Skip to content

Commit 9138df0

Browse files
committed
ci: name the module the Windows coverage crash dies in
The first round of diagnostics placed the fault: the crash log holds `open pid=...` and `entering pytest.main` and nothing else, so the process dies while pytest loads its initial conftests -- before collection, and far before any test runs. What it cannot say is which module. Three additions, none of which touch machine state: PYTHONVERBOSE makes the interpreter announce every import as it begins one, so the last line written before the process dies names the module it died in. Its output is stderr, so the session is redirected to a file rather than left in the step log: a crash takes the tail of a pipe with it, while each unbuffered write to a file has already reached the OS. The Windows event log carries an Application Error record for every process the OS kills, naming the faulting module and the offset within it, and it does so whether or not crash dumps are configured. That is the one fact no Python-level log can produce. Reading it needs no registry key set and nothing restored afterwards, which a LocalDumps setup would have on a machine shared with other jobs. PYTHONFAULTHANDLER has so far written nothing at all, which leaves two very different readings: the handler produced a traceback the dying process failed to flush, or it never ran because the fault happened outside any Python frame. run_pytest_with_stack.py now points faulthandler at its own unbuffered file when PYTEST_FAULTLOG is set, so still empty there means the handler genuinely had nothing to say. The three logs upload together as an artifact. No behaviour change: the step already carries continue-on-error, and every addition is inert on a run that does not crash. Signed-off-by: Rui Luo <ruluo@nvidia.com>
1 parent 69822ac commit 9138df0

2 files changed

Lines changed: 88 additions & 2 deletions

File tree

‎.github/workflows/coverage.yml‎

Lines changed: 71 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -390,17 +390,28 @@ jobs:
390390
# Only this step crashes. cuda.bindings runs the same wrapper, the same
391391
# 8 MB thread and the same coverage wheel, and finishes normally, so the
392392
# extra diagnostics are scoped to cuda.core alone.
393+
#
394+
# The crash log narrows the fault to "somewhere before pytest collects".
395+
# PYTHONVERBOSE narrows it to a module: the interpreter announces every
396+
# import as it begins one, so the last line written before the process
397+
# dies names the module it died in. That output is stderr, which is why
398+
# the session is redirected to a file rather than left in the step log --
399+
# a crash takes the tail of a pipe with it, while each unbuffered write
400+
# to a file has already reached the OS.
393401
- name: Run cuda.core tests (with 8MB stack)
394402
continue-on-error: true
395403
env:
396404
PYTEST_CRASHLOG: ${{ github.workspace }}/crashlog-core.txt
405+
PYTEST_FAULTLOG: ${{ github.workspace }}/faulthandler-core.txt
406+
PYTHONVERBOSE: "1"
397407
run: |
398408
"$GITHUB_WORKSPACE/.venv/Scripts/python" "$GITHUB_WORKSPACE/ci/tools/run_pytest_with_stack.py" \
399409
--isolate \
400410
--cwd "${{ steps.install-root.outputs.INSTALL_ROOT }}" \
401411
-v --cov=./cuda --cov-append --cov-context=test \
402412
--cov-config="$GITHUB_WORKSPACE/.coveragerc" \
403-
"$GITHUB_WORKSPACE/cuda_core/tests"
413+
"$GITHUB_WORKSPACE/cuda_core/tests" \
414+
> "$GITHUB_WORKSPACE/verbose-core.txt" 2>&1
404415
405416
# The step above has been exiting 139 on every run, and Git Bash reports
406417
# that as a bare "Segmentation fault". The crash log's final line says
@@ -411,12 +422,70 @@ jobs:
411422
run: |
412423
log="$GITHUB_WORKSPACE/crashlog-core.txt"
413424
if [ -f "$log" ]; then
414-
echo "=== $(wc -l < "$log") lines, last 20 ==="
425+
echo "=== crashlog: $(wc -l < "$log") lines, last 20 ==="
415426
tail -n 20 "$log"
416427
else
417428
echo "No crash log: died before pytest started."
418429
fi
419430
431+
echo ""
432+
fault="$GITHUB_WORKSPACE/faulthandler-core.txt"
433+
if [ -s "$fault" ]; then
434+
echo "=== faulthandler ==="
435+
cat "$fault"
436+
else
437+
# Empty is itself a result: the fault never reached Python's
438+
# handler, which points at native code running outside any Python
439+
# frame -- an extension module's initialisation, for instance.
440+
echo "=== faulthandler: no output ==="
441+
fi
442+
443+
echo ""
444+
verbose="$GITHUB_WORKSPACE/verbose-core.txt"
445+
if [ -f "$verbose" ]; then
446+
echo "=== PYTHONVERBOSE: $(wc -l < "$verbose") lines, last 40 ==="
447+
tail -n 40 "$verbose"
448+
fi
449+
450+
# Windows records an Application Error entry for every process it kills,
451+
# naming the faulting module and the offset within it, and it does so
452+
# whether or not crash dumps are configured. That is the one fact no
453+
# Python-level log can produce, and reading the log changes no machine
454+
# state, so nothing has to be restored afterwards.
455+
- name: Read the crash record from the Windows event log
456+
if: always()
457+
continue-on-error: true
458+
shell: powershell
459+
run: |
460+
$events = Get-WinEvent -FilterHashtable @{
461+
LogName = 'Application'
462+
ProviderName = 'Application Error', 'Windows Error Reporting'
463+
StartTime = (Get-Date).AddHours(-2)
464+
} -ErrorAction SilentlyContinue
465+
if (-not $events) {
466+
Write-Output "No Application Error / WER records in the last 2 hours."
467+
Write-Output "(If this image keeps no such log, a minidump is the fallback.)"
468+
exit 0
469+
}
470+
foreach ($e in $events | Select-Object -First 15) {
471+
Write-Output "===== $($e.TimeCreated.ToString('u')) $($e.ProviderName) Id=$($e.Id) ====="
472+
Write-Output $e.Message
473+
Write-Output ""
474+
}
475+
476+
- name: Upload crash diagnostics
477+
if: always()
478+
continue-on-error: true
479+
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
480+
with:
481+
name: crash-diagnostics-windows
482+
path: |
483+
crashlog-core.txt
484+
faulthandler-core.txt
485+
verbose-core.txt
486+
retention-days: 7
487+
if-no-files-found: warn
488+
420489
- name: Copy Windows coverage file to workspace
421490
run: |
422491
cp "${{ steps.install-root.outputs.INSTALL_ROOT }}/.coverage" "$GITHUB_WORKSPACE/.coverage.windows"

‎ci/tools/run_pytest_with_stack.py‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -95,6 +95,23 @@ def main():
9595
if args.isolate:
9696
sys.exit(_run_isolated(args.stack_mb, args.cwd, pytest_args))
9797

98+
# PYTHONFAULTHANDLER writes to stderr, and the nightly's stderr has so far
99+
# carried nothing at all -- which leaves two very different readings: the
100+
# handler produced a traceback that the dying process failed to flush, or
101+
# it never ran because the fault happened outside any Python frame. A
102+
# dedicated unbuffered file separates them: still empty here means the
103+
# handler genuinely had nothing to say. The handle is deliberately kept
104+
# alive for the life of the process; faulthandler writes through its fd.
105+
if os.environ.get("PYTEST_FAULTLOG"):
106+
import faulthandler
107+
108+
try:
109+
_fault_fp = open(os.environ["PYTEST_FAULTLOG"], "w", buffering=1) # noqa: SIM115
110+
except OSError as exc:
111+
print(f"[stack-wrapper] cannot open fault log: {exc}", file=sys.stderr, flush=True)
112+
else:
113+
faulthandler.enable(file=_fault_fp, all_threads=True)
114+
98115
plugins = []
99116
if os.environ.get("PYTEST_CRASHLOG"):
100117
# Load the progress logger from this script's directory, which is not

0 commit comments

Comments
 (0)