Skip to content

Auto-retry fuzz-daily runs killed by hosted-runner shutdown (#3471) - #3472

Merged
mcdonc merged 3 commits into
mainfrom
i3471-fuzz-daily
Oct 8, 2026
Merged

mcdonc merged 3 commits into
mainfrom
i3471-fuzz-daily

Conversation

@mcdonc

@mcdonc mcdonc commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Summary

Closes #3471.

Adds fuzz-daily-retry.yml, a small companion workflow that automatically re-runs the daily fuzz job when the failure is the hosted-runner shutdown signature from the issue — 6 of the 10 daily runs between Sep 10 and Sep 19, 2026 died with minutes of total silence, then "The runner has received a shutdown signal" and step exit 143, leaving the daily fuzz signal red with no actionable product cause.

How it works:

  • Triggers on workflow_run completions of "Fuzz API: daily run". A retry job inside fuzz-daily.yml itself cannot work: rerunFailedJobs requires the run to be completed, and the run is not complete while one of its own jobs is still executing.
  • Re-runs the fuzz job only on the kill signature: the fuzz job failed and its if: always() "Upload server log" step ended skipped — a runner shutdown skips post-steps (verified against the Sep 10–19 failure logs). A genuine fuzz-anomaly failure runs the upload step and is never retried, because a rerun draws a fresh random seed and could mask a real finding with a passing attempt.
  • Bounds retries with run_attempt < 3 (original attempt plus two reruns); after that the failure stands for human triage.
  • Grants exactly actions: write for the rerunFailedJobs call.

scripts/tests/test_fuzz_daily_retry.py pins the couplings that would otherwise drift silently: the trigger name matching fuzz-daily.yml's name: key, the attempt bound, the job/step names the signature keys on, and the narrow permission grant.

Testing

  • python -m pytest scripts/tests -v — 168 passed (4 new).
  • pre-commit run xenon --all-files — passed.
  • actionlint .github/workflows/fuzz-daily-retry.yml — clean.

CI-only change with no operator-visible effect, so no changelog entry (matches #3416).

Rider: file-sink suspend fix (#3551)

CI on this PR hit the Python 3.14.8 stdlib change twice (3.14.8 swallows WatchedFileHandler reopen failures into handleError), which had disabled RotationSafeFileHandler's suspend-on-failure path entirely. emit() now runs the reopen check inside its own try; a regression test installs the 3.14.8 swallowing shape to pin the behavior on any toolchain. Closes #3551.

@github-actions github-actions Bot added the backport/2.0 Merge also backports the squash commit to stable/2.0 (#3361) label Sep 19, 2026
@mcdonc

mcdonc commented Sep 19, 2026

Copy link
Copy Markdown
Owner Author

Fresh-eyes review

Reviewer verdict: APPROVE. The important item and nits 1, 2, and 4 are addressed in 234f73b; nit 3 is covered by the new 'deliberate gap' comment.

Review: PR #3472 — auto-retry fuzz-daily runner kills

I verified the load-bearing GitHub Actions semantics against the docs and real-world precedent before judging the code: rerunFailedJobs is permitted on schedule-triggered completed runs (docs limit it only to ≤30 days old and ≤50 reruns); workflow_run completed does re-fire per rerun attempt with an incremented run_attempt (this exact bounded pattern exists in other repos, and gptme#2777 shows the per-attempt re-fire); a job-level if can read github.event.workflow_run.*; actions: write is the documented minimal grant and implies the read needed for listJobsForWorkflowRun; github.paginate handles >100 jobs. The attempt arithmetic is correct: attempts 1 and 2 trigger a rerun, attempt 3 does not — exactly "original plus two reruns." Duplicate event delivery is serialized by the concurrency group and absorbed by the 422 path. actions/github-script@v9.0.0 exists (the repo has no prior github-script usage; tag-pinning matches the checkout@v7 style). actionlint, yamllint --strict, ruff, and the xenon gate all pass on the new files; the new test passes.

Blocking issues

None found.

Important issues

  1. 422 is swallowed with a possibly-wrong message — fuzz-daily-retry.yml:87-93. rerunFailedJobs returns 422 for any refusal (not only "already re-running": non-rerunnable state, past caps), and every such case is logged as run already re-running and the job exits green. This mechanism's whole design goal is to not fail silently, so log the response body or err.message inside the 422 branch too. Likelihood is low (the gate bounds attempts to ≤3 same-day), but it is a one-line fix.

Nits / questions

  1. Attempt bound not pinned — scripts/tests/test_fuzz_daily_retry.py:85 asserts only "run_attempt <". Someone loosening the policy from 3 to 99 attempts passes the suite. Assert "run_attempt < 3" if two-retries-then-triage is the policy.
  2. Wrong mechanism in the comments — fuzz-daily-retry.yml:8-9 (and the test docstring) say "a runner shutdown skips post-steps." The "Upload server log" step is an ordinary step, not a post step: the VM is gone, so remaining steps never start and GitHub reports them skipped. The signature is right; the explanation is not. A future maintainer extending the heuristic from "post-steps" reasoning would go wrong.
  3. Kill during the upload step is missed — if the runner dies in the seconds after the fuzzer fails genuinely (or during upload on a healthy run), the upload step concludes failure/cancelled, not skipped, so no retry fires. That is the fail-safe direction (red X stays visible, no masking), and the window is seconds against a ~25-min job — fine, but worth a comment saying it is deliberate.
  4. workflows scalar hazard — test_fuzz_daily_retry.py:74: if workflows: were ever rewritten to a bare string, in silently becomes a substring match (false positive). Add isinstance(trigger["workflows"], list).
  5. Inherent, no action needed: the retry workflow creates a run entry (job skipped) for every daily completion, successes included — unavoidable with workflow_run; and the retry job itself can be runner-killed, in which case that day's retry is lost and the failure stays visible for triage.

Conventions check: PR title carries (#3471), body has Closes #3471, no changelog entry needed (CI-only churn, policy says skip), test_ci_path_filters.py needs no entry (no pull_request trigger), and the test's file-contract style matches test_ci_path_filters.py.

Verdict

APPROVE — the mechanism is sound and every GitHub-Actions assumption it rests on checks out; the one Important item is a diagnosability fix, not a correctness bug.

mcdonc added 2 commits October 8, 2026 10:51
6 of the 10 daily fuzz runs between Sep 10 and Sep 19, 2026 died the
same way: minutes of total silence from both server and fuzzer, then
"The runner has received a shutdown signal" and step exit 143 — the
hosted ubuntu-latest VM reclaimed mid-fuzz, well under the 40-minute
job timeout and before the fuzzer's own deadline. The daily fuzz
signal was red ~60% of the time with no actionable product cause.

fuzz-daily-retry.yml listens for completions of "Fuzz API: daily run"
and re-runs the fuzz job when the kill signature is present: the fuzz
job failed and its if: always() "Upload server log" step ended
skipped (a runner shutdown skips post-steps). A genuine
fuzz-anomaly failure runs the upload step and is never retried — a
rerun draws a fresh random seed and could mask a real finding with a
passing attempt. The gate allows the original attempt plus two reruns
(run_attempt < 3); after that the failure stands for triage.

A retry job inside fuzz-daily.yml cannot work: rerunFailedJobs
requires a completed run, and the run is not complete while one of its
own jobs is executing — hence the separate workflow_run-triggered
workflow, with actions: write scoped to exactly that call.

scripts/tests/test_fuzz_daily_retry.py pins the couplings that would
otherwise drift silently: the trigger name matching fuzz-daily.yml's
name: key, the attempt bound, the job/step names the signature keys
on, and the narrow permission grant.
- The 422 catch logs err.message and the API's message, so any rerun
  refusal that is not a duplicate-delivery race is visible in the retry
  run's log instead of a fixed 'run already re-running'.
- The contract test pins 'run_attempt < 3' exactly and requires
  workflow_run.workflows to stay a list (a bare string would turn the
  membership check into a substring match).
- Comments and the test docstring describe the skip mechanism
  accurately: the runner VM is gone with the job, so every remaining
  step — the if: always() upload step included — never starts and is
  reported skipped; and the inverse gap (runner dies during the upload
  step: failure/cancelled, no retry, red run stays visible) is stated
  as deliberate.
Python 3.14.8 wraps WatchedFileHandler.emit's reopenIfNeeded in its
own try/except and routes failures to handleError, so a hostile log
path (file replaced by a directory, rotated volume gone) no longer
raises out of super().emit(). RotationSafeFileHandler's suspend path
therefore never engaged on 3.14.8: _sink_broken stayed False, the
one-shot suspension warning was never emitted, and every record
retried the failing reopen with a stderr traceback — the exact
behavior the suspend design exists to prevent. CI hit this as a
deterministic failure of test_reopen_failure_suspends_sink_not_the_
call_site (Python 3.14.8 on runners vs 3.14.7 in the devenv).

emit() now calls reopenIfNeeded() itself, inside its own try, and
delegates to FileHandler.emit directly (the stdlib's check inside
WatchedFileHandler.emit becomes a no-op after a successful reopen,
and on failure ours raises first). The recursion-contract test moves
its patch target from WatchedFileHandler.emit to FileHandler.emit —
the watched wrapper is no longer on the call path — and a new test
installs the 3.14.8 swallowing shape of the stdlib emit to pin the
suspend behavior on any toolchain.
@mcdonc
mcdonc merged commit 0d7490a into main Oct 8, 2026
11 checks passed
@mcdonc
mcdonc deleted the i3471-fuzz-daily branch October 8, 2026 16:31
@mcdonc

mcdonc commented Oct 8, 2026

Copy link
Copy Markdown
Owner Author

Created backport PR for stable/2.0:

Please cherry-pick the changes locally and resolve any conflicts.

git fetch origin backport-3472-to-stable/2.0
git worktree add --checkout .worktree/backport-3472-to-stable/2.0 backport-3472-to-stable/2.0
cd .worktree/backport-3472-to-stable/2.0
git reset --hard HEAD^
git cherry-pick -x 0d7490a92cd7210da08a88d91463074b0af86435

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport/2.0 Merge also backports the squash commit to stable/2.0 (#3361)

Projects

None yet

1 participant