Skip to content

fmtk e2e: restart_app recovers an already-dead app (#3469) - #3470

Closed
mcdonc wants to merge 2 commits into
mainfrom
i3469-fmtk-restart-dead-app
Closed

mcdonc wants to merge 2 commits into
mainfrom
i3469-fmtk-restart-dead-app

Conversation

@mcdonc

@mcdonc mcdonc commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Summary

The nightly flows suite has failed every run since Sep 10 with a constant
signature: the whole test_network.py module dies in a "No Flutter isolate
found" cascade. test_auth.py ends by booting the app at a deep link, which
leaves an instance that dies shortly after (documented in that scenario's
own comment); the first network scenario then calls
Harness.restart_app() — the designated recovery — and dies inside the
restart's own pre-restart drain, because app_errors() fails on the
already-dead app before the stop/relaunch can run. The failure then cascades
through the module (details and the 10-run evidence trail in the issue).

The fix: the pre-restart drain treats the gone-isolate state as an empty
drain — an app with no isolate has no error-monitor window left to launder —
and restart_app() clears the wedge marker before relaunching (the marker
describes the outgoing instance). Both are the same trade
FlutterRun.recover_from_wedge already documents: the alternative is every
subsequent test failing on a dead debug connection (#3231). Real app errors
and every other drain failure still fail the restart, so the
anti-laundering guarantee is intact.

Adds a pin in the core suite's test_harness_recovery.py: patch
app_errors to raise the synthetic gone-isolate failure (the same idiom
the wedge legs use — killing Chrome for real on CI takes the VM service
down with it, a different dead state), then restart_app() must relaunch
a drivable app with the marker cleared.

Closes #3469.

Validation

  • Dispatched fmtk-e2e-flows on this branch (run 35440973393): the
    network module passed for the first time since Sep 9 — the constant
    10-night cascade is gone. The three remaining failures in that run are
    the pre-existing intermittent class (admin Handle list-render probe on
    slow first boots; an SSO landing), separate from this fix.
  • Dispatched fmtk-e2e-core (run 35440971339) caught the first version of
    the new pin killing Chrome for real (it poisoned the session); the
    reworked pin fabricates the state synthetically and the re-dispatched
    core run (35442089582) is green: 18 passed, including the new pin.
  • Stubbed-logic check of the drain against both gone-isolate failure shapes
    (envelope-less with the mark in stderr; envelope-not-ok with the mark in
    the message), a non-gone fmtk failure (still raises), real app errors
    (still raise — anti-laundering intact), and at_path behavior
    (unchanged).
  • ruff + pre-commit clean; new code at xenon rank A (the harness file sits
    outside the graded set).

The flows nightly lost its whole network module ten runs running:
test_auth.py ends by booting the app at a deep link, which leaves an
instance that dies shortly after, and the first network scenario's
restart_app() died in its own pre-restart drain (app_errors() fails
with 'No Flutter isolate found') before the stop/relaunch could run.

The drain now treats the gone-isolate state as an empty drain — an app
with no isolate has no error-monitor window left to launder (the same
trade recover_from_wedge documents) — and the restart clears the wedge
marker, which describes the outgoing instance. Real app errors and
other drain failures still fail the restart.

Pinned live in test_harness_recovery.py (core suite): kill Chrome, then
restart_app() must relaunch a drivable app.
@github-actions github-actions Bot added the backport/2.0 Merge also backports the squash commit to stable/2.0 (#3361) label Sep 19, 2026
Killing Chrome for real takes the VM service down with it on CI
(connection refused) — a different dead state than the nightly's
(service answers, no isolate) — and the first dispatch run's scenario
poisoned the rest of the core session with it. Raise the synthetic
gone error from app_errors instead (the same idiom the wedge legs
use) and keep the real stop/relaunch leg of the pin.
@mcdonc

mcdonc commented Oct 8, 2026

Copy link
Copy Markdown
Owner Author

Closing as superseded: #3539 ("fmtk e2e: restart_app tolerates a dead isolate in its pre-restart drain (#3469)") landed the same fix while this PR was open, and issue #3469 is already closed by it. Main's version covers the same behavior this PR implemented — the pre-restart drain treats a gone-isolate app as an empty drain (nothing left to launder), the restart proceeds, real drain errors and other failures still raise, and the fabricated dead-drain scenario pins it without killing Chrome.

The rebase onto current main conflicts in fmtkharness.py / test_harness_recovery.py purely on the two competing implementations (drain_errors_before_restart vs drain_app_errors), with no unique behavior left on this branch to preserve.

@mcdonc mcdonc closed this Oct 8, 2026
@mcdonc
mcdonc deleted the i3469-fmtk-restart-dead-app branch October 8, 2026 16:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport/2.0 Merge also backports the squash commit to stable/2.0 (#3361)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fmtk e2e (flows): restart_app's pre-restart drain fails on an already-dead app — network module cascades every night

1 participant