Measured harness ledgerPublic result
GPT-5.6 LunaApple Carry — GPT-5.6 Luna Low
Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.
Low reasoningHeadline result
- Workflow cost
- $16.09
- Wall-clock
- Not recorded
- Processed tokens
- Not recorded
- Record state
- judge_verified_decision_loop_sweep_independently_replayed
Public summary
GPT-5.6 Luna Low judge_verified_decision_loop_sweep_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $16.09 Recorded provider cost.
Run identity and stack
- Result ID: apple-carry-xarm7-openroboto-gpt-5.6-luna-low-decision-loop-codex-2026-09-20
- Technical model: GPT-5.6 Luna
- Provider: OpenAI Codex (ChatGPT subscription)
- Client: codex exec -m gpt-5.6-luna -c model_reasoning_effort=low --output-schema, one process per decision
- Stack: OpenAI Codex (ChatGPT subscription)
- Stack: codex exec -m gpt-5.6-luna -c model_reasoning_effort=low --output-schema, one process per decision
- Stack: Technical model/configuration: GPT-5.6 Luna
- Stack: Embodied manipulation fixture
- Stack: Harness apple-carry prompt
Validation evidence
- Path: artifacts/apple-carry-xarm7-openroboto-gpt-5.6-luna-low-decision-loop-codex-2026-09-20/validation/judge-replay.public.json
- SHA-256: 66f793c08c9a17c6cbb88bd4d892323bbcb7cd6d07541ae328530aea6de1b518
- Validator SHA-256: 5facf7bb27541fca9a60dbaaded91151a3e4f9962419bfd55877bd0b49db5d81
Recorded caveats
- NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason.
- 20/20 PASS. $0 charged, API-equivalent $16.093187.
- A superseded seed-0 directory (seed00-superseded-call-accounting) is disclosed and not scored: the wrapper counted a retried attempt in api_calls but not model_requests.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.