Measured harness ledgerPublic result
GPT-5.6 Luna

Apple Carry — GPT-5.6 Luna Low

Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.

Low reasoningHeadline result
Workflow cost
$9.59
Wall-clock
Not recorded
Processed tokens
Not recorded
Record state
judge_verified_decision_loop_sweep_independently_replayed
Public summary

GPT-5.6 Luna Low judge_verified_decision_loop_sweep_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $9.59 Not an invoice: $0 charged. API-equivalent is at the harness on-record Luna rates..

Run identity and stack
  • Result ID: apple-carry-fr3-gripper-gpt-5.6-luna-low-decision-loop-codex-2026-09-20
  • Technical model: GPT-5.6 Luna
  • Provider: OpenAI Codex (ChatGPT subscription)
  • Client: codex exec -m gpt-5.6-luna -c model_reasoning_effort=low --output-schema, one process per decision
  • Stack: OpenAI Codex (ChatGPT subscription)
  • Stack: codex exec -m gpt-5.6-luna -c model_reasoning_effort=low --output-schema, one process per decision
  • Stack: Technical model/configuration: GPT-5.6 Luna
  • Stack: Embodied manipulation fixture
  • Stack: Harness apple-carry prompt
Validation evidence
  • Path: artifacts/apple-carry-fr3-gripper-gpt-5.6-luna-low-decision-loop-codex-2026-09-20/validation/judge-replay.public.json
  • SHA-256: fe17f409635827415b58e07d17f064cfa77a62b7eb958a7c625e622a785749a8
  • Validator SHA-256: e92e63eb3a4c358a423e88d14badd90f1fddbabbf97d0e55857a29d14d939ae9
Recorded caveats
  • NOT COMPARABLE WITH THE CODING-AGENT COHORTS. Every coding-agent result in this harness came from an agent with a shell, the fixture workspace, hours of wall clock and unlimited self-play against the simulator, submitting a controller it had already debugged. This is a decision-loop run: the model saw one structured observation at a time, answered two typed multiple-choice questions per control cycle (one intent, then the per-channel directions), and never wrote code, never ran the simulator, never saw an image and never read the judge. The executor that performed the increments is HARNESS-AUTHORED, not model-authored. Cycles, cost, wall clock and outcome mean different things in the two tracks and must never be put in one league table. cohort_eligible is false for exactly this reason.
  • TWENTY SEEDS, ONE EPISODE EACH.
  • Cost $0 charged. API-equivalent is an estimate at harness Luna rates, never mixed into cost_usd.
  • Seeds 1 and 13 FAIL stage 0 on malformed answers (choice not argmax), not on a physical failure.
  • The superseded xArm7 seed-0 call-accounting run is disclosed in the xarm ledger, not here.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console