Measured harness ledgerPublic result
GPT-6 AstraDex Cube Turn — GPT-6 Astra Low
Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.
Low reasoningHeadline result
- Workflow cost
- $1.08
- Wall-clock
- 3m 01s wall-clock
- Processed tokens
- 603K processed
- Record state
- judge_pass_independently_replayed
Public summary
GPT-6 Astra Low judge_pass_independently_replayed ledger: 3m 01s wall-clock, 603K processed, and $1.08 Standard API-equivalent estimate from the Codex wrapper receipt, not a subscription invoice: this was a ChatGPT-subscription run, $0 was charged and Codex emits no cost field..
Run identity and stack
- Result ID: dex-cube-turn-mujoco-gpt-6-astra-low-codex-2026-09-21
- Technical model: GPT-6 Astra
- Provider: OpenAI Codex
- Client: codex-sub wrapper (detached, headless, -l dexcube-astra-low -m gpt-6-astra -e low -s danger-full-access -t 7200), Codex CLI 0.154.0; ran to its own end_turn
- Stack: OpenAI Codex
- Stack: codex-sub wrapper (detached, headless, -l dexcube-astra-low -m gpt-6-astra -e low -s danger-full-access -t 7200), Codex CLI 0.154.0; ran to its own end_turn
- Stack: Technical model/configuration: GPT-6 Astra
- Stack: MuJoCo dexterous-hand fixture
- Stack: Harness cube-turn prompt
- Stack: Requested tool profile: python-mujoco-headless
Cost basis
- actual marginal subscription charge usd: $0.00 USD.
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-cube-turn-mujoco-gpt-6-astra-low-codex-2026-09-21/evidence/model-supplied/ctrl.npy
- SHA-256: 4988e6fa29a4bde5405d727e5fcc4b3f521a29254c82fc0b94fe5a5f62a2da3f
Validation evidence
- Result: PASS
- Path: artifacts/dex-cube-turn-mujoco-gpt-6-astra-low-codex-2026-09-21/validation/judge-replay.public.json
- SHA-256: 60137ec9c317bfd503fc882e67253e03d5b613d72a9ae6b0e636e54310477a73
- Validator SHA-256: 5f7750274d345db503f3320931fedc68d0d3d01573dd5a88052910b9213d7203
Recorded caveats
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- Output tokens include hidden reasoning, visible prose and tool-call JSON; cache reads are discounted, so processed-token volume overstates effective cost.
- This was a ChatGPT-subscription Codex run, so there is no dollar receipt: $0 was charged and Codex emits no cost field. The dollar figure is an API-equivalent estimate at the GPT-6 Astra list rates already on record in this harness, not an invoice.
- provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
- The raw Codex receipts (events.jsonl, meta.json, prompt and final-message files) carry the session id, the full reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
- Tier-1 solutions exploit the pre-positioned D grasp; tier 1 does not test finger placement, which is why the stage-2 task exists. Do not compare the tiers.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.