Measured harness ledgerPublic result
GPT-6 Astra

Claw Unlock — GPT-6 Astra Max

Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.

Max reasoningHeadline result
Workflow cost
$14.36
Wall-clock
53m 05s wall-clock
Processed tokens
8.19M processed
Record state
judge_pass_independently_replayed
Public summary

GPT-6 Astra Max judge_pass_independently_replayed ledger: 53m 05s wall-clock, 8.19M processed, and $14.36 Standard API-equivalent estimate from the Codex wrapper receipt, not a subscription invoice: this was a ChatGPT-subscription run, $0 was charged and Codex emits no cost field..

Run identity and stack
  • Result ID: claw-unlock-geometry-gpt-6-astra-max-codex-2026-09-19
  • Technical model: GPT-6 Astra
  • Provider: OpenAI Codex
  • Client: codex-sub wrapper (detached, headless, -l claw-astra -m gpt-6-astra -e max -s danger-full-access -t 10800), Codex CLI 0.154.0; ran to its own end_turn
  • Stack: OpenAI Codex
  • Stack: codex-sub wrapper (detached, headless, -l claw-astra -m gpt-6-astra -e max -s danger-full-access -t 10800), Codex CLI 0.154.0; ran to its own end_turn
  • Stack: Technical model/configuration: GPT-6 Astra
  • Stack: python-fcl geometry fixture
  • Stack: Harness claw-unlock prompt
  • Stack: Requested tool profile: python-fcl-geometry-headless
Cost basis
  • actual marginal subscription charge usd: $0.00 USD.
Primary artifact integrity
  • Kind: se3-rigid-body-path-polyline
  • Path: artifacts/claw-unlock-geometry-gpt-6-astra-max-codex-2026-09-19/evidence/model-supplied/path.npy
  • SHA-256: f7eab98c6fad074a27b14fb6fc7cc35c33856ba22326ee932cae2bb34c6debab
Validation evidence
  • Result: PASS
  • Path: artifacts/claw-unlock-geometry-gpt-6-astra-max-codex-2026-09-19/validation/judge-replay.public.json
  • SHA-256: 79846a26faabd4bf00e177132e5e114c010e76beadf1c643ec9e48646661899f
  • Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips for this result are NOT committed; they are pinned in provenance/clips.public.json by name, sha256, byte count, frame count and duration. Both clips were RE-RENDERED on 2026-09-20 with a route-independent OPTIMISED carry after the first render derived its two-arm carry from the published route, which is not Astra's route; the stock clawRig camera is used unchanged. PRESENTATION ONLY - the judged data is unchanged, the judged relative pose is preserved to 3.3e-16 per frame, and nothing in the dressing affects scoring. The known cosmetics of the optimised carry are listed under presentation.known_cosmetics.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Output tokens include hidden reasoning, visible prose and tool-call JSON; cache reads are discounted, so processed-token volume overstates effective cost.
  • This was a ChatGPT-subscription Codex run, so there is no dollar receipt: $0 was charged and Codex emits no cost field. The dollar figure is an API-equivalent estimate at the GPT-6 Astra list rates already on record in this harness, not an invoice.
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
  • The raw Codex receipts (events.jsonl, meta.json, prompt and final-message files) carry the session id, the full reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • The given harness/checker exposes the judge's own scoring in-process, so the solver could self-score without the operator; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
  • minClearance is reported, never gating: every known solution to this puzzle rides the contact surface, so the figure says little about quality.
  • The two claw meshes carry no established redistribution licence (see given/ATTRIBUTION.md); no licence is claimed for them and they will be removed on request.
  • The solver workspace carried a later revision of the non-scored given/ATTRIBUTION.md than the committed fixture: it gained a paragraph attributing the operator-side presentation layer. ATTRIBUTION.md is documentation, is not hashed into scene_sha256 and is not an input to the judge; every scored input matched the committed fixture byte for byte and the replay reproduced the verdict byte for byte.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • the two re-rendered presentation clips and their sha256 pins (a follow-up commit)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console