Measured harness ledgerPublic result
GPT-6 SolClaw Unlock — GPT-6 Sol Max
Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.
Max reasoningHeadline result
- Workflow cost
- $14.06
- Wall-clock
- 108m 35s wall-clock
- Processed tokens
- 58.32M processed
- Record state
- judge_fail_independently_replayed
Public summary
GPT-6 Sol Max judge_fail_independently_replayed ledger: 108m 35s wall-clock, 58.32M processed, and $14.06 Standard API-equivalent estimate from the Codex wrapper receipt, not a subscription invoice: this was a ChatGPT-subscription run, $0 was charged and Codex emits no cost field..
Run identity and stack
- Result ID: claw-unlock-geometry-gpt-6-sol-max-codex-2026-09-22
- Technical model: GPT-6 Sol
- Provider: OpenAI Codex
- Client: codex-sub wrapper (detached, headless, -l claw-sol -m gpt-6-sol -e max -s danger-full-access), Codex CLI 0.155.1; ran to its own end_turn
- Stack: OpenAI Codex
- Stack: codex-sub wrapper (detached, headless, -l claw-sol -m gpt-6-sol -e max -s danger-full-access), Codex CLI 0.155.1; ran to its own end_turn
- Stack: Technical model/configuration: GPT-6 Sol
- Stack: python-fcl geometry fixture
- Stack: Harness claw-unlock prompt
- Stack: Requested tool profile: python-fcl-headless
Cost basis
- actual marginal subscription charge usd: $0.00 USD.
Primary artifact integrity
- Kind: se3-path
- Path: artifacts/claw-unlock-geometry-gpt-6-sol-max-codex-2026-09-22/evidence/model-supplied/path.npy
- SHA-256: 8844b5dec4823f8e25ea18dd976c251fddef53b7c23e52afc5ae90c89653fcd5
Validation evidence
- Result: PASS
- Path: artifacts/claw-unlock-geometry-gpt-6-sol-max-codex-2026-09-22/validation/judge-replay.public.json
- SHA-256: 8b6bd5ae35d160d9ce37c8075a3fbd60f825d3ddb2537defa315e09bdd7d2a62
- Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- This was a ChatGPT-subscription Codex run, so there is no dollar receipt: $0 was charged and Codex emits no cost field. API-equivalent is $14.06 at the GPT-6 Sol list rates (corrected 2026-09-24; the original registration wrongly stated no rate was on record).
- provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
- The raw Codex receipts (events.jsonl, meta.json, prompt and final-message files) carry the session id, the full reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
- A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods (US), 62.8 + 144.6 min, $0.77 + $1.78 + $0.65 aborted batch = $3.21 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
- Nudge disclosure: none.
- The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
- Claw clips are dressed ALOHA 2 presentation (optimised carry, stock clawRig camera). PRESENTATION ONLY. Nothing judged changes.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.