Measured harness ledgerPublic result
Claude Opus 5.5

Claw Unlock — Claude Opus 5.5 Max

Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.

Max reasoningHeadline result
Workflow cost
$17.39
Wall-clock
93m 17s wall-clock
Processed tokens
43.03M processed
Record state
judge_fail_independently_replayed
Public summary

Claude Opus 5.5 Max judge_fail_independently_replayed ledger: 93m 17s wall-clock, 43.03M processed, and $17.39 recorded total_cost_usd from workspace FINAL.json (modelUsage.claude-opus-5-5, costBasis=list, Anthropic list price). Charged to the Claude subscription ($0 charged)..

Run identity and stack
  • Result ID: claw-unlock-geometry-opus-5.5-max-2026-09-22
  • Technical model: Claude Opus 5.5
  • Provider: Anthropic Claude Code
  • Client: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Anthropic Claude Code
  • Stack: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Technical model/configuration: Claude Opus 5.5
  • Stack: python-fcl geometry fixture
  • Stack: Harness claw-unlock prompt
  • Stack: Requested tool profile: python-fcl-geometry-headless
Primary artifact integrity
  • Kind: se3-rigid-body-path-polyline
  • Path: artifacts/claw-unlock-geometry-opus-5.5-max-2026-09-22/evidence/model-supplied/path.npy
  • SHA-256: d3e100c8a12f1731dacfec76284e6ff2a21cfb8fe1135aff9a51ecfdf662e3d9
Validation evidence
  • Result: PASS
  • Path: artifacts/claw-unlock-geometry-opus-5.5-max-2026-09-22/validation/judge-replay.public.json
  • SHA-256: a4dbde5d1084e7dfcea1a714ac7b7cc738528c9703f0100ac950cf9d9bce8d8c
  • Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Cost is the recorded list-price total_cost_usd from workspace FINAL.json (costBasis=list). Charged to the Claude subscription ($0 charged).
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The raw Claude CLI receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EU-CZ-1, 105.6 + 65.4 min, $1.30 + $0.81 = $2.11 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
  • Nudge disclosure: none.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
  • The published clips are dressed (ALOHA 2 presentation rig); the judged replay in evidence/model-supplied/replay/ is the judge's own two-body export. Nothing judged changes.
  • Clips are not committed for this task.
  • The given checker runs the judge's own judge() in-process, so the solver could self-score exactly without the operator; wall clock and cost are NOT comparable with the physics tasks in this suite.
  • Environment caveat: solver noted host load average up to ~325 from concurrent jobs.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console