Measured harness ledgerPublic result
Claude Opus 5.5

Dex Cube Turn — Claude Opus 5.5 Max

Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.

Max reasoningHeadline result
Workflow cost
$2.20
Wall-clock
9m 41s wall-clock
Processed tokens
1.97M processed
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5.5 Max judge_pass_independently_replayed ledger: 9m 41s wall-clock, 1.97M processed, and $2.20 recorded total_cost_usd from workspace FINAL.json (modelUsage.claude-opus-5-5, costBasis=list, Anthropic list price). Charged to the Claude subscription ($0 charged)..

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-opus-5.5-max-2026-09-22
  • Technical model: Claude Opus 5.5
  • Provider: Anthropic Claude Code
  • Client: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Anthropic Claude Code
  • Stack: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Technical model/configuration: Claude Opus 5.5
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-opus-5.5-max-2026-09-22/evidence/model-supplied/ctrl.npy
  • SHA-256: 51d848958861d13d773eab76bee5491f687dab7e40848a3ada7328588cd07617
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-cube-turn-mujoco-opus-5.5-max-2026-09-22/validation/judge-replay.public.json
  • SHA-256: 8c28db47a0fcb79e43f5b512a649306853d2febc0d8cf83d0abe396331f0ee23
  • Validator SHA-256: 5f7750274d345db503f3320931fedc68d0d3d01573dd5a88052910b9213d7203
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Cost is the recorded list-price total_cost_usd from workspace FINAL.json (costBasis=list). Charged to the Claude subscription ($0 charged).
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The raw Claude CLI receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EU-CZ-1, 105.6 + 65.4 min, $1.30 + $0.81 = $2.11 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
  • Nudge disclosure: none.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console