Measured harness ledgerPublic result
Claude Opus 5

Dex Cube — 3 Move — Claude Opus 5 Max

Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.

Max reasoningHeadline result
Workflow cost
$22.55
Wall-clock
1h 42m 42s (38.7 min to the self-stop + 64 min after the single resume) wall-clock
Processed tokens
Not recorded
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5 Max judge_pass_independently_replayed ledger: 1h 42m 42s (38.7 min to the self-stop + 64 min after the single resume) wall-clock, Not recorded, and $22.55 API-equivalent estimate from session token counts at Anthropic list price.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-stage-2-three-moves-opus-5-max
  • Technical model: Claude Opus 5
  • Provider: Anthropic Claude Code
  • Client: claude -p --model claude-opus-5 --effort max (headless subagent), resumed once
  • Stack: Anthropic Claude Code
  • Stack: claude -p --model claude-opus-5 --effort max (headless subagent), resumed once
  • Stack: Technical model/configuration: Claude Opus 5
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn stage-2 prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-opus-5-max/evidence/model-supplied/ctrl.npy
  • SHA-256: 040c23f0ed05cf4070c866376cbc926c8ecd897b55380be4a214bc5d98e3656c
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-opus-5-max/validation/judge-replay.public.json
  • SHA-256: e9cf7f713c487176358d3ad4feb13e5f02b2d65a180c85b23b45fccdb13b7e35
  • Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
  • The judge measures task completion only (facelet state, alignment, contact pads, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/orchestrator-results.md is the orchestrator's own result table copied verbatim and retains the operator's workspace paths.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Opus cost is an API-list-price equivalent from token counts, not a subscription charge.
  • Cache-read and cache-write counts are recorded rounded (26.1M / 712k) because the exact figures were not archived; fresh input is not reported.
  • The run was resumed once by the orchestrator after a premature self-stop (see protocol.nudge_disclosure); the wall clock and cost cover both segments.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • exact cache-read/cache-write/fresh-input counts
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console