Measured harness ledgerPublic result
Claude Opus 5

Dex Cube Turn — Claude Opus 5 Max

Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.

Max reasoningHeadline result
Workflow cost
$5.83
Wall-clock
20m 36s wall-clock
Processed tokens
5.86M processed
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5 Max judge_pass_independently_replayed ledger: 20m 36s wall-clock, 5.86M processed, and $5.83 API-equivalent estimate from session token counts at Anthropic list price.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-opus-5-max
  • Technical model: Claude Opus 5
  • Provider: Anthropic Claude Code
  • Client: claude -p --model claude-opus-5 --effort max (headless subagent)
  • Stack: Anthropic Claude Code
  • Stack: claude -p --model claude-opus-5 --effort max (headless subagent)
  • Stack: Technical model/configuration: Claude Opus 5
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-opus-5-max/evidence/model-supplied/ctrl.npy
  • SHA-256: ac81f1e375c66a7b08d803d3eb6f6aa7c50b4b06320855340805e665b027f03c
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-cube-turn-mujoco-opus-5-max/validation/judge-replay.public.json
  • SHA-256: 1af5fb4000e633fd56728361a0ae5216d785015247c46bdb20bf14c3c84a8354
  • Validator SHA-256: 8425328aed969dfd11cc1e5107b1e0c83b888c59675bdf2ce86c4bc2878dcee9
Recorded caveats
  • The judge measures task completion only (facelet state, alignment, contact pads, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/orchestrator-results.md is the orchestrator's own result table copied verbatim and retains the operator's workspace paths.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Opus cost is an API-list-price equivalent from token counts, not a subscription charge.
  • Both tier-1 cohort solutions exploit the pre-positioned D grasp; tier 1 does not test finger placement, which is why the stage-2 task exists. Do not compare the tiers.
  • Reconstructed attempts are this run's own progression, recovered and judged after the fact by the archive operator. They are unscored evidence: they do not change this run's verdict or cohort standing and are excluded from every comparison, including any of them that the judge passes.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console