Measured harness ledgerPublic result
Claude Opus 5

Eiffel Tower Drawing — Claude Opus 5 Max

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

Max reasoningHeadline result
Workflow cost
$5.51
Wall-clock
15m 54s wall-clock
Processed tokens
5.11M processed
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5 Max judge_pass_independently_replayed ledger: 15m 54s wall-clock, 5.11M processed, and $5.51 API-equivalent estimate from session token counts at Anthropic list price (input $5, output $25, cache read $0.50, cache write $6.25 per MTok).

Run identity and stack
  • Result ID: dex-draw-mujoco-opus-5-max
  • Technical model: Claude Opus 5
  • Provider: Anthropic Claude Code
  • Client: claude -p --model opus --permission-mode bypassPermissions (headless subagent, max effort), one shot, no nudges
  • Stack: Anthropic Claude Code
  • Stack: claude -p --model opus --permission-mode bypassPermissions (headless subagent, max effort), one shot, no nudges
  • Stack: Technical model/configuration: Claude Opus 5
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-opus-5-max/evidence/model-supplied/ctrl.npy
  • SHA-256: 8ad346115b19782b48808a8ee5d5b44fce463c9be38670c9b667bd698f257b4b
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-draw-mujoco-opus-5-max/validation/judge-replay.public.json
  • SHA-256: dea8c64888e69904c0d2d351f3b697b25335e047bf5f1f0836446b284c963a9d
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • The judge measures task completion only (contact-gated ink against the given polyline, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export and ink stream with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/gold-results.md is the orchestrator's own result file copied verbatim.
  • usedThreeFingers / threeFingerFraction are reported, never gating.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • The given harness exposes the judge's own simulate/rasterise/score pipeline, so the solver could self-score in-process; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
  • Opus cost is an API-list-price equivalent from token counts, not a subscription charge; fresh input was not reported separately.
  • Three judge attempts were recorded, all PASS tier 3; the first already passed and the later two refined the follower.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • fresh-input token count
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console