Measured harness ledgerPublic result
Claude Opus 5Eiffel Tower Drawing — Claude Opus 5 Max
Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.
Max reasoningHeadline result
- Workflow cost
- $5.51
- Wall-clock
- 15m 54s wall-clock
- Processed tokens
- 5.11M processed
- Record state
- judge_pass_independently_replayed
Public summary
Claude Opus 5 Max judge_pass_independently_replayed ledger: 15m 54s wall-clock, 5.11M processed, and $5.51 API-equivalent estimate from session token counts at Anthropic list price (input $5, output $25, cache read $0.50, cache write $6.25 per MTok).
Run identity and stack
- Result ID: dex-draw-mujoco-opus-5-max
- Technical model: Claude Opus 5
- Provider: Anthropic Claude Code
- Client: claude -p --model opus --permission-mode bypassPermissions (headless subagent, max effort), one shot, no nudges
- Stack: Anthropic Claude Code
- Stack: claude -p --model opus --permission-mode bypassPermissions (headless subagent, max effort), one shot, no nudges
- Stack: Technical model/configuration: Claude Opus 5
- Stack: MuJoCo pencil-drawing fixture
- Stack: Harness drawing prompt
- Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-draw-mujoco-opus-5-max/evidence/model-supplied/ctrl.npy
- SHA-256: 8ad346115b19782b48808a8ee5d5b44fce463c9be38670c9b667bd698f257b4b
Validation evidence
- Result: PASS
- Path: artifacts/dex-draw-mujoco-opus-5-max/validation/judge-replay.public.json
- SHA-256: dea8c64888e69904c0d2d351f3b697b25335e047bf5f1f0836446b284c963a9d
- Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
- The judge measures task completion only (contact-gated ink against the given polyline, runtime limits); there is no visual quality or blind evaluation for this task.
- The replay clip is rendered by the operator's viewer from the judge's trajectory export and ink stream with a fixed camera; it is presentation evidence, not a measurement.
- provenance/gold-results.md is the orchestrator's own result file copied verbatim.
- usedThreeFingers / threeFingerFraction are reported, never gating.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- The given harness exposes the judge's own simulate/rasterise/score pipeline, so the solver could self-score in-process; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
- Opus cost is an API-list-price equivalent from token counts, not a subscription charge; fresh input was not reported separately.
- Three judge attempts were recorded, all PASS tier 3; the first already passed and the later two refined the follower.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- fresh-input token count
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.