Measured harness ledgerPublic result
Claude Opus 5.5

Eiffel Tower Drawing — Claude Opus 5.5 Max

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

Max reasoningHeadline result
Workflow cost
$2.01
Wall-clock
7m 43s wall-clock
Processed tokens
1.74M processed
Record state
judge_pass_independently_replayed
Public summary

Claude Opus 5.5 Max judge_pass_independently_replayed ledger: 7m 43s wall-clock, 1.74M processed, and $2.01 recorded total_cost_usd from workspace FINAL.json (modelUsage.claude-opus-5-5, costBasis=list, Anthropic list price). Charged to the Claude subscription ($0 charged)..

Run identity and stack
  • Result ID: dex-draw-mujoco-opus-5.5-max-2026-09-22
  • Technical model: Claude Opus 5.5
  • Provider: Anthropic Claude Code
  • Client: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Anthropic Claude Code
  • Stack: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
  • Stack: Technical model/configuration: Claude Opus 5.5
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-opus-5.5-max-2026-09-22/evidence/model-supplied/ctrl.npy
  • SHA-256: c77417d416769740512f1c348aa6cec8782a5c0cf6e83f147656c421c8482686
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-draw-mujoco-opus-5.5-max-2026-09-22/validation/judge-replay.public.json
  • SHA-256: 7e47ba27a3fa4f1c92a3a25e37f4d85f2ae90025b55b620caec5ba2371a9afda
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Cost is the recorded list-price total_cost_usd from workspace FINAL.json (costBasis=list). Charged to the Claude subscription ($0 charged).
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The raw Claude CLI receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EU-CZ-1, 105.6 + 65.4 min, $1.30 + $0.81 = $2.11 total. Two-pass render: frames 1–265 (pod B, EU-CZ-1; rsync failed, no space left on device) then frames 266–1422 (RTX 4090 Secure US-NC-1, 83.3 min, $1.03). FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
  • Nudge disclosure: none.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console