Measured harness ledgerPublic result
Claude Opus 5

Off-road Mud Game — Claude Opus 5 Max

Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.

Max reasoningHeadline result
Workflow cost
$124.35
Wall-clock
2h 48m 08.6s wall-clock
Processed tokens
149.70M processed
Record state
partial_token_timing_source_build_ledger
Public summary

Claude Opus 5 Max partial_token_timing_source_build_ledger ledger: 2h 48m 08.6s wall-clock, 149.70M processed, and $124.35 Official Anthropic API-list-price equivalent, not an itemized Claude Code subscription charge.

Run identity and stack
  • Result ID: off-road-driving-game-threejs-opus-5-max
  • Technical model: claude-opus-5
  • Provider: Anthropic Claude Code
  • Stack: Anthropic Claude Code
  • Stack: Technical model/configuration: claude-opus-5
  • Stack: Three.js vehicle workflow
  • Stack: Harness v1 off-road prompt
Cost basis
  • All-5-minute cache-write alternative: $113.74.
Primary artifact integrity
  • Kind: Fresh Vite `dist/` directory archive
  • Path: artifacts/off-road-driving-game-threejs-opus-5-max/production-build.tar.gz
  • SHA-256: 81e47b14a92cb622a929eca09fa9146ccb861696d542a52879202f0c1b88026c
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, package installation, browser automation, WASM initialization, and model latency; it is not model-only compute.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
  • Cache-read tokens are discounted, so the total-processed figure materially overstates effective cost.
  • The cache-write TTL mix is not independently reconstructible from the transcript; the primary uses the source receipt's one-hour declaration and records the all-five-minute alternative.
  • The retained `.verify` scripts and captures are model-authored output. Their console results are not supplied and they were not independently rerun, so they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
  • No standardized final gameplay capture or blind-evaluation record is supplied.
Visible evidence gaps
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • multi-minute sustained-drive and generation-seam receipt
  • four-corner suspension-articulation receipt
  • standardized final gameplay capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console