Measured harness ledgerPublic result
Claude Opus 5CARVE — Claude Opus 5 Max
Build a playable Babylon.js WebGPU snowboarding game with the supplied rider asset, carved terrain, tricks, obstacles, audio, and a complete downhill run.
Max reasoningHeadline result
- Workflow cost
- $57.60
- Wall-clock
- 5h 39m 17.2s wall-clock
- Processed tokens
- 80.04M processed
- Record state
- partial_root_session_metrics_source_and_build_ledger
Public summary
Claude Opus 5 Max partial_root_session_metrics_source_and_build_ledger ledger: 5h 39m 17.2s wall-clock, 80.04M processed, and $57.60 Root-session token-only standard API-equivalent estimate, not the actual Claude Code subscription charge or a full all-agent campaign cost.
Run identity and stack
- Result ID: carve-webgpu-snowboarding-opus-5-max
- Technical model: claude-opus-5
- Provider: Anthropic Claude Code
- Stack: Anthropic Claude Code
- Stack: Technical model/configuration: claude-opus-5
- Stack: Babylon.js WebGPU
- Stack: Supplied snowboarder fixture
- Stack: Harness v1 CARVE prompt
Cost basis
- No separately priced tool or non-token cost is included because no verifiable usage counts were supplied.
- The supplied receipt declares a one-hour prompt-cache TTL, so the $10/MTok cache-creation rate is applied.
- All-5-minute cache-write alternative: $54.36.
Primary artifact integrity
- Kind: interactive-webgpu-snowboarding-game
- Path: artifacts/carve-webgpu-snowboarding-opus-5-max/source/src/main.js
- SHA-256: 2fa99fe97a730b47fc71aea295ce5a436de1d0e60893d5ed7941c30c73981b2a
Validation evidence
- Path: artifacts/carve-webgpu-snowboarding-opus-5-max/validation/build-and-source-check.public.json
- SHA-256: f86d2b1628d54127f836b67a72876ef3ab3ee2cbf9ddbf3ee829c6e6781aa99c
Recorded caveats
- The root task was one genuine benchmark request, but the model used nine background subagents. Their usage is separately billed and absent from the supplied root receipt, so the published tokens and $57.60 estimate are root-session-only rather than all-agent campaign totals.
- Wall-clock is end-to-end workflow latency including tool execution, rendering, and background-agent waits, not model-only compute.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
- The one-hour cache-write rate is based on the receipt's declared TTL; the source receipt also supplies the $54.3608 alternative if its cache writes were actually five-minute writes.
- The cost is an API-list-price equivalent, not an actual subscription-backed Claude Code charge.
- The production build passed, but no independent browser acceptance run, local FPS measurement, hardware/viewport receipt, audio behavior test, or blind evaluation was supplied.
- The source describes the rider GLB as Meshy-generated. This ledger preserves its required input without making an independent ownership or redistribution claim.
Visible evidence gaps
- all-agent subagent token and cost receipts
- independent WebGPU browser acceptance run
- timed-run, checkpoint, wipeout/respawn, and run-summary checks
- persistent snow deformation and re-ridable track check
- audio response check
- 1080p browser, GPU, viewport, quality-preset, and local FPS receipt
- independent final-capture receipt
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
