Measured harness ledgerPublic result
Claude Opus 5Explorable space-flight game — Claude Opus 5 Max
Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.
Max reasoningHeadline result
- Workflow cost
- $65.22
- Wall-clock
- 57m 55.5s wall-clock
- Processed tokens
- 84.33M processed
- Record state
- partial_token_timing_artifact_build_ledger
Public summary
Claude Opus 5 Max partial_token_timing_artifact_build_ledger ledger: 57m 55.5s wall-clock, 84.33M processed, and $65.22 First-party Claude API list-price equivalent, not a subscription invoice.
Run identity and stack
- Result ID: space-flight-game-opus-5-max
- Technical model: claude-opus-5
- Provider: Anthropic Claude Code
- Stack: Anthropic Claude Code
- Stack: Technical model/configuration: claude-opus-5
- Stack: Three.js / Vite game
- Stack: Generated spaceship assets
- Stack: Harness v1 space-flight prompt
Cost basis
- The session declares a one-hour prompt-cache TTL; the source does not record TTL per entry.
- All-5-minute cache-write alternative: $60.89.
Primary artifact integrity
- Kind: vite-threejs-space-flight-source-and-production-build-entry-point
- Path: artifacts/space-flight-game-opus-5-max/source/src/main.js
- SHA-256: 4beccede1b410533b569aa49f9fff07d131bb27eee8df5536e0e2a822415f533
Recorded caveats
- The source session span includes a 5h 07m 48.2s user-idle gap. The result row uses the source-reported 57m 55.5s active build window instead.
- Output tokens include hidden reasoning, code, and tool-call JSON, not only visible prose.
- Cache reads dominate processed-token volume but are billed at 0.1× base input, so processed tokens substantially overstate cost.
- The primary $65.22 estimate uses a session-declared one-hour cache TTL; the all-five-minute alternative is $60.89.
- The archive build passed, but no browser-runtime, local FPS, hardware, final capture, or blind-evaluation receipt was supplied.
Visible evidence gaps
- browser runtime verification
- local FPS and hardware receipt
- final capture with viewport/browser evidence
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
