Measured harness ledgerPublic result
Claude Fable 5JRPG boss battle — Claude Fable 5 Max
Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.
Max reasoningHeadline result
- Workflow cost
- $123.42
- Wall-clock
- 50:48.1 wall-clock
- Processed tokens
- 76.28M processed
- Record state
- partial_token_timing_artifact_validation_ledger
Public summary
Claude Fable 5 Max partial_token_timing_artifact_validation_ledger ledger: 50:48.1 wall-clock, 76.28M processed, and $123.42 API-equivalent estimate with exact cache-write TTL accounting, not a subscription invoice.
Run identity and stack
- Result ID: jrpg-boss-battle-fable-5-max
- Technical model: claude-fable-5
- Provider: Anthropic Claude Code
- Stack: Anthropic Claude Code
- Stack: Technical model/configuration: claude-fable-5
- Stack: Three.js / Vite game
- Stack: Supplied GLB fixture bundle
- Stack: Harness v1 JRPG validator
- Stack: Requested tool profile: codex-cli-imagegen
Cost basis
- The supplied transcript reports the exact 5-minute and 1-hour cache-creation split, so no cache-TTL assumption is required.
Primary artifact integrity
- Kind: interactive-threejs-jrpg-boss-battle
- Path: artifacts/jrpg-boss-battle-fable-5-max/source/src/main.js
- SHA-256: 800bdc502aed7190cf59dcdb8025e7f729eed50113d3bacd55ee99302ca9d41d
Validation evidence
- Result: PASS, 41/41 checks
- Verification: The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.
- Path: artifacts/jrpg-boss-battle-fable-5-max/validation/scorecard.json
- SHA-256: 316fd215980bd44411a95eeeae333d808a7f68c6f01a26f1c968e197d6cea3eb
- Validator SHA-256: b571ff1895fd6a61d98ccb4a9f7ca4ecefc961abfc3f6b77e23742bddd921ab8
Recorded caveats
- Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool execution and idle waits.
- Output tokens include hidden reasoning, code, and tool-call JSON.
- Cache reads are billed at a 90% discount to fresh input, so total processed tokens substantially overstate cost.
- The supplied snapshot includes the metrics turns themselves.
- The exact browser version and local hardware identity were not supplied.
- The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
- The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
- The validator establishes functional completion but does not replace blind visual-quality evaluation.
Visible evidence gaps
- browser and local-hardware identity
- independent validator rerun
- model tool-use transcript or disclosure
- blind-evaluation record
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Kimi K3 Launch 002 · v1
