Measured harness ledgerPublic result
Claude Opus 5

JRPG boss battle — Claude Opus 5 Max

Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.

Max reasoningHeadline result
Workflow cost
$48.95
Wall-clock
1h 15m 49.9s wall-clock
Processed tokens
49.93M processed
Record state
partial_token_timing_artifact_validation_ledger
Public summary

Claude Opus 5 Max partial_token_timing_artifact_validation_ledger ledger: 1h 15m 49.9s wall-clock, 49.93M processed, and $48.95 First-party Claude API list-price equivalent, not a Claude Code subscription invoice.

Run identity and stack
  • Result ID: jrpg-boss-battle-opus-5-max
  • Technical model: claude-opus-5
  • Provider: Anthropic Claude Code
  • Stack: Anthropic Claude Code
  • Stack: Technical model/configuration: claude-opus-5
  • Stack: Three.js / Vite game
  • Stack: Supplied GLB fixture bundle
  • Stack: Harness v1 JRPG validator
  • Stack: Requested tool profile: codex-cli-imagegen
Cost basis
  • The source receipt reports a one-hour prompt-cache TTL, so the primary calculation uses the one-hour cache-write rate.
  • All-5-minute cache-write alternative: $44.41.
Primary artifact integrity
  • Kind: interactive-threejs-jrpg-boss-battle
  • Path: artifacts/jrpg-boss-battle-opus-5-max/source/src/main.js
  • SHA-256: 4547b0712a38f5927a9b76c0f242dabd4931b5bd42c5342f37a0be858b56dcb0
Validation evidence
  • Result: PASS, 41/41 checks
  • Verification: Freshly rerun during archival on 2026-07-30; scorecard duration 61.646 seconds.
  • Path: artifacts/jrpg-boss-battle-opus-5-max/validation/scorecard.json
  • SHA-256: dfc831b73db3626a23a0d678b594acd4c7db3f01bdc2aac99b178e41dc2f7977
  • Validator SHA-256: b571ff1895fd6a61d98ccb4a9f7ca4ecefc961abfc3f6b77e23742bddd921ab8
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, browser runs, image generation, and human-side idle time; it is not model-only compute.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
  • The exact browser version and local hardware identity were not supplied.
  • The metrics receipt reports a subscription-backed Claude Code session; the displayed amount is API-equivalent rather than an itemized cash charge.
Visible evidence gaps
  • browser and local-hardware identity
  • independent blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console