Measured harness ledgerPublic result
Claude Opus 5Red Sands v1 — Claude Opus 5 Max
Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.
Max reasoningHeadline result
- Workflow cost
- $129.47
- Wall-clock
- 1h 30m 54.0s wall-clock
- Processed tokens
- 193.93M processed
- Record state
- artifact_operator_static_checks_ledger_no_playthrough_verification
Public summary
Claude Opus 5 Max artifact_operator_static_checks_ledger_no_playthrough_verification ledger: 1h 30m 54.0s wall-clock, 193.93M processed, and $129.47 First-party Claude API list-price equivalent for the measured token mix; not an itemized subscription charge.
Run identity and stack
- Result ID: red-sands-quest-threejs-opus-5-max
- Technical model: claude-opus-5
- Provider: Provider not separately recorded
- Client: Claude Code
- Stack: Claude Code
- Stack: Technical model/configuration: claude-opus-5
- Stack: Three.js quest-authoring fixture
- Stack: Harness v1 frozen prompt
Cost basis
- The transcript does not record cache TTL per request; the one-hour rate is taken from this session's stated prompt-cache configuration.
- 98.6% of tokens processed were cache reads, costing 73.9% of the bill. The 194M total-processed figure badly overstates cost.
- All-5-minute cache-write alternative: $121.57.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/red-sands-quest-threejs-opus-5-max/source/src/quest/quests/shallow-ground.js
- SHA-256: a58fa9c55f36230e557f187137422fb053dada614025185cc96c1a287c095046
Validation evidence
- Path: artifacts/red-sands-quest-threejs-opus-5-max/verification/operator-smoke-report.json
- SHA-256: 976c6e73f16c248a1a91891bedf6a13e36117bba4fd19935100374dd001e7ef4
Recorded caveats
- This task has NO automated completion gate, by design. Nothing in this ledger establishes that the quest plays start to finish.
- No operator playthrough exists — the verification browser could not boot the game, because a backgrounded pane throttles the requestAnimationFrame-driven ~10 s terrain generation.
- npm run smoke runs under ?capture=1, which never constructs the quest system; a green smoke says nothing about the quest.
- Everything about how the quest plays rests on the model's own account in SHALLOW-GROUND.md.
- The archived smoke report and capture are the operator's rerun, which overwrote the model's earlier copy.
- The voice audio was not listened to; delivery, count, and casting are confirmed, performance is not.
- Wall-clock is end-to-end latency including about fifteen world-generation rebuilds, ElevenLabs synthesis, and npm install — not model-only compute.
- Output tokens include hidden reasoning, code, and tool-call JSON.
- Cache reads bill at 0.1x input, so the 194M total-processed figure badly overstates cost.
- This is a quest-authorship task over a fixed game, so its token and cost figures are not comparable to the archive's from-scratch build rows.
- The archive is a delta over the fixture, not a standalone runnable tree.
- No quality score and no blind-evaluation record exist — which for this task is where the entire measurement lives.
Visible evidence gaps
- a human playthrough of the quest, start to finish
- blind-evaluation record — the primary instrument for this task
- mechanical completion evidence that the quest starts, advances, completes, and fires its voice lines in play
- a second candidate to compare against
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Red Sands v1 · Builder projects v1
