Measured harness ledgerPublic result
DeepSeek V4 FlashJRPG boss battle — DeepSeek V4 Flash Max
Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.
Max reasoningHeadline result
- Workflow cost
- $0.07
- Wall-clock
- 32m 0.3s wall-clock
- Processed tokens
- 9.23M processed
- Record state
- partial_token_timing_artifact_validation_ledger
Public summary
DeepSeek V4 Flash Max partial_token_timing_artifact_validation_ledger ledger: 32m 0.3s wall-clock, 9.23M processed, and $0.07 API-equivalent usage accounting, not an itemized subscription cash charge.
Run identity and stack
- Result ID: jrpg-boss-battle-deepseek-v4-flash-max
- Technical model: DeepSeek V4 Flash
- Provider: OpenCode Go
- Stack: OpenCode Go
- Stack: Technical model/configuration: DeepSeek V4 Flash
- Stack: Three.js / Vite game
- Stack: Supplied GLB fixture bundle
- Stack: Harness v1 JRPG validator
Cost basis
- Visible output and reasoning are both output-priced. OpenCode Go subscription pricing is not divided into a marginal cash cost for this run.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/jrpg-boss-battle-deepseek-v4-flash-max/source/src/main.js
- SHA-256: 9092d1a9c31afb184af7683cdc918143b7486f115962992f1eb9980334f65f26
Recorded caveats
- Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle gaps; it is not model-only compute.
- OpenCode splits one agentic turn into multiple assistant/API messages.
- Cache-read tokens are deeply discounted, so total processed tokens overstate cost.
- The public record excludes private OpenCode session/export data; their source hashes and the stable benchmark-window hash are retained.
- The supplied validator scorecard reports PASS 41/41 and 60.2 FPS, but it was not independently rerun during archival.
- No final RTX Pro 6000 render or blind-evaluation record was supplied.
Visible evidence gaps
- independent validator rerun
- final RTX Pro 6000 render or capture metadata
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
