Measured harness ledgerPublic result
DeepSeek V4 ProFighter Game v1 — DeepSeek V4 Pro Max
Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.
Max reasoningHeadline result
- Workflow cost
- $0.35
- Wall-clock
- 53m 59.574s wall-clock
- Processed tokens
- 35.95M processed
- Record state
- one_shot_metrics_and_artifact_ledger_partial_independent_runtime_evidence
Public summary
DeepSeek V4 Pro Max one_shot_metrics_and_artifact_ledger_partial_independent_runtime_evidence ledger: 53m 59.574s wall-clock, 35.95M processed, and $0.35 OpenCode-recorded usage accounting, exactly recomputed from the supplied OpenCode Go rate snapshot; not an itemized subscription cash charge.
Run identity and stack
- Result ID: courtyard-arena-blender-mcp-deepseek-v4-pro-max
- Technical model: DeepSeek V4 Pro
- Provider: OpenCode Go
- Client: OpenCode 1.18.17
- Stack: OpenCode Go
- Stack: OpenCode 1.18.17
- Stack: Technical model/configuration: DeepSeek V4 Pro
- Stack: Blender MCP
- Stack: Three.js fighting-game fixture
- Stack: Harness v1 frozen prompt
Cost basis
- OpenCode Go is a subscription path; the marginal cash charge for this run is not itemized.
Primary artifact integrity
- Kind: rebuilt Blender-exported arena geometry
- Path: artifacts/courtyard-arena-blender-mcp-deepseek-v4-pro-max/source/web/models/arena.glb
- SHA-256: 18df1d9235185b09a597fa57e7381294ff3b271a0cab39f3d16d83b6508eae4c
Recorded caveats
- Wall-clock is end-to-end workflow latency, including tool execution, Blender work, browser checks, and waits; it is not model-only compute.
- One benchmark user turn contains 176 model requests, including three child explore-agent requests.
- Output tokens include visible output plus reasoning output. Cache reads dominate total processed tokens and are deeply discounted, so total processed tokens overstate cost.
- Two tool calls reported errors; they are retained in the metrics rather than hidden.
- The candidate input-package boundary is supported by the supplied package, but the private path-redacted transcript cannot independently prove every tool-access path.
- The final GLB was independently reopened in Blender, but the static-server browser check reached only the asset-loading screen. The game captures are model-supplied, not an independent completed-match replay.
- No standardized FPS or hitch receipt, blind evaluation, or comparative quality score is archived.
- This is an environment-improvement result over a fixed starter; its metrics are not comparable to the harness's from-scratch creative builds.
Visible evidence gaps
- independent completed-match replay and console receipt
- standardized FPS and hitch measurement
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
