Measured harness ledgerPublic result
DeepSeek V4 Pro

Fighter Game v1 — DeepSeek V4 Pro Max

Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.

Max reasoningHeadline result
Workflow cost
$0.35
Wall-clock
53m 59.574s wall-clock
Processed tokens
35.95M processed
Record state
one_shot_metrics_and_artifact_ledger_partial_independent_runtime_evidence
Public summary

DeepSeek V4 Pro Max one_shot_metrics_and_artifact_ledger_partial_independent_runtime_evidence ledger: 53m 59.574s wall-clock, 35.95M processed, and $0.35 OpenCode-recorded usage accounting, exactly recomputed from the supplied OpenCode Go rate snapshot; not an itemized subscription cash charge.

Run identity and stack
  • Result ID: courtyard-arena-blender-mcp-deepseek-v4-pro-max
  • Technical model: DeepSeek V4 Pro
  • Provider: OpenCode Go
  • Client: OpenCode 1.18.17
  • Stack: OpenCode Go
  • Stack: OpenCode 1.18.17
  • Stack: Technical model/configuration: DeepSeek V4 Pro
  • Stack: Blender MCP
  • Stack: Three.js fighting-game fixture
  • Stack: Harness v1 frozen prompt
Cost basis
  • OpenCode Go is a subscription path; the marginal cash charge for this run is not itemized.
Primary artifact integrity
  • Kind: rebuilt Blender-exported arena geometry
  • Path: artifacts/courtyard-arena-blender-mcp-deepseek-v4-pro-max/source/web/models/arena.glb
  • SHA-256: 18df1d9235185b09a597fa57e7381294ff3b271a0cab39f3d16d83b6508eae4c
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, Blender work, browser checks, and waits; it is not model-only compute.
  • One benchmark user turn contains 176 model requests, including three child explore-agent requests.
  • Output tokens include visible output plus reasoning output. Cache reads dominate total processed tokens and are deeply discounted, so total processed tokens overstate cost.
  • Two tool calls reported errors; they are retained in the metrics rather than hidden.
  • The candidate input-package boundary is supported by the supplied package, but the private path-redacted transcript cannot independently prove every tool-access path.
  • The final GLB was independently reopened in Blender, but the static-server browser check reached only the asset-loading screen. The game captures are model-supplied, not an independent completed-match replay.
  • No standardized FPS or hitch receipt, blind evaluation, or comparative quality score is archived.
  • This is an environment-improvement result over a fixed starter; its metrics are not comparable to the harness's from-scratch creative builds.
Visible evidence gaps
  • independent completed-match replay and console receipt
  • standardized FPS and hitch measurement
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console