Versioned benchmark collection
Fighter Game v1
Harness v1 · release candidate blocked · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- Fighter Game v1
Model coverage
- GPT-5.6 Sol · Max
- Kimi K3 · Max
- Claude Opus 5 · Max
- Qwen3.8 Max · xhigh
Evidence available
- Four selected measured run ledgers
- Exact frozen prompt identity
- Per-run arena-geometry artifact identities
- Exact clean GPT-5.6 Sol rerun with superseded contaminated attempt excluded
Still missing
- Redistribution and commercial-use authority for third-party starter assets
- Blind evaluation and scored side-by-side review
- Standardized FPS and hitch receipts
- Repeat-run reliability evidence
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01GPT-5.6 SolMax reasoning · Fighter Game v1$10.74 estimated API-list-price equivalentrecorded estimate31m 38.6s wall-clockEnd-to-end workflowOpen 02Kimi K3Max reasoning · Fighter Game v1≥$2.91 estimated API-list-price equivalent (lower bound)recorded lower bound31m 15.6s wall-clockEnd-to-end workflowOpen 03Claude Opus 5Max reasoning · Fighter Game v1≥$104.53 estimated API-list-price equivalent (lower bound)recorded lower bound1h 23m 18.7s wall-clockEnd-to-end workflowOpen 04Qwen3.8 Maxxhigh reasoning · Fighter Game v1¥147.93 estimated API-list-price equivalent (upper bound)recorded upper bound1h 17m 10.2s wall-clockEnd-to-end workflowOpen
