Versioned benchmark collection

Fighter Game v1

Harness v1 · release candidate blocked · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Arena-only preview from the clean Fighter Game v1 result
Arena-only preview from the clean Fighter Game v1 result
4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Fighter Game v1

Model coverage

  • GPT-5.6 Sol · Max
  • Kimi K3 · Max
  • Claude Opus 5 · Max
  • Qwen3.8 Max · xhigh

Evidence available

  • Four selected measured run ledgers
  • Exact frozen prompt identity
  • Per-run arena-geometry artifact identities
  • Exact clean GPT-5.6 Sol rerun with superseded contaminated attempt excluded

Still missing

  • Redistribution and commercial-use authority for third-party starter assets
  • Blind evaluation and scored side-by-side review
  • Standardized FPS and hitch receipts
  • Repeat-run reliability evidence
Public receipts

4 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.