PUBLISHED TESTS

The scoreboard grows one test at a time.

Start with the recorded outputs and receipts, then inspect any available editorial and live blind matchups. No synthetic overall score is added after the fact.

Published comparisons
4
Recorded runs
135
Newest first

Published comparisons

Launch 003

I Gave Qwen3.8 Max, Kimi K3, GPT-5.6 and Fable 5 the Same 9 Game Tests

Kimi K3 was more consistent; Qwen's visual-agent workflow was not reliable enough to win overall.

Review scope
9 tests · 42 recorded runs
Reproduction path
Not released
Launch 002

Kimi K3 vs GPT-5.6 SOL vs Claude Fable 5: 9 Game Tests

Fable 5 sets the visual ceiling. Kimi K3 wins on value. GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.

Review scope
9 tests · 33 recorded runs
Play it blind
7 matchups live now
Reproduction path
Project + prompt available
Tech Review 001

GPT-5.6 vs Claude Fable 5: Same Prompts, 7 AI Game-Dev Tests

Fable sets the visual ceiling. Sol Ultra wins the value argument.

Review scope
7 tests · 28 recorded runs
Play it blind
5 matchups live now
Reproduction path
Prompt only

Each listed result is one recorded attempt. Reliability is not established. Costs retain the estimate type and caveats recorded in the corresponding ledger.