PUBLISHED TESTS

The scoreboard grows one test at a time.

Start with the recorded outputs and receipts, then inspect any available editorial and live blind matchups. No synthetic overall score is added after the fact.

Published comparisons
3
Featured run entries
103
Newest first

Published comparisons

Launch 002

Kimi K3 vs GPT-5.6 SOL vs Claude Fable 5: 9 Game Tests

Fable 5 sets the visual ceiling. Kimi K3 wins on value. GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.

Review scope
9 tests · 33 recorded runs
Blind campaign
7 live pairs
Builder package
Project Pack v1 + Research Pack v1 · 25 ready · 5 reference-only · 3 partial · 0 excluded
Tech Review 001

GPT-5.6 vs Claude Fable 5: Same Prompts, 7 AI Game-Dev Tests

Fable sets the visual ceiling. Sol Ultra wins the value argument.

Review scope
7 tests · 28 recorded runs
Blind campaign
5 live pairs
Builder package
Tech Review 001 · v1 · verified

Each listed result is one recorded attempt. Reliability is not established. Costs retain the estimate type and caveats recorded in the corresponding ledger.