Start with the recorded outputs and receipts, then inspect any available editorial and live blind matchups. No synthetic overall score is added after the fact.
Published comparisons
4
Recorded runs
135
Newest first
Published comparisons
Launch 004
Claude Opus 5 across nine game-building tests
Opus 5 delivered the strongest detail in several rounds. It also made the efficiency argument against itself.
Kimi K3 vs GPT-5.6 SOL vs Claude Fable 5: 9 Game Tests
Fable 5 sets the visual ceiling. Kimi K3 wins on value. GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.
Each listed result is one recorded attempt. Reliability is not established. Costs retain the estimate type and caveats recorded in the corresponding ledger.