Fable sets the strongest visual ceiling across this published comparison.
Kimi K3 vs GPT-5.6 Sol vs Claude Fable 5
Fable 5 sets the visual ceiling. Kimi K3 wins on value.
Nine frozen graphics and game-development tasks. One request per run. No rescue prompts.
GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.
Same frozen target and prompt within each test · disclosed provider-specific stacks
Single recorded runs. Reliability is not established.
7 real blind matchups are ready now.
- 9 testsfrozen graphics + game tasks
- 33featured run records
- One requestper run
- Disclosedprovider-specific stacks
- API-equivalentcost receipts
- 7 liveblind matchups
Three useful answers. No synthetic ranking.
Harry’s editorial verdict separates visual ceiling, recorded value, and the practical technical middle instead of manufacturing an overall score.
Kimi wins on value across the recorded attempts, with material counterexamples kept visible.
Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.
No percentage is published before the declared unique-voter minimum is met.
These are single recorded attempts, not repeat-run reliability evidence.
One request is not one API call.
- One user request per attempt, with no human follow-up or manual correction after submission.
- The same frozen target and prompt within each test, using disclosed provider-specific stacks.
- Single recorded attempts with their cost labels, caveats, and missing evidence kept attached.
Read the full seven-point methodology
- Each attempt began with one user request and had no human follow-up or manual correction after submission.
- The same frozen target and prompt were used within each test, with disclosed provider-specific stacks.
- Those stacks could use disclosed tools, internal retries, and sub-agents. This is not a claim of one model or API call.
- Most Kimi attempts used Kimi Code CLI; Fable used Anthropic Claude Code; GPT-5.6 used OpenAI Codex. Neo-Gothic Storm City used OpenCode Go / Moonshot AI, exactly as its ledger reports.
- Costs are recorded API-equivalent estimates or first-party API estimates as labelled in each ledger—not subscription invoices.
- Cost, time, tokens, artifacts, validation, and missing evidence come from the pinned harness ledgers.
- Every result is a single recorded attempt. Missing evidence and lower-bound coverage remain attached to the corresponding row.
The MacBook Kimi cost is ≥$10.85. One API request lacks a paired usage record, so the actual total may be higher. The Kimi Code session used an Allegretto subscription; this is token-arithmetic equivalence, not a per-run cash charge. Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.
Sponsors can support RemakeBench. They cannot buy benchmark outcomes, rankings, blind votes, or editorial verdicts.
Inspect the output, then open the ledger.
Each pair shows the recorded Kimi and Fable output. The receipt table keeps every featured GPT-5.6 configuration and its material caveats visible.
Neo-Gothic Storm City
A storm-lashed neo-gothic city shader judged on architecture, atmosphere, lighting, water, and motion.
Fable Max is the visual-quality winner. Kimi beats the displayed Sol results on water and lighting, but takes much longer.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxBase model | OpenCode Go / Moonshot AI | $2.30API-equivalent estimate | 38:29.074 wall-clock | partial token timing and artifact ledger | Moonshot publishes one output rate; Kimi K3 reasoning and visible output are both billed at that rate. No cache-write rate was needed because cache-write tokens are zero. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxBase model | Anthropic Claude Code | $11.39API-equivalent estimate | 28:48 wall-clock | partial token timing and artifact ledger | Unlike most Claude receipts in this set, the supplied deduplicated ledger separately identifies both 5-minute and 1-hour cache-write tokens. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighBase model | OpenAI Codex | $1.96API-equivalent estimate | 8:49.5 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraBase model | OpenAI Codex | $2.84API-equivalent estimate | 22:10.9 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 47 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
MacBook-Class Cinematic
A polished product-ad shot with a modeled laptop, industrial detail, intentional materials, lighting, and camera movement.
Fable Max is best overall. Kimi preserves the MacBook-class chassis shape better than Sol xhigh but gets the hinge wrong. The archived validator logs report 49/49 checks for the compared scenes; those checks do not establish equal visual quality.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | ≥$10.85API-equivalent estimate | 1:26:58.2 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | One API request lacks a paired usage record, so the actual total may be higher. The Kimi Code session used an Allegretto subscription; this is token-arithmetic equivalence, not a per-run cash charge. Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher. Evidence still missing (5)
| Open result |
| Claude Fable 5 MaxTool-assisted | Anthropic Claude Code | $75.18API-equivalent estimate | 1:01:04.8 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | The primary total uses the official 5-minute cache-write rate because the supplied transcript aggregates cache writes without recording TTL. All-1-hour cache-write alternative: $81.28. Primary cost assumes 5-minute cache writes because the supplied usage does not record TTLs. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighTool-assisted | OpenAI Codex | $7.67API-equivalent estimate | 55:49.9 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | Requests with more than 272,000 input tokens use long-context pricing; all 81 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
JRPG Boss Battle
An interactive Three.js boss battle using supplied assets, required combat beats, animation, effects, sound controls, and a validator-ready loop.
Fable Max creates the more cinematic battle and broader effects. Kimi is the value winner: Fable cost 43.1× as much in this recorded run. The archived scorecards report 41/41 functional checks for all three; they do not erase the visible quality differences or replace an independent rerun.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI / Moonshot AI | $2.86API-equivalent estimate | 41:20.2 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | The Kimi Code CLI session used an Allegretto subscription. This is token-arithmetic equivalence, not a per-run cash charge. Evidence still missing (3)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $123.42API-equivalent estimate | 50:48.1 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | The supplied transcript reports the exact 5-minute and 1-hour cache-creation split, so no cache-TTL assumption is required. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraFull production stack | OpenAI Codex | $9.98API-equivalent estimate | 25:51.8 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | Cache-creation is zero because the Codex transcript schema does not expose a cache-write token field; no cache-write amount was inferred. Priority service-tier pricing. Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context. No cache-write amount was inferred from the supplied transcript schema. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
Campfire Under a Starry Night
A polished interactive Three.js campfire judged on fire, light, environmental detail, and controllable camera motion.
Fable Max has the strongest fire and environment; Fable Medium is fastest; Kimi is the value result at $0.45.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI | $0.45API-equivalent estimate | 14:19.3 wall-clock | partial token timing and source ledger | The Kimi Code session used an Allegretto subscription. This is token-arithmetic equivalence, not a per-run cash charge. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $11.17API-equivalent estimate | 12:48.9 wall-clock | partial token timing and source ledger | The supplied receipt applies standard Anthropic cache-rate multipliers because cache rates were not itemized on the fetched official page. Evidence still missing (4)
| Open result |
| Claude Fable 5 MediumFull production stack | Anthropic Claude Code | $5.12estimated API cost | 4:24.6 wall-clock | partial token timing and source ledger | Primary cost uses 5-minute cache writes. All-1-hour cache-write alternative: $6.57. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol xhighFull production stack | OpenAI Codex | $2.26API-equivalent estimate | 28:52.6 wall-clock | partial token timing and source ledger | Requests with more than 272,000 input tokens use long-context pricing; all 44 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
Space Flight Game
A playable browser space-flight game with responsive controls, a coherent environment, lighting, assets, and a game loop.
Fable Max has the visual ceiling. Kimi beats the displayed Sol Ultra result visually and costs $3.25, but its workflow takes about an hour.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI / Moonshot AI | $3.25API-equivalent estimate | 1:04:11 wall-clock | partial token timing and source ledger | The Kimi Code CLI session used an Allegretto subscription. This is token-arithmetic equivalence, not a per-run cash charge. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $151.55API-equivalent estimate | 1:05:19.2 wall-clock | partial token timing and source ledger | If every cache write used the official one-hour rate instead, the total would be $167.36. Primary cost uses 5-minute cache writes. All-1-hour cache-write alternative: $167.36. No deduplication audit was supplied for this usage receipt; the token and cost figures retain the supplied basis. Evidence still missing (4)
| Open result |
| Claude Fable 5 MediumFull production stack | Anthropic Claude Code | $22.85estimated API cost | 18:37.7 wall-clock | partial token timing and source ledger | Primary cost uses 5-minute cache writes. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraFull production stack | OpenAI Codex | $17.94API-equivalent estimate | 37:44.2 wall-clock | partial token timing and source ledger | The supplied metrics name gpt-5.6-sol; that model ID is used instead of the surrounding GPT 5.5 label. Evidence still missing (4)
| Open result |
Shrine Village
An inspectable voxel shrine-village scene judged on composition, architecture, landscaping, water, and detail.
Kimi wins the editorial comparison against Sol Ultra on bridge placement, the lamp, and water. It gets surprisingly close to Fable Medium for far less money, but takes longer.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.23API-equivalent estimate | 35:53.5 wall-clock | partial token timing and artifact ledger | Kimi Code ran on an Allegretto subscription ($39/month); the marginal per-run cash cost was not itemized. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $20.28estimated API cost | 18:37.5 wall-clock | partial token timing and artifact ledger | Primary cost uses 5-minute cache writes. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $1.98API-equivalent estimate | 20:36.2 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 29 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
Infinite Cathedral
An infinite stained-glass cathedral corridor shader judged on depth, architectural repetition, light, reflections, and motion.
Fable Max wins visual quality. Sol Ultra is the practical choice in this use case. Kimi lacks architectural detail and floor reflections and needs significant handholding despite its low price.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxBase model | Kimi Code CLI | $0.92API-equivalent estimate | 21:54.1 wall-clock | partial token timing and artifact ledger | Kimi Code ran on an Allegretto subscription ($39/month); the marginal per-run cash cost was not itemized. Evidence still missing (5)
| Open result |
| Claude Fable 5 MaxBase model | Anthropic Claude Code | $43.85API-equivalent estimate | 31:13.8 wall-clock | partial token timing and artifact ledger | Primary cost uses 1-hour cache writes (supplied). All-5-minute cache-write alternative: $38.66. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighBase model | OpenAI Codex | $4.10API-equivalent estimate | 24:19.8 wall-clock | partial token timing and source ledger | Requests with more than 272,000 input tokens use long-context pricing; all 57 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol UltraBase model | OpenAI Codex | $5.61API-equivalent estimate | 38:22.4 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 72 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
Oasis Outpost
An inspectable voxel oasis outpost judged on settlement readability, terrain, vegetation, water, and environmental detail.
Kimi beats the displayed GPT-5.6 outputs on quality and gives Fable the only close visual competition. Fable cost 7.8× as much as Kimi in this recorded pair; Kimi is slower but looks steerable with iteration.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.72API-equivalent estimate | 39:27.9 wall-clock | partial token timing and artifact ledger | Kimi Code ran on an Allegretto subscription ($39/month); the marginal per-run cash cost was not itemized. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $21.23estimated API cost | 19:23.7 wall-clock | partial token timing and artifact ledger | Primary cost uses 5-minute cache writes. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $5.51API-equivalent estimate | 16:30.9 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
| GPT-5.6 Terra UltraTool-assisted | OpenAI Codex | $1.33API-equivalent estimate | 28:14.2 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 21 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
Jungle Temple
An inspectable high-density voxel jungle-temple diorama created at the requested fixed grid scale.
Fable Medium wins visual quality and speed. Kimi is inconsistent here and takes 53:37.7, making this the clearest counterexample to its value story.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.60API-equivalent estimate | 53:37.7 wall-clock | partial token timing and artifact ledger | Kimi Code ran on an Allegretto subscription ($39/month); the marginal per-run cash cost was not itemized. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $10.10estimated API cost | 12:04.2 wall-clock | partial token timing and artifact ledger | Primary cost uses 5-minute cache writes. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $3.23API-equivalent estimate | 26:39.7 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 42 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
| GPT-5.6 Terra UltraTool-assisted | OpenAI Codex | $1.20API-equivalent estimate | 16:43.2 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 19 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
Reproduce a published build, then change it.
Start from released Blender, browser and shader projects with their exact versioned prompts. Reproduce the published baseline, modify it for another model or direction, and preserve the output as your own comparison. Technical status: 25 ready · 5 reference-only · 3 partial · 0 excluded.
Keep judging the builds with us.
Join the free RemakeBench Discord after you have inspected the verdict, videos, and public receipts.
