Fable sets the strongest visual ceiling across this published comparison.
Kimi K3 vs GPT-5.6 Sol vs Claude Fable 5
Fable 5 sets the visual ceiling. Kimi K3 wins on value.
Nine frozen graphics and game-development tasks. One request per run. No rescue prompts.
GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.
Same frozen target and prompt within each test · disclosed provider-specific stacks
Single recorded runs. Reliability is not established.
7 real blind matchups are ready now.
- 9 testsfrozen graphics + game tasks
- 33featured run records
- One requestper run
- Disclosedprovider-specific stacks
- API-equivalentcost receipts
- 7 liveblind matchups
Three useful answers. No synthetic ranking.
Harry’s editorial verdict separates visual ceiling, recorded value, and the practical technical middle instead of manufacturing an overall score.
Kimi wins on value across the recorded attempts, with material counterexamples kept visible.
Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.
No percentage is published before the declared unique-voter minimum is met.
These are single recorded attempts, not repeat-run reliability evidence.
One request is not one API call.
- One user request per attempt, with no human follow-up or manual correction after submission.
- The same frozen target and prompt within each test, using disclosed provider-specific stacks.
- Single recorded attempts with their cost labels, caveats, and missing evidence kept attached.
Read the full seven-point methodology
- Each attempt began with one user request and had no human follow-up or manual correction after submission.
- The same frozen target and prompt were used within each test, with disclosed provider-specific stacks.
- Those stacks could use disclosed tools, internal retries, and sub-agents. This is not a claim of one model or API call.
- Most Kimi attempts used Kimi Code CLI; Fable used Anthropic Claude Code; GPT-5.6 used OpenAI Codex. Neo-Gothic Storm City used OpenCode Go / Moonshot AI, exactly as its ledger reports.
- Costs are recorded API-equivalent estimates or first-party API estimates as labelled in each ledger—not subscription invoices.
- Cost, time, tokens, artifacts, validation, and missing evidence come from the pinned harness ledgers.
- Every result is a single recorded attempt. Missing evidence and lower-bound coverage remain attached to the corresponding row.
The MacBook Kimi cost is ≥$10.85. Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.
Sponsors can support RemakeBench. They cannot buy benchmark outcomes, rankings, blind votes, or editorial verdicts.
Inspect the output, then open the ledger.
Each pair shows the recorded Kimi and Fable output. The receipt table keeps every featured GPT-5.6 configuration and its material caveats visible.
Neo-Gothic Storm City
A storm-lashed neo-gothic city shader judged on architecture, atmosphere, lighting, water, and motion.
Fable Max is the visual-quality winner. Kimi beats the displayed Sol results on water and lighting, but takes much longer.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxBase model | OpenCode Go / Moonshot AI | $2.30API-equivalent estimate | 38:29.074 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end workflow latency including tool execution, not model-only compute. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxBase model | Anthropic Claude Code | $11.39API-equivalent estimate | 28:48 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency including user idle time between turns, not model-only compute. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighBase model | OpenAI Codex | $1.96API-equivalent estimate | 8:49.5 wall-clock | partial token timing and artifact ledger | Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraBase model | OpenAI Codex | $2.84API-equivalent estimate | 22:10.9 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
MacBook-Class Cinematic
A polished product-ad shot with a modeled laptop, industrial detail, intentional materials, lighting, and camera movement.
Fable Max is best overall. Kimi preserves the MacBook-class chassis shape better than Sol xhigh but gets the hinge wrong. The archived validator logs report 49/49 checks for the compared scenes; those checks do not establish equal visual quality.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | ≥$10.85API-equivalent estimate | 1:26:58.2 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher. Evidence still missing (5)
| Open result |
| Claude Fable 5 MaxTool-assisted | Anthropic Claude Code | $75.18API-equivalent estimate | 1:01:04.8 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | All-1-hour cache-write alternative: $81.28. Primary cost assumes 5-minute cache writes because the supplied usage does not record TTLs. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighTool-assisted | OpenAI Codex | $7.67API-equivalent estimate | 55:49.9 wall-clock | partial token timing artifact validation ledgerPASS, 49/49 checks | Requests with more than 272,000 input tokens use long-context pricing; all 81 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (3)
| Open result |
JRPG Boss Battle
An interactive Three.js boss battle using supplied assets, required combat beats, animation, effects, sound controls, and a validator-ready loop.
Fable Max creates the more cinematic battle and broader effects. Kimi is the value winner: Fable cost 43.1× as much in this recorded run. The archived scorecards report 41/41 functional checks for all three; they do not erase the visible quality differences or replace an independent rerun.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI / Moonshot AI | $2.86API-equivalent estimate | 41:20.2 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step. Evidence still missing (3)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $123.42API-equivalent estimate | 50:48.1 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraFull production stack | OpenAI Codex | $9.98API-equivalent estimate | 25:51.8 wall-clock | partial token timing artifact validation ledgerPASS, 41/41 checks | Priority service-tier pricing. Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context. No cache-write amount was inferred from the supplied transcript schema. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
Campfire Under a Starry Night
A polished interactive Three.js campfire judged on fire, light, environmental detail, and controllable camera motion.
Fable Max has the strongest fire and environment; Fable Medium is fastest; Kimi is the value result at $0.45.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI | $0.45API-equivalent estimate | 14:19.3 wall-clock | partial token timing and source ledger | Wall-clock is end-to-end workflow latency and includes tool execution, browser checks, approvals, and idle time. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $11.17API-equivalent estimate | 12:48.9 wall-clock | partial token timing and source ledger | Wall-clock is end-to-end latency, including tool execution, idle gaps, and human think time; it is not model-only compute. Evidence still missing (4)
| Open result |
| Claude Fable 5 MediumFull production stack | Anthropic Claude Code | $5.12estimated API cost | 4:24.6 wall-clock | partial token timing and source ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol xhighFull production stack | OpenAI Codex | $2.26API-equivalent estimate | 28:52.6 wall-clock | partial token timing and source ledger | Requests with more than 272,000 input tokens use long-context pricing; all 44 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (4)
| Open result |
Space Flight Game
A playable browser space-flight game with responsive controls, a coherent environment, lighting, assets, and a game loop.
Fable Max has the visual ceiling. Kimi beats the displayed Sol Ultra result visually and costs $3.25, but its workflow takes about an hour.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxFull production stack | Kimi Code CLI / Moonshot AI | $3.25API-equivalent estimate | 1:04:11 wall-clock | partial token timing and source ledger | Wall-clock is end-to-end workflow latency, including tool execution, manual-approval waits, browser playtests, and idle time. Evidence still missing (4)
| Open result |
| Claude Fable 5 MaxFull production stack | Anthropic Claude Code | $151.55API-equivalent estimate | 1:05:19.2 wall-clock | partial token timing and source ledger | Primary cost uses 5-minute cache writes. All-1-hour cache-write alternative: $167.36. No deduplication audit was supplied for this usage receipt; the token and cost figures retain the supplied basis. Evidence still missing (4)
| Open result |
| Claude Fable 5 MediumFull production stack | Anthropic Claude Code | $22.85API-equivalent estimate | 18:37.7 wall-clock | partial token timing and source ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (4)
| Open result |
| GPT-5.6 Sol UltraFull production stack | OpenAI Codex | $17.94API-equivalent estimate | 37:44.2 wall-clock | partial token timing and source ledger | Single supplied pre-request snapshot; repeat runs not attached Evidence still missing (4)
| Open result |
Shrine Village
An inspectable voxel shrine-village scene judged on composition, architecture, landscaping, water, and detail.
Kimi wins the editorial comparison against Sol Ultra on bridge placement, the lamp, and water. It gets surprisingly close to Fable Medium for far less money, but takes longer.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.23API-equivalent estimate | 35:53.5 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $20.28estimated API cost | 18:37.5 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $1.98API-equivalent estimate | 20:36.2 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
Infinite Cathedral
An infinite stained-glass cathedral corridor shader judged on depth, architectural repetition, light, reflections, and motion.
Fable Max wins visual quality. Sol Ultra is the practical choice in this use case. Kimi lacks architectural detail and floor reflections and needs significant handholding despite its low price.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxBase model | Kimi Code CLI | $0.92API-equivalent estimate | 21:54.1 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end workflow latency and includes tools, waits, and idle time. Evidence still missing (5)
| Open result |
| Claude Fable 5 MaxBase model | Anthropic Claude Code | $43.85API-equivalent estimate | 31:13.8 wall-clock | partial token timing and artifact ledger | Primary cost uses 1-hour cache writes (supplied). All-5-minute cache-write alternative: $38.66. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol xhighBase model | OpenAI Codex | $4.10API-equivalent estimate | 24:19.8 wall-clock | partial token timing and source ledger | Requests with more than 272,000 input tokens use long-context pricing; all 57 supplied calls were short-context. Separately priced tools and non-token services are excluded. Evidence still missing (5)
| Open result |
| GPT-5.6 Sol UltraBase model | OpenAI Codex | $5.61API-equivalent estimate | 38:22.4 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
Oasis Outpost
An inspectable voxel oasis outpost judged on settlement readability, terrain, vegetation, water, and environmental detail.
Kimi beats the displayed GPT-5.6 outputs on quality and gives Fable the only close visual competition. Fable cost 7.8× as much as Kimi in this recorded pair; Kimi is slower but looks steerable with iteration.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.72API-equivalent estimate | 39:27.9 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $21.23estimated API cost | 19:23.7 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $5.51API-equivalent estimate | 16:30.9 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Terra UltraTool-assisted | OpenAI Codex | $1.33API-equivalent estimate | 28:14.2 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
Jungle Temple
An inspectable high-density voxel jungle-temple diorama created at the requested fixed grid scale.
Fable Medium wins visual quality and speed. Kimi is inconsistent here and takes 53:37.7, making this the clearest counterexample to its value story.


| Configuration | Provider / client | Workflow cost | Wall-clock | Record | Caveats | Open result |
|---|---|---|---|---|---|---|
| Kimi K3 MaxTool-assisted | Kimi Code CLI | $2.60API-equivalent estimate | 53:37.7 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time. Evidence still missing (5)
| Open result |
| Claude Fable 5 MediumTool-assisted | Anthropic Claude Code | $10.10estimated API cost | 12:04.2 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Sol UltraTool-assisted | OpenAI Codex | $3.23API-equivalent estimate | 26:39.7 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
| GPT-5.6 Terra UltraTool-assisted | OpenAI Codex | $1.20API-equivalent estimate | 16:43.2 wall-clock | partial token timing and artifact ledger | Wall-clock is end-to-end latency, not model-only compute. Evidence still missing (3)
| Open result |
Get the projects behind this test.
Open 33 reviewed project outputs as a full Project Pack or nine task-sized downloads: editable sanitized Blender scenes, browser projects, shaders, permitted assets, exact prompts, configurations, and receipts. Customer status: 25 ready · 5 reference-only · 3 partial · 0 excluded. Reference-only and partial projects retain their exact prerequisites and limitations.
Keep judging the builds with us.
Join the free RemakeBench Discord after you have inspected the verdict, videos, and public receipts.
