REMAKEBENCH · KIMI K3 LAUNCH TEST

Kimi K3 vs GPT-5.6 Sol vs Claude Fable 5

Fable 5 sets the visual ceiling. Kimi K3 wins on value.

Nine frozen graphics and game-development tasks. One request per run. No rescue prompts.

GPT-5.6 Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.

Same frozen target and prompt within each test · disclosed provider-specific stacks

Single recorded runs. Reliability is not established.

7 real blind matchups are ready now.

  • 9 testsfrozen graphics + game tasks
  • 33featured run records
  • One requestper run
  • Disclosedprovider-specific stacks
  • API-equivalentcost receipts
  • 7 liveblind matchups
Published editorial verdict

Three useful answers. No synthetic ranking.

Harry’s editorial verdict separates visual ceiling, recorded value, and the practical technical middle instead of manufacturing an overall score.

Visual ceilingClaude Fable 5

Fable sets the strongest visual ceiling across this published comparison.

Recorded valueKimi K3

Kimi wins on value across the recorded attempts, with material counterexamples kept visible.

Practical middleGPT-5.6 Sol Ultra

Sol Ultra is the practical middle when technical coherence matters without paying Fable’s full cost.

Community preferenceCollecting blind votes

No percentage is published before the declared unique-voter minimum is met.

ReliabilityNot established

These are single recorded attempts, not repeat-run reliability evidence.

Methodology

One request is not one API call.

  1. One user request per attempt, with no human follow-up or manual correction after submission.
  2. The same frozen target and prompt within each test, using disclosed provider-specific stacks.
  3. Single recorded attempts with their cost labels, caveats, and missing evidence kept attached.
Read the full seven-point methodology
  1. Each attempt began with one user request and had no human follow-up or manual correction after submission.
  2. The same frozen target and prompt were used within each test, with disclosed provider-specific stacks.
  3. Those stacks could use disclosed tools, internal retries, and sub-agents. This is not a claim of one model or API call.
  4. Most Kimi attempts used Kimi Code CLI; Fable used Anthropic Claude Code; GPT-5.6 used OpenAI Codex. Neo-Gothic Storm City used OpenCode Go / Moonshot AI, exactly as its ledger reports.
  5. Costs are recorded API-equivalent estimates or first-party API estimates as labelled in each ledger—not subscription invoices.
  6. Cost, time, tokens, artifacts, validation, and missing evidence come from the pinned harness ledgers.
  7. Every result is a single recorded attempt. Missing evidence and lower-bound coverage remain attached to the corresponding row.

The MacBook Kimi cost is ≥$10.85. Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.

Sponsors can support RemakeBench. They cannot buy benchmark outcomes, rankings, blind votes, or editorial verdicts.

Nine frozen targets · 33 recorded results

Inspect the output, then open the ledger.

Each pair shows the recorded Kimi and Fable output. The receipt table keeps every featured GPT-5.6 configuration and its material caveats visible.

Test 01 · Shader

Neo-Gothic Storm City

A storm-lashed neo-gothic city shader judged on architecture, atmosphere, lighting, water, and motion.

Editorial assessment

Fable Max is the visual-quality winner. Kimi beats the displayed Sol results on water and lighting, but takes much longer.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
Neo-Gothic Storm City recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxBase modelOpenCode Go / Moonshot AI$2.30API-equivalent estimate38:29.074 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end workflow latency including tool execution, not model-only compute.

Evidence still missing (4)
  • initial local FPS
  • optimized local FPS
  • final RTX Pro 6000 1080p60 render or capture metadata
  • blind-evaluation record
Open result
Claude Fable 5 MaxBase modelAnthropic Claude Code$11.39API-equivalent estimate28:48 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency including user idle time between turns, not model-only compute.

Evidence still missing (5)
  • TWIGL compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
GPT-5.6 Sol xhighBase modelOpenAI Codex$1.96API-equivalent estimate8:49.5 wall-clockpartial token timing and artifact ledger

Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context. Separately priced tools and non-token services are excluded.

Evidence still missing (4)
  • initial local FPS
  • optimized local FPS
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraBase modelOpenAI Codex$2.84API-equivalent estimate22:10.9 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • initial local FPS
  • optimized local FPS
  • blind-evaluation record
Open result
Evidence policy

Single recorded attempts; each row preserves its own missing-evidence disclosure.

Test 02 · Blender

MacBook-Class Cinematic

A polished product-ad shot with a modeled laptop, industrial detail, intentional materials, lighting, and camera movement.

Editorial assessment

Fable Max is best overall. Kimi preserves the MacBook-class chassis shape better than Sol xhigh but gets the hinge wrong. The archived validator logs report 49/49 checks for the compared scenes; those checks do not establish equal visual quality.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
MacBook-Class Cinematic recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxTool-assistedKimi Code CLI≥$10.85API-equivalent estimate1:26:58.2 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checks

Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.

Evidence still missing (5)
  • metrics reconciliation for the one unpaired API request
  • independent headless validator rerun
  • render-hardware identity
  • Kimi tool-use transcript or disclosure
  • blind-evaluation record
Open result
Claude Fable 5 MaxTool-assistedAnthropic Claude Code$75.18API-equivalent estimate1:01:04.8 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checks

All-1-hour cache-write alternative: $81.28. Primary cost assumes 5-minute cache writes because the supplied usage does not record TTLs.

Evidence still missing (5)
  • independent headless validator rerun
  • render-hardware identity
  • Blender MCP tool-use transcript or disclosure
  • final capture or rubric stills
  • blind-evaluation record
Open result
GPT-5.6 Sol xhighTool-assistedOpenAI Codex$7.67API-equivalent estimate55:49.9 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checks

Requests with more than 272,000 input tokens use long-context pricing; all 81 supplied calls were short-context. Separately priced tools and non-token services are excluded.

Evidence still missing (3)
  • render-hardware identity
  • Blender MCP tool-use transcript or disclosure
  • blind-evaluation record
Open result
Test 03 · Browser game

JRPG Boss Battle

An interactive Three.js boss battle using supplied assets, required combat beats, animation, effects, sound controls, and a validator-ready loop.

Editorial assessment

Fable Max creates the more cinematic battle and broader effects. Kimi is the value winner: Fable cost 43.1× as much in this recorded run. The archived scorecards report 41/41 functional checks for all three; they do not erase the visible quality differences or replace an independent rerun.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
JRPG Boss Battle recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxFull production stackKimi Code CLI / Moonshot AI$2.86API-equivalent estimate41:20.2 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checks

The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.

Evidence still missing (3)
  • browser and local-hardware identity
  • independent validator rerun
  • blind-evaluation record
Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$123.42API-equivalent estimate50:48.1 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checks

The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.

Evidence still missing (4)
  • browser and local-hardware identity
  • independent validator rerun
  • model tool-use transcript or disclosure
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraFull production stackOpenAI Codex$9.98API-equivalent estimate25:51.8 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checks

Priority service-tier pricing. Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context. No cache-write amount was inferred from the supplied transcript schema. Separately priced tools and non-token services are excluded.

Evidence still missing (4)
  • browser and local-hardware identity
  • independent validator rerun
  • model tool-use transcript or disclosure
  • blind-evaluation record
Open result
Test 04 · Interactive scene

Campfire Under a Starry Night

A polished interactive Three.js campfire judged on fire, light, environmental detail, and controllable camera motion.

Editorial assessment

Fable Max has the strongest fire and environment; Fable Medium is fastest; Kimi is the value result at $0.45.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
Campfire Under a Starry Night recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxFull production stackKimi Code CLI$0.45API-equivalent estimate14:19.3 wall-clockpartial token timing and source ledger

Wall-clock is end-to-end workflow latency and includes tool execution, browser checks, approvals, and idle time.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$11.17API-equivalent estimate12:48.9 wall-clockpartial token timing and source ledger

Wall-clock is end-to-end latency, including tool execution, idle gaps, and human think time; it is not model-only compute.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Open result
Claude Fable 5 MediumFull production stackAnthropic Claude Code$5.12estimated API cost4:24.6 wall-clockpartial token timing and source ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • final capture
  • browser and hardware environment
  • blind-evaluation record
Open result
GPT-5.6 Sol xhighFull production stackOpenAI Codex$2.26API-equivalent estimate28:52.6 wall-clockpartial token timing and source ledger

Requests with more than 272,000 input tokens use long-context pricing; all 44 supplied calls were short-context. Separately priced tools and non-token services are excluded.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Open result
Test 05 · Browser game

Space Flight Game

A playable browser space-flight game with responsive controls, a coherent environment, lighting, assets, and a game loop.

Editorial assessment

Fable Max has the visual ceiling. Kimi beats the displayed Sol Ultra result visually and costs $3.25, but its workflow takes about an hour.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
Space Flight Game recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxFull production stackKimi Code CLI / Moonshot AI$3.25API-equivalent estimate1:04:11 wall-clockpartial token timing and source ledger

Wall-clock is end-to-end workflow latency, including tool execution, manual-approval waits, browser playtests, and idle time.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$151.55API-equivalent estimate1:05:19.2 wall-clockpartial token timing and source ledger

Primary cost uses 5-minute cache writes. All-1-hour cache-write alternative: $167.36. No deduplication audit was supplied for this usage receipt; the token and cost figures retain the supplied basis.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Open result
Claude Fable 5 MediumFull production stackAnthropic Claude Code$22.85API-equivalent estimate18:37.7 wall-clockpartial token timing and source ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (4)
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraFull production stackOpenAI Codex$17.94API-equivalent estimate37:44.2 wall-clockpartial token timing and source ledger

Single supplied pre-request snapshot; repeat runs not attached

Evidence still missing (4)
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Open result
Test 06 · Blender

Shrine Village

An inspectable voxel shrine-village scene judged on composition, architecture, landscaping, water, and detail.

Editorial assessment

Kimi wins the editorial comparison against Sol Ultra on bridge placement, the lamp, and water. It gets surprisingly close to Fable Medium for far less money, but takes longer.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
Shrine Village recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxTool-assistedKimi Code CLI$2.23API-equivalent estimate35:53.5 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.

Evidence still missing (5)
  • independent Blender scene and build-script validation
  • final capture metadata including renderer, hardware, and resolution
  • initial and optimized local performance measurements
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$20.28estimated API cost18:37.5 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$1.98API-equivalent estimate20:36.2 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
Test 07 · Shader

Infinite Cathedral

An infinite stained-glass cathedral corridor shader judged on depth, architectural repetition, light, reflections, and motion.

Editorial assessment

Fable Max wins visual quality. Sol Ultra is the practical choice in this use case. Kimi lacks architectural detail and floor reflections and needs significant handholding despite its low price.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Max standardized benchmark capture poster
Claude Fable 5 Max
Infinite Cathedral recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxBase modelKimi Code CLI$0.92API-equivalent estimate21:54.1 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end workflow latency and includes tools, waits, and idle time.

Evidence still missing (5)
  • Shadertoy or WebGL2 compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
Claude Fable 5 MaxBase modelAnthropic Claude Code$43.85API-equivalent estimate31:13.8 wall-clockpartial token timing and artifact ledger

Primary cost uses 1-hour cache writes (supplied). All-5-minute cache-write alternative: $38.66.

Evidence still missing (5)
  • Shadertoy or WebGL2 compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
GPT-5.6 Sol xhighBase modelOpenAI Codex$4.10API-equivalent estimate24:19.8 wall-clockpartial token timing and source ledger

Requests with more than 272,000 input tokens use long-context pricing; all 57 supplied calls were short-context. Separately priced tools and non-token services are excluded.

Evidence still missing (5)
  • Shadertoy compatibility adapter or direct Shadertoy source
  • initial local FPS
  • optimized local FPS
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraBase modelOpenAI Codex$5.61API-equivalent estimate38:22.4 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • initial local FPS
  • optimized local FPS
  • blind-evaluation record
Open result
Evidence policy

Single recorded attempts; each row preserves its own missing-evidence disclosure.

Test 08 · Blender

Oasis Outpost

An inspectable voxel oasis outpost judged on settlement readability, terrain, vegetation, water, and environmental detail.

Editorial assessment

Kimi beats the displayed GPT-5.6 outputs on quality and gives Fable the only close visual competition. Fable cost 7.8× as much as Kimi in this recorded pair; Kimi is slower but looks steerable with iteration.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
Oasis Outpost recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxTool-assistedKimi Code CLI$2.72API-equivalent estimate39:27.9 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.

Evidence still missing (5)
  • independent Blender scene and build-script validation
  • final capture metadata including renderer, hardware, and resolution
  • initial and optimized local performance measurements
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$21.23estimated API cost19:23.7 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$5.51API-equivalent estimate16:30.9 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Terra UltraTool-assistedOpenAI Codex$1.33API-equivalent estimate28:14.2 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
Test 09 · Blender

Jungle Temple

An inspectable high-density voxel jungle-temple diorama created at the requested fixed grid scale.

Editorial assessment

Fable Medium wins visual quality and speed. Kimi is inconsistent here and takes 53:37.7, making this the clearest counterexample to its value story.

Kimi K3 Max standardized benchmark capture poster
Kimi K3 Max
Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
Jungle Temple recorded cost and time receipts
ConfigurationProvider / clientWorkflow costWall-clockRecordCaveatsOpen result
Kimi K3 MaxTool-assistedKimi Code CLI$2.60API-equivalent estimate53:37.7 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.

Evidence still missing (5)
  • independent Blender scene and build-script validation
  • final capture metadata including renderer, hardware, and resolution
  • initial and optimized local performance measurements
  • RTX Pro 6000 final render
  • blind-evaluation record
Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$10.10estimated API cost12:04.2 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$3.23API-equivalent estimate26:39.7 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
GPT-5.6 Terra UltraTool-assistedOpenAI Codex$1.20API-equivalent estimate16:43.2 wall-clockpartial token timing and artifact ledger

Wall-clock is end-to-end latency, not model-only compute.

Evidence still missing (3)
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Open result
See all 10 public Kimi K3 records →
Builder

Get the projects behind this test.

Open 33 reviewed project outputs as a full Project Pack or nine task-sized downloads: editable sanitized Blender scenes, browser projects, shaders, permitted assets, exact prompts, configurations, and receipts. Customer status: 25 ready · 5 reference-only · 3 partial · 0 excluded. Reference-only and partial projects retain their exact prerequisites and limitations.

Verified member Discord

Keep judging the builds with us.

Join the free RemakeBench Discord after you have inspected the verdict, videos, and public receipts.