RemakeBench Tech Review 001

GPT-5.6 vs Claude Fable 5

Same frozen task prompt. One request per model. Seven workflows. Every recorded dollar.

Same target. Frozen prompt. Disclosed stack. Honest result.

Lower total cost
GPT-5.6 Sol Ultra$43.97 API-equivalent estimate
Faster completion
Claude Fable 5 Medium97:00.9 total workflow time

No synthetic overall winner. Quality and community preference remain separate evidence categories.

  • 7benchmark prompts
  • 28recorded runs
  • 4configurations
  • Disclosedestimated API costs
  • 5 activeof seven blind matchups
Builder research package

Audit the workflow behind this result.

Get the 7 exact frozen prompts, 28 available ledgers, four configurations, and cost/time exports from Tech Review 001.

Headline receipts

Verdict first. Receipts immediately after.

Cost and time are earned from seven public run ledgers per headline configuration. Missing quality, preference, and repeat evidence stays missing.

Lower total costSol Ultra

$43.97 Sol API-equivalent estimate versus $107.84 Fable combined estimate (API-equivalent estimate + estimated API cost). Fable cost 2.45× as much.

Faster total completionFable Medium

97:00.9 versus 185:12.7. Sol took 1.91× as long.

Editorial qualityFable sets the visual ceiling. Sol Ultra wins the value argument.

Fable produced some of the strongest high-end environments in the suite, especially Storm City and Space Flight. Sol Ultra won Cathedral and often came close for $43.97 versus Fable’s $107.84, while taking 1.91× as long. With one recorded run per configuration per test, reliability is not established.

Community preferenceCollecting blind votes

No percentage appears before a real minimum sample exists.

ReliabilityNot established — single recorded runs

Repeat-run evidence has not been attached.

Watch the 8:51 technical verdict →
Methodology

One request. No rescue pass.

  1. One user request per run, with no human follow-up or manual correction after submission.
  2. The benchmark target and frozen task prompt matched within each round.
  3. Provider-specific agent stacks could use their disclosed tools and internal workflows; the stacks were not identical.
  4. Costs are estimates based on recorded token usage and each record’s cited pricing basis—not subscription invoices.
  5. Published benchmark versions stay immutable, and missing evidence remains visible.

Sponsors can support RemakeBench. They cannot buy benchmark outcomes, rankings, blind votes, or editorial verdicts.

Open the public methodology and evidence →
Seven frozen targets

Every round, every recorded configuration.

Fable Medium and Sol Ultra lead the comparison. Terra Ultra and Luna extra-high remain visible as capability and value references.

Round 01 · Shader

Infinite Cathedral Corridor

Create an infinite stained-glass cathedral corridor shader with convincing depth, architectural repetition, light, and motion.

Recorded comparisonFable Medium cost 3.4× as muchFable Medium finished fasterEditorial assessment

Sol Ultra is the quality winner. Terra is the value surprise, while Luna shows the capability floor.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Post-hoc reproduction

Post-hoc RTX PRO 6000 evidence verified

1920×1080 · 60 fps · 900 frames

Source hash matched · hardware and renderer attested
Deterministic offline capture + separate sustained-performance test

Claude Fable 5 Medium1510.21 FPSindependently recomputed aggregate
Uncapped trials
1512.39 · 1510.32 · 1507.91 FPS
Frame time
p50 0.656 ms · p95 0.731 ms · p99 0.755 ms
Trial variability
0.148% sample CV
Sustained 60 fps
Pass · 0 late frames · 0 deadline misses
Classification
qualifying_posthoc_reproduction

Three post-hoc performance trials of one generated artifact; model-generation sample n=1.

GPT-5.6 Sol Ultra387.61 FPSindependently recomputed aggregate
Uncapped trials
387.75 · 387.62 · 387.46 FPS
Frame time
p50 2.586 ms · p95 2.631 ms · p99 2.651 ms
Trial variability
0.037% sample CV
Sustained 60 fps
Pass · 0 late frames · 0 deadline misses
Classification
qualifying_posthoc_reproduction

Three post-hoc performance trials of one generated artifact; model-generation sample n=1.

Fable produced 3.90× Sol’s measured post-hoc EGL throughput on this task.The raw FPS number is not a visual-quality verdict, and these two opposing task results do not establish either model as generally more performant.

Post-hoc uncapped headless EGL/OpenGL throughput on an NVIDIA RTX PRO 6000 at 1920×1080. This is not the original local FPS or native TWIGL/ShaderToy browser performance.

Claude Fable 5 · Medium

Fable Medium

partial token timing and artifact ledger
Workflow cost
$19.17 estimated API cost
Workflow time
15:38.1 wall-clock
Processed tokens
6.95M
Reasoning
Medium

Anthropic Claude Code · ShaderToy fragment shader · Harness v1 infinite cathedral prompt

User-supplied report using official Anthropic API rates; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and artifact ledger
Workflow cost
$5.61 API-equivalent estimate
Workflow time
38:22.4 wall-clock
Processed tokens
5.78M
Reasoning
Ultra

OpenAI Codex · ShaderToy fragment shader · Harness v1 infinite cathedral prompt

API-equivalent estimate, not a subscription invoice

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$0.55API-equivalent estimate · 20:39.8 wall-clockLuna extra-high$0.21API-equivalent estimate · 9:18.2 wall-clock
Evidence still missing

initial local FPS · optimized local FPS · blind-evaluation record

Round 02 · Shader

Neo-Gothic Storm City

Create a storm-lashed neo-gothic city shader with legible architecture, atmosphere, lighting, water, and motion.

Recorded comparisonFable Medium cost 3.2× as muchFable Medium finished fasterEditorial assessment

Fable wins this round. Its ocean, fog and lightning create the more convincing storm, even if some towers fracture under closer inspection.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Post-hoc reproduction

Post-hoc RTX PRO 6000 evidence verified

1920×1080 · 60 fps · 900 frames

Source hash matched · hardware and renderer attested
Deterministic offline capture + separate sustained-performance test

Claude Fable 5 Medium821.75 FPSindependently recomputed aggregate
Uncapped trials
822.00 · 821.85 · 821.41 FPS
Frame time
p50 1.217 ms · p95 1.322 ms · p99 1.379 ms
Trial variability
0.037% sample CV
Sustained 60 fps
Pass · 0 late frames · 0 deadline misses
Classification
qualifying_posthoc_reproduction

Three post-hoc performance trials of one generated artifact; model-generation sample n=1.

GPT-5.6 Sol Ultra2045.20 FPSindependently recomputed aggregate
Uncapped trials
2049.47 · 2045.13 · 2041.02 FPS
Frame time
p50 0.491 ms · p95 0.546 ms · p99 0.565 ms
Trial variability
0.207% sample CV
Sustained 60 fps
Pass · 0 late frames · 0 deadline misses
Classification
qualifying_posthoc_reproduction

Three post-hoc performance trials of one generated artifact; model-generation sample n=1.

Sol produced 2.49× Fable’s measured post-hoc EGL throughput on this task.The raw FPS number is not a visual-quality verdict, and these two opposing task results do not establish either model as generally more performant.

Post-hoc uncapped headless EGL/OpenGL throughput on an NVIDIA RTX PRO 6000 at 1920×1080. This is not the original local FPS or native TWIGL/ShaderToy browser performance.

Claude Fable 5 · Medium

Fable Medium

partial token timing and artifact ledger
Workflow cost
$9.09 estimated API cost
Workflow time
8:15 wall-clock
Processed tokens
1.93M
Reasoning
Medium

Anthropic Claude Code · TWIGL fragment shader · Harness v1 neo-gothic storm city prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and artifact ledger
Workflow cost
$2.84 API-equivalent estimate
Workflow time
22:10.9 wall-clock
Processed tokens
2.81M
Reasoning
Ultra

OpenAI Codex · TWIGL classic fragment shader · Harness v1 neo-gothic storm city prompt

API-equivalent estimate, not a subscription invoice

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$3.27API-equivalent estimate · 16:35.1 wall-clockLuna extra-high$0.21API-equivalent estimate · 8:46.5 wall-clock
Evidence still missing

raw Claude Code meta.json · initial local FPS · optimized local FPS · blind-evaluation record

Round 03 · Browser game

Space Flight Game

Build a playable browser space-flight game with responsive controls, a coherent environment, lighting, assets, and a game loop.

Recorded comparisonFable Medium cost 1.3× as muchFable Medium finished fasterEditorial assessment

Fable creates the more expansive environment and stronger visual detail. Terra has the best flight controls of the four.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Claude Fable 5 · Medium

Fable Medium

partial token timing and source ledger
Workflow cost
$22.85 API-equivalent estimate
Workflow time
18:37.7 wall-clock
Processed tokens
13.96M
Reasoning
Medium

Anthropic Claude Code · Three.js source project · Generated GLB spaceship asset · Harness v1 space-flight prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and source ledger
Workflow cost
$17.94 API-equivalent estimate
Workflow time
37:44.2 wall-clock
Processed tokens
25.9M
Reasoning
Ultra

OpenAI Codex · Three.js / app-router source bundle · Generated GLB spaceship asset · Harness v1 space-flight prompt

User-supplied API-equivalent pre-request snapshot

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$1.44API-equivalent estimate · 16:51.4 wall-clockLuna extra-high$0.33API-equivalent estimate · 10:55.9 wall-clock
Evidence still missing

browser and hardware environment · local FPS · final capture · blind-evaluation record

Vote on this campaign
Round 04 · Interactive scene

Campfire Under a Starry Night

Build a polished interactive Three.js campfire with convincing fire, light, environmental detail, and controllable camera motion.

Recorded comparisonSol Ultra cost 1.3× as muchFable Medium finished fasterEditorial assessment

Sol cost more than Fable. Luna was cheapest and included more controls; no clean quality winner was claimed.

Claude Fable 5 Medium benchmark output capture
Claude Fable 5 Medium
GPT-5.6 Sol Ultra benchmark output capture
GPT-5.6 Sol Ultra
Claude Fable 5 · Medium

Fable Medium

partial token timing and source ledger
Workflow cost
$5.12 estimated API cost
Workflow time
4:24.6 wall-clock
Processed tokens
1.14M
Reasoning
Medium

Anthropic Claude Code · Three.js interactive scene · Harness v1 campfire prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and source ledger
Workflow cost
$6.86 API-equivalent estimate
Workflow time
23:08.4 wall-clock
Processed tokens
8.72M
Reasoning
Ultra

OpenAI Codex · Three.js source project · Harness v1 campfire prompt

API-equivalent estimate, not a subscription invoice

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$1.86API-equivalent estimate · 24:28.8 wall-clockLuna extra-high$0.56API-equivalent estimate · 8:40.2 wall-clock
Evidence still missing

final capture · browser and hardware environment · blind-evaluation record · local FPS

Vote on this campaign
Round 05 · Blender

Jungle Temple

Create an inspectable high-density voxel jungle-temple diorama in Blender at the requested fixed grid scale.

Recorded comparisonFable Medium cost 3.1× as muchFable Medium finished fasterEditorial assessment

Fable creates the more complete and coherent scene. Sol has the preferred bridge design, while Terra and Luna expose the capability gap.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Claude Fable 5 · Medium

Fable Medium

partial token timing and artifact ledger
Workflow cost
$10.10 estimated API cost
Workflow time
12:04.2 wall-clock
Processed tokens
2.52M
Reasoning
Medium

Anthropic Claude Code · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 jungle temple prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and artifact ledger
Workflow cost
$3.23 API-equivalent estimate
Workflow time
26:39.7 wall-clock
Processed tokens
2.56M
Reasoning
Ultra

OpenAI Codex · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 jungle temple prompt

API-equivalent estimate, not a subscription invoice

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$1.20API-equivalent estimate · 16:43.2 wall-clockLuna extra-high$0.37API-equivalent estimate · 10:55.9 wall-clock
Evidence still missing

Blender and render environment · final capture metadata · blind-evaluation record

Vote on this campaign
Round 06 · Blender

Shrine Village

Create an inspectable voxel shrine-village scene in Blender with coherent composition, architecture, landscaping, and detail.

Recorded comparisonFable Medium cost 10.2× as muchFable Medium finished fasterEditorial assessment

Sol overperforms on value at $1.98 versus Fable’s $20.28, but floating roof artifacts and an unusable bridge prevent a clean quality win.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Claude Fable 5 · Medium

Fable Medium

partial token timing and artifact ledger
Workflow cost
$20.28 estimated API cost
Workflow time
18:37.5 wall-clock
Processed tokens
5.94M
Reasoning
Medium

Anthropic Claude Code · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 shrine village prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and artifact ledger
Workflow cost
$1.98 API-equivalent estimate
Workflow time
20:36.2 wall-clock
Processed tokens
1.53M
Reasoning
Ultra

OpenAI Codex · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 shrine village prompt

API-equivalent estimate, not a subscription invoice

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$1.30API-equivalent estimate · 19:30.7 wall-clockLuna extra-high$0.44API-equivalent estimate · 14:02.1 wall-clock
Evidence still missing

Blender and render environment · final capture metadata · blind-evaluation record

Vote on this campaign
Round 07 · Blender

Oasis Outpost

Create an inspectable voxel oasis outpost in Blender with a readable settlement, terrain, vegetation, water, and environmental detail.

Recorded comparisonFable Medium cost 3.9× as muchSol Ultra finished fasterEditorial assessment

Fable creates the broadest environment. Terra preserves a complete, readable oasis at $1.33; no clean quality winner was claimed.

Claude Fable 5 Medium standardized benchmark capture poster
Claude Fable 5 Medium
GPT-5.6 Sol Ultra standardized benchmark capture poster
GPT-5.6 Sol Ultra
Claude Fable 5 · Medium

Fable Medium

partial token timing and artifact ledger
Workflow cost
$21.23 estimated API cost
Workflow time
19:23.7 wall-clock
Processed tokens
6.57M
Reasoning
Medium

Anthropic Claude Code · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 oasis outpost prompt

Anthropic first-party Claude API, standard global pricing; 5-minute cache-write assumption

GPT-5.6 Sol · Ultra

Sol Ultra

partial token timing and artifact ledger
Workflow cost
$5.51 API-equivalent estimate
Workflow time
16:30.9 wall-clock
Processed tokens
1.85M
Reasoning
Ultra

OpenAI Codex · Blender MCP · 0.05-unit voxel grid · Inspectable .blend artifact · Harness v1 oasis outpost prompt

Priority API short-context rates as an API-equivalent scenario because the supplied transcript observed service_tier=priority; this is not proof of an API invoice.

Additional references

Same target, disclosed alternative configurations.

Terra Ultra$1.33API-equivalent estimate · 28:14.2 wall-clockLuna extra-high$0.34API-equivalent estimate · 9:43.6 wall-clock
Evidence still missing

Blender and render environment · final capture metadata · blind-evaluation record

Vote on this campaign
5 of seven blind matchups are live

Judge the outputs before you see the model.

Explore each real build for at least 15 seconds. Sign in only when you record the vote. Identity, cost, tokens, time, and evidence reveal afterward.

Builder research package

Take the exact prompts, available ledgers, configurations, and exports with you.

Get prompts + ledgers