RemakeBench · Launch 004 · published comparison

Claude Opus 5 across nine game-building tests

Opus 5 delivered the strongest detail in several rounds. It also made the efficiency argument against itself.

Across this suite, Opus 5 often spent more time and API-equivalent value to reach stronger visual detail. Shrine Exploration shows the upside of a seven-hour long-horizon run—and the animation, character and collision debt that current models still leave behind.

Single recorded attempts. Reliability is not established.

Workflow disclosure

This is a workflow comparison, not a bare-model test. Claude Opus 5 and Claude Fable 5 ran through Claude Code; GPT-5.6 Sol ran through Codex; Kimi K3 ran through Kimi CLI. Each benchmark run used one frozen request, with no human follow-up or manual correction after submission. The agents could use their disclosed tools, internal retries and subagents where supported.

The exact Builder components are included for active Builder members.

  • 9episode rounds
  • 32exact records
  • 14new hash-bound clips
  • 1frozen request rule
  • Rawtime, token and cost receipts
The public answer

Nine frozen tests, in the script’s order.

Every row preserves its exact configuration, measured values, artifact state, validation scope and permanent record. Media appears only when its result identity is exact.

Round 01 · 3 results

Off-road Mud Game

One-request procedural off-road driving, vehicle physics, suspension, terrain streaming, mud feedback and multiple cameras.

Owner assessment

Neither Opus 5 Max nor GPT-5.6 Sol Max produced a clean one-shot result. Opus has invisible tires and doors, an obstructed cockpit view and a tendency to catapult, though its forest and wind animation are strong. Sol finished faster and cheaper with the more planted vehicle and glossy mud. Kimi keeps its doors and tires but also catapults and obstructs the cockpit view.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
GPT-5.6 Sol · Max presentation recording poster
GPT-5.6 Sol · MaxLaunch 004 episode capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxLaunch 004 episode capture · audio removed
Off-road Mud Game exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codeoff-road-driving-game-threejs-opus-5-max2h 48m 08.6s149,702,730$124.35Official Anthropic API-list-price equivalent, not an itemized Claude Code subscription charge
1 cost qualifier
  • All-5-minute cache-write alternative: $113.74.
Fresh Vite `dist/` directory archiveRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency, including tool execution, package installation, browser automation, WASM initialization, and model latency; it is not model-only compute.Open result
GPT-5.6 Sol · MaxOpenAI Codexoff-road-driving-game-threejs-gpt-5.6-sol-max1h 2m 40.9s25,940,954$19.10Token-only standard API-equivalent estimate, not the actual Codex subscription charge
1 cost qualifier
  • The source receipt applies long-context rates only when a call exceeds 272,000 input tokens. It reports zero such calls.
source project package manifestRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency, including tools and idle gaps; it is not model-only compute time.Open result
Kimi K3 · MaxKimi Code CLI / Moonshot AIoff-road-driving-game-threejs-kimi-k3-max3h 50m 07.6s64,179,371$28.72API-equivalent usage accounting, not an itemized subscription cash chargefresh vite dist directory archiveRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, package installation, browser automation, WASM initialization, and model latency; it is not model-only compute.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
  • Cache-read tokens are discounted, so the total-processed figure materially overstates effective cost.
  • The cache-write TTL mix is not independently reconstructible from the transcript; the primary uses the source receipt's one-hour declaration and records the all-five-minute alternative.
  • The retained `.verify` scripts and captures are model-authored output. Their console results are not supplied and they were not independently rerun, so they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
  • No standardized final gameplay capture or blind-evaluation record is supplied.
Still missing
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • multi-minute sustained-drive and generation-seam receipt
  • four-corner suspension-articulation receipt
  • standardized final gameplay capture
  • blind-evaluation record

GPT-5.6 Sol · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, including tools and idle gaps; it is not model-only compute time.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
  • The generated production output is deliberately not archived because its server-side build files contain a build-time prerender secret. The complete committed source required to regenerate it is archived instead.
  • The production build passes, but the supplied npm test command fails two stale starter-scaffold assertions. No independent browser acceptance, sustained-drive, local-FPS, seam, or suspension-articulation receipt is archived.
  • No final gameplay capture or blind evaluation is supplied.
Still missing
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • multi-minute sustained-drive and generation-seam receipt
  • four-corner suspension-articulation receipt
  • final gameplay capture
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
  • One user prompt produced 355 billable Kimi model requests across one main and three child agents.
  • Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • No before/after /usage quota evidence was supplied, so membership-quota consumption cannot be measured from this run.
  • The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
  • The retained `.verify` scripts and captures are model-authored output and were not independently rerun; they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
  • No standardized final gameplay capture or blind-evaluation record is supplied.
Still missing
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • multi-minute sustained-drive and generation-seam receipt
  • four-corner suspension-articulation receipt
  • standardized final gameplay capture
  • blind-evaluation record
Round 02 · 4 results

Space Flight Stage 2 — Landfall

A second frozen request that extends each Stage 1 space-flight project with seamless landing, takeoff, an explorable surface and a fixed character asset.

Owner assessment

Opus and Fable both created appealing but different environments. Opus supports flexible crash landing but flickers and makes the valid landing region hard to find; Fable automates landing but launches the ship facing down. Sol's seamless planet is much smaller and cheaper. Kimi's landing is the hardest to complete and distorts the character, although its environment has a distinct visual appeal.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Max presentation recording poster
Claude Fable 5 · MaxLaunch 004 episode capture · audio removed
GPT-5.6 Sol · Ultra presentation recording poster
GPT-5.6 Sol · UltraLaunch 004 episode capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxLaunch 004 episode capture · audio removed
Space Flight Stage 2 — Landfall exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codespace-flight-game-threejs-stage-2-landfall-opus-5-max2h 31m 33.5s207,637,757$133.83Claude API list-price equivalent, not an itemized Claude Code charge
1 cost qualifier
  • All-5-minute cache-write alternative: $129.37.
source project package manifestRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.Open result
Claude Fable 5 · MaxAnthropic Claude Codespace-flight-game-threejs-stage-2-landfall-fable-5-max2h 05m 20.4s193,034,691$280.14Claude API list-price equivalent, not an itemized Claude Code charge
1 cost qualifier
  • All-1-hour cache-write alternative: $308.12.
vite production build tarballRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.Open result
GPT-5.6 Sol · UltraOpenAI Codexspace-flight-game-threejs-stage-2-landfall-gpt-5.6-sol-ultra1h 29m 57.7s29,319,597$20.84Official OpenAI standard API-list-price equivalent; not an itemized Codex Pro subscription charge
1 cost qualifier
  • Separately priced tools and non-token services are excluded.
Fresh vinext `dist/` directory archiveRecord status: partial token timing source build static acceptance ledgerNo validation result attachedWall-clock is end-to-end workflow latency, including tools and idle time, not model-only compute.Open result
Kimi K3 · MaxKimi Code CLI / Moonshot AIspace-flight-game-threejs-stage-2-landfall-kimi-k3-max2h 23m 18.1s39,986,997$15.58API-equivalent usage accounting, not an itemized subscription cash chargesource project package manifestRecord status: partial token timing source build ledgerNo validation result attachedWall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • The archived Vite build passes, but no independent full-loop browser acceptance run, local FPS/transition-hitch receipt, final capture, or blind evaluation is archived.
  • The one-hour cache-write assumption is explicitly reported by the source receipt but cannot be independently reconstructed from the sanitized token totals.
Still missing
  • independent uninterrupted landing-to-orbit acceptance run
  • local FPS and transition-hitch receipt with hardware details
  • final viewport or browser capture
  • blind-evaluation record

Claude Fable 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • The source receipt does not distinguish cache-write TTLs. The primary estimate follows its stated five-minute assumption; the all-one-hour alternative is included separately.
  • The archive build passes, but no independent full-loop browser acceptance run, local FPS/transition-hitch receipt, final capture, or blind evaluation is archived.
Still missing
  • independent uninterrupted landing-to-orbit acceptance run
  • local FPS and transition-hitch receipt with hardware details
  • final viewport or browser capture
  • blind-evaluation record

GPT-5.6 Sol · Ultra

Material caveats
  • Wall-clock is end-to-end workflow latency, including tools and idle time, not model-only compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a Codex Pro subscription-backed session; separately priced tools and non-token services are excluded.
  • The acceptance evidence is model-authored static/server-rendered testing plus a fresh build and lint run. No independent full-loop browser playthrough, FPS/hitch receipt, or final runtime capture is supplied.
  • The supplied social preview is retained for provenance but is not used as a final render or quality score.
  • No blind-evaluation record is supplied.
Still missing
  • independent uninterrupted orbit-to-surface-to-orbit acceptance run
  • local FPS and transition-hitch receipt with hardware details
  • final browser or viewport capture
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
  • One user prompt produced 202 billable Kimi model requests.
  • Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • The source receipt reports no before/after /usage quota evidence; it cannot quantify the subscription quota consumed by this run.
  • The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
  • The archive build passes, but no independent browser acceptance run, local FPS/hardware measurement, final capture, or blind evaluation is archived.
Still missing
  • independent uninterrupted landing-to-orbit acceptance run without test fallback
  • local FPS and hardware receipt, including transition hitch measurement
  • final capture with viewport/browser evidence
  • blind-evaluation record
Round 03 · 4 results

Neo-Gothic Flooded City

A procedural neo-gothic storm-city shader balancing architecture, atmosphere, floodwater, reflections and motion.

Owner assessment

Opus took substantially longer and cost more than Fable, with more structural detail. Both satisfy the neo-gothic aesthetic: Fable's floodwater is more violent, while Opus's water looks more realistic and reflective. The final preference is a taste call, not a universal winner.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Max presentation recording poster
Claude Fable 5 · MaxExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.

Neo-Gothic Flooded City exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codeneo-gothic-storm-city-opus-5-max1h 52m 04.4s23,105,513$19.55API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: The session declares a one-hour prompt-cache TTL; the source does not record TTL per entry.
  • All-5-minute cache-write alternative: $17.94.
twigl fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency including tool execution, headless browser renders, validation, web fetches, and user think time; it is not model-only compute.Open result
Claude Fable 5 · MaxAnthropic Claude Codeneo-gothic-storm-city-fable-5-max28:484,770,667$11.39API-equivalent estimate, not a subscription invoicetwigl fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency including user idle time between turns, not model-only compute.Open result
GPT-5.6 Sol · XHighOpenAI Codexneo-gothic-storm-city-gpt-5.6-sol-xhigh8:49.51,618,640$1.96API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
twigl classic webgl1 fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxOpenCode Go / Moonshot AIneo-gothic-storm-city-kimi-k3-max38:29.0743,715,570$2.30API-equivalent usage accounting, not an itemized subscription cash chargetwigl classic webgl1 fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end workflow latency including tool execution, not model-only compute.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end latency including tool execution, headless browser renders, validation, web fetches, and user think time; it is not model-only compute.
  • Output tokens include hidden reasoning, tool-call JSON, and shader code, not only visible prose.
  • Cache reads are roughly 96% of processed tokens but bill at 0.1× base input, so processed tokens substantially overstate cost.
  • The primary $19.55 estimate applies the session's declared one-hour cache TTL; an all-five-minute cache-write alternative is $17.94.
  • The source receipt is a mid-session snapshot and excludes the later turn that wrote the receipt.
  • The shader was archived without a fresh TWIGL compile or browser run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
  • TWIGL compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record

Claude Fable 5 · Max

Material caveats
  • Wall-clock is end-to-end latency including user idle time between turns, not model-only compute.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache reads are billed at a deep discount, so total processed tokens overstate cost.
  • Raw JSONL per-content-block totals double-counted multi-block Claude Code messages. This ledger uses the supplied deduplicated billing basis instead.
  • Fable 5 Max was confirmed by the user after initial archival; the raw metrics receipt itself identifies only the base model.
  • The shader was archived without a fresh TWIGL compile or browser run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
  • TWIGL compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record

GPT-5.6 Sol · XHigh

Material caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Static compilation is not a substitute for Twigl runtime compatibility or measured FPS.
Still missing
  • initial local FPS
  • optimized local FPS
  • final capture metadata
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency including tool execution, not model-only compute.
  • OpenCode splits one agentic turn into multiple assistant/API messages (51 here).
  • Reasoning tokens and visible output tokens are recorded separately; the table's output-token figure is their combined total.
  • Cache-read tokens are heavily discounted, so total processed tokens overstate cost.
  • OpenCode Go ledger cost is API-equivalent usage accounting, not an itemized cash charge.
  • The bootstrap extraction is version-specific to OpenCode 1.18.3 until validated on more benchmark runs.
  • The raw metrics ledger records the session window but not an artifact hash; this result ledger supplies the SHA-256 binding to the archived shader.
Still missing
  • initial local FPS
  • optimized local FPS
  • final RTX Pro 6000 1080p60 render or capture metadata
  • blind-evaluation record
Round 04 · 4 results

Infinite Cathedral

A procedural stained-glass cathedral corridor shader with depth, repeated architecture, lighting and continuous motion.

Owner assessment

Opus produced one of the strongest cathedral results, especially in interior detail, but omitted pillars, exaggerated window light and let distant hanging lights glow from ceiling to floor. It also spent far longer than Fable. Opus is visually stronger than the shown Sol result, while Kimi trails the frontier models on this test.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Max presentation recording poster
Claude Fable 5 · MaxExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.

Infinite Cathedral exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codeinfinite-cathedral-shadertoy-opus-5-max1h 41m 22.5s32,282,622$29.25API-equivalent estimate, not a subscription invoice
1 cost qualifier
  • Cache-write basis: Exact TTL split supplied from transcript usage payload: 116,747 one-hour cache-write tokens and 1,286,290 five-minute cache-write tokens.
shadertoy fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock time is end-to-end latency, not model-only compute time; it includes shader compiles, browser renders, screenshot round-trips, and user think time.Open result
Claude Fable 5 · MaxAnthropic Claude Codeinfinite-cathedral-shadertoy-fable-5-max31:13.815,524,794$43.85API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: 1-hour cache writes (supplied)
  • All-5-minute cache-write alternative: $38.66.
shadertoy fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock time is end-to-end latency, not model-only compute time.Open result
GPT-5.6 Sol · XHighOpenAI Codexinfinite-cathedral-shadertoy-gpt-5.6-sol-xhigh24:19.83,913,289$4.10API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Requests with more than 272,000 input tokens use long-context pricing; all 57 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
webgl2 fragment shader source moduleRecord status: partial token timing and source ledgerNo validation result attachedWall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxKimi Code CLIinfinite-cathedral-shadertoy-kimi-k3-max21:54.1338,412$0.92API-equivalent estimate, not an itemized subscription cash chargeshadertoy fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end workflow latency and includes tools, waits, and idle time.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock time is end-to-end latency, not model-only compute time; it includes shader compiles, browser renders, screenshot round-trips, and user think time.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache-read tokens are billed at a deep discount, so total processed tokens overstate effective cost.
  • Pricing was verified live in the supplied receipt, and the cache-write TTL split is recorded exactly in the corrected usage payload.
  • The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
  • Shadertoy or WebGL2 compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record

Claude Fable 5 · Max

Material caveats
  • Wall-clock time is end-to-end latency, not model-only compute time.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache-read tokens are billed at a deep discount, so total processed tokens overstate effective cost.
  • The supplied 1-hour cache-write TTL is used for the primary estimate.
  • The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
  • Shadertoy or WebGL2 compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record

GPT-5.6 Sol · XHigh

Material caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • The submitted artifact is a WebGL2 application rather than a direct Shadertoy paste.
Still missing
  • Shadertoy compatibility adapter or direct Shadertoy source
  • initial local FPS
  • optimized local FPS
  • final capture metadata
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency and includes tools, waits, and idle time.
  • Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate effective cost.
  • The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
  • The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
  • Shadertoy or WebGL2 compile/run receipt
  • initial local FPS
  • optimized local FPS
  • RTX Pro 6000 final render
  • blind-evaluation record
Round 05 · 4 results

MacBook Pro

A validator-ready Blender product-ad scene with modeled industrial detail, materials, lighting and camera movement.

Owner assessment

The same specification still produces materially different products. Opus took only six minutes longer than Fable in the recorded runs while costing much less; Kimi remains especially competitive when cost is considered. The owner leaves the visual preference to the audience.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Max presentation recording poster
Claude Fable 5 · MaxExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.

MacBook Pro exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Code with Blender MCPmacbook-cinematic-blender-opus-5-max1h 07m 02.5s35,897,343$27.40API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: The session declares a one-hour prompt-cache TTL; the source does not record TTL per entry.
  • All-5-minute cache-write alternative: $26.03.
blender cinematic product ad sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end workflow latency, not model-only compute; the source receipt records Blender renders, boolean evaluation, and two headless Blender launches within this window.Open result
Claude Fable 5 · MaxAnthropic Claude Codemacbook-cinematic-blender-fable-5-max1:01:04.835,176,594$75.18API-equivalent estimate, not a subscription invoice
1 cost qualifier
  • All-1-hour cache-write alternative: $81.28.
blender cinematic product ad sceneRecord status: partial token timing artifact validation ledgerPASS, 49/49 checksWall-clock is end-to-end workflow latency, not model-only compute; the supplied record says Blender-side Cycles renders, boolean evaluation, and BVH sweeps account for a substantial portion of the session time.Open result
GPT-5.6 Sol · XHighOpenAI Codexmacbook-cinematic-blender-gpt-5.6-sol-xhigh55:49.97,464,477$7.67API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Requests with more than 272,000 input tokens use long-context pricing; all 81 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
blender cinematic product ad sceneRecord status: partial token timing artifact validation ledgerPASS, 49/49 checksWall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxKimi Code CLImacbook-cinematic-blender-kimi-k3-max1:26:58.2≥26,649,794$10.85+API-equivalent lower bound over paired usage records, not an itemized subscription cash charge
1 cost qualifier
  • Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.
blender cinematic product ad sceneRecord status: partial token timing artifact validation ledgerPASS, 49/49 checksWall-clock is end-to-end workflow latency and includes tool execution, Blender renders, approval waits, and idle time.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; the source receipt records Blender renders, boolean evaluation, and two headless Blender launches within this window.
  • Output tokens include hidden reasoning, generated Python, and tool-call JSON, not only visible prose.
  • Cache reads are about 98% of processed-token volume but bill at 0.1× base input, so processed tokens substantially overstate cost.
  • The primary $27.40 estimate applies the session's declared one-hour cache TTL; an all-five-minute cache-write alternative is $26.03.
  • The supplied receipt reports PASS 49/49, but no validator log was supplied and the archival environment cannot perform another Blender re-run.
  • No final 4K film render, render-hardware identity, Blender MCP tool-use transcript, or blind-evaluation record was supplied.
Still missing
  • retained validator log
  • final 4K film render and renderer identity
  • blind-evaluation record

Claude Fable 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; the supplied record says Blender-side Cycles renders, boolean evaluation, and BVH sweeps account for a substantial portion of the session time.
  • Output tokens include hidden reasoning, tool-bound code, and tool-call JSON, not just visible prose.
  • Cache reads are billed at a deep discount, so 35.2M processed tokens substantially overstate cost; most processed tokens were cache reads.
  • The supplied metrics ledger reports the base model claude-fable-5. The Max reasoning configuration is user-specified metadata.
  • The cache-creation aggregate does not identify TTL. If every cache write used the official one-hour rate, the total would be $81.28 rather than $75.18.
  • The supplied validation log reports PASS 49/49, but it was not independently rerun during archival because Blender is unavailable in this environment.
  • No final 4K film render, render-hardware identity, Blender MCP tool-use transcript, or blind-evaluation record was supplied.
Still missing
  • independent headless validator rerun
  • render-hardware identity
  • Blender MCP tool-use transcript or disclosure
  • final capture or rubric stills
  • blind-evaluation record

GPT-5.6 Sol · XHigh

Material caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • The metrics ledger does not identify the actual render GPU or preserve a Blender MCP tool-call transcript; the requested tool profile is recorded but not independently attested.
  • The 1267x712 Cycles stills are audit evidence; the task intentionally does not require executing the configured 4K final animation.
Still missing
  • render-hardware identity
  • Blender MCP tool-use transcript or disclosure
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency and includes tool execution, Blender renders, approval waits, and idle time.
  • One user prompt produced 123 Kimi model requests, but only 122 have reconciled usage records.
  • Kimi Code CLI 0.26.0 exposed no authoritative separate reasoning-token field; output tokens include provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate cost.
  • The supplied scene validator reports PASS 49/49, but it was not independently rerun during archival because Blender is unavailable in this environment.
  • The private session ZIP named in the receipt was not present in the supplied project folder and is intentionally not published.
  • No /usage screenshot, render-hardware identity, or blind-evaluation record was supplied.
Still missing
  • metrics reconciliation for the one unpaired API request
  • independent headless validator rerun
  • render-hardware identity
  • Kimi tool-use transcript or disclosure
  • blind-evaluation record
Round 06 · 4 results

JRPG Boss Battle

A supplied-asset Three.js boss battle with cinematic camera beats, combat state, magic effects, audio controls and validator-ready outcomes.

Owner assessment

Opus is a solid result and much cheaper than Fable, with an appealing floor and cinematic movement, but its generated backdrop cuts off at the top. Its magic composition is stronger than the shown Sol Ultra shortcut. Kimi meets the effects and game contract without notable bugs, but lacks the regular JRPG camera movement; its cost stands out.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Max presentation recording poster
Claude Fable 5 · MaxExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · Ultra. Their measured records remain below; no title-matched substitute was used.

JRPG Boss Battle exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codejrpg-boss-battle-opus-5-max1h 15m 49.9s49,933,667$48.95First-party Claude API list-price equivalent, not a Claude Code subscription invoice
2 cost qualifiers
  • Cache-write basis: The source receipt reports a one-hour prompt-cache TTL, so the primary calculation uses the one-hour cache-write rate.
  • All-5-minute cache-write alternative: $44.41.
interactive threejs jrpg boss battleRecord status: partial token timing artifact validation ledgerPASS, 41/41 checksWall-clock is end-to-end workflow latency, including tool execution, browser runs, image generation, and human-side idle time; it is not model-only compute.Open result
Claude Fable 5 · MaxAnthropic Claude Codejrpg-boss-battle-fable-5-max50:48.176,281,510$123.42API-equivalent estimate with exact cache-write TTL accounting, not a subscription invoiceinteractive threejs jrpg boss battleRecord status: partial token timing artifact validation ledgerPASS, 41/41 checksWall-clock is end-to-end workflow latency, not model-only compute; it includes tool execution and idle waits.Open result
GPT-5.6 Sol · UltraOpenAI Codexjrpg-boss-battle-gpt-5.6-sol-ultra25:51.85,046,213$9.98API-equivalent estimate, not a subscription invoice
4 cost qualifiers
  • priority service-tier pricing.
  • Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context.
  • No cache-write amount was inferred from the supplied transcript schema.
  • Separately priced tools and non-token services are excluded.
interactive threejs jrpg boss battleRecord status: partial token timing artifact validation ledgerPASS, 41/41 checksWall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxKimi Code CLI / Moonshot AIjrpg-boss-battle-kimi-k3-max41:20.25,769,491$2.86API-equivalent usage accounting, not an itemized subscription cash chargeinteractive threejs jrpg boss battleRecord status: partial token timing artifact validation ledgerPASS, 41/41 checksWall-clock includes tools, installs, browser checks, approval waits, and idle time.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, browser runs, image generation, and human-side idle time; it is not model-only compute.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
  • The exact browser version and local hardware identity were not supplied.
  • The metrics receipt reports a subscription-backed Claude Code session; the displayed amount is API-equivalent rather than an itemized cash charge.
Still missing
  • browser and local-hardware identity
  • independent blind-evaluation record

Claude Fable 5 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool execution and idle waits.
  • Output tokens include hidden reasoning, code, and tool-call JSON.
  • Cache reads are billed at a 90% discount to fresh input, so total processed tokens substantially overstate cost.
  • The supplied snapshot includes the metrics turns themselves.
  • The exact browser version and local hardware identity were not supplied.
  • The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
  • The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
  • browser and local-hardware identity
  • independent validator rerun
  • model tool-use transcript or disclosure
  • blind-evaluation record

GPT-5.6 Sol · Ultra

Material caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent Priority-tier estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • The exact browser version and local hardware identity were not supplied.
  • The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
  • browser and local-hardware identity
  • independent validator rerun
  • model tool-use transcript or disclosure
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock includes tools, installs, browser checks, approval waits, and idle time.
  • One user turn produced 62 billable Kimi model requests.
  • Kimi Code CLI 0.27.0 exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate cost.
  • Allegretto does not itemize a per-run cash charge; the $39 monthly subscription is not divided by this run.
  • The exact browser version and local hardware identity were not supplied.
  • The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
  • The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
  • browser and local-hardware identity
  • independent validator rerun
  • blind-evaluation record
Round 07 · 4 results

Jungle Temple

A high-voxel-density Blender jungle-temple diorama with coherent scene composition, inspectable geometry and fixed voxel scale.

Owner assessment

Opus produced a coherent, highly detailed scene with no major visual flaw, at materially higher cost than Fable. Both Opus and Sol use a large bridge, but Sol falls behind in scene detail; Kimi's scene is comparatively sparse.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Medium presentation recording poster
Claude Fable 5 · MediumExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · Ultra. Their measured records remain below; no title-matched substitute was used.

Jungle Temple exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codejungle-temple-opus-5-max52m 24.4s44,730,735$40.58First-party Claude API list-price equivalent, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: The source receipt states a one-hour prompt-cache TTL; the primary calculation uses the one-hour cache-write rate.
  • All-5-minute cache-write alternative: $38.23.
blender sceneRecord status: partial token timing artifact scene inspection ledgerNo validation result attachedWall-clock is end-to-end latency including Blender MCP tool execution, renders, file I/O, and waits; it is not model-only compute time.Open result
Claude Fable 5 · MediumAnthropic Claude Codejungle-temple-fable-5-medium12:04.22,519,008$10.10First-party API-list-price-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: 5-minute cache writes
  • Pricing basis: Anthropic first-party Claude API, standard global pricing.
blender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.Open result
GPT-5.6 Sol · UltraOpenAI Codexjungle-temple-gpt-5.6-sol-ultra26:39.72,555,331$3.23API-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Requests with more than 272,000 input tokens use long-context pricing; all 42 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
blender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxKimi Code CLIjungle-temple-kimi-k3-max53:37.74,180,898$2.60API-equivalent estimate, not an itemized subscription cash chargeblender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end latency including Blender MCP tool execution, renders, file I/O, and waits; it is not model-only compute time.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON.
  • Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
  • The displayed $40.58 workflow estimate includes the receipt's $0.01 web-search charge; the token-only estimate is $40.57.
  • The scene was read-only inspected with Blender 5.1.2, but no initial/optimized FPS measurement, render hardware receipt, RTX Pro 6000 render, or blind evaluation was supplied.
Still missing
  • initial and optimized local performance measurements
  • render hardware receipt
  • RTX Pro 6000 final render
  • blind-evaluation record

Claude Fable 5 · Medium

Material caveats
  • Wall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.
  • Output tokens include hidden reasoning, code, and tool-call JSON, not just visible text.
  • Cache-read tokens are billed at a deep discount, so total processed tokens greatly overstate effective cost.
  • The primary total uses the supplied five-minute cache-write rate; a one-hour cache write would increase the cost.
  • Blender and render environment, final capture metadata, and blind-evaluation evidence have not been supplied.
Still missing
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record

GPT-5.6 Sol · Ultra

Material caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Subagent logs are excluded to avoid double-counting inherited parent context.
Still missing
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.
  • Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate effective cost.
  • The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
  • The archived preview was visually inspected, but Blender is unavailable in the archival environment, so the scene and build log were not independently rerun.
  • No render-hardware identity, initial/optimized FPS measurement, RTX final render, or blind-evaluation record was supplied.
Still missing
  • independent Blender scene and build-script validation
  • final capture metadata including renderer, hardware, and resolution
  • initial and optimized local performance measurements
  • RTX Pro 6000 final render
  • blind-evaluation record
Round 08 · 4 results

Oasis Outpost

A high-voxel-density Blender oasis diorama testing scene population, architectural detail, pattern, shade and landscape composition.

Owner assessment

Opus packs the most detail into roofs, walls and terrain, perhaps making the landscape too dynamic, and also takes the longest. Sol wins on completion time and cost. Kimi works at a smaller scale but populates its scene well and is impressive for the price.

Claude Opus 5 · Max presentation recording poster
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Claude Fable 5 · Medium presentation recording poster
Claude Fable 5 · MediumExact earlier comparator capture · audio removed
Kimi K3 · Max presentation recording poster
Kimi K3 · MaxExact earlier comparator capture · audio removed

No exact public episode capture is bound for GPT-5.6 Sol · Ultra. Their measured records remain below; no title-matched substitute was used.

Oasis Outpost exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 · MaxAnthropic Claude Codeoasis-outpost-opus-5-max49m 10.1s27,763,841$34.68First-party Claude API list-price equivalent, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: Exact source-reported split: 1,544,808 five-minute cache-write tokens and 38,024 one-hour cache-write tokens.
  • All-5-minute cache-write alternative: $34.54.
blender high voxel density dioramaRecord status: partial token timing artifact scene inspection ledgerNo validation result attachedWall-clock is end-to-end latency, including Blender MCP tool execution, render time, file I/O, and waits; it is not model-only compute time.Open result
Claude Fable 5 · MediumAnthropic Claude Codeoasis-outpost-fable-5-medium19:23.76,565,527$21.23First-party API-list-price-equivalent estimate, not a subscription invoice
2 cost qualifiers
  • Cache-write basis: 5-minute cache writes
  • Pricing basis: Anthropic first-party Claude API, standard global pricing.
blender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.Open result
GPT-5.6 Sol · UltraOpenAI Codexoasis-outpost-gpt-5.6-sol-ultra16:30.91,850,193$5.51API-equivalent estimate, not a subscription invoice
3 cost qualifiers
  • Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context.
  • Pricing basis: Priority API short-context rates as an API-equivalent scenario because the supplied transcript observed service_tier=priority; this is not proof of an API invoice..
  • Separately priced tools and non-token services are excluded.
blender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.Open result
Kimi K3 · MaxKimi Code CLIoasis-outpost-kimi-k3-max39:27.94,390,406$2.72API-equivalent estimate, not an itemized subscription cash chargeblender sceneRecord status: partial token timing and artifact ledgerNo validation result attachedWall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 · Max

Material caveats
  • Wall-clock is end-to-end latency, including Blender MCP tool execution, render time, file I/O, and waits; it is not model-only compute time.
  • Output tokens include hidden reasoning, generated code, and tool-call JSON.
  • Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
  • The source receipt flags self-measurement drift because it reads an active transcript that can grow while being measured.
  • The render was visually inspected during archival as a coherent scene, but no visual-quality score or blind comparison has been assigned.
  • No initial/optimized FPS measurement, render hardware receipt, or RTX Pro 6000 render was supplied.
Still missing
  • initial and optimized local performance measurements
  • render hardware receipt
  • RTX Pro 6000 final render
  • blind-evaluation record

Claude Fable 5 · Medium

Material caveats
  • Wall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.
  • Output tokens include hidden reasoning, code, and tool-call JSON, not just visible text.
  • Cache-read tokens are billed at a deep discount, so total processed tokens greatly overstate effective cost.
  • The primary total uses the supplied five-minute cache-write rate; a one-hour cache write would increase the cost.
  • Blender and render environment, final capture metadata, and blind-evaluation evidence have not been supplied.
Still missing
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record

GPT-5.6 Sol · Ultra

Material caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • The supplied record observed Priority service tier, so this estimate uses Priority rates and is not directly price-comparable with Standard-tier estimates.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Subagent logs are excluded to avoid double-counting inherited parent context.
Still missing
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record

Kimi K3 · Max

Material caveats
  • Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.
  • Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate effective cost.
  • The receipt records one tool error, even though token reconciliation and final benchmark completion passed.
  • The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
  • The archived preview was visually inspected, but Blender is unavailable in the archival environment, so the scene and build script were not independently rerun.
  • No render-hardware identity, initial/optimized FPS measurement, RTX final render, or blind-evaluation record was supplied.
Still missing
  • independent Blender scene and build-script validation
  • final capture metadata including renderer, hardware, and resolution
  • initial and optimized local performance measurements
  • RTX Pro 6000 final render
  • blind-evaluation record
Round 09 · 1 result

Shrine Exploration

A seven-hour one-request Unity long-horizon build combining environment, character creation, rigging, animation, audio, collision and release output.

Owner assessment

Shrine Exploration demonstrates the visual upside of a long-horizon Opus run, especially for the environment. It also exposes major animation, character-quality and collision debt. Only Opus has been tested so far, visual acceptance failed at 14/40, and the $204.17+ figure remains a parent-session API list-price lower bound because seven subagents lack a priced token mix.

Claude Opus 5 Max presentation recording poster
Claude Opus 5 MaxLaunch 004 episode capture · audio removed
Shrine Exploration exact result receipts
ConfigurationWall-clockProcessed tokensQualified costArtifact / validationMaterial caveatResult link
Claude Opus 5 MaxAnthropic Claude Codeshrine-approach-unity-opus-5-max7h 41m 36.5s327,042,040 parent + 1,888,337 subagent$204.17+Parent-session Claude API list-price equivalent plus an unknown additional subagent amount; not a Claude Code invoice, subscription charge, or exact complete run cost.
3 cost qualifiers
  • Only the parent session has the token mix required for exact API list-price arithmetic.
  • Seven separately billed subagents used 1,888,337 tokens, but their input, cache and output split is unavailable.
  • The exact complete run cost is unknown and higher than $204.17.
active public benchmark recordRecord status: verified visual acceptance failedVisual acceptance failed · 14/40Visual acceptance failed at 14/40, but that failure does not suppress the benchmark result or eligible downloads.Open result
Open the full caveat and missing-evidence ledger

Claude Opus 5 Max

Material caveats
  • Visual acceptance failed at 14/40, but that failure does not suppress the benchmark result or eligible downloads.
  • The strongest evidence is environmental; animation, character quality and collision remain material weaknesses.
  • The $204.17+ figure is a parent-session API list-price lower bound, not a Claude Code invoice or complete run cost.
  • Platform evidence is target-specific and must not be copied from macOS to WebGL or Windows.
  • This single run does not establish a reliability rate.
Still missing
  • public player hosting and signed-out read-back verification
  • WebGL gameplay, physical-input and audible-output verification
  • Windows runtime, gameplay, physical-input and audio verification
  • complete subagent token mix and exact total API-equivalent cost
  • repeat-run reliability evidence
Shrine Exploration deep dive

Seven hours of upside—and visible debt.

The benchmark stays public even though visual acceptance failed. The environment is the strongest evidence; animation, character quality and collision remain material weaknesses.

Review 112/ 40
Review 214/ 40
Review 314/ 40
Visual acceptance failed14 / 40

The failed visual gate does not suppress the result, its Library listing or download eligibility.

Wall-clock
7h 41m 36.5s
Parent processed tokens
327,042,040
Subagent tokens
1,888,337
Cost boundary
$204.17+

Parent-session Claude API list-price equivalent plus an unknown additional subagent amount; not a Claude Code invoice, subscription charge, or exact complete run cost. Seven separately billed subagents are recorded, but their input, cache and output token mix is unavailable. The exact total is unknown and higher.

Fixed benchmark input

Get the exact Shrine Exploration build.

Builder includes the complete Unity project, its fixed 53-bone samurai benchmark fixture, six animation clips, katana and saya, all redistributable Shrine assets, pipeline tools, reusable skills and the exact frozen prompt.

The Shrine run remained one request. The samurai was a separately completed AI-generated fixture—produced iteratively with Opus 5, Meshy and Blender—and supplied unchanged. Its production time, tokens and cost are not counted in the Shrine run receipt.

SAMURAI CHARACTER FIXTURE V1
  • 53-bone rig
  • 6 authored animation clips
  • Katana + saya
  • 5 PBR texture maps
  • 15 files · 80,410,681 bytes
Fixed input · not contestant output

Builder fulfillment verified: all 15 fixture members are present; exact bytes or privacy-sanitized path equivalents are hash-bound to the approved Unity source.

macOS ARM64

shrine-approach-v1-macos-arm64.app.zip

Bytes
190,194,015
SHA-256
f045ee61380dfe18d144d13827758cd80f89bc066f52be50b684b76b0078dd81
Evidence
operator-attested
Trust
ad-hoc signed, not notarized
Gameplay, physical input, footsteps and ambience are operator-attested on macOS.

Gameplay, physical input, footsteps and ambience are operator-attested. The app is ad-hoc signed, is not notarized, and macOS may require right-click → Open.

Public artifact · Hosted · read-back passed
WebGL

shrine-approach-v1-webgl.zip

Bytes
157,918,695
SHA-256
aa518976838751c4a9a1bc32d5357c369ab2c36929d99f3e63760bc07f4def4b
Evidence
rendered startup only
Trust
browser-sandboxed, unsigned
Browser startup verified. Gameplay, physical input and audible output are not yet verified on WebGL.

Rendered startup, WebGL 2.0, and zero recorded console errors are verified. Gameplay, physical input and audible output are not yet verified on WebGL.

Public artifact · Hosted · read-back passed
Windows x86-64

shrine-approach-v1-windows-x86_64.zip

Bytes
198,547,544
SHA-256
3c07e2df476e7650a3d0cbf2fa2c53e22b62ad496d5e6351f01fb64f1a86602e
Evidence
PE32+ GUI x86-64 structure only
Trust
unsigned; Windows runtime untested
Windows runtime, gameplay, physical input and audio remain unverified.

The genuine PE32+ x86-64 player was structurally verified but not launched on Windows. Gameplay, physical input and audio remain unverified, and the executable is unsigned.

Public artifact · Hosted · read-back passed
Public evidence and receipts

Claims you can inspect without joining.

This immutable public projection binds the active benchmark record, player limitations, licence classes, exclusions, checksums and missing-evidence states. The exact prompt intentionally preserves 7 home-relative path examples. Local paths are preserved because this is the exact submitted prompt. The portable fixture binding resolves the supplied Samurai package for reproduction. Activated from owner approval and exact storage read-back. Ordinary-customer preactivation retrieval was waived and not performed; no customer-retrieval receipt exists.

Prompt identity7f557c5a9cbdcd6d33565cc7c9262f63e4055abc8d507f3f2d704e6a396550ed

Frozen at harness commit b464e86bf89f45f36350e97d92c6fe65ce10051c.

Release manifest5b84c152334af952c202735c0d03eb50f5d0916002ddad7bff7924bea3d85e92

Final receipt 19e8ad57ecfa557410dfb6f50bcddf9287c3ba7783d1d22ce2839c7b1d28bb20.

Reliability boundaryOne recorded Shrine run

This single run does not establish a reliability rate.

Builder package

Shrine Exploration v1 — Builder Pack

Open the Unity project, inspect and remix the complete redistributable scene, reuse the pipeline tools and skills, and rerun the exact frozen prompt against another model stack.

Activated · owner-approved, storage-verified3 task-sized components

All three exact, hash-bound components are available with an active Builder entitlement.

Unity source members
322
Reusable skills
4
Pipeline documents
8
Pipeline tools
6
Exact prompts
1
Included benchmark fixture

Get the exact Shrine Exploration build.

Builder includes the complete Unity project, its fixed 53-bone samurai benchmark fixture, six animation clips, katana and saya, all redistributable Shrine assets, pipeline tools, reusable skills and the exact frozen prompt.

The Shrine run remained one request. The samurai was a separately completed AI-generated fixture—produced iteratively with Opus 5, Meshy and Blender—and supplied unchanged. Its production time, tokens and cost are not counted in the Shrine run receipt.

SAMURAI CHARACTER FIXTURE V1
  • 53-bone rig
  • 6 authored animation clips
  • Katana + saya
  • 5 PBR texture maps
  • 15 files · 80,410,681 bytes
Fixed input · not contestant output

Builder fulfillment verified: all 15 fixture members are present; exact bytes or privacy-sanitized path equivalents are hash-bound to the approved Unity source.

unity source

shrine-approach-v1-publication-builder-unity-source.zip

Files
322
Bytes
223,688,953
SHA-256
9008477c0124a1665027a8b009b1e1410dfd04f35b8d0f3c1479f7e7aecaebef
Available to Builder
pipeline skills

shrine-approach-v1-pipeline-skills-v1.zip

Files
26
Bytes
409,186
SHA-256
790061ca557f8f6ef8d953ebc6f191b0f06869dd0194bfa511bd01f9415b0731
Available to Builder
prompt receipts

shrine-approach-v1-prompt-public-receipts-v1.zip

Files
9
Bytes
29,300
SHA-256
d84d238975f956c97bec9fee4bbad2d086982062d080db914d93306689c43993
Available to Builder
Real free preview

Read the exact frozen prompt and public receipts now.

Local paths are preserved because this is the exact submitted prompt. The portable fixture binding resolves the supplied Samurai package for reproduction. Active Builder members can retrieve the protected Unity source, pipeline tools and reusable skills through short-lived download grants.

The Builder offer includes all three exact components; the contextual continuation preserves this collection.

Named exclusions

Five private concept references stay out.

  • reference/A_hero_view.pngbfff498e9929f0bdfc27ec65cd3dbc21490bd2996894358801bc092c3ae587fc
  • reference/B_overview_aerial.pngaef9a04beed4f1fdd2b657d04d5f762efd4f4ace6a817f8769647a888765a52b
  • reference/C_sheet_a_hero_reverse_overview.png32754deccfb381d599eb6030f5646629779faefe07e96746688451abd402b1d0
  • reference/D_sheet_b_elevations_details_materials.pngdb2051803a5ec3c5cf3a90fb78bdf16f024e15814056795dde827a20ffd966a9
  • reference/E_detail_eave_paving_materials.pngbea2b890469573dcf4a8e23178956e1bb7692a62eb355df089f2394b0fc5ad0e

Builder includes all redistributable project assets—not the five private references above.

Activation evidence

Activated from owner approval and exact storage read-back. Ordinary-customer preactivation retrieval was waived and not performed; no customer-retrieval receipt exists. The immutable prepared-candidate receipts remain available as historical evidence.

Contextual pricing, auth and checkout

Continue without losing the Shrine package.

The continuation preserves the exact Shrine collection, campaign and safe internal return path through pricing, authentication and checkout. Entitled customers return to the collection for the activated component downloads.

Continue to Builder pricing
After the proof and package

Keep testing the boundary with us.

Join the member Discord after you have inspected the nine tests, public receipts and current Builder activation boundary.

Join the RemakeBench Discord