Opus 5 delivered the strongest detail in several rounds. It also made the efficiency argument against itself.
Across this suite, Opus 5 often spent more time and API-equivalent value to reach stronger visual detail. Shrine Exploration shows the upside of a seven-hour long-horizon run—and the animation, character and collision debt that current models still leave behind.
Single recorded attempts. Reliability is not established.
Workflow disclosure
This is a workflow comparison, not a bare-model test. Claude Opus 5 and Claude Fable 5 ran through Claude Code; GPT-5.6 Sol ran through Codex; Kimi K3 ran through Kimi CLI. Each benchmark run used one frozen request, with no human follow-up or manual correction after submission. The agents could use their disclosed tools, internal retries and subagents where supported.
The exact Builder components are included for active Builder members.
9episode rounds
32exact records
14new hash-bound clips
1frozen request rule
Rawtime, token and cost receipts
The public answer
Nine frozen tests, in the script’s order.
Every row preserves its exact configuration, measured values, artifact state, validation scope and permanent record. Media appears only when its result identity is exact.
Neither Opus 5 Max nor GPT-5.6 Sol Max produced a clean one-shot result. Opus has invisible tires and doors, an obstructed cockpit view and a tendency to catapult, though its forest and wind animation are strong. Sol finished faster and cheaper with the more planted vehicle and glossy mud. Kimi keeps its doors and tires but also catapults and obstructs the cockpit view.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
GPT-5.6 Sol · MaxLaunch 004 episode capture · audio removed
Kimi K3 · MaxLaunch 004 episode capture · audio removed
Off-road Mud Game exact result receipts
Configuration
Wall-clock
Processed tokens
Qualified cost
Artifact / validation
Material caveat
Result link
Claude Opus 5 · MaxAnthropic Claude Codeoff-road-driving-game-threejs-opus-5-max
2h 48m 08.6s
149,702,730
$124.35Official Anthropic API-list-price equivalent, not an itemized Claude Code subscription charge1 cost qualifier
All-5-minute cache-write alternative: $113.74.
Fresh Vite `dist/` directory archiveRecord status: partial token timing source build ledgerNo validation result attached
Wall-clock is end-to-end workflow latency, including tool execution, package installation, browser automation, WASM initialization, and model latency; it is not model-only compute.
Wall-clock is end-to-end workflow latency, including tool execution, package installation, browser automation, WASM initialization, and model latency; it is not model-only compute.
Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
Cache-read tokens are discounted, so the total-processed figure materially overstates effective cost.
The cache-write TTL mix is not independently reconstructible from the transcript; the primary uses the source receipt's one-hour declaration and records the all-five-minute alternative.
The retained `.verify` scripts and captures are model-authored output. Their console results are not supplied and they were not independently rerun, so they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
No standardized final gameplay capture or blind-evaluation record is supplied.
Still missing
independent browser acceptance run
local FPS, browser, viewport, and hardware receipt
multi-minute sustained-drive and generation-seam receipt
four-corner suspension-articulation receipt
standardized final gameplay capture
blind-evaluation record
GPT-5.6 Sol · Max
Material caveats
Wall-clock is end-to-end workflow latency, including tools and idle gaps; it is not model-only compute time.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
The generated production output is deliberately not archived because its server-side build files contain a build-time prerender secret. The complete committed source required to regenerate it is archived instead.
The production build passes, but the supplied npm test command fails two stale starter-scaffold assertions. No independent browser acceptance, sustained-drive, local-FPS, seam, or suspension-articulation receipt is archived.
No final gameplay capture or blind evaluation is supplied.
Still missing
independent browser acceptance run
local FPS, browser, viewport, and hardware receipt
multi-minute sustained-drive and generation-seam receipt
four-corner suspension-articulation receipt
final gameplay capture
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
One user prompt produced 355 billable Kimi model requests across one main and three child agents.
Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache-read tokens are discounted, so total processed tokens overstate cost.
No before/after /usage quota evidence was supplied, so membership-quota consumption cannot be measured from this run.
The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
The retained `.verify` scripts and captures are model-authored output and were not independently rerun; they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
No standardized final gameplay capture or blind-evaluation record is supplied.
Still missing
independent browser acceptance run
local FPS, browser, viewport, and hardware receipt
multi-minute sustained-drive and generation-seam receipt
four-corner suspension-articulation receipt
standardized final gameplay capture
blind-evaluation record
Round 02 · 4 results
Space Flight Stage 2 — Landfall
A second frozen request that extends each Stage 1 space-flight project with seamless landing, takeoff, an explorable surface and a fixed character asset.
Owner assessment
Opus and Fable both created appealing but different environments. Opus supports flexible crash landing but flickers and makes the valid landing region hard to find; Fable automates landing but launches the ship facing down. Sol's seamless planet is much smaller and cheaper. Kimi's landing is the hardest to complete and distorts the character, although its environment has a distinct visual appeal.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Wall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.
Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
Cache-read tokens are discounted, so total processed tokens overstate cost.
The archived Vite build passes, but no independent full-loop browser acceptance run, local FPS/transition-hitch receipt, final capture, or blind evaluation is archived.
The one-hour cache-write assumption is explicitly reported by the source receipt but cannot be independently reconstructed from the sanitized token totals.
Still missing
independent uninterrupted landing-to-orbit acceptance run
local FPS and transition-hitch receipt with hardware details
final viewport or browser capture
blind-evaluation record
Claude Fable 5 · Max
Material caveats
Wall-clock is end-to-end workflow latency, including tools, browser automation, builds, and idle time; it is not model-only compute.
Output tokens include hidden reasoning, generated code, and tool-call JSON, not only visible prose.
Cache-read tokens are discounted, so total processed tokens overstate cost.
The source receipt does not distinguish cache-write TTLs. The primary estimate follows its stated five-minute assumption; the all-one-hour alternative is included separately.
The archive build passes, but no independent full-loop browser acceptance run, local FPS/transition-hitch receipt, final capture, or blind evaluation is archived.
Still missing
independent uninterrupted landing-to-orbit acceptance run
local FPS and transition-hitch receipt with hardware details
final viewport or browser capture
blind-evaluation record
GPT-5.6 Sol · Ultra
Material caveats
Wall-clock is end-to-end workflow latency, including tools and idle time, not model-only compute.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cache-read input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a Codex Pro subscription-backed session; separately priced tools and non-token services are excluded.
The acceptance evidence is model-authored static/server-rendered testing plus a fresh build and lint run. No independent full-loop browser playthrough, FPS/hitch receipt, or final runtime capture is supplied.
The supplied social preview is retained for provenance but is not used as a final render or quality score.
No blind-evaluation record is supplied.
Still missing
independent uninterrupted orbit-to-surface-to-orbit acceptance run
local FPS and transition-hitch receipt with hardware details
final browser or viewport capture
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
One user prompt produced 202 billable Kimi model requests.
Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache-read tokens are discounted, so total processed tokens overstate cost.
The source receipt reports no before/after /usage quota evidence; it cannot quantify the subscription quota consumed by this run.
The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
The archive build passes, but no independent browser acceptance run, local FPS/hardware measurement, final capture, or blind evaluation is archived.
Still missing
independent uninterrupted landing-to-orbit acceptance run without test fallback
local FPS and hardware receipt, including transition hitch measurement
final capture with viewport/browser evidence
blind-evaluation record
Round 03 · 4 results
Neo-Gothic Flooded City
A procedural neo-gothic storm-city shader balancing architecture, atmosphere, floodwater, reflections and motion.
Owner assessment
Opus took substantially longer and cost more than Fable, with more structural detail. Both satisfy the neo-gothic aesthetic: Fable's floodwater is more violent, while Opus's water looks more realistic and reflective. The final preference is a taste call, not a universal winner.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Kimi K3 · MaxExact earlier comparator capture · audio removed
No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.
Neo-Gothic Flooded City exact result receipts
Configuration
Wall-clock
Processed tokens
Qualified cost
Artifact / validation
Material caveat
Result link
Claude Opus 5 · MaxAnthropic Claude Codeneo-gothic-storm-city-opus-5-max
1h 52m 04.4s
23,105,513
$19.55API-equivalent estimate, not a subscription invoice2 cost qualifiers
Cache-write basis: The session declares a one-hour prompt-cache TTL; the source does not record TTL per entry.
All-5-minute cache-write alternative: $17.94.
twigl fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attached
Wall-clock is end-to-end latency including tool execution, headless browser renders, validation, web fetches, and user think time; it is not model-only compute.
Wall-clock is end-to-end latency including tool execution, headless browser renders, validation, web fetches, and user think time; it is not model-only compute.
Output tokens include hidden reasoning, tool-call JSON, and shader code, not only visible prose.
Cache reads are roughly 96% of processed tokens but bill at 0.1× base input, so processed tokens substantially overstate cost.
The primary $19.55 estimate applies the session's declared one-hour cache TTL; an all-five-minute cache-write alternative is $17.94.
The source receipt is a mid-session snapshot and excludes the later turn that wrote the receipt.
The shader was archived without a fresh TWIGL compile or browser run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
TWIGL compile/run receipt
initial local FPS
optimized local FPS
RTX Pro 6000 final render
blind-evaluation record
Claude Fable 5 · Max
Material caveats
Wall-clock is end-to-end latency including user idle time between turns, not model-only compute.
Output tokens include hidden reasoning, code, and tool-call JSON.
Cache reads are billed at a deep discount, so total processed tokens overstate cost.
Raw JSONL per-content-block totals double-counted multi-block Claude Code messages. This ledger uses the supplied deduplicated billing basis instead.
Fable 5 Max was confirmed by the user after initial archival; the raw metrics receipt itself identifies only the base model.
The shader was archived without a fresh TWIGL compile or browser run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
TWIGL compile/run receipt
initial local FPS
optimized local FPS
RTX Pro 6000 final render
blind-evaluation record
GPT-5.6 Sol · XHigh
Material caveats
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
Separately priced tools and non-token services are excluded.
Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
Static compilation is not a substitute for Twigl runtime compatibility or measured FPS.
Still missing
initial local FPS
optimized local FPS
final capture metadata
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency including tool execution, not model-only compute.
OpenCode splits one agentic turn into multiple assistant/API messages (51 here).
Reasoning tokens and visible output tokens are recorded separately; the table's output-token figure is their combined total.
Cache-read tokens are heavily discounted, so total processed tokens overstate cost.
OpenCode Go ledger cost is API-equivalent usage accounting, not an itemized cash charge.
The bootstrap extraction is version-specific to OpenCode 1.18.3 until validated on more benchmark runs.
The raw metrics ledger records the session window but not an artifact hash; this result ledger supplies the SHA-256 binding to the archived shader.
Still missing
initial local FPS
optimized local FPS
final RTX Pro 6000 1080p60 render or capture metadata
blind-evaluation record
Round 04 · 4 results
Infinite Cathedral
A procedural stained-glass cathedral corridor shader with depth, repeated architecture, lighting and continuous motion.
Owner assessment
Opus produced one of the strongest cathedral results, especially in interior detail, but omitted pillars, exaggerated window light and let distant hanging lights glow from ceiling to floor. It also spent far longer than Fable. Opus is visually stronger than the shown Sol result, while Kimi trails the frontier models on this test.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Kimi K3 · MaxExact earlier comparator capture · audio removed
No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.
Infinite Cathedral exact result receipts
Configuration
Wall-clock
Processed tokens
Qualified cost
Artifact / validation
Material caveat
Result link
Claude Opus 5 · MaxAnthropic Claude Codeinfinite-cathedral-shadertoy-opus-5-max
1h 41m 22.5s
32,282,622
$29.25API-equivalent estimate, not a subscription invoice1 cost qualifier
Cache-write basis: Exact TTL split supplied from transcript usage payload: 116,747 one-hour cache-write tokens and 1,286,290 five-minute cache-write tokens.
shadertoy fragment shaderRecord status: partial token timing and artifact ledgerNo validation result attached
Wall-clock time is end-to-end latency, not model-only compute time; it includes shader compiles, browser renders, screenshot round-trips, and user think time.
Wall-clock time is end-to-end latency, not model-only compute time; it includes shader compiles, browser renders, screenshot round-trips, and user think time.
Output tokens include hidden reasoning, code, and tool-call JSON.
Cache-read tokens are billed at a deep discount, so total processed tokens overstate effective cost.
Pricing was verified live in the supplied receipt, and the cache-write TTL split is recorded exactly in the corrected usage payload.
The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
Shadertoy or WebGL2 compile/run receipt
initial local FPS
optimized local FPS
RTX Pro 6000 final render
blind-evaluation record
Claude Fable 5 · Max
Material caveats
Wall-clock time is end-to-end latency, not model-only compute time.
Output tokens include hidden reasoning, code, and tool-call JSON.
Cache-read tokens are billed at a deep discount, so total processed tokens overstate effective cost.
The supplied 1-hour cache-write TTL is used for the primary estimate.
The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
Shadertoy or WebGL2 compile/run receipt
initial local FPS
optimized local FPS
RTX Pro 6000 final render
blind-evaluation record
GPT-5.6 Sol · XHigh
Material caveats
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
Separately priced tools and non-token services are excluded.
Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
The submitted artifact is a WebGL2 application rather than a direct Shadertoy paste.
Still missing
Shadertoy compatibility adapter or direct Shadertoy source
initial local FPS
optimized local FPS
final capture metadata
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency and includes tools, waits, and idle time.
Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache reads are discounted, so total processed tokens overstate effective cost.
The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
The shader was archived without a fresh Shadertoy/WebGL2 run in this environment; initial local FPS, optimized local FPS, RTX final render, and blind evaluation remain unavailable.
Still missing
Shadertoy or WebGL2 compile/run receipt
initial local FPS
optimized local FPS
RTX Pro 6000 final render
blind-evaluation record
Round 05 · 4 results
MacBook Pro
A validator-ready Blender product-ad scene with modeled industrial detail, materials, lighting and camera movement.
Owner assessment
The same specification still produces materially different products. Opus took only six minutes longer than Fable in the recorded runs while costing much less; Kimi remains especially competitive when cost is considered. The owner leaves the visual preference to the audience.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Kimi K3 · MaxExact earlier comparator capture · audio removed
No exact public episode capture is bound for GPT-5.6 Sol · XHigh. Their measured records remain below; no title-matched substitute was used.
MacBook Pro exact result receipts
Configuration
Wall-clock
Processed tokens
Qualified cost
Artifact / validation
Material caveat
Result link
Claude Opus 5 · MaxAnthropic Claude Code with Blender MCPmacbook-cinematic-blender-opus-5-max
1h 07m 02.5s
35,897,343
$27.40API-equivalent estimate, not a subscription invoice2 cost qualifiers
Cache-write basis: The session declares a one-hour prompt-cache TTL; the source does not record TTL per entry.
All-5-minute cache-write alternative: $26.03.
blender cinematic product ad sceneRecord status: partial token timing and artifact ledgerNo validation result attached
Wall-clock is end-to-end workflow latency, not model-only compute; the source receipt records Blender renders, boolean evaluation, and two headless Blender launches within this window.
Wall-clock is end-to-end workflow latency, not model-only compute; the supplied record says Blender-side Cycles renders, boolean evaluation, and BVH sweeps account for a substantial portion of the session time.
Wall-clock is end-to-end workflow latency, not model-only compute; the source receipt records Blender renders, boolean evaluation, and two headless Blender launches within this window.
Output tokens include hidden reasoning, generated Python, and tool-call JSON, not only visible prose.
Cache reads are about 98% of processed-token volume but bill at 0.1× base input, so processed tokens substantially overstate cost.
The primary $27.40 estimate applies the session's declared one-hour cache TTL; an all-five-minute cache-write alternative is $26.03.
The supplied receipt reports PASS 49/49, but no validator log was supplied and the archival environment cannot perform another Blender re-run.
No final 4K film render, render-hardware identity, Blender MCP tool-use transcript, or blind-evaluation record was supplied.
Still missing
retained validator log
final 4K film render and renderer identity
blind-evaluation record
Claude Fable 5 · Max
Material caveats
Wall-clock is end-to-end workflow latency, not model-only compute; the supplied record says Blender-side Cycles renders, boolean evaluation, and BVH sweeps account for a substantial portion of the session time.
Output tokens include hidden reasoning, tool-bound code, and tool-call JSON, not just visible prose.
Cache reads are billed at a deep discount, so 35.2M processed tokens substantially overstate cost; most processed tokens were cache reads.
The supplied metrics ledger reports the base model claude-fable-5. The Max reasoning configuration is user-specified metadata.
The cache-creation aggregate does not identify TTL. If every cache write used the official one-hour rate, the total would be $81.28 rather than $75.18.
The supplied validation log reports PASS 49/49, but it was not independently rerun during archival because Blender is unavailable in this environment.
No final 4K film render, render-hardware identity, Blender MCP tool-use transcript, or blind-evaluation record was supplied.
Still missing
independent headless validator rerun
render-hardware identity
Blender MCP tool-use transcript or disclosure
final capture or rubric stills
blind-evaluation record
GPT-5.6 Sol · XHigh
Material caveats
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
Separately priced tools and non-token services are excluded.
Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
The metrics ledger does not identify the actual render GPU or preserve a Blender MCP tool-call transcript; the requested tool profile is recorded but not independently attested.
The 1267x712 Cycles stills are audit evidence; the task intentionally does not require executing the configured 4K final animation.
Still missing
render-hardware identity
Blender MCP tool-use transcript or disclosure
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency and includes tool execution, Blender renders, approval waits, and idle time.
One user prompt produced 123 Kimi model requests, but only 122 have reconciled usage records.
Kimi Code CLI 0.26.0 exposed no authoritative separate reasoning-token field; output tokens include provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache reads are discounted, so total processed tokens overstate cost.
The supplied scene validator reports PASS 49/49, but it was not independently rerun during archival because Blender is unavailable in this environment.
The private session ZIP named in the receipt was not present in the supplied project folder and is intentionally not published.
No /usage screenshot, render-hardware identity, or blind-evaluation record was supplied.
Still missing
metrics reconciliation for the one unpaired API request
independent headless validator rerun
render-hardware identity
Kimi tool-use transcript or disclosure
blind-evaluation record
Round 06 · 4 results
JRPG Boss Battle
A supplied-asset Three.js boss battle with cinematic camera beats, combat state, magic effects, audio controls and validator-ready outcomes.
Owner assessment
Opus is a solid result and much cheaper than Fable, with an appealing floor and cinematic movement, but its generated backdrop cuts off at the top. Its magic composition is stronger than the shown Sol Ultra shortcut. Kimi meets the effects and game contract without notable bugs, but lacks the regular JRPG camera movement; its cost stands out.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Wall-clock is end-to-end workflow latency, including tool execution, browser runs, image generation, and human-side idle time; it is not model-only compute.
Wall-clock is end-to-end workflow latency, including tool execution, browser runs, image generation, and human-side idle time; it is not model-only compute.
Output tokens include hidden reasoning, code, and tool-call JSON.
Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
The validator establishes functional completion but does not replace blind visual-quality evaluation.
The exact browser version and local hardware identity were not supplied.
The metrics receipt reports a subscription-backed Claude Code session; the displayed amount is API-equivalent rather than an itemized cash charge.
Still missing
browser and local-hardware identity
independent blind-evaluation record
Claude Fable 5 · Max
Material caveats
Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool execution and idle waits.
Output tokens include hidden reasoning, code, and tool-call JSON.
Cache reads are billed at a 90% discount to fresh input, so total processed tokens substantially overstate cost.
The supplied snapshot includes the metrics turns themselves.
The exact browser version and local hardware identity were not supplied.
The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
browser and local-hardware identity
independent validator rerun
model tool-use transcript or disclosure
blind-evaluation record
GPT-5.6 Sol · Ultra
Material caveats
Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent Priority-tier estimate rather than the actual charge for a subscription-backed Codex session.
Separately priced tools and non-token services are excluded.
The exact browser version and local hardware identity were not supplied.
The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
browser and local-hardware identity
independent validator rerun
model tool-use transcript or disclosure
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock includes tools, installs, browser checks, approval waits, and idle time.
One user turn produced 62 billable Kimi model requests.
Kimi Code CLI 0.27.0 exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache reads are discounted, so total processed tokens overstate cost.
Allegretto does not itemize a per-run cash charge; the $39 monthly subscription is not divided by this run.
The exact browser version and local hardware identity were not supplied.
The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
The validator establishes functional completion but does not replace blind visual-quality evaluation.
Still missing
browser and local-hardware identity
independent validator rerun
blind-evaluation record
Round 07 · 4 results
Jungle Temple
A high-voxel-density Blender jungle-temple diorama with coherent scene composition, inspectable geometry and fixed voxel scale.
Owner assessment
Opus produced a coherent, highly detailed scene with no major visual flaw, at materially higher cost than Fable. Both Opus and Sol use a large bridge, but Sol falls behind in scene detail; Kimi's scene is comparatively sparse.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
Wall-clock is end-to-end latency including Blender MCP tool execution, renders, file I/O, and waits; it is not model-only compute time.
Output tokens include hidden reasoning, generated code, and tool-call JSON.
Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
The displayed $40.58 workflow estimate includes the receipt's $0.01 web-search charge; the token-only estimate is $40.57.
The scene was read-only inspected with Blender 5.1.2, but no initial/optimized FPS measurement, render hardware receipt, RTX Pro 6000 render, or blind evaluation was supplied.
Still missing
initial and optimized local performance measurements
render hardware receipt
RTX Pro 6000 final render
blind-evaluation record
Claude Fable 5 · Medium
Material caveats
Wall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.
Output tokens include hidden reasoning, code, and tool-call JSON, not just visible text.
Cache-read tokens are billed at a deep discount, so total processed tokens greatly overstate effective cost.
The primary total uses the supplied five-minute cache-write rate; a one-hour cache write would increase the cost.
Blender and render environment, final capture metadata, and blind-evaluation evidence have not been supplied.
Still missing
Blender and render environment
final capture metadata
blind-evaluation record
GPT-5.6 Sol · Ultra
Material caveats
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
Separately priced tools and non-token services are excluded.
Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
Subagent logs are excluded to avoid double-counting inherited parent context.
Still missing
Blender and render environment
final capture metadata
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.
Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache reads are discounted, so total processed tokens overstate effective cost.
The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
The archived preview was visually inspected, but Blender is unavailable in the archival environment, so the scene and build log were not independently rerun.
No render-hardware identity, initial/optimized FPS measurement, RTX final render, or blind-evaluation record was supplied.
Still missing
independent Blender scene and build-script validation
final capture metadata including renderer, hardware, and resolution
initial and optimized local performance measurements
RTX Pro 6000 final render
blind-evaluation record
Round 08 · 4 results
Oasis Outpost
A high-voxel-density Blender oasis diorama testing scene population, architectural detail, pattern, shade and landscape composition.
Owner assessment
Opus packs the most detail into roofs, walls and terrain, perhaps making the landscape too dynamic, and also takes the longest. Sol wins on completion time and cost. Kimi works at a smaller scale but populates its scene well and is impressive for the price.
Claude Opus 5 · MaxLaunch 004 episode capture · audio removed
GPT-5.6 Sol · UltraOpenAI Codexoasis-outpost-gpt-5.6-sol-ultra
16:30.9
1,850,193
$5.51API-equivalent estimate, not a subscription invoice3 cost qualifiers
Requests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context.
Pricing basis: Priority API short-context rates as an API-equivalent scenario because the supplied transcript observed service_tier=priority; this is not proof of an API invoice..
Separately priced tools and non-token services are excluded.
blender sceneRecord status: partial token timing and artifact ledgerNo validation result attached
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Wall-clock is end-to-end latency, including Blender MCP tool execution, render time, file I/O, and waits; it is not model-only compute time.
Output tokens include hidden reasoning, generated code, and tool-call JSON.
Cache reads are billed at a deep discount, so total processed tokens overstate effective cost.
The source receipt flags self-measurement drift because it reads an active transcript that can grow while being measured.
The render was visually inspected during archival as a coherent scene, but no visual-quality score or blind comparison has been assigned.
No initial/optimized FPS measurement, render hardware receipt, or RTX Pro 6000 render was supplied.
Still missing
initial and optimized local performance measurements
render hardware receipt
RTX Pro 6000 final render
blind-evaluation record
Claude Fable 5 · Medium
Material caveats
Wall-clock is end-to-end latency including tool execution and user think-time, not model-only compute.
Output tokens include hidden reasoning, code, and tool-call JSON, not just visible text.
Cache-read tokens are billed at a deep discount, so total processed tokens greatly overstate effective cost.
The primary total uses the supplied five-minute cache-write rate; a one-hour cache write would increase the cost.
Blender and render environment, final capture metadata, and blind-evaluation evidence have not been supplied.
Still missing
Blender and render environment
final capture metadata
blind-evaluation record
GPT-5.6 Sol · Ultra
Material caveats
Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
Cached input is deeply discounted, so total processed tokens overstate cost.
This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
The supplied record observed Priority service tier, so this estimate uses Priority rates and is not directly price-comparable with Standard-tier estimates.
Separately priced tools and non-token services are excluded.
Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
Subagent logs are excluded to avoid double-counting inherited parent context.
Still missing
Blender and render environment
final capture metadata
blind-evaluation record
Kimi K3 · Max
Material caveats
Wall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.
Kimi Code CLI 0.26.0 does not expose separate reasoning-token counts; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
Cache reads are discounted, so total processed tokens overstate effective cost.
The receipt records one tool error, even though token reconciliation and final benchmark completion passed.
The raw private session ZIP has not been published; the stable benchmark-window receipt and its hash are retained.
The archived preview was visually inspected, but Blender is unavailable in the archival environment, so the scene and build script were not independently rerun.
No render-hardware identity, initial/optimized FPS measurement, RTX final render, or blind-evaluation record was supplied.
Still missing
independent Blender scene and build-script validation
final capture metadata including renderer, hardware, and resolution
initial and optimized local performance measurements
RTX Pro 6000 final render
blind-evaluation record
Round 09 · 1 result
Shrine Exploration
A seven-hour one-request Unity long-horizon build combining environment, character creation, rigging, animation, audio, collision and release output.
Owner assessment
Shrine Exploration demonstrates the visual upside of a long-horizon Opus run, especially for the environment. It also exposes major animation, character-quality and collision debt. Only Opus has been tested so far, visual acceptance failed at 14/40, and the $204.17+ figure remains a parent-session API list-price lower bound because seven subagents lack a priced token mix.
Claude Opus 5 MaxLaunch 004 episode capture · audio removed
Shrine Exploration exact result receipts
Configuration
Wall-clock
Processed tokens
Qualified cost
Artifact / validation
Material caveat
Result link
Claude Opus 5 MaxAnthropic Claude Codeshrine-approach-unity-opus-5-max
7h 41m 36.5s
327,042,040 parent + 1,888,337 subagent
$204.17+Parent-session Claude API list-price equivalent plus an unknown additional subagent amount; not a Claude Code invoice, subscription charge, or exact complete run cost.3 cost qualifiers
Only the parent session has the token mix required for exact API list-price arithmetic.
Seven separately billed subagents used 1,888,337 tokens, but their input, cache and output split is unavailable.
The exact complete run cost is unknown and higher than $204.17.
active public benchmark recordRecord status: verified visual acceptance failedVisual acceptance failed · 14/40
Visual acceptance failed at 14/40, but that failure does not suppress the benchmark result or eligible downloads.
Visual acceptance failed at 14/40, but that failure does not suppress the benchmark result or eligible downloads.
The strongest evidence is environmental; animation, character quality and collision remain material weaknesses.
The $204.17+ figure is a parent-session API list-price lower bound, not a Claude Code invoice or complete run cost.
Platform evidence is target-specific and must not be copied from macOS to WebGL or Windows.
This single run does not establish a reliability rate.
Still missing
public player hosting and signed-out read-back verification
WebGL gameplay, physical-input and audible-output verification
Windows runtime, gameplay, physical-input and audio verification
complete subagent token mix and exact total API-equivalent cost
repeat-run reliability evidence
Shrine Exploration deep dive
Seven hours of upside—and visible debt.
The benchmark stays public even though visual acceptance failed. The environment is the strongest evidence; animation, character quality and collision remain material weaknesses.
Review 112/ 40
Review 214/ 40
Review 314/ 40
Visual acceptance failed14 / 40
The failed visual gate does not suppress the result, its Library listing or download eligibility.
Wall-clock
7h 41m 36.5s
Parent processed tokens
327,042,040
Subagent tokens
1,888,337
Cost boundary
$204.17+
Parent-session Claude API list-price equivalent plus an unknown additional subagent amount; not a Claude Code invoice, subscription charge, or exact complete run cost. Seven separately billed subagents are recorded, but their input, cache and output token mix is unavailable. The exact total is unknown and higher.
Fixed benchmark input
Get the exact Shrine Exploration build.
Builder includes the complete Unity project, its fixed 53-bone samurai benchmark fixture, six animation clips, katana and saya, all redistributable Shrine assets, pipeline tools, reusable skills and the exact frozen prompt.
The Shrine run remained one request. The samurai was a separately completed AI-generated fixture—produced iteratively with Opus 5, Meshy and Blender—and supplied unchanged. Its production time, tokens and cost are not counted in the Shrine run receipt.
Builder fulfillment verified: all 15 fixture members are present; exact bytes or privacy-sanitized path equivalents are hash-bound to the approved Unity source.
Gameplay, physical input, footsteps and ambience are operator-attested on macOS.
Gameplay, physical input, footsteps and ambience are operator-attested. The app is ad-hoc signed, is not notarized, and macOS may require right-click → Open.
Windows runtime, gameplay, physical input and audio remain unverified.
The genuine PE32+ x86-64 player was structurally verified but not launched on Windows. Gameplay, physical input and audio remain unverified, and the executable is unsigned.
This immutable public projection binds the active benchmark record, player limitations, licence classes, exclusions, checksums and missing-evidence states. The exact prompt intentionally preserves 7 home-relative path examples. Local paths are preserved because this is the exact submitted prompt. The portable fixture binding resolves the supplied Samurai package for reproduction. Activated from owner approval and exact storage read-back. Ordinary-customer preactivation retrieval was waived and not performed; no customer-retrieval receipt exists.
Final receipt 19e8ad57ecfa557410dfb6f50bcddf9287c3ba7783d1d22ce2839c7b1d28bb20.
Reliability boundaryOne recorded Shrine run
This single run does not establish a reliability rate.
Builder package
Shrine Exploration v1 — Builder Pack
Open the Unity project, inspect and remix the complete redistributable scene, reuse the pipeline tools and skills, and rerun the exact frozen prompt against another model stack.
All three exact, hash-bound components are available with an active Builder entitlement.
Unity source members
322
Reusable skills
4
Pipeline documents
8
Pipeline tools
6
Exact prompts
1
Included benchmark fixture
Get the exact Shrine Exploration build.
Builder includes the complete Unity project, its fixed 53-bone samurai benchmark fixture, six animation clips, katana and saya, all redistributable Shrine assets, pipeline tools, reusable skills and the exact frozen prompt.
The Shrine run remained one request. The samurai was a separately completed AI-generated fixture—produced iteratively with Opus 5, Meshy and Blender—and supplied unchanged. Its production time, tokens and cost are not counted in the Shrine run receipt.
Builder fulfillment verified: all 15 fixture members are present; exact bytes or privacy-sanitized path equivalents are hash-bound to the approved Unity source.
Read the exact frozen prompt and public receipts now.
Local paths are preserved because this is the exact submitted prompt. The portable fixture binding resolves the supplied Samurai package for reproduction. Active Builder members can retrieve the protected Unity source, pipeline tools and reusable skills through short-lived download grants.
Builder includes all redistributable project assets—not the five private references above.
Activation evidence
Activated from owner approval and exact storage read-back. Ordinary-customer preactivation retrieval was waived and not performed; no customer-retrieval receipt exists. The immutable prepared-candidate receipts remain available as historical evidence.
Contextual pricing, auth and checkout
Continue without losing the Shrine package.
The continuation preserves the exact Shrine collection, campaign and safe internal return path through pricing, authentication and checkout. Entitled customers return to the collection for the activated component downloads.