REMAKEBENCH · LAUNCH 003

I Gave Qwen3.8 Max, Kimi K3, GPT-5.6 and Fable 5 the Same 9 Game Tests

Kimi K3 was more consistent; Qwen's visual-agent workflow was not reliable enough to win overall.

Across nine frozen tests, Qwen showed its best work in Campfire and JRPG, but Kimi produced more coherent results across the suite. Qwen's biggest failure was operational: once visual review began, the agent repeatedly struggled to stop.

Single recorded attempts. Reliability is not established.

Blind assignments are unavailable, so the primary action safely falls back to the nine public receipts.

  • 9featured Qwen rounds
  • 42derived comparison records
  • 6 preparedcandidate blind pairs
  • Publishedowner editorial active
  • No scoreno rank or winner
  • Receiptscost · time · tokens
MacBook result · independent axes

49/49 PASS does not erase agent non-termination.

The validator-passing artifact and the measured agent completion are reported independently. No visual quality score was assigned, and this result is excluded from blind assignment.

Objective validator
49/49 PASS
Time to objective gate
38m 10.5s
Passing artifact saved
40m 48.4s
Independent validator rerun
49/49 PASS
Model-declared final response
absent
Measured agent completion
no
Visual quality score
not assigned
Blind-vote eligibility
no
Full-attempt processed tokens
18,433,707
Full-attempt qualified API-list-price equivalent
$29.2650895
Post-objective usage
11,238,597 tokens · $17.9743255
Later render preservation
fix9 was not saved over the validator-passing artifact
Nine Qwen outputs · 42 derived comparison records

Inspect the artifacts behind the verdict.

Every Qwen row is resolved by exact result ID from the pinned harness import. Dollar values are qualified third-party API-list-price equivalents, not Alibaba charges.

Test 01 · Shader

Neo-Gothic Storm City

Real-time shader composition, architectural depth, atmosphere, water, lighting, and motion.

Qwen ledger statuspartial token timing artifact validation ledgerEditorial assessment

Qwen finished in 7m 56.8s, but missed both the buildings and the neo-Gothic target. Kimi's slower result was more legible and more faithful to the brief.

Qwen3.8 Max Preview · Neo-Gothic Storm City presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing artifact validation ledger
Wall-clock
7m 56.8s
Processed tokens
107,839
Qualified API-list-price equivalent
$0.278487
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
twigl webgl fragment shader
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (6)
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 28,947 thought tokens are a subset of the 33,351 output tokens.
  • No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
  • The $0.28 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The preceding failed Cathedral attempt counts against account quota but is not a benchmark result because no output artifact was preserved.
  • No post-run quota screenshots, local FPS measurement, optimization pass, RTX Pro 6000 final render, final capture, or blind-evaluation record were supplied.
Evidence still missing (4)
  • initial local FPS measurement receipt
  • optimized local FPS measurement receipt
  • RTX Pro 6000 1080p60 final capture
  • blind-evaluation record
Neo-Gothic Storm City measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$0.278487Qualified API-list-price equivalent7m 56.8spartial token timing artifact validation ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxBase modelOpenCode Go / Moonshot AI$2.30API-equivalent estimate38:29.074 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end workflow latency including tool execution, not model-only compute.Open result
Claude Fable 5 MaxBase modelAnthropic Claude Code$11.39API-equivalent estimate28:48 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency including user idle time between turns, not model-only compute.Open result
GPT-5.6 Sol xhighBase modelOpenAI Codex$1.96API-equivalent estimate8:49.5 wall-clockpartial token timing and artifact ledgerRequests with more than 272,000 input tokens use long-context pricing; all 30 supplied calls were short-context. Separately priced tools and non-token services are excluded.Open result
GPT-5.6 Sol UltraBase modelOpenAI Codex$2.84API-equivalent estimate22:10.9 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 02 · Blender

MacBook-Class Cinematic

Blender MCP product modeling, scene structure, materials, camera animation, deterministic validation, and agent-loop termination.

Qwen ledger statuspartial objective pass agent nonterminationEditorial assessment

Qwen reached the objective 49/49 gate in 38m 10.5s and saved a passing artifact, but its visual-review process did not terminate and the presentation fell short of the other models.

Qwen3.8 Max Preview · MacBook-Class Cinematic presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial objective pass agent nontermination
Wall-clock
7h 00m 56.4s
Processed tokens
18,433,707
Qualified API-list-price equivalent
$29.2650895
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.1
Artifact classification
blender objective gate artifact
Validator
49/49 PASS
Runtime
See ledger caveats
Material caveats (5)
  • This is a dual result: the core construction objective passed, but the agent workflow did not terminate normally. It is visible as PARTIAL, not as a completed or blind-vote-ready benchmark result.
  • The 38m 10.5s number is time to the independent 49/49 objective gate; the full agent attempt lasted 7h 00m 56.4s including the quota pause.
  • 61.42% of the API-equivalent cost was incurred after the objective gate had already passed.
  • Output includes model-accounted thoughts; the thoughts figure is a subset of output and is not additive.
  • The API-equivalent estimate is not Alibaba pricing, provider cost, or an itemized subscription charge.
Evidence still missing (0)

No additional missing-evidence list was supplied by this ledger.

MacBook-Class Cinematic measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.1Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.1$29.2650895Qualified API-list-price equivalent7h 00m 56.4spartial objective pass agent nontermination49/49 PASSSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxTool-assistedKimi Code CLI≥$10.85API-equivalent estimate1:26:58.2 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checksLower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.Open result
Claude Fable 5 MaxTool-assistedAnthropic Claude Code$75.18API-equivalent estimate1:01:04.8 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checksAll-1-hour cache-write alternative: $81.28. Primary cost assumes 5-minute cache writes because the supplied usage does not record TTLs.Open result
GPT-5.6 Sol xhighTool-assistedOpenAI Codex$7.67API-equivalent estimate55:49.9 wall-clockpartial token timing artifact validation ledgerPASS, 49/49 checksRequests with more than 272,000 input tokens use long-context pricing; all 81 supplied calls were short-context. Separately priced tools and non-token services are excluded.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 03 · Browser game

JRPG Boss Battle

Three.js game construction with supplied assets, combat state, animation, audio controls, validation, and runtime inspection.

Qwen ledger statuspublished artifact validation runtime ledgerEditorial assessment

Qwen delivered a complete 41/41 battle with working contact, animation, effects, and a strong floor and backdrop, although it took the longest of the four workflows.

Qwen3.8 Max Preview · JRPG Boss Battle presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
published artifact validation runtime ledger
Wall-clock
58m 40.6s
Processed tokens
9,636,912
Qualified API-list-price equivalent
$15.1479165
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.1
Artifact classification
vite threejs jrpg source and production build entry point
Validator
PASS — 41/41
Runtime
PASS
Material caveats (5)
  • Wall-clock is end-to-end workflow latency, not model-only compute time.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 123,945 thought tokens are a subset of 197,871 output tokens.
  • The independent 119.99 FPS sample is Apple M1 Pro / ANGLE Metal evidence, not a standardized cross-machine comparison.
  • The shared 275-Credit batch cannot be apportioned among Voxel Maps, the intervening activity, the failed JRPG attempt, and this successful retry.
  • No formal RemakeBench quality score or blind-voting result has been recorded.
Evidence still missing (0)

No additional missing-evidence list was supplied by this ledger.

JRPG Boss Battle measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.1Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.1$15.1479165Qualified API-list-price equivalent58m 40.6spublished artifact validation runtime ledgerPASS — 41/41Single recorded attempt. Reliability is not established.Open result
Kimi K3 MaxFull production stackKimi Code CLI / Moonshot AI$2.86API-equivalent estimate41:20.2 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checksThe supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$123.42API-equivalent estimate50:48.1 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checksThe supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.Open result
GPT-5.6 Sol UltraFull production stackOpenAI Codex$9.98API-equivalent estimate25:51.8 wall-clockpartial token timing artifact validation ledgerPASS, 41/41 checksPriority service-tier pricing. Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context. No cache-write amount was inferred from the supplied transcript schema. Separately priced tools and non-token services are excluded.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 04 · Interactive scene

Campfire Under a Starry Night

Interactive Three.js scene construction, fire, lighting, environment detail, and controllable camera motion.

Qwen ledger statuspartial token timing and source ledgerEditorial assessment

Qwen produced a stronger fire effect than Kimi and finished in roughly half the time, but its surrounding trees and environment remained crude.

Qwen3.8 Max Preview · Campfire Under a Starry Night presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing and source ledger
Wall-clock
6m 0.9s
Processed tokens
139,142
Qualified API-list-price equivalent
$0.299006
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
interactive threejs campfire scene
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (7)
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 18,213 thought tokens are a subset of the 25,798 output tokens.
  • Cache-read tokens are discounted or plan-accounted differently from fresh input, so processed-token totals are not a cost proxy.
  • The $0.30 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The account-meter delta includes an intervening nonbenchmark Blender MCP probe, and its modeled campfire credit allocation is explicitly non-exact.
  • The archived HTML uses external jsDelivr imports, so it is not self-contained offline.
  • Runtime verification is reported by the archived receipt; a fresh local browser run, local FPS measurement, final capture, and blind evaluation were not supplied.
Evidence still missing (4)
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Campfire Under a Starry Night measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$0.299006Qualified API-list-price equivalent6m 0.9spartial token timing and source ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxFull production stackKimi Code CLI$0.45API-equivalent estimate14:19.3 wall-clockpartial token timing and source ledgerWall-clock is end-to-end workflow latency and includes tool execution, browser checks, approvals, and idle time.Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$11.17API-equivalent estimate12:48.9 wall-clockpartial token timing and source ledgerWall-clock is end-to-end latency, including tool execution, idle gaps, and human think time; it is not model-only compute.Open result
Claude Fable 5 MediumFull production stackAnthropic Claude Code$5.12estimated API cost4:24.6 wall-clockpartial token timing and source ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Sol xhighFull production stackOpenAI Codex$2.26API-equivalent estimate28:52.6 wall-clockpartial token timing and source ledgerRequests with more than 272,000 input tokens use long-context pricing; all 44 supplied calls were short-context. Separately priced tools and non-token services are excluded.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 05 · Browser game

Space Flight Game

Long-horizon browser game construction, controls, assets, camera orientation, runtime checks, and visual self-review.

Qwen ledger statuspartial token timing runtime artifact visual self verification ledgerEditorial assessment

After three attempts and a repaired vision path, Qwen produced a runnable but stylistically inconsistent result that took over two hours. Kimi's CLI workflow was materially more stable.

Qwen3.8 Max Preview · Space Flight Game presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing runtime artifact visual self verification ledger
Wall-clock
2h 4m 17.3s
Processed tokens
29,903,941
Qualified API-list-price equivalent
$45.894351
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
threejs space flight game source entry point
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (7)
  • Wall-clock is end-to-end workflow latency, not model-only compute time.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 185,832 thought tokens are a subset of 296,697 output tokens.
  • The actual captures are 1920×993; 1920×1080 was the requested Chrome window size.
  • The 120–123 FPS values are HUD-reported on Apple M1 Pro / ANGLE Metal, not standardized cross-machine performance benchmarks.
  • The image-isolation architecture succeeded but exceeded its intended agent-call budget, making this successful retry unusually expensive.
  • No task-isolated Token Plan Credit charge can be assigned because the meter combines attempts 2 and 3.
  • Formal RemakeBench quality scoring, standardized RTX capture, and blind preference remain unassigned.
Evidence still missing (5)
  • bounded visual-inspector retry with six or fewer inspector calls and three or fewer capture/edit cycles
  • standardized local FPS measurement receipt
  • standardized RTX capture
  • formal RemakeBench quality scoring
  • blind-evaluation record
Space Flight Game measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$45.894351Qualified API-list-price equivalent2h 4m 17.3spartial token timing runtime artifact visual self verification ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxFull production stackKimi Code CLI / Moonshot AI$3.25API-equivalent estimate1:04:11 wall-clockpartial token timing and source ledgerWall-clock is end-to-end workflow latency, including tool execution, manual-approval waits, browser playtests, and idle time.Open result
Claude Fable 5 MaxFull production stackAnthropic Claude Code$151.55API-equivalent estimate1:05:19.2 wall-clockpartial token timing and source ledgerPrimary cost uses 5-minute cache writes. All-1-hour cache-write alternative: $167.36. No deduplication audit was supplied for this usage receipt; the token and cost figures retain the supplied basis.Open result
Claude Fable 5 MediumFull production stackAnthropic Claude Code$22.85API-equivalent estimate18:37.7 wall-clockpartial token timing and source ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Sol UltraFull production stackOpenAI Codex$17.94API-equivalent estimate37:44.2 wall-clockpartial token timing and source ledgerSingle supplied pre-request snapshot; repeat runs not attachedOpen result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 06 · Blender

Shrine Village

Blender MCP voxel-diorama construction, composition, architecture, landscaping, water, and inspectable artifact creation.

Qwen ledger statuspartial token timing artifact validation ledgerEditorial assessment

Kimi won on scene coherence. Qwen had the right asset family and stronger water than GPT-5.6 Sol, but overexposure, random bridges, skewed paths, and choppy stairs weakened the full composition.

Qwen3.8 Max Preview · Shrine Village presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing artifact validation ledger
Wall-clock
30m 59.8s
Processed tokens
2,357,522
Qualified API-list-price equivalent
$3.754501
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
blender scene
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (8)
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 32,656 thought tokens are a subset of the 62,348 output tokens.
  • One Blender execute-code failure was recovered. The result is retained because the final scene completed and passed independent load checks.
  • No clean standalone source script was saved; the preserved execute-code transcript is source evidence only and has not been verified as standalone replayable.
  • No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
  • The $3.75 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The post-run first-party quota evidence covers four activities. It supports only aggregate capacity consumption and cannot assign Shrine a task-only Credit amount.
  • No blind evaluation was supplied.
Evidence still missing (1)
  • blind-evaluation record
Shrine Village measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$3.754501Qualified API-list-price equivalent30m 59.8spartial token timing artifact validation ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxTool-assistedKimi Code CLI$2.23API-equivalent estimate35:53.5 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$20.28estimated API cost18:37.5 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$1.98API-equivalent estimate20:36.2 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 07 · Shader

Infinite Cathedral

ShaderToy corridor generation, stained-glass lighting, repetition, depth, image transport, and runtime inspection.

Qwen ledger statuspublished artifact runtime ledgerEditorial assessment

Kimi beat Qwen on lighting and interior detail, but both results looked primitive beside the GPT-5.6 Sol and Fable 5 generations.

Qwen3.8 Max Preview · Infinite Cathedral presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
published artifact runtime ledger
Wall-clock
25m 9.2s
Processed tokens
3,409,324
Qualified API-list-price equivalent
$5.324798
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
shadertoy webgl fragment shader
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (7)
  • Wall-clock is end-to-end workflow latency, not model-only compute time.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 36,650 thought tokens are a subset of 60,232 output tokens.
  • The actual captures are 1920×993; 1920×1080 was the requested Chrome window size.
  • The independent 19.87 FPS sample is Apple M1 Pro / ANGLE Metal evidence, not a standardized cross-machine performance comparison.
  • Task-required initial and optimized local FPS receipts and an RTX Pro 6000 1080p60 final were not supplied.
  • The 28-Credit estimate (conservative range 25–32) is a reset-aware, after-only Token Plan capacity observation rather than a cash charge or invoice-grade measurement.
  • No formal RemakeBench quality score or blind-voting result has been recorded; the retained independent review is nonbinding provenance.
Evidence still missing (0)

No additional missing-evidence list was supplied by this ledger.

Infinite Cathedral measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$5.324798Qualified API-list-price equivalent25m 9.2spublished artifact runtime ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxBase modelKimi Code CLI$0.92API-equivalent estimate21:54.1 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end workflow latency and includes tools, waits, and idle time.Open result
Claude Fable 5 MaxBase modelAnthropic Claude Code$43.85API-equivalent estimate31:13.8 wall-clockpartial token timing and artifact ledgerPrimary cost uses 1-hour cache writes (supplied). All-5-minute cache-write alternative: $38.66.Open result
GPT-5.6 Sol xhighBase modelOpenAI Codex$4.10API-equivalent estimate24:19.8 wall-clockpartial token timing and source ledgerRequests with more than 272,000 input tokens use long-context pricing; all 57 supplied calls were short-context. Separately priced tools and non-token services are excluded.Open result
GPT-5.6 Sol UltraBase modelOpenAI Codex$5.61API-equivalent estimate38:22.4 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 08 · Blender

Oasis Outpost

Blender MCP voxel-diorama construction, settlement readability, terrain, vegetation, water, and scene preservation.

Qwen ledger statuspartial token timing artifact validation ledgerEditorial assessment

Qwen produced a usable outpost in roughly the same time as Kimi, but Kimi's scene had better attention to detail and won the comparison despite Qwen's fixable overexposure.

Qwen3.8 Max Preview · Oasis Outpost presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing artifact validation ledger
Wall-clock
34m 48.5s
Processed tokens
2,916,981
Qualified API-list-price equivalent
$4.698861
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
blender scene
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (7)
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 70,221 thought tokens are a subset of the 92,397 output tokens.
  • The post-generation preservation and 38.939-second final render are excluded from the model wall-clock. The model did not save a .blend or render before its final response.
  • No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
  • The $4.70 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The before-meter predates a failed Cathedral attempt and successful Gothic City task; no post-run quota screenshot isolates this Oasis run.
  • No blind evaluation was supplied.
Evidence still missing (1)
  • blind-evaluation record
Oasis Outpost measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$4.698861Qualified API-list-price equivalent34m 48.5spartial token timing artifact validation ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxTool-assistedKimi Code CLI$2.72API-equivalent estimate39:27.9 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$21.23estimated API cost19:23.7 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$5.51API-equivalent estimate16:30.9 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Terra UltraTool-assistedOpenAI Codex$1.33API-equivalent estimate28:14.2 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Test 09 · Blender

Jungle Temple

Blender MCP high-density voxel construction, architectural readability, vegetation, and inspectable artifact creation.

Qwen ledger statuspartial token timing artifact validation ledgerEditorial assessment

Kimi was the clear winner. Qwen took 48m 49.9s yet produced the least coherent temple result and left much less to iterate on.

Qwen3.8 Max Preview · Jungle Temple presentation recording poster
Qwen3.8 Max Preview · presentation recording · not standardized FPS evidence
Exact status
partial token timing artifact validation ledger
Wall-clock
48m 49.9s
Processed tokens
4,347,869
Qualified API-list-price equivalent
$6.769691
Provider / client
Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0
Artifact classification
blender scene
Validator
No validator result attached
Runtime
See ledger caveats
Material caveats (8)
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 45,797 thought tokens are a subset of the 70,825 output tokens.
  • The post-generation 3.931-second render is excluded from model wall-clock.
  • No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
  • The $6.77 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The final capture has a pale, low-contrast palette and frames the temple lower-right rather than the prompt's upper-right requirement; no quality score is assigned.
  • The task crossed a five-hour quota reset boundary, so five-hour percentages cannot be subtracted. The weekly 12% to 16% change supports a nominal 100 Credits with a whole-percent display range of 75–125, not an invoice-grade exact charge.
  • No blind evaluation was supplied.
Evidence still missing (1)
  • blind-evaluation record
Jungle Temple measured comparison records
ConfigurationProvider / clientWorkflow costWall-clockRecordEvidence noteOpen result
Qwen3.8 Max PreviewQwen Code 0.20.0Alibaba ModelStudio Token Plan Personal Lite, Singapore · Qwen Code CLI 0.20.0$6.769691Qualified API-list-price equivalent48m 49.9spartial token timing artifact validation ledgerSingle recorded attempt. Reliability is not established.Open result
Kimi K3 MaxTool-assistedKimi Code CLI$2.60API-equivalent estimate53:37.7 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end workflow latency and includes tools, Blender renders, approval waits, and idle time.Open result
Claude Fable 5 MediumTool-assistedAnthropic Claude Code$10.10estimated API cost12:04.2 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Sol UltraTool-assistedOpenAI Codex$3.23API-equivalent estimate26:39.7 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
GPT-5.6 Terra UltraTool-assistedOpenAI Codex$1.20API-equivalent estimate16:43.2 wall-clockpartial token timing and artifact ledgerWall-clock is end-to-end latency, not model-only compute.Open result
Evidence boundary

The presentation video does not establish FPS or hardware facts absent from the ledger.

Vision and reliability chronology

Failed attempts are evidence, not extra scored rounds.

Metal fixed visibility. Disposable inspectors fixed image transport. Only the harness can enforce when the visual loop stops.

  1. Space Flight · attempt 1

    Runnable artifact retained with an orientation failure.

  2. Space Flight · attempt 2

    Five parent screenshots produced 24.80 MiB of parent context and HTTP 413.

  3. Space Flight · attempt 3

    Isolated visual inspection reached a runnable artifact, but used 26 visual-inspector contexts before the later hard cap existed.

  4. JRPG · attempt 1

    Qwen Code 0.20.0 screenshot compaction disabled mandatory thinking and returned HTTP 400.

  5. JRPG · attempt 2

    Qwen Code 0.20.1 reached 41/41 and passed runtime verification, while exceeding the intended visual budget.

  6. MacBook · attempt 1

    Ten parent images accumulated before an HTTP 413 failure.

  7. MacBook · attempt 2

    An incorrect camera +Z direction produced four uniform gray 16-bit renders and data_inspection_failed.

  8. MacBook · attempt 3

    The artifact reached 49/49, then the run used 20 inspectors, 19 direct parent image reads, 77 quota errors, correction renders through fix9, and ended without a final response.

The MacBook attempt demonstrated no reliable operational path to termination. It is not a claim that the model literally would have run forever.

Capacity receipt

Subscription Credits and API-equivalent dollars remain separate.

  1. The actual attempts used Alibaba ModelStudio Token Plan Personal Lite in Singapore.
  2. Each displayed dollar figure is a qualified NanoGPT API-list-price-equivalent scenario for the measured token mix.
  3. The values are not Alibaba provider cost, invoices, subscription charges, or cash spent.
Verified member Discord

Keep judging the builds with us.

Continue to the verified-member Discord after sign-in and discuss the published evidence.