Shader / WebGLPick a benchmark. Build your version.
Browse every released Builder test individually. Get its exact prompt and any released projects, Harness workflows and skills.
12 benchmark tests included
Exact prompts for every test, plus the released source projects, RemakeBench Harness and skills attached to each one.
Shader / WebGL
Shader / WebGL
Three.js / browserSpace-Flight Game
- Prompt
- 4 projects
3 ready + 1 reference-only
The GPT-5.6 Sol Ultra pre-request snapshot is source-only and labelled reference-only. Bundled ship-model copies retain CC BY 4.0 attribution and modification notices.
Open test materials
Three.js / browserCampfire Under a Starry Night
- Prompt
- 4 projects
3 ready + 1 reference-only
GPT-5.6 Sol xhigh is a source-only reference capture; inspect its per-output status before use.
Open test materials
Blender
Blender
Blender
BlenderMacBook Cinematic
- Prompt
- 3 projects
3 partial projects
Customer scenes are privacy-sanitized and unbranded; supplied reference stills, checkpoint files, branding, keyboard legends, and system-font-derived material are excluded.
Open test materials
Three.js / browserJRPG Boss Battle
- Prompt
- 3 projects
3 reference-only projects
Supplied GLB fixtures and target-video media are excluded. Each project includes asset-reconstruction instructions and is labelled reference-only.
Open test materials
Three.js / browserRed Sands v1
- Prompt
- 4 projects
4 ready projects
Release smoke proves clean installation, production build, and the fixture smoke route; it does not prove a human completed each authored quest.
Open test materials
Three.js / browserFighter Game v1
- Prompt
- 4 projects
4 ready projects
Release smoke proves clean extraction, checksum verification, clean installation, production build, built-project start, and browser runtime boot; it is not a blind comparative evaluation or quality score.
Open test materials
UnityShrine Exploration
- Prompt
- Unity project
- Harness
- 4 skills
Completed Unity source project
Open test materialsResults, costs, tokens, workflow time, run ledgers, methodology, evidence, and known gaps remain free to inspect.
Latest published comparisons
Inspect the published outcomes first. Evidence details and known gaps stay accessible below each comparison.
Launch 004 · public comparisonClaude Opus 5 across nine game-building tests
Off-road Mud Game · Space Flight Stage 2 — Landfall · Neo-Gothic Flooded City · Infinite Cathedral · MacBook Pro · JRPG Boss Battle · Jungle Temple · Oasis Outpost · Shrine Exploration
Evidence available & known gaps
Evidence available
- Cost, workflow time, token, stack, artifact, validation, and caveat ledgers
- Nine frozen tests with 32 exact featured records
- Fourteen hash-bound Launch 004 presentation recordings
- Public Shrine receipts, disclosed prompt privacy findings, platform limitations, and exclusion proof
- Published owner benchmark assessment
Still missing
- Companion video episode metadata (deferred after benchmark publication)
- Repeat-run reliability evidence
- Complete Shrine subagent token mix and exact total cost
- WebGL gameplay, physical-input, and audible-output verification
- Windows runtime, gameplay, physical-input, and audio verification
Launch 003 · public comparisonI Gave Qwen3.8 Max, Kimi K3, GPT-5.6 and Fable 5 the Same 9 Game Tests
Neo-Gothic Storm City · MacBook-Class Cinematic · JRPG Boss Battle · Campfire Under a Starry Night · Space Flight Game · Shrine Village · Infinite Cathedral · Oasis Outpost · Jungle Temple
Evidence available & known gaps
Evidence available
- Cost, workflow time, token, stack, artifact, and caveat ledgers
- Nine frozen Qwen tests with a 42-record comparison pool
- Nine public Qwen media recordings and linked artifacts
- Published owner editorial assessment and official video
Still missing
- Minimum community preference sample
- Repeat-run reliability evidence
- A completed Qwen MacBook agent run
Launch 002 · public comparisonKimi K3 vs GPT-5.6 SOL vs Claude Fable 5: 9 Game Tests
Neo-Gothic Storm City · MacBook-Class Cinematic · JRPG Boss Battle · Campfire Under a Starry Night · Space Flight Game · Shrine Village · Infinite Cathedral · Oasis Outpost · Jungle Temple
Evidence available & known gaps
Evidence available
- Cost, workflow time, token, stack, and artifact ledgers
- Nine frozen tests with 33 featured run records
- Seven anonymous Kimi-versus-Fable blind matchups
- Published editorial assessment and official video
Still missing
- Minimum community preference sample
- Repeat-run reliability evidence
- One reconciled Kimi MacBook API request
Launch 001 · public comparisonGPT-5.6 vs Claude Fable 5: Same Prompts, 7 AI Game-Dev Tests
Infinite Cathedral Corridor · Neo-Gothic Storm City · Space Flight Game · Campfire Under a Starry Night · Jungle Temple · Shrine Village · Oasis Outpost
Evidence available & known gaps
Evidence available
- Cost, workflow time, token, stack, and artifact ledgers
- Seven standardized Fable-versus-Sol comparisons
- Anonymous blind exploration; sign-in only to record a vote
- Published editorial assessment and official video
Still missing
- Minimum community preference sample
- Repeat-run reliability evidence
Versioned benchmark collections
Stable scopes over every measured run. Open one to see which models were tested, how each attempt went, and where the public evidence still has gaps.
Harness v1 · Builder projects v1Red Sands v1
Red Sands v1
Evidence available & known gaps
Evidence available
- Four selected measured run ledgers
- Exact frozen prompt identity
- Per-run quest source and generated voice artifact identities
- Four reconstructed projects with clean install, production build, and browser smoke receipts
Still missing
- Human start-to-finish playthrough for every run
- Blind evaluation and quality scoring
- Repeat-run reliability evidence
Harness v1 · Builder projects v1Fighter Game v1
Fighter Game v1
Evidence available & known gaps
Evidence available
- Four selected measured run ledgers
- Exact frozen prompt identity
- Per-run arena-geometry artifact identities
- Exact clean GPT-5.6 Sol rerun with superseded contaminated attempt excluded
Still missing
- Blind evaluation and scored side-by-side review
- Standardized FPS and hitch receipts
- Repeat-run reliability evidence
Harness v1Space Flight v1
Explorable space-flight game
Evidence available & known gaps
Evidence available
- Eight measured run ledgers
- Public browser builds
- Cost, token, and workflow timing
Still missing
- Browser/hardware environment
- Local FPS
- Final capture metadata
- Minimum blind-vote sample
Harness v2Space Flight v2
Space Flight Stage 2 — Landfall
Evidence available & known gaps
Evidence available
- Four measured run ledgers
- Four public presentation recordings
- Cost, token, workflow timing, and artifact receipts
Still missing
- Repeat-run reliability evidence
- Standardized gameplay and physical-input verification
- Comparable local FPS measurements
Harness v1Shader Frontier v1
Infinite cathedral corridor shader · Neo-gothic storm city shader
Evidence available & known gaps
Evidence available
- Nineteen measured run ledgers
- Shader artifacts with SHA-256
- Cost, token, and workflow timing
Still missing
- Initial and optimized local FPS
- RTX Pro 6000 final-capture metadata
- Minimum blind-vote sample
Harness v1 · Capture protocol v1Blender Voxel Worlds v1
High-voxel-density Jungle Temple diorama · High-voxel-density Oasis Outpost diorama · High-voxel-density Shrine Village diorama
Evidence available & known gaps
Evidence available
- Twenty-four measured run ledgers
- Primary artifact SHA-256 records
- Standardized turntable media
Still missing
- Recorded Blender/render environment for original runs
- Repeat-run evidence
- Minimum blind-vote sample
Harness v1Campfire Interactive Scene v1
Campfire under the stars
Evidence available & known gaps
Evidence available
- Seven measured run ledgers
- Public browser builds
- Cost, token, and workflow timing
Still missing
- Browser/hardware environment
- Local FPS for several runs
- Final capture metadata
- Minimum blind-vote sample
Harness v1MacBook Cinematic v1
MacBook-class cinematic ad scene
Evidence available & known gaps
Evidence available
- Five measured run ledgers
- Public recorded outputs
- Cost, token, and workflow timing
Still missing
- One reconciled Kimi API request
- Repeat-run reliability evidence
- Minimum blind-vote sample
Harness v1JRPG Boss Battle v1
JRPG boss battle
Evidence available & known gaps
Evidence available
- Five measured run ledgers
- Public recorded outputs
- Cost, token, and workflow timing
Still missing
- Repeat-run reliability evidence
- Minimum blind-vote sample
Harness v1Interactive Voxel Viewer v1
Interactive Three.js voxel map viewer
Evidence available & known gaps
Evidence available
- Seven measured run ledgers
- Primary artifact integrity records
- Cost, token, and workflow timing
Still missing
- Standardized public capture set
- Repeat-run reliability evidence
- Minimum blind-vote sample
Result v1 · public evidenceShrine Exploration v1
Shrine Exploration
Evidence available & known gaps
Evidence available
- One measured long-horizon result record
- 12 → 14 → 14 / 40 visual score history
- $204.17+ parent-session API list-price lower-bound receipt
- Three hosted public-player archives with streamed read-back and a WebGL startup route
- Target-specific Builder-source artifact identities
Still missing
- Complete subagent token mix and exact total cost
- Repeat-run reliability evidence
Discuss the result. Build the next one.
Join after signing in. Builder adds released reproduction components and private Builder Lab; it is not required for the community Discord or public benchmark evidence.
