Measured harness ledgerPublic result
GPT-5.6 SolFighter Game v1 — GPT-5.6 Sol Max
Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.
Max reasoningHeadline result
- Workflow cost
- $10.74
- Wall-clock
- 31m 38.6s wall-clock
- Processed tokens
- 15.61M processed
- Record state
- artifact_independent_boot_ledger_closed_book_verified
Public summary
GPT-5.6 Sol Max clean closed-book rerun artifact_independent_boot_ledger_closed_book_verified ledger: 31m 38.6s wall-clock, 15.61M processed, and $10.74 API-equivalent estimate from official OpenAI API Standard pricing for the measured token mix; NOT the actual subscription-backed Codex charge.
Run identity and stack
- Result ID: courtyard-arena-blender-mcp-gpt-5.6-sol-max
- Technical model: gpt-5.6-sol
- Provider: Provider not separately recorded
- Client: Codex
- Stack: Codex
- Stack: Technical model/configuration: gpt-5.6-sol
- Stack: Blender MCP
- Stack: Three.js fighting-game fixture
- Stack: Harness v1 frozen prompt
Cost basis
- Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/courtyard-arena-blender-mcp-gpt-5.6-sol-max-clean-rerun/source/web/models/arena.glb
- SHA-256: b3fc2884cee5666002320bbe5cb368ffb6872a2c23aaf251a3b29c73c9e863fa
Validation evidence
- Result: PASS
Recorded caveats
- The metrics receipt records 2 user turns where the protocol is 1 — disclosed, not disqualifying.
- Cache-creation is recorded as zero because Codex logs expose no cache-write count; it is a logging absence, not a measured zero.
- Wall-clock is end-to-end workflow latency including Blender builds and browser checks, not model-only compute.
- Output tokens include the 20,303-token reasoning subset, code, and tool-call JSON.
- Cache reads are discounted, so the 15.6M total-processed figure overstates cost.
- The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and excludes tool services.
- The deliverable captures were recovered from the run's own output directory, not from the handback folder.
- This is an environment-improvement task over a fixed starter, so its figures are not comparable to the archive's from-scratch build rows.
- The archive is a delta over the committed fixture, not a standalone runnable tree.
- No frame-rate receipt exists — neither the model nor the operator recorded a valid FPS sample.
- No blind evaluation and no scored side-by-side against the reference run or the other candidates are recorded.
- The metrics receipt records two user turns where the protocol is one fixed prompt — the same deviation the earlier attempt showed. Disclosed, not disqualifying.
Visible evidence gaps
- blind-evaluation record
- standardized FPS and hitch receipt under sustained play
- scored side-by-side against the reference run's judging anchors
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Fighter Game v1 · Builder projects v1
