Measured harness ledgerPublic result
GPT-5.6 Sol

Fighter Game v1 — GPT-5.6 Sol Max

Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.

Max reasoningHeadline result
Workflow cost
$10.74
Wall-clock
31m 38.6s wall-clock
Processed tokens
15.61M processed
Record state
artifact_independent_boot_ledger_closed_book_verified
Public summary

GPT-5.6 Sol Max clean closed-book rerun artifact_independent_boot_ledger_closed_book_verified ledger: 31m 38.6s wall-clock, 15.61M processed, and $10.74 API-equivalent estimate from official OpenAI API Standard pricing for the measured token mix; NOT the actual subscription-backed Codex charge.

Run identity and stack
  • Result ID: courtyard-arena-blender-mcp-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: Provider not separately recorded
  • Client: Codex
  • Stack: Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Blender MCP
  • Stack: Three.js fighting-game fixture
  • Stack: Harness v1 frozen prompt
Cost basis
  • Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/courtyard-arena-blender-mcp-gpt-5.6-sol-max-clean-rerun/source/web/models/arena.glb
  • SHA-256: b3fc2884cee5666002320bbe5cb368ffb6872a2c23aaf251a3b29c73c9e863fa
Validation evidence
  • Result: PASS
Recorded caveats
  • The metrics receipt records 2 user turns where the protocol is 1 — disclosed, not disqualifying.
  • Cache-creation is recorded as zero because Codex logs expose no cache-write count; it is a logging absence, not a measured zero.
  • Wall-clock is end-to-end workflow latency including Blender builds and browser checks, not model-only compute.
  • Output tokens include the 20,303-token reasoning subset, code, and tool-call JSON.
  • Cache reads are discounted, so the 15.6M total-processed figure overstates cost.
  • The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and excludes tool services.
  • The deliverable captures were recovered from the run's own output directory, not from the handback folder.
  • This is an environment-improvement task over a fixed starter, so its figures are not comparable to the archive's from-scratch build rows.
  • The archive is a delta over the committed fixture, not a standalone runnable tree.
  • No frame-rate receipt exists — neither the model nor the operator recorded a valid FPS sample.
  • No blind evaluation and no scored side-by-side against the reference run or the other candidates are recorded.
  • The metrics receipt records two user turns where the protocol is one fixed prompt — the same deviation the earlier attempt showed. Disclosed, not disqualifying.
Visible evidence gaps
  • blind-evaluation record
  • standardized FPS and hitch receipt under sustained play
  • scored side-by-side against the reference run's judging anchors
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Fighter Game v1 · Builder projects v1
RemakeBenchResearch console