Measured harness ledgerPublic result
GPT-5.6 Sol

JRPG boss battle — GPT-5.6 Sol Ultra

Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.

Ultra reasoningHeadline result
Workflow cost
$9.98
Wall-clock
25:51.8 wall-clock
Processed tokens
5.05M processed
Record state
partial_token_timing_artifact_validation_ledger
Public summary

GPT-5.6 Sol Ultra partial_token_timing_artifact_validation_ledger ledger: 25:51.8 wall-clock, 5.05M processed, and $9.98 API-equivalent estimate, not a subscription invoice.

Run identity and stack
  • Result ID: jrpg-boss-battle-gpt-5.6-sol-ultra
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js / Vite game
  • Stack: Supplied GLB fixture bundle
  • Stack: Harness v1 JRPG validator
  • Stack: Requested tool profile: codex-cli-imagegen
Cost basis
  • Cache-creation is zero because the Codex transcript schema does not expose a cache-write token field; no cache-write amount was inferred.
  • Priority service-tier pricing.
  • Requests with more than 272,000 input tokens use long-context pricing; all 55 supplied calls were short-context.
  • No cache-write amount was inferred from the supplied transcript schema.
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: interactive-threejs-jrpg-boss-battle
  • Path: artifacts/jrpg-boss-battle-gpt-5.6-sol-ultra/source/src/main.js
  • SHA-256: 50a64fd104936a58f185b62fcc71d082f753ecd88714c95964cc5d799f5a9e54
Validation evidence
  • Result: PASS, 41/41 checks
  • Verification: The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.
  • Path: artifacts/jrpg-boss-battle-gpt-5.6-sol-ultra/validation/scorecard.json
  • SHA-256: e90442b5074dac429a96457cf4c898ad2e82aab3daceca59d562e2b7ca9fd8c9
  • Validator SHA-256: b571ff1895fd6a61d98ccb4a9f7ca4ecefc961abfc3f6b77e23742bddd921ab8
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent Priority-tier estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • The exact browser version and local hardware identity were not supplied.
  • The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
Visible evidence gaps
  • browser and local-hardware identity
  • independent validator rerun
  • model tool-use transcript or disclosure
  • blind-evaluation record
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console