Measured harness ledgerPublic result
GPT-5.6 Sol

Explorable space-flight game — GPT-5.6 Sol Ultra

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Ultra reasoningHeadline result
Workflow cost
$17.94
Wall-clock
37:44.2 wall-clock
Processed tokens
25.90M processed
Record state
partial_token_timing_and_source_ledger
Public summary

GPT-5.6 Sol Ultra partial_token_timing_and_source_ledger ledger: 37:44.2 wall-clock, 25.90M processed, and $17.94 user-supplied API-equivalent pre-request snapshot.

Run identity and stack
  • Result ID: space-flight-game-gpt-5.6-sol-ultra-pre-request-snapshot
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Primary artifact integrity
  • Kind: threejs-space-flight-game-primary-component
  • Path: artifacts/space-flight-game-gpt-5.6-sol-ultra/source/app/game/HeliosDrift.tsx
  • SHA-256: ea4a13f7873e1ef004ca05d5c88263cf3b50f5b7bf7aca196824d3f0388a4e93
Recorded caveats
  • The supplied metrics name gpt-5.6-sol; that model ID is used instead of the surrounding GPT 5.5 label.
  • This is a pre-request snapshot. The bundled post-generation meta.json represents a later, larger session and is not substituted for the supplied snapshot.
  • The supplied cost applies standard short-context rates to the aggregate tokens. Per-call long-context usage was not supplied, so no context-bucket adjustment has been inferred.
  • Wall-clock is end-to-end latency, not model-only compute. Cache-read tokens are discounted, so total processed tokens overstate cost as uncached volume.
Visible evidence gaps
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Builder test available

This result is part of Builder tests. Open them for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Tech Review 001 · v1
  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console