Measured harness ledgerPublic result
GPT-5.6 LunaExplorable space-flight game — GPT-5.6 Luna Max
Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.
Max reasoningHeadline result
- Workflow cost
- $0.25
- Wall-clock
- 22m 48s wall-clock
- Processed tokens
- 7.05M processed
- Record state
- partial_token_timing_source_and_build_ledger
Public summary
GPT-5.6 Luna Max partial_token_timing_source_and_build_ledger ledger: 22m 48s wall-clock, 7.05M processed, and $0.25 API-equivalent estimate, not a subscription invoice.
Run identity and stack
- Result ID: space-flight-game-gpt-5.6-luna-max
- Technical model: gpt-5.6-luna
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-5.6-luna
- Stack: Three.js / Vite game
- Stack: Generated spaceship assets
- Stack: Harness v1 space-flight prompt
Cost basis
- Cache-creation is zero because the Codex transcript schema did not expose a cache-write field; no cache-write amount was inferred.
- Default service-tier pricing.
- Prompts with more than 272,000 input tokens use long-context pricing; all 82 supplied calls were short-context.
- No cache-write amount was inferred from the supplied transcript schema.
- Separately priced tools and non-token services are excluded.
Primary artifact integrity
- Kind: interactive-threejs-space-flight-game
- Path: artifacts/space-flight-game-gpt-5.6-luna-max/source/app/page.tsx
- SHA-256: ae64b2a75324215ee2d9e821e5211a4d4109be07791bdfd5eececf6050836db3
Recorded caveats
- Wall-clock is end-to-end workflow latency, not model-only compute; it includes tool time and idle gaps between user turns.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cached input is deeply discounted, so total processed tokens overstate cost.
- This is an API-equivalent default-tier estimate rather than the actual charge for a subscription-backed Codex session.
- Separately priced tools and non-token services are excluded.
- The archive build check does not establish interactive frame rate or visual quality.
Visible evidence gaps
- browser and local-hardware identity
- local FPS
- final capture
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
