Measured harness ledgerPublic result
GPT-6 Astra

Explorable space-flight game — GPT-6 Astra Max

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Max reasoningHeadline result
Workflow cost
$20.53
Wall-clock
63m 51.9s (qualified snapshot span) wall-clock
Processed tokens
14.07M processed
Record state
partial_post_task_metrics_snapshot_with_build_and_live_load
Public summary

GPT-6 Astra Max partial_post_task_metrics_snapshot_with_build_and_live_load ledger: 63m 51.9s (qualified snapshot span) wall-clock, 14.07M processed, and $20.53 Standard API-equivalent estimate from a qualified cumulative root-session snapshot, not a subscription invoice.

Run identity and stack
  • Result ID: space-flight-game-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Recorded caveats
  • The supplied token and cost totals are cumulative root-session snapshot values that include the beginning of the metrics follow-up; they are not a clean isolated generation-turn score.
  • The raw parser did not recognize response_item user records, so the displayed wall-clock is a schema-aware recovered span rather than the parser's raw duration.
  • Wall-clock includes tools and idle gaps, not model-only compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates cost.
  • Supplied CDP input and FPS evidence is retained, but the archive operator only independently rebuilt and live-loaded the game; no fresh keyboard/mouse flight replay was run.
  • Vinext's build IDs and chunk hashes are non-deterministic, so the successful clean build does not byte-match the submitted production tree.
  • Blind visual evaluation remains outstanding.
Visible evidence gaps
  • clean isolated generation-turn token and timing receipt
  • fresh independent keyboard and mouse flight replay
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console