Measured harness ledgerPublic result
GPT-6 Astra

Off-road Mud Game — GPT-6 Astra Max

Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.

Max reasoningHeadline result
Workflow cost
$274.71
Wall-clock
13h 57m 41.1s wall-clock
Processed tokens
195.96M processed
Record state
partial_cumulative_root_snapshot_with_clean_build_and_model_supplied_runtime_evidence
Public summary

GPT-6 Astra Max partial_cumulative_root_snapshot_with_clean_build_and_model_supplied_runtime_evidence ledger: 13h 57m 41.1s wall-clock, 195.96M processed, and $274.71 Standard API-equivalent estimate from the supplied root Codex receipt, not an actual subscription charge or all-agent spend..

Run identity and stack
  • Result ID: off-road-driving-game-threejs-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Three.js vehicle workflow
  • Stack: Harness v1 off-road prompt
Primary artifact integrity
  • Kind: interactive-threejs-off-road-driving-game
  • Path: artifacts/off-road-driving-game-threejs-gpt-6-astra-max/source/src/main.js
  • SHA-256: 9debf02c80a8f4534a9837d1ac9514f515ef3beeb7768509cff4dc4a4b9b8fbf
Recorded caveats
  • The cost, tokens, and wall-clock are a cumulative root-transcript snapshot through the metrics request, not a strict completed-generation-only measurement.
  • Subagent logs are excluded to avoid inherited-context double counting, so the displayed figures are not a complete all-agent spend total.
  • Wall-clock includes tools and idle gaps; output includes hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates effective cost.
  • The archive operator reproduced a clean build but did not independently drive the game, measure FPS, or judge visual or 3D-asset quality. Submitted runtime and visual evidence is labelled model-supplied.
  • The submitted driving receipt reports 31.95 FPS on an Apple M1 Pro Balanced setting, not 60 FPS or a standardized desktop-GPU result.
  • The submitted report documents an explicit eight-round limit and remaining visual limitations.
  • No blind evaluation is recorded.
Visible evidence gaps
  • strict completed-generation-only all-agent token and timing ledger
  • independent browser gameplay replay through real input
  • standardized fixed-hardware FPS measurement
  • independent visual and 3D asset-quality acceptance
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console