Measured harness ledgerPublic result
GPT-6 Astra

Space Flight Stage 2 — Landfall — GPT-6 Astra Max

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

Max reasoningHeadline result
Workflow cost
$188.63
Wall-clock
Not recorded
Processed tokens
130.54M processed
Record state
partial_cumulative_stage1_and_landfall_root_metrics_clean_build_and_model_supplied_stage2_evidence
Public summary

GPT-6 Astra Max partial_cumulative_stage1_and_landfall_root_metrics_clean_build_and_model_supplied_stage2_evidence ledger: Not recorded wall-clock, 130.54M processed, and $188.63 Official OpenAI Standard API-list-price equivalent for the cumulative Stage 1 plus Landfall root-task receipt; not a Stage-2-only cost or an itemized Codex subscription charge.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-gpt-6-astra-max
  • Technical model: gpt-6-astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-6-astra
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Cost basis
  • Prompts above 272K input tokens use long-context pricing for the full request; all 1,101 supplied calls were short-context.
Recorded caveats
  • No valid wall-clock timing is available in the supplied receipt; no duration is reported here.
  • The published token counts and API-equivalent total are cumulative across Stage 1 and Landfall, not a Stage-2-only delta; subagent logs are excluded.
  • Output tokens include hidden reasoning, visible prose and code, and tool-call JSON.
  • Cache-read input is deeply discounted, so processed-token volume overstates cost.
  • The clean build verifies archive reproducibility, not gameplay behavior. The candidate's RC14 reports are model-supplied evidence, not independent verification.
  • The current source receipt does not byte-pin the provided Stage 1 baseline.
Visible evidence gaps
  • Stage-2-only token, timing, and all-agent measurement receipt
  • byte-pinned Stage 1 input baseline
  • independent real-input uninterrupted orbit-to-surface-to-orbit replay
  • independent repeatability replay
  • standardized FPS and transition-hitch receipt with hardware details
  • repair or explicitly accepted waiver for the close-ladder N camera-framing failure
  • surface-quality acceptance above the submitted minimum-3 failure
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console