Measured harness ledgerPublic result
GPT-5.6 Luna

Space Flight Stage 2 — Landfall — GPT-5.6 Luna Max

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

Max reasoningHeadline result
Workflow cost
$0.69
Wall-clock
1h 28m 30.7s wall-clock
Processed tokens
23.34M processed
Record state
partial_token_timing_production_build_and_supplied_acceptance_evidence_ledger
Public summary

GPT-5.6 Luna Max partial_token_timing_production_build_and_supplied_acceptance_evidence_ledger ledger: 1h 28m 30.7s wall-clock, 23.34M processed, and $0.69 Official OpenAI standard API-list-price equivalent; not an itemized ChatGPT Pro/Codex subscription charge.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-gpt-5.6-luna-max
  • Technical model: gpt-5.6-luna
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-luna
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Cost basis
  • Prompts above 272K input tokens use long-context pricing for the full request; all 146 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: user-supplied-final-vinext-production-build
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-gpt-5.6-luna-max/production-build.tar.gz
  • SHA-256: 050e47f28b369bc1a61c05760a814c0c36df58b81322705f8d00ea21f41b90e9
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tools and idle time between two user turns, not model-only compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a Codex Pro subscription-backed session; separately priced tools and non-token services are excluded.
  • The published production build is user-supplied. It was statically inspected but not rebuilt or rerun by the archive operator.
  • The supplied QA log demonstrates one recorded descent, surface exploration, and return to orbit, but the verifier timed out before its second-landing assertion. Repeatability therefore remains inconclusive.
  • No local FPS, transition-hitch receipt, hardware identity, independent final runtime capture, or blind-evaluation record is supplied.
Visible evidence gaps
  • one-user-turn benchmark protocol
  • independent uninterrupted orbit-to-surface-to-orbit acceptance run
  • repeatable second landing acceptance run
  • local FPS and transition-hitch receipt with hardware details
  • final browser or viewport capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console