Measured harness ledgerPublic result
GPT-5.6 Sol

Space Flight Stage 2 — Landfall — GPT-5.6 Sol Ultra

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

Ultra reasoningHeadline result
Workflow cost
$20.84
Wall-clock
1h 29m 57.7s wall-clock
Processed tokens
29.32M processed
Record state
partial_token_timing_source_build_static_acceptance_ledger
Public summary

GPT-5.6 Sol Ultra partial_token_timing_source_build_static_acceptance_ledger ledger: 1h 29m 57.7s wall-clock, 29.32M processed, and $20.84 Official OpenAI standard API-list-price equivalent; not an itemized Codex Pro subscription charge.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-gpt-5.6-sol-ultra
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Cost basis
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: Fresh vinext `dist/` directory archive
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-gpt-5.6-sol-ultra/production-build.tar.gz
  • SHA-256: 300e43fadb1e7941f166ca03138508fb7b60b5a60e0270f21f6602813f17ff53
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tools and idle time, not model-only compute.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a Codex Pro subscription-backed session; separately priced tools and non-token services are excluded.
  • The acceptance evidence is model-authored static/server-rendered testing plus a fresh build and lint run. No independent full-loop browser playthrough, FPS/hitch receipt, or final runtime capture is supplied.
  • The supplied social preview is retained for provenance but is not used as a final render or quality score.
  • No blind-evaluation record is supplied.
Visible evidence gaps
  • independent uninterrupted orbit-to-surface-to-orbit acceptance run
  • local FPS and transition-hitch receipt with hardware details
  • final browser or viewport capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console