Measured harness ledgerPublic result
GPT-5.6 Sol

F-117 Stealth Jet — GPT-5.6 Sol Max

Build an interactive Three.js F-117 stealth-jet experience under the pinned Harness task and fixture contract.

Max reasoningHeadline result
Workflow cost
$4.51
Wall-clock
15m 04.6s wall-clock
Processed tokens
5.32M processed
Record state
partial_two_user_turn_token_timing_source_build_ledger
Public summary

GPT-5.6 Sol Max partial_two_user_turn_token_timing_source_build_ledger ledger: 15m 04.6s wall-clock, 5.32M processed, and $4.51 Token-only standard API-equivalent estimate, not the actual Codex subscription charge.

Run identity and stack
  • Result ID: f117-stealth-jet-threejs-gpt-5.6-sol-max
  • Technical model: gpt-5.6-sol
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-sol
  • Stack: Three.js flight experience
  • Stack: Harness v1 F-117 prompt
Cost basis
  • Prompts with more than 272,000 input tokens use the published long-context rate for the whole request. The receipt reports zero calls over that threshold.
  • No separately priced tool or non-token cost is included because no verifiable usage counts were supplied.
Primary artifact integrity
  • Kind: interactive-threejs-f117-model-component
  • Path: artifacts/f117-stealth-jet-threejs-gpt-5.6-sol-max/source/app/NighthawkExperience.tsx
  • SHA-256: 34d264adbf0b9aee37870d9f65dfddf3db7d1e7c67f4695a771d2cc599a6b2ed
Validation evidence
  • Path: artifacts/f117-stealth-jet-threejs-gpt-5.6-sol-max/validation/build-and-test.public.json
  • SHA-256: f44396a620958ce02c26d3581c9af74a6ee3c987dba5c21913695f332459cb0e
Recorded caveats
  • The root-task receipt records two user turns where the frozen benchmark protocol expects one; it is disclosed rather than treated as a clean one-shot result.
  • Wall-clock is end-to-end workflow latency including tools and idle gaps, not model-only compute time.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens substantially overstate cost.
  • Cache creation is recorded as zero because the transcript schema did not expose a cache-write field; no estimate was inferred.
  • The cost is an API-list-price equivalent, not the actual subscription-backed Codex charge, and separately priced tools are excluded.
  • The production build passed, but the bundled starter tests are stale against the delivered app and fail. No independent browser acceptance run was available.
  • No local FPS measurement, final capture, hardware/viewport receipt, or blind evaluation was supplied.
  • The supplied root-task receipt reports two user turns but does not identify the purpose or content of each. This is therefore not presented as a clean one-shot run.
Visible evidence gaps
  • clean one-user-turn rerun
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console