Measured harness ledgerPublic result
GPT-5.6 Luna

Campfire under the stars — GPT-5.6 Luna Max

Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.

Max reasoningHeadline result
Workflow cost
$0.06
Wall-clock
17m 11.6s wall-clock
Processed tokens
595K processed
Record state
partial_token_timing_and_artifact_ledger
Public summary

GPT-5.6 Luna Max partial_token_timing_and_artifact_ledger ledger: 17m 11.6s wall-clock, 595K processed, and $0.06 API-equivalent estimate, not a subscription invoice.

Run identity and stack
  • Result ID: campfire-threejs-gpt-5.6-luna-max
  • Technical model: gpt-5.6-luna
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-luna
  • Stack: Three.js interactive scene
  • Stack: Harness v1 campfire prompt
Cost basis
  • Requests with more than 272,000 input tokens use long-context pricing; all 14 observed calls were short-context.
Primary artifact integrity
  • Kind: single-file-interactive-threejs-campfire
  • Path: artifacts/campfire-threejs-gpt-5.6-luna-max/index.html
  • SHA-256: e0512f429bbf88d6b812b1c9ab9e5e6e78857f75b6781f97255d932aed9b4aaa
Recorded caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • The post-generation metrics request is included in the recorded wall-clock and token total.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • The scene depends on the Three.js r160 CDN and has not been independently browser-tested during archival.
  • No local FPS measurement, browser/hardware environment, final capture, or blind evaluation is archived.
Visible evidence gaps
  • browser and hardware environment
  • local FPS
  • final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console