Measured harness ledgerPublic result
GPT-6 AstraCampfire under the stars — GPT-6 Astra Max
Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.
Max reasoningHeadline result
- Workflow cost
- $15.49
- Wall-clock
- Not recorded
- Processed tokens
- 11.38M processed
- Record state
- partial_qualified_root_task_snapshot_with_build_and_live_load
Public summary
GPT-6 Astra Max partial_qualified_root_task_snapshot_with_build_and_live_load ledger: Not recorded wall-clock, 11.38M processed, and $15.49 Standard API-equivalent estimate from a qualified root-task snapshot, not a subscription invoice.
Run identity and stack
- Result ID: campfire-threejs-gpt-6-astra-max
- Technical model: GPT-6 Astra
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: GPT-6 Astra
- Stack: Three.js interactive scene
- Stack: Harness v1 campfire prompt
Recorded caveats
- The supplied usage and cost totals are a qualified root-task snapshot rather than a clean isolated generation-turn score.
- The unchanged parser did not recognize response_item user records, so raw parser timing is null. The displayed 51m 30.5s is a supplementary prompt-to-final-usage span, not model-only compute and not paired with isolated token accounting.
- Wall-clock includes tools and idle gaps between user turns.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cache reads are discounted, so processed-token volume overstates effective cost.
- Cache creation is reported as zero because the transcript schema exposes no cache-write field; it is not proof that no cache write was billed.
- The public source hosting configuration has its deployment project identifier replaced with null; no application source or asset was changed.
- The archive operator observed the live page and controls but did not execute fresh direct keyboard, pointer, sound, pause, or atmosphere interactions.
- No final capture, browser/hardware/viewport record, local FPS measurement, or blind evaluation is archived.
Visible evidence gaps
- clean isolated generation-turn token and timing receipt
- final capture with browser, hardware, viewport, and local FPS
- fresh independent input replay
- lint-clean source or documented lint remediation
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.