Measured harness ledgerPublic result
GPT-6 Astra

Campfire under the stars — GPT-6 Astra Max

Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.

Max reasoningHeadline result
Workflow cost
$15.49
Wall-clock
Not recorded
Processed tokens
11.38M processed
Record state
partial_qualified_root_task_snapshot_with_build_and_live_load
Public summary

GPT-6 Astra Max partial_qualified_root_task_snapshot_with_build_and_live_load ledger: Not recorded wall-clock, 11.38M processed, and $15.49 Standard API-equivalent estimate from a qualified root-task snapshot, not a subscription invoice.

Run identity and stack
  • Result ID: campfire-threejs-gpt-6-astra-max
  • Technical model: GPT-6 Astra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: GPT-6 Astra
  • Stack: Three.js interactive scene
  • Stack: Harness v1 campfire prompt
Recorded caveats
  • The supplied usage and cost totals are a qualified root-task snapshot rather than a clean isolated generation-turn score.
  • The unchanged parser did not recognize response_item user records, so raw parser timing is null. The displayed 51m 30.5s is a supplementary prompt-to-final-usage span, not model-only compute and not paired with isolated token accounting.
  • Wall-clock includes tools and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so processed-token volume overstates effective cost.
  • Cache creation is reported as zero because the transcript schema exposes no cache-write field; it is not proof that no cache write was billed.
  • The public source hosting configuration has its deployment project identifier replaced with null; no application source or asset was changed.
  • The archive operator observed the live page and controls but did not execute fresh direct keyboard, pointer, sound, pause, or atmosphere interactions.
  • No final capture, browser/hardware/viewport record, local FPS measurement, or blind evaluation is archived.
Visible evidence gaps
  • clean isolated generation-turn token and timing receipt
  • final capture with browser, hardware, viewport, and local FPS
  • fresh independent input replay
  • lint-clean source or documented lint remediation
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console