Measured harness ledgerPublic result
GPT-5.6 Luna

Fighter Game v1 — GPT-5.6 Luna Max

Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.

Max reasoningHeadline result
Workflow cost
$0.42
Wall-clock
33m 29.7s wall-clock
Processed tokens
14.98M processed
Record state
partial_multi_user_turn_token_timing_artifact_ledger
Public summary

GPT-5.6 Luna Max partial_multi_user_turn_token_timing_artifact_ledger ledger: 33m 29.7s wall-clock, 14.98M processed, and $0.42 API-equivalent estimate from official OpenAI Standard pricing; not the actual subscription-backed Codex charge.

Run identity and stack
  • Result ID: courtyard-arena-blender-mcp-gpt-5.6-luna-max
  • Technical model: gpt-5.6-luna
  • Provider: Provider not separately recorded
  • Client: Codex
  • Stack: Codex
  • Stack: Technical model/configuration: gpt-5.6-luna
  • Stack: Blender MCP
  • Stack: Three.js fighting-game fixture
  • Stack: Harness v1 frozen prompt
Cost basis
  • Prompts above 272K input tokens price at 2x input and 1.5x output for the whole request. No call crossed the threshold.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/courtyard-arena-blender-mcp-gpt-5.6-luna-max/artifact/arena_polished.blend
  • SHA-256: 35b5b9f3efbf014f7cfb5348cc843529134969c92ea7f9c661f30088814651eb
Recorded caveats
  • Two user turns are recorded where the fixture protocol is one fixed prompt.
  • The candidate input boundary is unverified, so this result makes no closed-book or blind-vote-eligibility claim.
  • Wall-clock is end-to-end workflow latency, not model-only compute.
  • Output includes hidden reasoning, generated code, and tool-call JSON.
  • Cache reads are discounted, so total tokens processed overstate cost.
  • Cache-creation is a transcript-schema absence, not a measured zero.
  • The public cost is an API-equivalent estimate, not a subscription invoice; tools and non-token services are excluded.
  • No independent Blender reopen, browser replay, FPS/hitch measurement, browser-console receipt, or blind evaluation is recorded.
Visible evidence gaps
  • verified closed-book candidate input boundary
  • independent runtime replay and console receipt
  • standardized FPS and hitch receipt
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console