Versioned benchmark collection

Campfire Interactive Scene v1

Harness v1 · 8 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Campfire benchmark output in the Campfire Interactive Scene v1 collection
Campfire benchmark output in the Campfire Interactive Scene v1 collection
8Measured runs
1Frozen targets
6Model families
8Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Campfire under the stars

Model coverage

  • Claude Fable 5 · Max
  • Claude Fable 5 · Medium
  • GPT-5.6 Luna · xhigh
  • GPT-5.6 Sol · Ultra
  • GPT-5.6 Sol · xhigh
  • GPT-5.6 Terra · Ultra
  • Kimi K3 · Max
  • Qwen3.8 Max Preview · Unreported

Evidence available

  • Seven measured run ledgers
  • Public browser builds
  • Cost, token, and workflow timing

Still missing

  • Browser/hardware environment
  • Local FPS for several runs
  • Final capture metadata
  • Minimum blind-vote sample
Public receipts

8 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.

01Claude Fable 5Max reasoning · Campfire under the stars$11.17 API-equivalent estimate, not a subscription invoicerecorded estimate12:48.9 wall-clockEnd-to-end workflowOpen 02Claude Fable 5Medium reasoning · Campfire under the stars$5.12 Anthropic first-party Claude API, standard global pricingrecorded estimate4:24.6 wall-clockEnd-to-end workflowOpen 03GPT-5.6 Lunaxhigh reasoning · Campfire under the stars$0.56 API-equivalent estimate, not a subscription invoicerecorded estimate8:40.2 wall-clockEnd-to-end workflowOpen 04GPT-5.6 SolUltra reasoning · Campfire under the stars$6.86 API-equivalent estimate, not a subscription invoicerecorded estimate23:08.4 wall-clockEnd-to-end workflowOpen 05GPT-5.6 Solxhigh reasoning · Campfire under the stars$2.26 API-equivalent estimate, not a subscription invoicerecorded estimate28:52.6 wall-clockEnd-to-end workflowOpen 06GPT-5.6 TerraUltra reasoning · Campfire under the stars$1.86 API-equivalent estimate, not a subscription invoicerecorded estimate24:28.8 wall-clockEnd-to-end workflowOpen 07Kimi K3Max reasoning · Campfire under the stars$0.45 API-equivalent usage accounting, not an itemized subscription cash chargerecorded estimate14:19.3 wall-clockEnd-to-end workflowOpen 08Qwen3.8 Max PreviewUnreported reasoning · Campfire under the stars$0.30 Third-party NanoGPT API-list-price equivalent for the measured token mix; not an Alibaba Token Plan cash chargerecorded estimate6m 0.9s wall-clockEnd-to-end workflowOpen