Versioned benchmark collection

Campfire Interactive Scene v1

Harness v1 · 8 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Campfire benchmark output in the Campfire Interactive Scene v1 collection
Campfire benchmark output in the Campfire Interactive Scene v1 collection
8Measured runs
1Frozen targets
6Model families
8Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Campfire Under a Starry Night

Model coverage

  • Claude Fable 5 · Medium
  • GPT-5.6 Luna · xhigh
  • GPT-5.6 Sol · Ultra
  • GPT-5.6 Terra · Ultra
  • Claude Fable 5 · Max
  • GPT-5.6 Sol · xhigh
  • Kimi K3 · Max
  • Qwen3.8 Max Preview · Unreported

Evidence available

  • Seven measured run ledgers
  • Public browser builds
  • Cost, token, and workflow timing

Still missing

  • Browser/hardware environment
  • Local FPS for several runs
  • Final capture metadata
  • Minimum blind-vote sample
Public receipts

8 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.