Versioned benchmark collection

Interactive Voxel Viewer v1

Harness v1 · 8 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

RemakeBench voxel collection artwork
RemakeBench voxel collection artwork
8Measured runs
1Frozen targets
6Model families
8Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • Interactive Voxel Map Viewer

Model coverage

  • Claude Fable 5 · Medium
  • GPT-5.6 Sol · Ultra
  • GPT-5.6 Terra · Ultra
  • GPT-5.6 Luna · xhigh
  • Claude Fable 5 · Max
  • GPT-5.6 Sol · xhigh
  • Kimi K3 · Max
  • Qwen3.8 Max Preview · Unreported

Evidence available

  • Seven measured run ledgers
  • Primary artifact integrity records
  • Cost, token, and workflow timing

Still missing

  • Standardized public capture set
  • Repeat-run reliability evidence
  • Minimum blind-vote sample
Public receipts

8 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.