Versioned benchmark collection
MacBook Cinematic v1
Harness v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- MacBook-Class Cinematic
Model coverage
- Claude Fable 5 · Max
- GPT-5.6 Sol · xhigh
- Kimi K3 · Max
- Qwen3.8 Max Preview · Mandatory thinking enabled
Evidence available
- Three measured run ledgers
- Public recorded outputs
- Cost, token, and workflow timing
Still missing
- One reconciled Kimi API request
- Repeat-run reliability evidence
- Minimum blind-vote sample
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01Claude Fable 5Max reasoning · MacBook-Class Cinematic$75.18 estimated API-equivalent costrecorded estimate1:01:04.8 wall-clockEnd-to-end workflowOpen 02GPT-5.6 Solxhigh reasoning · MacBook-Class Cinematic$7.67 estimated API-equivalent costrecorded estimate55:49.9 wall-clockEnd-to-end workflowOpen 03Kimi K3Max reasoning · MacBook-Class Cinematic≥$10.85 estimated API-equivalent cost (lower bound)recorded lower bound1:26:58.2 wall-clockEnd-to-end workflowOpen 04Qwen3.8 Max PreviewMandatory thinking enabled reasoning · MacBook-Class Cinematic$29.27 estimated API-equivalent costrecorded estimate7h 00m 56.4s wall-clockEnd-to-end workflowOpen
