Measured harness ledgerPublic result
Kimi K3Red Sands v1 — Kimi K3 Unreported
Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.
Unreported reasoningHeadline result
- Workflow cost
- ≥$5.68
- Wall-clock
- 1h 0m 6.6s (measured window only) wall-clock
- Processed tokens
- 15.13M processed
- Record state
- artifact_model_playthrough_complete_ledger_partial_window_metrics
Public summary
Kimi K3 Unreported artifact_model_playthrough_complete_ledger_partial_window_metrics ledger: 1h 0m 6.6s (measured window only) wall-clock, 15.13M processed, and ≥$5.68 Independently recomputed API-list-price equivalent for the measured token mix; not an itemized subscription charge (lower bound).
Run identity and stack
- Result ID: red-sands-quest-threejs-kimi-k3-max
- Technical model: k3
- Provider: openai-compatible
- Client: Kimi Code CLI 0.34.0
- Stack: openai-compatible
- Stack: Kimi Code CLI 0.34.0
- Stack: Technical model/configuration: k3
- Stack: Three.js quest-authoring fixture
- Stack: Harness v1 frozen prompt
Cost basis
- MEASURED WINDOW ONLY — roughly the final third of the task. See measurement_window_caveat.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/red-sands-quest-threejs-kimi-k3-max/source/src/quest/quests/an-honest-man.js
- SHA-256: d0bdf4e5b7e5ec259f3f352940aa5e279c5b514327b64db33b72b3c97144c8da
Recorded caveats
- THE METRICS COVER ROUGHLY THE FINAL THIRD OF THE TASK. The receipt discloses this itself: the window anchors at a mid-task 'continue' message at 03:55:57Z, while the quest prompt was at 02:03:00Z. Measured wall-clock is 3,606.6 s against a 10,383 s task span — about 34.7%. Tokens, timing, and cost are lower bounds.
- This task has NO automated completion gate, by design.
- No operator playthrough exists — the verification browser cannot boot the game, because a backgrounded pane throttles the rAF-driven ~10 s terrain generation.
- npm run smoke runs under ?capture=1, which never constructs the quest system.
- The completion evidence is the model's own scripted playthrough record, not an operator reproduction.
- The prompt's required handback note was not delivered; no stated-gaps section exists in the submission.
- Kimi Code does not emit a separate reasoning-token count, so reasoning is folded into the 25,710 output tokens.
- Cache reads are discounted, so total processed tokens overstate cost.
- The Allegretto plan does not itemize a per-run cash charge.
- No quota evidence exists — there is no /usage capture in the session.
- The voice audio was not listened to; delivery, count, and casting are confirmed, performance is not.
- This is a quest-authorship task over a fixed game, so its figures are not comparable to the archive's from-scratch build rows.
- The archive is a delta over the fixture, not a standalone runnable tree.
- No quality score and no blind-evaluation record exist — which for this task is where the entire measurement lives.
Visible evidence gaps
- a whole-task metrics receipt anchored at the first benchmark user turn
- a human playthrough of the quest, start to finish
- blind-evaluation record — the primary instrument for this task
- the handback note the prompt asks for
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Red Sands v1 · Builder projects v1
