Measured harness ledgerPublic result
Kimi K3The Last of Us Main Menu — Kimi K3 Max
Recreate the pinned The Last of Us-style main menu reference as an interactive browser experience.
Max reasoningHeadline result
- Workflow cost
- Not recorded
- Wall-clock
- 43m 37s wall-clock
- Processed tokens
- 11.55M processed
- Record state
- verified_exact_metrics_runtime_replay_no_blind_evaluation
Public summary
Kimi K3 Max verified_exact_metrics_runtime_replay_no_blind_evaluation ledger: 43m 37s wall-clock, 11.55M processed, and Not recorded.
Run identity and stack
- Result ID: last-of-us-main-menu-browser-kimi-k3-max
- Technical model: k3
- Provider: Kimi Code CLI / Moonshot AI
- Stack: Kimi Code CLI / Moonshot AI
- Stack: Technical model/configuration: k3
- Stack: Browser recreation
- Stack: Pinned visual reference
- Stack: Harness v1 menu prompt
Cost basis
- No measured or estimated run cost is recorded.
Primary artifact integrity
- Kind: interactive-browser-title-menu
- Path: artifacts/last-of-us-main-menu-browser-kimi-k3-max/source/index.html
- SHA-256: 1ad0473aae1d28c70341c443c81b230b49572102e6b7c19b438d52a7b8ac3df7
Validation evidence
- Result: PASS
Recorded caveats
- Wall-clock time includes tool execution, installation, rendering, visual iteration, and browser verification; it is not model-only compute time.
- Output tokens include provider-accounted thinking, visible prose/code, and tool-call JSON because the Kimi wire format has no separate reasoning-token field.
- The ¥35.88 value is an API-equivalent estimate from official K3 rates, not an actual marginal Kimi Code subscription charge.
- npm audit reported three high-severity dependency advisories in the frozen candidate dependency graph; the benchmark archive preserves the candidate output rather than silently upgrading it.
- Independent replay verifies runtime behavior and performance on one machine/browser, not visual-fidelity scoring or cross-device performance.
Visible evidence gaps
- blind visual-fidelity evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
