Measured harness ledgerPublic result
Kimi K3ECHO Main Menu — Kimi K3 Max
Recreate the pinned ECHO-style cinematic main menu reference as an interactive browser experience.
Max reasoningHeadline result
- Workflow cost
- Not recorded
- Wall-clock
- 1h 49m 29s wall-clock
- Processed tokens
- 32.57M processed
- Record state
- verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation
Public summary
Kimi K3 Max verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation ledger: 1h 49m 29s wall-clock, 32.57M processed, and Not recorded.
Run identity and stack
- Result ID: echo-main-menu-browser-kimi-k3-max
- Technical model: k3
- Provider: Kimi Code CLI / Moonshot AI
- Stack: Kimi Code CLI / Moonshot AI
- Stack: Technical model/configuration: k3
- Stack: Browser recreation
- Stack: Pinned visual references
- Stack: Harness v1 ECHO prompt
Cost basis
- No measured or estimated run cost is recorded.
Primary artifact integrity
- Kind: interactive-browser-title-menu
- Path: artifacts/echo-main-menu-browser-kimi-k3-max/source/index.html
- SHA-256: 6c4cfbf6ab45ec3912c5151189315ee9096eadcc9f1305b38d422ee6073a55c9
Validation evidence
- Result: PASS
Recorded caveats
- The run is one preserved candidate session but not one uninterrupted runner process; all quota-resume segments are disclosed.
- Chronological wall-clock includes provider quota wait time; active runner wall-clock excludes gaps between invocations but includes tools and failed quota probes.
- One user-visible benchmark contained 163 attempted requests, of which 155 completed with usage and eight ended at the provider quota boundary.
- Output tokens include provider-accounted thinking, visible prose/code, and tool-call JSON because the Kimi wire format has no separate reasoning-token field.
- Cache-read tokens are discounted; total processed tokens therefore overstate API-equivalent cost.
- The $15.14 value is an API-equivalent estimate from current official K3 rates, not an actual marginal Kimi Code subscription charge.
- Independent replay verifies runtime behavior and performance on one machine/browser, not visual-fidelity scoring or cross-device performance.
- The initial run and first continuation reached Kimi's five-hour quota before end_turn. Six later same-session probes were rejected with HTTP 403 before generation. The exact preserved conversation and workspace were resumed after quota reset, and the final continuation ran until Kimi returned end_turn.
Visible evidence gaps
- blind visual-fidelity evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
