Measured harness ledgerPublic result
Kimi K3

ECHO Main Menu — Kimi K3 Max

Recreate the pinned ECHO-style cinematic main menu reference as an interactive browser experience.

Max reasoningHeadline result
Workflow cost
Not recorded
Wall-clock
1h 49m 29s wall-clock
Processed tokens
32.57M processed
Record state
verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation
Public summary

Kimi K3 Max verified_exact_metrics_runtime_replay_qualified_protocol_deviation_no_blind_evaluation ledger: 1h 49m 29s wall-clock, 32.57M processed, and Not recorded.

Run identity and stack
  • Result ID: echo-main-menu-browser-kimi-k3-max
  • Technical model: k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: k3
  • Stack: Browser recreation
  • Stack: Pinned visual references
  • Stack: Harness v1 ECHO prompt
Cost basis
  • No measured or estimated run cost is recorded.
Primary artifact integrity
  • Kind: interactive-browser-title-menu
  • Path: artifacts/echo-main-menu-browser-kimi-k3-max/source/index.html
  • SHA-256: 6c4cfbf6ab45ec3912c5151189315ee9096eadcc9f1305b38d422ee6073a55c9
Validation evidence
  • Result: PASS
Recorded caveats
  • The run is one preserved candidate session but not one uninterrupted runner process; all quota-resume segments are disclosed.
  • Chronological wall-clock includes provider quota wait time; active runner wall-clock excludes gaps between invocations but includes tools and failed quota probes.
  • One user-visible benchmark contained 163 attempted requests, of which 155 completed with usage and eight ended at the provider quota boundary.
  • Output tokens include provider-accounted thinking, visible prose/code, and tool-call JSON because the Kimi wire format has no separate reasoning-token field.
  • Cache-read tokens are discounted; total processed tokens therefore overstate API-equivalent cost.
  • The $15.14 value is an API-equivalent estimate from current official K3 rates, not an actual marginal Kimi Code subscription charge.
  • Independent replay verifies runtime behavior and performance on one machine/browser, not visual-fidelity scoring or cross-device performance.
  • The initial run and first continuation reached Kimi's five-hour quota before end_turn. Six later same-session probes were rejected with HTTP 403 before generation. The exact preserved conversation and workspace were resumed after quota reset, and the final continuation ran until Kimi returned end_turn.
Visible evidence gaps
  • blind visual-fidelity evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console