Measured harness ledgerPublic result
Kimi K3

Off-road Mud Game — Kimi K3 Max

Build a one-request procedural off-road driving game with vehicle physics, independent suspension, streamed terrain, mud feedback, and multiple cameras.

Max reasoningHeadline result
Workflow cost
$28.72
Wall-clock
3h 50m 07.6s wall-clock
Processed tokens
64.18M processed
Record state
partial_token_timing_source_build_ledger
Public summary

Kimi K3 Max partial_token_timing_source_build_ledger ledger: 3h 50m 07.6s wall-clock, 64.18M processed, and $28.72 API-equivalent usage accounting, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: off-road-driving-game-threejs-kimi-k3-max
  • Technical model: kimi-k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: kimi-k3
  • Stack: Three.js vehicle workflow
  • Stack: Harness v1 off-road prompt
Primary artifact integrity
  • Kind: fresh-vite-dist-directory-archive
  • Path: artifacts/off-road-driving-game-threejs-kimi-k3-max/production-build.tar.gz
  • SHA-256: 1a7f8858deb97fb7e086fd21d8d0bc46e9df18a04017ae7d96559d45d74c74a8
Recorded caveats
  • Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
  • One user prompt produced 355 billable Kimi model requests across one main and three child agents.
  • Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • No before/after /usage quota evidence was supplied, so membership-quota consumption cannot be measured from this run.
  • The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
  • The retained `.verify` scripts and captures are model-authored output and were not independently rerun; they do not satisfy the harness's browser, FPS, seam, sustained-drive, or four-corner articulation measurement requirements.
  • No standardized final gameplay capture or blind-evaluation record is supplied.
Visible evidence gaps
  • independent browser acceptance run
  • local FPS, browser, viewport, and hardware receipt
  • multi-minute sustained-drive and generation-seam receipt
  • four-corner suspension-articulation receipt
  • standardized final gameplay capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console