Measured harness ledgerPublic result
Kimi K3

Space Flight Stage 2 — Landfall — Kimi K3 Max

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

Max reasoningHeadline result
Workflow cost
$15.58
Wall-clock
2h 23m 18.1s wall-clock
Processed tokens
39.99M processed
Record state
partial_token_timing_source_build_ledger
Public summary

Kimi K3 Max partial_token_timing_source_build_ledger ledger: 2h 23m 18.1s wall-clock, 39.99M processed, and $15.58 API-equivalent usage accounting, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-kimi-k3-max
  • Technical model: kimi-k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: kimi-k3
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Primary artifact integrity
  • Kind: vite-threejs-stage-2-landfall-source-and-production-build-entry-point
  • Path: artifacts/space-flight-game-threejs-stage-2-landfall-kimi-k3-max/source/src/main.js
  • SHA-256: cc8dcf91a4d98d2f7885d4839254bf9234775b74b4f7d35199ca5002337f064d
Recorded caveats
  • Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
  • One user prompt produced 202 billable Kimi model requests.
  • Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • The source receipt reports no before/after /usage quota evidence; it cannot quantify the subscription quota consumed by this run.
  • The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
  • The archive build passes, but no independent browser acceptance run, local FPS/hardware measurement, final capture, or blind evaluation is archived.
Visible evidence gaps
  • independent uninterrupted landing-to-orbit acceptance run without test fallback
  • local FPS and hardware receipt, including transition hitch measurement
  • final capture with viewport/browser evidence
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console