Measured harness ledgerPublic result
Kimi K3Space Flight Stage 2 — Landfall — Kimi K3 Max
Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.
Max reasoningHeadline result
- Workflow cost
- $15.58
- Wall-clock
- 2h 23m 18.1s wall-clock
- Processed tokens
- 39.99M processed
- Record state
- partial_token_timing_source_build_ledger
Public summary
Kimi K3 Max partial_token_timing_source_build_ledger ledger: 2h 23m 18.1s wall-clock, 39.99M processed, and $15.58 API-equivalent usage accounting, not an itemized subscription cash charge.
Run identity and stack
- Result ID: space-flight-game-threejs-stage-2-landfall-kimi-k3-max
- Technical model: kimi-k3
- Provider: Kimi Code CLI / Moonshot AI
- Stack: Kimi Code CLI / Moonshot AI
- Stack: Technical model/configuration: kimi-k3
- Stack: Three.js Stage 2 workflow
- Stack: Harness v1 Landfall prompt
Primary artifact integrity
- Kind: vite-threejs-stage-2-landfall-source-and-production-build-entry-point
- Path: artifacts/space-flight-game-threejs-stage-2-landfall-kimi-k3-max/source/src/main.js
- SHA-256: cc8dcf91a4d98d2f7885d4839254bf9234775b74b4f7d35199ca5002337f064d
Recorded caveats
- Wall-clock is end-to-end workflow latency including tool execution, installs, browser checks, and idle time; it is not model-only compute.
- One user prompt produced 202 billable Kimi model requests.
- Kimi's observed wire format exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
- Cache-read tokens are discounted, so total processed tokens overstate cost.
- The source receipt reports no before/after /usage quota evidence; it cannot quantify the subscription quota consumed by this run.
- The private Kimi session ZIP is intentionally omitted because it contains prompts, hidden reasoning, local paths, and tool output.
- The archive build passes, but no independent browser acceptance run, local FPS/hardware measurement, final capture, or blind evaluation is archived.
Visible evidence gaps
- independent uninterrupted landing-to-orbit acceptance run without test fallback
- local FPS and hardware receipt, including transition hitch measurement
- final capture with viewport/browser evidence
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
