Measured harness ledgerPublic result
Kimi K3

Neon Gunner — Kimi K3 Max

Build a playable destructible neon run-and-gun game in the browser under the pinned Harness task contract.

Max reasoningHeadline result
Workflow cost
Not recorded
Wall-clock
5378.1s wall-clock
Processed tokens
Not recorded
Record state
partial_quota_interrupted_model_runtime_verified
Public summary

Kimi K3 Max partial_quota_interrupted_model_runtime_verified ledger: 5378.1s wall-clock, Not recorded, and Not recorded.

Run identity and stack
  • Result ID: neon-gunner-canvas-kimi-k3-max
  • Technical model: k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: k3
  • Stack: Browser canvas game
  • Stack: Harness v1 Neon Gunner prompt
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/neon-gunner-canvas-kimi-k3-max/source/index.html
  • SHA-256: e986e047ba50dee502ae9c9cb32a76c5760b28c426dc66321d56d366bbf4109d
Validation evidence
  • Result: RUNTIME_PASS_MODEL_QUOTA_INTERRUPTED
  • Path: artifacts/neon-gunner-canvas-kimi-k3-max/verification/summary.public.json
  • SHA-256: c263162193013d25968415fe1b7e5a195132c7576a214944a6eb99f1fe7571dd
  • Validator SHA-256: fa5ec13e9ef2f5f96949283077b7c6278bf9c9f14032817c21743ab973d06090
Recorded caveats
  • Kimi hit quota after passing the fixture and before final polish, its README, and final completion response.
  • 29 usage records reconcile with 29 completed steps; the 30th attempted request has no successful usage record.
  • Operator independently completed gameplay, destruction and FPS checks without modifying generated runtime source.
  • The successful real-input run took about 53 seconds, shorter than the requested 75–90 seconds.
  • Operator polling continued after the quota event; that idle time is excluded from the 5378.123-second work duration.
  • No blind evaluation, independent audio listening, or current API-equivalent cost is claimed.
Visible evidence gaps
  • model normal completion / final verification response
  • blind-evaluation record
  • verified API-equivalent cost
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console