Measured harness ledgerPublic result
Claude Opus 5.5Space Flight Stage 2 — Landfall — Claude Opus 5.5 Unreported
Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.
Unreported reasoningHeadline result
- Workflow cost
- $33.62
- Wall-clock
- Not recorded
- Processed tokens
- Not recorded
- Record state
- complete_by_model_internal_verdict_with_input_only_verification_gap
Public summary
Claude Opus 5.5 Unreported complete_by_model_internal_verdict_with_input_only_verification_gap ledger: Not recorded wall-clock, Not recorded, and $33.62 API-equivalent list price, not a marginal subscription cash charge.
Run identity and stack
- Result ID: space-flight-game-threejs-stage-2-landfall-opus-5.5-high
- Technical model: Claude Opus 5.5
- Provider: Anthropic first-party Max subscription
- Client: Claude Code 2.1.280
- Stack: Anthropic first-party Max subscription
- Stack: Claude Code 2.1.280
- Stack: Technical model/configuration: Claude Opus 5.5
- Stack: Three.js Stage 2 workflow
- Stack: Harness v1 Landfall prompt
Recorded caveats
- The model's three-cycle browser script directly set ship orientation and on-foot camera-forward for aim, while using keyboard movement and action keys. It is not an ordinary-input-only full-loop proof.
- Independent keyboard play verified launch, ship movement, and help controls, not the complete orbit-to-surface-to-orbit loop. Real mouse capture and physical gamepad were not tested.
- Final frame timings were captured on a contended local Mac. Earlier runs had slow frames, and the final sampled cadence is not a universal hardware FPS guarantee.
- An unrelated benchmark briefly replaced the default dev-server port; the affected trace was excluded and final checks ran against an isolated listener without altering model source.
- The canonical same-session transcript metrics prompt overcounted duplicate rows and mispriced one-hour cache writes. The generation-only Claude CLI receipt and official one-hour rate determine the published API-equivalent estimate; this is not an itemized Max subscription charge.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.