Measured harness ledgerPublic result
Claude Opus 5.5

Space Flight Stage 2 — Landfall — Claude Opus 5.5 Unreported

Extend the Stage 1 space-flight project with a seamless orbit-to-surface round trip, explorable terrain, takeoff, landing, and a fixed character asset using one frozen follow-on request.

Unreported reasoningHeadline result
Workflow cost
$33.62
Wall-clock
Not recorded
Processed tokens
Not recorded
Record state
complete_by_model_internal_verdict_with_input_only_verification_gap
Public summary

Claude Opus 5.5 Unreported complete_by_model_internal_verdict_with_input_only_verification_gap ledger: Not recorded wall-clock, Not recorded, and $33.62 API-equivalent list price, not a marginal subscription cash charge.

Run identity and stack
  • Result ID: space-flight-game-threejs-stage-2-landfall-opus-5.5-high
  • Technical model: Claude Opus 5.5
  • Provider: Anthropic first-party Max subscription
  • Client: Claude Code 2.1.280
  • Stack: Anthropic first-party Max subscription
  • Stack: Claude Code 2.1.280
  • Stack: Technical model/configuration: Claude Opus 5.5
  • Stack: Three.js Stage 2 workflow
  • Stack: Harness v1 Landfall prompt
Recorded caveats
  • The model's three-cycle browser script directly set ship orientation and on-foot camera-forward for aim, while using keyboard movement and action keys. It is not an ordinary-input-only full-loop proof.
  • Independent keyboard play verified launch, ship movement, and help controls, not the complete orbit-to-surface-to-orbit loop. Real mouse capture and physical gamepad were not tested.
  • Final frame timings were captured on a contended local Mac. Earlier runs had slow frames, and the final sampled cadence is not a universal hardware FPS guarantee.
  • An unrelated benchmark briefly replaced the default dev-server port; the affected trace was excluded and final checks ran against an isolated listener without altering model source.
  • The canonical same-session transcript metrics prompt overcounted duplicate rows and mispriced one-hour cache writes. The generation-only Claude CLI receipt and official one-hour rate determine the published API-equivalent estimate; this is not an itemized Max subscription charge.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console