Measured harness ledgerPublic result
Claude Opus 5.5Apple Carry — Claude Opus 5.5 Max
Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.
Max reasoningHeadline result
- Workflow cost
- $2.00
- Wall-clock
- 25m 29s wall-clock
- Processed tokens
- 1.85M processed
- Record state
- judge_pass_independently_replayed
Public summary
Claude Opus 5.5 Max judge_pass_independently_replayed ledger: 25m 29s wall-clock, 1.85M processed, and $2.00 recorded total_cost_usd from workspace FINAL.json (modelUsage.claude-opus-5-5, costBasis=list, Anthropic list price). Charged to the Claude subscription ($0 charged)..
Run identity and stack
- Result ID: apple-carry-fr3-gripper-opus-5.5-max-2026-09-22
- Technical model: Claude Opus 5.5
- Provider: Anthropic Claude Code
- Client: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
- Stack: Anthropic Claude Code
- Stack: Claude CLI 2.1.280, claude -p --model claude-opus-5-5 --effort max --dangerously-skip-permissions --output-format json; 3-at-a-time scheduler; ran to its own end_turn
- Stack: Technical model/configuration: Claude Opus 5.5
- Stack: Embodied manipulation fixture
- Stack: Harness apple-carry prompt
- Stack: Requested tool profile: python-superdex-headless
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/apple-carry-fr3-gripper-opus-5.5-max-2026-09-22/evidence/model-supplied/seed00/robot.glb
- SHA-256: 009696c95e3e3515677b7a65a065c195eb6f34bc72cc453b80c05eb185bbe1a1
Validation evidence
- Result: PASS
- Path: artifacts/apple-carry-fr3-gripper-opus-5.5-max-2026-09-22/validation/judge-replay.public.json
- SHA-256: 859d6a02a553f35d294a3e653c670b9ea1274c2f327734a74cb922a31c0b5a47
Recorded caveats
- This IS a coding-agent cohort run on apple-carry Track A (fr3-gripper). It is not comparable with decision-loop rows.
- Five seeds in one session, not twenty. Wilson interval is wide.
- The published progression clip is seed 0 plus a 0.5 s freeze: it holds fewer segments than the five distinct archived seed logs.
- The compositor burned in 'Final: FAIL' on the progression clip; the judged result is PASS 5/5 stage 3.
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- Cost is the recorded list-price total_cost_usd from workspace FINAL.json (costBasis=list). Charged to the Claude subscription ($0 charged).
- provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
- The raw Claude CLI receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
- A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EU-CZ-1, 105.6 + 65.4 min, $1.30 + $0.81 = $2.11 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
- Nudge disclosure: none.
- The replay clips are presentation evidence, not a measurement, and are not committed.
- Scope note: solver disclosed writing throwaway scripts to the session scratchpad before moving them into the workspace.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.