Measured harness ledgerPublic result
Grok 4.7Apple Carry — Grok 4.7 xhigh
Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.
xhigh reasoningHeadline result
- Workflow cost
- $2.48
- Wall-clock
- 37m 36s wall-clock
- Processed tokens
- 3.07M processed
- Record state
- judge_pass_independently_replayed
Public summary
Grok 4.7 xhigh judge_pass_independently_replayed ledger: 37m 36s wall-clock, 3.07M processed, and $2.48 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: apple-carry-fr3-gripper-grok-4.7-xhigh-2026-09-22
- Technical model: Grok 4.7
- Provider: xAI Grok
- Client: grok-sub wrapper (detached, headless, sandbox off), grok-4.7 at xhigh; ran to its own end_turn
- Stack: xAI Grok
- Stack: grok-sub wrapper (detached, headless, sandbox off), grok-4.7 at xhigh; ran to its own end_turn
- Stack: Technical model/configuration: Grok 4.7
- Stack: Embodied manipulation fixture
- Stack: Harness apple-carry prompt
- Stack: Requested tool profile: python-superdex-headless
Cost basis
- usage cost reported by xAI for the session (grok-sub receipt line, confirmed by the wrapper run's meta.json)
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/apple-carry-fr3-gripper-grok-4.7-xhigh-2026-09-22/evidence/model-supplied/seed00/robot.glb
- SHA-256: 009696c95e3e3515677b7a65a065c195eb6f34bc72cc453b80c05eb185bbe1a1
Validation evidence
- Result: PASS
- Path: artifacts/apple-carry-fr3-gripper-grok-4.7-xhigh-2026-09-22/validation/judge-replay.public.json
- SHA-256: b384f0a84d7ac7e423f1ac929794575b762c7efee1ad81d9ca0ee324e318aca4
Recorded caveats
- This IS a coding-agent cohort run on apple-carry Track A (fr3-gripper). It is not comparable with decision-loop rows.
- Five seeds in one session, not twenty. Wilson interval is wide.
- The published progression clip is seed 0 plus a 0.5 s freeze: it holds fewer segments than the five distinct archived seed logs.
- The compositor burned in 'Final: FAIL' on the progression clip; the judged result is PASS 5/5 stage 3.
- A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EUR-IS-2, 52 + 154 min, $0.64 + $1.90 + $0.25 aborted batch = $2.79 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
- The replay clips are presentation evidence, not a measurement, and are not committed.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs.
- Cost is the xAI-reported session total from the grok-sub receipt line.
- provenance/orchestrator-results.md is the task project's own result section copied with absolute local paths replaced.
- The raw grok-sub receipts carry the session id, the reasoning transcript and local paths; they are not committed.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.