Measured harness ledgerPublic result
Grok 4.6Apple Stem — Grok 4.6 xhigh
Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.
xhigh reasoningHeadline result
- Workflow cost
- Not recorded
- Wall-clock
- 1h 30m 00s wrapper wall (submission written at 51m 21s) wall-clock
- Processed tokens
- Not recorded
- Record state
- judge_pass_independently_replayed
Public summary
Grok 4.6 xhigh judge_pass_independently_replayed ledger: 1h 30m 00s wrapper wall (submission written at 51m 21s) wall-clock, Not recorded, and Not recorded.
Run identity and stack
- Result ID: dex-apple-stem-superdex-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok
- Client: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
- Stack: xAI Grok
- Stack: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
- Stack: Technical model/configuration: Grok 4.6
- Stack: SuperDex fingertip fixture
- Stack: Harness apple-stem prompt
- Stack: Requested tool profile: python-superdex-headless
Cost basis
- No measured or estimated run cost is recorded.
Primary artifact integrity
- Kind: open-loop-200hz-joint-target-log
- Path: artifacts/dex-apple-stem-superdex-grok-4.6-xhigh/evidence/model-supplied/ctrl.npy
- SHA-256: 83d1736388b28ed34a5220add4ff583f4713cc733258489bf4558f9c70ee4eae
Validation evidence
- Result: PASS
- Path: artifacts/dex-apple-stem-superdex-grok-4.6-xhigh/validation/judge-replay.public.json
- SHA-256: 0a4285cd2a70af0b9b548c33108b19bd981b0da620e15effb82bc0e23637e127
- Validator SHA-256: 8d6c638716c7c28701a7369792ab05ddadc72ecc4ad2b6183dfb04b4122b4647
Recorded caveats
- The judge measures task completion only (lift height, hold, stem-only contact, runtime limits); there is no visual quality or blind evaluation for this task.
- The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
- provenance/gold-results.md is the orchestrator's own result file copied verbatim.
- The solver did not see a verdict for its final log: its self-check was killed by the wrapper wall. The archived verdict.json is the operator's canonical judge run, independently reproduced by the archive operator.
- The wrapper reported $0 on its timeout path and the local session record carries no usage fields, so the xAI cost is unrecorded (not zero); comparable runs suggest $3-5; the authoritative figure is in the xAI console.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- Grok token counts are not available from the grok-sub receipt; only the xAI usage dollars, elapsed seconds and turn count were recorded.
- The Grok workspace carried the pre-Y-up-fix export_scene.py, so its archived robot.glb was written by that older exporter; trajectory bins and verdict are unaffected.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- token-count receipt (xAI usage dollars only)
- xAI usage cost for the session (wrapper timeout path recorded $0)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.