Measured harness ledgerPublic result
Gemini 3.8 FlashApple Stem — Gemini 3.8 Flash High
Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.
High reasoningHeadline result
- Workflow cost
- $5.39
- Wall-clock
- 69m 1s wall-clock
- Processed tokens
- 20.97M processed
- Record state
- judge_fail_independently_replayed
Public summary
Gemini 3.8 Flash High judge_fail_independently_replayed ledger: 69m 1s wall-clock, 20.97M processed, and $5.39 API-equivalent list price, not a marginal subscription cash charge.
Run identity and stack
- Result ID: dex-apple-stem-superdex-gemini-3.8-flash-high-antigravity-2026-09-16
- Technical model: Gemini 3.8 Flash
- Provider: Google Antigravity
- Client: Antigravity CLI (headless print mode)
- Stack: Google Antigravity
- Stack: Antigravity CLI (headless print mode)
- Stack: Technical model/configuration: Gemini 3.8 Flash
- Stack: SuperDex fingertip fixture
- Stack: Harness apple-stem prompt
- Stack: Requested tool profile: python-superdex-headless
Cost basis
- actual marginal subscription charge usd: $0.00 USD.
- Introductory 2026 API-equivalent alternative: $2.69.
Primary artifact integrity
- Kind: open-loop-200hz-joint-target-log
- Path: artifacts/dex-apple-stem-superdex-gemini-3.8-flash-high-antigravity-2026-09-16/evidence/model-supplied/ctrl.npy
- SHA-256: cb7c4ce43bd126eb69f22915b3a458c000aa9f5bfe357dbd951d574d10d2a740
Validation evidence
- Result: pass
- Path: artifacts/dex-apple-stem-superdex-gemini-3.8-flash-high-antigravity-2026-09-16/validation/judge-replay.public.json
- SHA-256: 84fbdaace1bc44f35e360138e4877be9132b8cdd8d46ad2103ca512add33c2da
- Validator SHA-256: 8d6c638716c7c28701a7369792ab05ddadc72ecc4ad2b6183dfb04b4122b4647
Recorded caveats
- The judge measures task completion only (lift height, hold, stem-only contact, runtime limits); there is no visual quality or blind evaluation for this task.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs and idle gaps; it is not model-only compute.
- The replay and progression clips are rendered by the operator's viewer from the judge's trajectory export with a fixed camera; they are presentation evidence, not a measurement.
- Solver-authored source is published byte for byte as the run left it. Several scripts contain the workspace's absolute path in a literal string; that is model output, not an operator receipt, and it is not rewritten. Operator provenance, validation and ledger files carry no local paths.
- Two dex-cube tier-2 runs shared the same Mac for the whole window. SuperDex simulation dominates per-iteration cost here, so every wall clock in this cohort carries CPU contention and plausibly cost the slower models real iterations.
- Cost is API-equivalent on two published bases (2026 introductory and standard); the Antigravity subscription's marginal charge was $0 and print mode exposes no quota delta.
- This run used the newer gemini-3.8-flash-high model in headless print mode rather than the operator's frozen Gemini 3.6 Flash TUI stack; disclosed as a stack deviation.
- The conversation id is withheld per precedent; the conversation database, the stream-json transcript and the CLI log are committed by hash only; provenance/run-notes.md and the sim-attempt records are sanitized copies.
- This is one of the first four runs on fixture revision v1.1. Verdicts stay comparable with the three v1 runs; only archive availability and the fixture-identity scene_sha256 differ.
- Three of the nineteen sim-attempt records are the judge's own re-simulations, not solver simulate() calls, because this task's validate.py does not set RB_NO_ARCHIVE=1 for itself. Counts that matter are stated separately in attempt_history.sim_attempts.
- The failure is one criterion away from a pass: the apple is lifted and held by the stem with no body contact, but the thumb touches r_thumb_distal/r_thumb_middle instead of r_thumb_pad, so only one distinct pad registers.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- /usage quota delta (unavailable in print mode)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.