Measured harness ledgerPublic result
Gemini 3.8 FlashDex Cube — 3 Move — Gemini 3.8 Flash High
Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.
High reasoningHeadline result
- Workflow cost
- $11.30
- Wall-clock
- 44m 30s wall-clock
- Processed tokens
- 30.97M processed
- Record state
- judge_fail_independently_replayed
Public summary
Gemini 3.8 Flash High judge_fail_independently_replayed ledger: 44m 30s wall-clock, 30.97M processed, and $11.30 API-equivalent list price, not a marginal subscription cash charge.
Run identity and stack
- Result ID: dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16
- Technical model: Gemini 3.8 Flash
- Provider: Google Antigravity
- Client: Antigravity CLI 1.2.4 (headless print mode)
- Stack: Google Antigravity
- Stack: Antigravity CLI 1.2.4 (headless print mode)
- Stack: Technical model/configuration: Gemini 3.8 Flash
- Stack: MuJoCo dexterous-hand fixture
- Stack: Harness cube-turn stage-2 prompt
- Stack: Requested tool profile: python-mujoco-headless
Cost basis
- actual marginal subscription charge usd: $0.00 USD.
- Introductory 2026 API-equivalent alternative: $5.65.
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16/evidence/model-supplied/ctrl.npy
- SHA-256: 31963cf8a88d5d0fd40f31d93041e6910f3f3970b5249861412bf3e01c35fb53
Validation evidence
- Result: pass
- Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16/validation/judge-replay.public.json
- SHA-256: 73a85124036b651ad6d4b77bdb75438a176df9ddb710e5020afc80010623fbd5
- Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
- The judge measures task completion only; there is no visual quality or blind evaluation for this task.
- This is a judge FAIL with partial credit: movesReached 1 of 3. The log is violation-free, so the failure is the two unattempted moves, not a safety breach.
- Wall-clock includes tools, simulation and judge runs, and idle gaps; it is not model-only compute. The solver used only about 30% of its 150-minute budget and ended on its own end_turn; it was never interrupted, nudged or resumed.
- Cost is API-equivalent on two published bases (introductory 2026 and standard); the subscription's marginal charge was $0 and print mode exposes no quota delta.
- This run used the newer gemini-3.8-flash-high model in headless print mode rather than the operator's frozen Gemini 3.6 Flash TUI stack; disclosed as a stack deviation.
- The conversation id is withheld per precedent; raw receipts, the stream-json transcript, the CLI log and the conversation database are committed by hash only; provenance/run-notes.md and the sim-attempt records are sanitized copies (conversation id and workspace paths replaced).
- This is the first stage-2 run on fixture revision v1.1 (harness sim-attempt archiving only). Verdicts remain comparable with the two v1 stage-2 runs, which have no harness sim-attempt archives; the judge file is identical across the revisions.
- The replay clips are rendered by the operator's viewer from the judge's trajectory exports with a fixed camera; they are presentation evidence, not measurements.
- There is no source/ directory because the solver left no scripts behind (see artifacts.source_note).
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- /usage quota delta (unavailable in print mode)
- solver-authored source files (the solver wrote none to disk; only its transcript holds the inline programs, and that is private)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.