Measured harness ledgerPublic result
Gemini 3.8 Flash

Dex Cube — 3 Move — Gemini 3.8 Flash High

Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.

High reasoningHeadline result
Workflow cost
$11.30
Wall-clock
44m 30s wall-clock
Processed tokens
30.97M processed
Record state
judge_fail_independently_replayed
Public summary

Gemini 3.8 Flash High judge_fail_independently_replayed ledger: 44m 30s wall-clock, 30.97M processed, and $11.30 API-equivalent list price, not a marginal subscription cash charge.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16
  • Technical model: Gemini 3.8 Flash
  • Provider: Google Antigravity
  • Client: Antigravity CLI 1.2.4 (headless print mode)
  • Stack: Google Antigravity
  • Stack: Antigravity CLI 1.2.4 (headless print mode)
  • Stack: Technical model/configuration: Gemini 3.8 Flash
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn stage-2 prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • actual marginal subscription charge usd: $0.00 USD.
  • Introductory 2026 API-equivalent alternative: $5.65.
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: 31963cf8a88d5d0fd40f31d93041e6910f3f3970b5249861412bf3e01c35fb53
Validation evidence
  • Result: pass
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-gemini-3.8-flash-high-antigravity-2026-09-16/validation/judge-replay.public.json
  • SHA-256: 73a85124036b651ad6d4b77bdb75438a176df9ddb710e5020afc80010623fbd5
  • Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • This is a judge FAIL with partial credit: movesReached 1 of 3. The log is violation-free, so the failure is the two unattempted moves, not a safety breach.
  • Wall-clock includes tools, simulation and judge runs, and idle gaps; it is not model-only compute. The solver used only about 30% of its 150-minute budget and ended on its own end_turn; it was never interrupted, nudged or resumed.
  • Cost is API-equivalent on two published bases (introductory 2026 and standard); the subscription's marginal charge was $0 and print mode exposes no quota delta.
  • This run used the newer gemini-3.8-flash-high model in headless print mode rather than the operator's frozen Gemini 3.6 Flash TUI stack; disclosed as a stack deviation.
  • The conversation id is withheld per precedent; raw receipts, the stream-json transcript, the CLI log and the conversation database are committed by hash only; provenance/run-notes.md and the sim-attempt records are sanitized copies (conversation id and workspace paths replaced).
  • This is the first stage-2 run on fixture revision v1.1 (harness sim-attempt archiving only). Verdicts remain comparable with the two v1 stage-2 runs, which have no harness sim-attempt archives; the judge file is identical across the revisions.
  • The replay clips are rendered by the operator's viewer from the judge's trajectory exports with a fixed camera; they are presentation evidence, not measurements.
  • There is no source/ directory because the solver left no scripts behind (see artifacts.source_note).
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • /usage quota delta (unavailable in print mode)
  • solver-authored source files (the solver wrote none to disk; only its transcript holds the inline programs, and that is private)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console