Measured harness ledgerPublic result
GPT-6 Astra

Apple Carry — GPT-6 Astra Max

Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.

Max reasoningHeadline result
Workflow cost
$2.68
Wall-clock
8m 00s wall-clock
Processed tokens
1.61M processed
Record state
judge_pass_independently_replayed
Public summary

GPT-6 Astra Max judge_pass_independently_replayed ledger: 8m 00s wall-clock, 1.61M processed, and $2.68 API-equivalent; subscription charge $0..

Run identity and stack
  • Result ID: apple-carry-fr3-gripper-gpt-6-astra-max-codex-2026-09-21
  • Technical model: GPT-6 Astra
  • Provider: OpenAI Codex
  • Client: codex-sub wrapper (detached, headless, -l dexapple-carry-astra -m gpt-6-astra -e max -s danger-full-access -t 7200)
  • Stack: OpenAI Codex
  • Stack: codex-sub wrapper (detached, headless, -l dexapple-carry-astra -m gpt-6-astra -e max -s danger-full-access -t 7200)
  • Stack: Technical model/configuration: GPT-6 Astra
  • Stack: Embodied manipulation fixture
  • Stack: Harness apple-carry prompt
Validation evidence
  • Path: artifacts/apple-carry-fr3-gripper-gpt-6-astra-max-codex-2026-09-21/validation/judge-replay.public.json
  • SHA-256: d732d54cb5d7b0f565ee8ebf8d6fb9c873f01edb0f1d0bc6e8f7016538443c44
  • Validator SHA-256: e92e63eb3a4c358a423e88d14badd90f1fddbabbf97d0e55857a29d14d939ae9
Recorded caveats
  • This IS a coding-agent cohort run (the first on apple-carry). It is still not comparable with decision-loop rows.
  • Five seeds in one session, not twenty. Wilson interval is wide.
  • $0 charged; $2.68 is API-equivalent at harness Astra rates.
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console