Measured harness ledgerPublic result
Qwen3.8 Flash

Apple Stem — Qwen3.8 Flash Highest

Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.

Highest reasoningHeadline result
Workflow cost
$0.24
Wall-clock
88m 21.1s wall-clock
Processed tokens
8.71M processed
Record state
failed_no_submission
Public summary

Qwen3.8 Flash Highest failed_no_submission ledger: 88m 21.1s wall-clock, 8.71M processed, and $0.24 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-apple-stem-superdex-qwen3.8-flash-max-opencode-openrouter-2026-09-16
  • Technical model: Qwen 3.8 Flash
  • Provider: OpenRouter
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenRouter
  • Stack: OpenCode 1.18.23
  • Stack: OpenCode → OpenRouter
  • Stack: Technical model/configuration: Qwen 3.8 Flash
  • Stack: Requested variant: max
  • Stack: SuperDex fingertip fixture
  • Stack: Harness apple-stem prompt
  • Stack: Requested tool profile: python-superdex-headless
Cost basis
  • actual marginal cash charged usd: $0.00 USD.
Validation evidence
  • Result: pass
Recorded caveats
  • This run produced no submission, so it has no verdict, no stage, no replay artefacts and no clips. Nothing was reconstructed after the fact.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs and idle gaps; it is not model-only compute.
  • Solver-authored source is published byte for byte as the run left it. Several scripts contain the workspace's absolute path in a literal string; that is model output, not an operator receipt, and it is not rewritten. Operator provenance, validation and ledger files carry no local paths.
  • Two dex-cube tier-2 runs shared the same Mac for the whole window. SuperDex simulation dominates per-iteration cost here, so every wall clock in this cohort carries CPU contention and plausibly cost the slower models real iterations.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • The 5400 s launcher alarm killed the run mid-request, so the metrics receipt is PARTIAL: the final root request finished as `tool-calls` rather than `stop` and the session ledger records variant `default` where `--variant max` was requested. Every token component reconciles with 0 mismatches; nothing was relabelled.
  • OpenRouter billed $0 (BYOK Alibaba key); the OpenCode-recorded $0.2397 is a catalog-rate estimate. The receipt's own billing block wrongly equates it with marginal cash because the extractor assumes pay-as-you-go credits.
  • No agent-written receipt bundle exists: the metrics turn aborted once with finish reason `length` (the invocation omitted --variant max, so the workspace opencode.json override did not apply) and the relaunch ran 40 minutes without emitting a stream event before its own alarm killed it. metrics/receipt.json comes from the canonical extractor run directly against the marker. Both failed metrics turns are archived privately; the counted generation run is one clean session and was never relaunched.
  • The pricing snapshot is the OpenCode catalogue entry carried over from the same day's cube run for the identical model, provider and workspace override; it reproduces this run's recorded total exactly.
  • ctrl_v6.npy has an mtime of 2026-09-16T21:56:02Z, after the 21:53:55Z alarm: a python process the agent had already started finished writing after the agent was killed. No model request exists after the alarm.
  • This is one of the first four runs on fixture revision v1.1; the revision is recorded from the workspace inputs because this run has no verdict.
  • Reconstructed attempts are the solver's own saved arrays judged after the fact by the archive operator. They are unscored evidence of progression: they do not change this run's verdict, stage or cohort standing and are excluded from every comparison.
Visible evidence gaps
  • a submission of any kind (the run produced none)
  • a judge verdict, replay artefacts and an evidence clip (all downstream of the missing submission)
  • the model's own NOTES.md (never written)
  • an agent-written metrics receipt bundle (only the canonical extractor output exists)
  • hardware and interpreter receipt from the solver session
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console