Choose a model. Inspect every run.
Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder adds exact prompts, released projects, the RemakeBench Harness and production skills for the tests in its library.
Order by average cost or average time. Run count stays visible for context.
Download for Google SheetsCSV18 models shown, ordered by average cost. 23 runs match the search.
Every run, cost × time.
Compare individual model runs without flattening them into an average. The lower-left region is lower-cost and faster.
Each dot is one public run. Cost runs left to right; wall-clock time runs bottom to top. Logarithmic spacing keeps the full range legible.
GLM 5.3 Flash — Infinite cathedral corridor shader lacks measured cost or time · GPT-5.6 Luna — Infinite cathedral corridor shader Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · DeepSeek V4 Flash — Infinite cathedral corridor shader lacks measured cost or time · Qwen3.8 Flash — Infinite cathedral corridor shader lacks measured cost or time · Qwen3.8 Max — Infinite cathedral corridor shader Non-USD cost; excluded from USD averages and the cost × time plot.
