Choose a model. Inspect every run.
Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder adds exact prompts, released projects, the RemakeBench Harness and production skills for the tests in its library.
Order by average cost or average time. Run count stays visible for context.
Download for Google SheetsCSV10 models shown, ordered by average cost. 10 runs match the search.
Every run, cost × time.
Compare individual model runs without flattening them into an average. The lower-left region is lower-cost and faster.
Each dot is one public run. Cost runs left to right; wall-clock time runs bottom to top. Logarithmic spacing keeps the full range legible.
GLM 5.3 Flash — Sekiro Main Menu lacks measured cost or time · GPT-5.6 Luna — Sekiro Main Menu Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · DeepSeek V4 Flash — Sekiro Main Menu lacks measured cost or time · Kimi K3 — Sekiro Main Menu Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · Qwen3.8 Flash — Sekiro Main Menu lacks measured cost or time
