Research index

Choose a model. Inspect every run.

Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder adds exact prompts, released projects, the RemakeBench Harness and production skills for the tests in its library.

Published model index

Order by average cost or average time. Run count stays visible for context.

Download for Google SheetsCSV
21 models · 274 runs

10 models shown, ordered by average cost. 10 runs match the search.

01DeepSeek V4 FlashOpenCode Go · Provider not separately recorded · opencode-goMaxVision exp24runs$0.2623/24 USD measured2h 3m13/24 measured02Qwen3.8 FlashProvider not separately recordedUnreported11runs$0.4311/11 USD measured0/11 measured03GPT-5.6 LunaOpenAI Codex · Provider not separately recordedMaxxhigh25runs$0.4914/25 USD measured23m 26s14/25 measured04GLM 5.3 FlashOpenCode Go · Provider not separately recorded · Unavailable · ZCode · opencode-goMax28runs$1.5124/28 USD measured1h 1m14/28 measured05Gemini 3.8 FlashGoogle · OpenRouter · UnavailableHighHighest14runs$2.8014/14 USD measured54m 25s14/14 measured06Grok 4.6Grok CLI subscription · xAI Grok Build · xAI via grok-subxhigh21runs$5.4418/21 USD measured29m 31s20/21 measured07Kimi K3Kimi Code · Kimi Code CLI · Kimi Code CLI / Moonshot AI · OpenCode Go / Moonshot AI · openai-compatibleMaxUnreported18runs≥$5.9615/18 USD measured1h 11m18/18 measured08GPT-5.6 SolOpenAI Codex · Provider not separately recordedMaxUltraxhigh26runs$8.4526/26 USD measured39m 15s26/26 measured09Claude Opus 5Anthropic · Anthropic Claude Code · Anthropic Claude Code with Blender MCP · Claude Code · Provider not separately recordedMax19runs≥$80.7619/19 USD measured2h 54m19/19 measured10Claude Fable 5.1Anthropic Claude CodeMax17runs$105.6717/17 USD measured1h 48m17/17 measured
Every public run

Every run, cost × time.

Compare individual model runs without flattening them into an average. The lower-left region is lower-cost and faster.

Run matrix5 measured runs on one field

Each dot is one public run. Cost runs left to right; wall-clock time runs bottom to top. Logarithmic spacing keeps the full range legible.

GLM 5.3 Flash1GPT-5.6 Sol1GPT-5.6 Luna1DeepSeek V4 Flash1Grok 4.61Claude Opus 51Kimi K31Claude Fable 5.11Gemini 3.8 Flash1Qwen3.8 Flash1
5 unplotted

GLM 5.3 FlashSekiro Main Menu lacks measured cost or time · GPT-5.6 LunaSekiro Main Menu Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · DeepSeek V4 FlashSekiro Main Menu lacks measured cost or time · Kimi K3Sekiro Main Menu Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · Qwen3.8 FlashSekiro Main Menu lacks measured cost or time

RemakeBenchResearch console