Model Leaderboard

Which DeepSWE model gives you the most performance per dollar — and where you are overpaying for the same score.

Best Models

The most cost-efficient model for each performance score, from all available models.

  1. claude-opus-5high
    72.8%±1.9
    $0.08/pt
  2. gpt-5-6-solhigh
    69.4%±1.4
    $0.05/pt
  3. gpt-5-6-lunamax
    67.2%±4.0
    $0.05/pt
  4. gemini-3-7-flashmedium
    65.5%±3.1
    $0.03/pt
  5. deepseek-v4-promax
    62.8%±6.3
    0.38¢/pt
  6. deepseek-v4-flashmax
    53.3%±3.6
    0.19¢/pt
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty
How it works

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Charts

Compare DeepSWE pass rate with mean cost. The top-right corner is the sweet spot: higher performance for less money.

15 selected models
Providers

Chart loads in the browser.

Comparison Table

15 selected models · sort by performance, price, or speed

Sort by
  • claude-opus-5max· Anthropic
    73.6%±3.9
    $0.16/pt
    claude-opus-5high
    72.8%±1.9
    $0.08/pt
    49% cheaper0.8% worse
  • gpt-5-6-solmax· OpenAI
    72.7%±2.8
    $0.12/pt
    claude-opus-5high
    72.8%±1.9
    $0.08/pt
    28% cheaper0.2% better
  • claude-fable-5xhigh· Anthropic
    69.9%±3.2
    $0.19/pt
    gpt-5-6-solhigh
    69.4%±1.4
    $0.05/pt
    74% cheaper0.5% worse
  • gpt-5-6-terramax· OpenAI
    69.6%±2.6
    $0.07/pt
    gpt-5-6-solhigh
    69.4%±1.4
    $0.05/pt
    30% cheaper0.2% worse
  • glm-5-3max· Zhipu
    69.0%±3.0
    $0.06/pt
    gpt-5-6-lunamax
    67.2%±4.0
    $0.05/pt
    24% cheaper1.8% worse
  • kimi-k3max· Moonshot
    68.5%±4.5
    $0.07/pt
    gpt-5-6-lunamax
    67.2%±4.0
    $0.05/pt
    35% cheaper1.3% worse
  • grok-4-6medium· xAI
    67.5%±2.3
    $0.05/pt
    gpt-5-6-lunamax
    67.2%±4.0
    $0.05/pt
    12% cheaper0.3% worse
  • gpt-5-6-lunamax· OpenAI
    67.2%±4.0
    $0.05/pt
    Best in class
  • gemini-3-7-flashmedium· Google
    65.5%±3.1
    $0.03/pt
    Best in class
  • deepseek-v4-promax· DeepSeek
    62.8%±6.3
    0.38¢/pt
    Best in class
  • qwen3-8-maxxhigh· Alibaba
    57.5%±2.7
    $0.06/pt
    deepseek-v4-promax
    62.8%±6.3
    0.38¢/pt
    94% cheaper5.4% better
  • muse-spark-1-2xhigh· Meta
    54.9%±2.1
    $0.07/pt
    deepseek-v4-promax
    62.8%±6.3
    0.38¢/pt
    93% cheaper8.0% better
  • claude-sonnet-5max· Anthropic
    53.8%±4.2
    $0.49/pt
    deepseek-v4-flashmax
    53.3%±3.6
    0.19¢/pt
    100% cheaper0.5% worse
  • deepseek-v4-flashmax· DeepSeek
    53.3%±3.6
    0.19¢/pt
    Best in class
  • kimi-k2-7-codedefault· Moonshot
    30.5%±0.5
    $0.09/pt
    deepseek-v4-flashmax
    53.3%±3.6
    0.19¢/pt
    96% cheaper22.8% better
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty
How it works

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Model selector

Choose the model and reasoning effort rows to compare.

15 of 44 selected
ProvidersAll providers shown

Alibaba

qwen3-8-max

Alibaba

Anthropic

claude-opus-5

Anthropic

claude-fable-5

Anthropic

claude-sonnet-5

Anthropic

DeepSeek

deepseek-v4-pro

DeepSeek

deepseek-v4-flash

DeepSeek

Google

gemini-3-7-flash

Google

Meta

muse-spark-1-2

Meta

Moonshot

kimi-k3

Moonshot

kimi-k2-7-code

Moonshot

OpenAI

gpt-5-6-sol

OpenAI

gpt-5-6-terra

OpenAI

gpt-5-6-luna

OpenAI

xAI

grok-4-6

xAI

Zhipu

glm-5-3

Zhipu