The top DeepSWE scores are close enough that cost changes the decision. GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 sit within half a percentage point in this snapshot. Their average task costs range from $2.36 to $11.84.
I’m keeping this page as a running comparison of frontier coding models. I want to see what a model solves, what the run costs, and whether the result holds up in the agent I would use.
Last source check . Datacurve’s leaderboard is dated . This is a curated snapshot.
DeepSWE at a glance
DeepSWE v1.1 has 113 tasks across 91 repositories and five languages. Datacurve writes original engineering tasks with functional and regression tests. The leaderboard uses mini-swe-agent across models, which makes it useful for comparing models under a shared harness.
Pass@1 measures success on one attempt per task, averaged over the evaluation. Average cost per task includes unsuccessful attempts. A model that spends more time on a failing task can still leave you with a larger bill.
COST / CAPABILITY
DeepSWE v1.1
Official runs · 113 tasks · mini-swe-agent · checked
All 70 published configurations across 28 official models, plus self-reported Muse Spark 1.3, Fable 5.1, and DeepSeek V4.1 Flash results.
SCORE ONLY · COST NOT REPORTED
Cost per DeepSWE task is unknown. Select an entry for its source and evaluation details. It has no position on the cost axis.
Hover, tap, or tab to a point. On narrow screens, scroll the chart horizontally or inspect a model below. With more than six models visible, labels identify frontier models, self-reported results, and the inspected model. Filter to six or fewer to label every model.
GPT-6 Astra [xhigh]
Official leaderboard result · On this selection’s official frontier
Harness · mini-swe-agent
- Pass@1
- 74.1% ±3 pp
- Cost / task
- $6.52
- Output tokens
- 30k
- Agent steps
- 29
Official frontier members in this selection are GPT-5.6 Luna [low], GPT-5.6 Luna [medium], GPT-5.6 Luna [high], GLM-5.3 Flash [max], GPT-5.6 Luna [max], Gemini 3.8 Flash [medium], Gemini 3.8 Flash [high], GPT-6 Astra [xhigh]. Pending results never change that frontier. Solid lines join published reasoning levels of the same model and evidence type. Values between points are not measurements.
Configuration table 76 results
| Model / effort | Pass@1 | USD / task | Out tokens | Steps | Evidence |
|---|---|---|---|---|---|
| GPT-6 Astra[xhigh] | 74.1%±3 pp | $6.52 | 30k | 29 | Official |
| Gemini 3.8 Flash[high] | 73.8%±1 pp | $2.36 | 143k | 166 | Official |
| Claude Opus 5[max] | 73.6%±4 pp | $11.84 | 118k | 99 | Official |
| GPT-6 Astra[high] | 73.2%±3 pp | $5.72 | 27k | 27 | Official |
| GPT-6 Astra[max] | 73.2%±1 pp | $12.37 | 61k | 28 | Official |
| Claude Opus 5[xhigh] | 73.2%±3 pp | $9.07 | 92k | 89 | Official |
| Claude Opus 5[high] | 72.8%±2 pp | $6.08 | 64k | 73 | Official |
| GPT-6 Astra[medium] | 72.8%±3 pp | $4.38 | 20k | 26 | Official |
| GPT-5.6 Sol[max] | 72.7%±3 pp | $6.46 | 60k | 61 | Official |
| Gemini 3.8 Flash[medium] | 71.0%±2 pp | $1.97 | 125k | 147 | Official |
| GPT-5.6 Sol[xhigh] | 70.7%±1 pp | $3.60 | 41k | 44 | Official |
| Claude Fable 5[xhigh] | 69.9%±3 pp | $13.41 | 80k | 68 | Official |
| Claude Fable 5[max] | 69.7%±4 pp | $21.63 | 119k | 88 | Official |
| GPT-5.6 Terra[max] | 69.6%±3 pp | $3.96 | 72k | 76 | Official |
| GPT-5.6 Sol[high] | 69.4%±1 pp | $2.66 | 28k | 37 | Official |
| GLM-5.3[max] | 69.0%±3 pp | $3.99 | 80k | 124 | Official |
| Claude Opus 5[medium] | 68.9%±1 pp | $3.29 | 37k | 52 | Official |
| Claude Fable 5[high] | 68.6%±1 pp | $9.18 | 57k | 59 | Official |
| Kimi K3[max] | 68.5%±5 pp | $4.65 | 81k | 98 | Official |
| Grok 4.6[medium] | 67.5%±2 pp | $3.45 | 50k | 70 | Official |
| GPT-5.6 Luna[max] | 67.2%±4 pp | $0.61 | 73k | 102 | Official |
| GPT-5.5[xhigh] | 67.0%±6 pp | $7.23 | 46k | 82 | Official |
| GPT-6 Astra[low] | 67.0%±1 pp | $2.19 | 11k | 20 | Official |
| Grok 4.6[xhigh] | 66.7%±2 pp | $5.50 | 71k | 87 | Official |
| Gemini 3.7 Flash[medium] | 65.5%±3 pp | $2.03 | 94k | 117 | Official |
| Claude Fable 5[medium] | 65.4%±4 pp | $6.09 | 40k | 48 | Official |
| Gemini 3.7 Flash[high] | 65.3%±2 pp | $2.18 | 107k | 125 | Official |
| Grok 4.6[high] | 65.2%±2 pp | $4.38 | 61k | 79 | Official |
| GPT-5.5[high] | 64.4%±3 pp | $5.10 | 31k | 62 | Official |
| GLM-5.3 Flash[max] | 63.4%±4 pp | $0.24 | 73k | 123 | Official |
| DeepSeek V4 Pro[max] | 62.8%±6 pp | $1.67 | 106k | 155 | Official |
| GPT-5.6 Sol[medium] | 61.1%±2 pp | $1.42 | 18k | 31 | Official |
| GPT-5.6 Terra[xhigh] | 60.2%±2 pp | $1.70 | 40k | 43 | Official |
| Claude Fable 5[low] | 59.6%±3 pp | $3.76 | 25k | 38 | Official |
| Claude Opus 4.8[max] | 59.0%±2 pp | $13.22 | 135k | 120 | Official |
| Claude Opus 5[low] | 58.1%±2 pp | $1.66 | 20k | 36 | Official |
| Qwen3.8 Max[xhigh] | 57.5%±3 pp | $3.73 | 95k | 111 | Official |
| GPT-5.6 Luna[xhigh] | 56.9%±2 pp | $0.31 | 45k | 71 | Official |
| Muse Spark 1.2[xhigh] | 54.9%±2 pp | $3.70 | 99k | 101 | Official |
| Claude Opus 4.8[xhigh] | 54.4%±4 pp | $8.01 | 86k | 95 | Official |
| GPT-5.5[medium] | 54.0%±3 pp | $2.75 | 20k | 46 | Official |
| Claude Sonnet 5[max] | 53.8%±4 pp | $26.40 | 214k | 268 | Official |
| Gemini 3.7 Flash[low] | 53.8%±3 pp | $1.83 | 73k | 130 | Official |
| GPT-5.6 Terra[high] | 53.8%±4 pp | $0.91 | 22k | 34 | Official |
| Grok 4.5[high] | 53.8%±2 pp | $2.42 | 36k | 61 | Official |
| DeepSeek V4 Flash[max] | 53.3%±4 pp | $0.46 | 108k | 153 | Official |
| Muse Spark 1.1[xhigh] | 53.3%±3 pp | $2.36 | 74k | 96 | Official |
| Claude Opus 4.8[high] | 51.8%±5 pp | $4.28 | 50k | 73 | Official |
| GPT-5.4[xhigh] | 51.8%±2 pp | $5.65 | 71k | 70 | Official |
| Claude Sonnet 5[xhigh] | 49.7%±3 pp | $11.89 | 121k | 186 | Official |
| Claude Opus 4.8[medium] | 48.7%±2 pp | $3.44 | 41k | 66 | Official |
| Claude Sonnet 5[high] | 48.2%±5 pp | $7.43 | 87k | 147 | Official |
| Gemini 3.6 Flash[high] | 46.7%±4 pp | $2.21 | 96k | 117 | Official |
| GPT-5.6 Sol[low] | 45.4%±2 pp | $0.82 | 11k | 23 | Official |
| GPT-5.6 Luna[high] | 44.2%±3 pp | $0.16 | 26k | 49 | Official |
| GLM-5.2[max] | 43.8%±2 pp | $3.92 | 78k | 129 | Official |
| Grok 4.6[low] | 41.6%±2 pp | $1.04 | 16k | 44 | Official |
| Claude Opus 4.8[low] | 40.8%±1 pp | $2.29 | 29k | 54 | Official |
| Claude Sonnet 5[medium] | 39.8%±3 pp | $4.08 | 57k | 108 | Official |
| GLM-5.2[high] | 36.3%±5 pp | $2.84 | 54k | 122 | Official |
| Gemini 3.5 Flash[high] | 36.1%±4 pp | $3.45 | 76k | 105 | Official |
| GPT-5.6 Terra[medium] | 35.1%±3 pp | $0.47 | 12k | 25 | Official |
| Kimi K2.7 Code[default] | 30.5%±1 pp | $2.82 | 59k | 149 | Official |
| Claude Sonnet 5[low] | 30.5%±1 pp | $2.19 | 36k | 77 | Official |
| Claude Sonnet 4.6[high] | 29.9%±4 pp | $5.52 | 76k | 134 | Official |
| GPT-5.5[low] | 27.0%±2 pp | $1.20 | 9k | 28 | Official |
| GPT-5.6 Terra[low] | 24.1%±1 pp | $0.34 | 9k | 21 | Official |
| Gemini 3.1 Pro Preview[high] | 11.7%±1 pp | $2.14 | 28k | 76 | Official |
| GPT-5.6 Luna[medium] | 11.3%±1 pp | $0.04 | 8k | 24 | Official |
| GPT-5.6 Luna[low] | 1.5%±1 pp | $0.01 | 3k | 12 | Official |
| Muse Spark 1.3 (Standard)[max] | 75.4%Range not reported | ~$2.75estimate | ~74kestimated | ~80estimated | Self-reportedPending review |
| Muse Spark 1.3 (Contributor)[max] | 75.4%Range not reported | ~$0.25estimate | ~74kestimated | ~80estimated | Self-reportedPending review |
| Claude Fable 5.1[max] | 67.4%Range not reported | Not reported | Not reported | Not reported | Self-reportedPending review |
| DeepSeek V4.1 Flash[low] | 66.0%Range not reported | ~$0.19estimate | ~38kestimated | ~95estimated | Self-reportedPending review |
| DeepSeek V4.1 Flash[high] | 72.0%Range not reported | ~$0.30estimate | ~60kestimated | ~120estimated | Self-reportedPending review |
| DeepSeek V4.1 Flash[max] | 74.2%Range not reported | ~$0.48estimate | ~95kestimated | ~140estimated | Self-reportedPending review |
The comparison includes all 70 configurations across the 28 models in Datacurve’s current v1.1 dataset, plus the Muse Spark 1.3, Fable 5.1, and DeepSeek V4.1 Flash claims. Muse has separate Standard and Contributor entries, each with its own model filter. Use the searchable model dropdown on the chart to choose a comparison, and the Options dropdown beside it for the scale, frontier, uncertainty, and self-reported toggles. Fable 5 and Luna are included at every published reasoning level. DeepSeek V4.1 Flash connects low, high, and max with the standard reasoning-series line; its hollow markers identify the self-reported points. Fable 5.1 appears in the score-only panel and configuration table because its DeepSWE cost is unknown.
The upper-right is the useful direction on this chart, as on Datacurve. Cost decreases toward the right. The default logarithmic cost scale keeps the roughly $0.01 to $26.40 range readable. Switch it off for a linear scale. The axes rescale to the visible models, so a narrow comparison uses the full plot area. Each solid curve connects the published reasoning levels of one model. Hover or select any point for that configuration’s score and cost. Small comparisons label every model. Larger comparisons label the frontier models, self-reported results, and whichever model you inspect.
A configuration is on the Pareto frontier when no other included configuration is at least as good on both cost and score and strictly better on one. Turn on “Show official frontier” to see it as a dashed line. It is computed from the visible official configurations and updates with the model filters. With all models selected, it uses the full 70-configuration snapshot. A configuration below this line could still be useful for a different workload.
What I take from the current results
Gemini 3.8 Flash at high effort is the cheapest of the three leading configurations here. Astra at xhigh buys a small increase in the point estimate for a higher task cost. Opus 5 at max costs more than either and has a slightly lower score in this run.
Luna at max reaches 67.2% for $0.61 per task at Datacurve’s corrected prices. GLM-5.3 Flash reaches 63.4% for $0.24. Both extend the frontier at lower budgets. Fable 5 peaks at 69.9% at xhigh for $13.41, so it sits below the cost/score frontier in this run.
The reported uncertainty ranges overlap. A 0.3-point gap between Astra and Gemini is too small to treat as an established quality advantage from these numbers alone. The chart uses the point estimates to calculate dominance. It does not test statistical significance.
Astra also uses far fewer output tokens and agent steps in this snapshot. Gemini’s lower bill comes with a longer trajectory. Steps are useful context, but wall-clock latency also depends on token throughput, tool execution, and provider delays. I would measure latency separately before choosing an interactive default.
Where DeepSeek V4.1 Flash fits
DeepSeek’s V4.1 Flash model card reports a continuously controllable reasoning effort from 1 to 100. For consistency with the rest of this chart, I label the three supplied evaluation points at 25, 75, and 100 as low, high, and max. This is an editorial chart mapping, not DeepSeek’s API terminology: its technical report maps the public API presets low, high, and max to 50, 75, and 100. The report does not tabulate an exact DeepSWE coordinate for effort 50, so I have not inferred a separate medium point from the figure. The displayed sweep rises from 66.0% at low to 72.0% at high and 74.2% at max. These remain self-reported points outside the official frontier until an independent run supplies comparable uncertainty and billing data.
At max, the reported 74.2% reaches the same band as Claude Opus 5 and Gemini 3.8 Flash. DeepSeek’s own comparison reports Opus 5 at 74.0%; the Datacurve snapshot plotted here has Opus at 73.6% and Gemini at 73.8%. The estimated V4.1 Flash bill is about $0.48 per task, roughly 4% of Opus 5’s measured $11.84 bill in this snapshot.
The low setting is the more consequential result for routine work. Its reported 66.0% already exceeds DeepSeek V4 Pro at max in the Datacurve snapshot, 62.8%, while the estimated 38k output tokens are about 64% below Pro’s measured 106k. The estimate is about $0.19 per task. High provides a middle point at 72.0%, about 60k output tokens, and roughly $0.30.
DeepSeek also reports Terminal-Bench 2.1 scores of 82.4%, 88.0%, and 90.6% at the chart’s low, high, and max points. Those numbers show the same monotonic effort tradeoff, but they do not enter the Terminal-Bench 4.0 table below. Benchmark versions and harnesses are part of the result.
DeepSeek’s launch documentation lists output at $0.60 per million tokens and uncached input at $0.155 per million, with lower cache-hit rates that vary by peak window. The chart’s cost, token, and step values are estimates based on the release report rather than measured Datacurve bills, so the points use hollow markers.
Where Fable 5.1 fits
Anthropic’s system card reports 67.4% on DeepSWE v1.1, averaged over five trials. Section 8.3 is on page 168. The standard configuration is adaptive thinking at max effort. The DeepSWE section does not specify the harness or provide cost, token usage, steps, or uncertainty.
Fable 5.1 is selectable alongside Fable 5. The chart keeps Anthropic’s result as a score-only entry because that run has no published cost. The result is self-reported and remains outside the official frontier.
Artificial Analysis provides a separate five-level effort sweep on its Intelligence Index v4.1.1. That index combines nine evaluations, including Terminal-Bench v2.1, SciCode, HLE, and GPQA Diamond. The scores and weighted task costs below belong to that suite.
| Effort | Intelligence Index score | USD / Intelligence Index task |
|---|---|---|
| max | 66 | $3.69 |
| xhigh | 65 | $2.65 |
| high | 62 | $1.43 |
| medium | 60 | $1.00 |
| low | 58 | $0.77 |
Pre-release evaluation with default server-side fallback. Model-page costs are shown above. The September 1 launch article recorded $3.76 at max and $2.72 at xhigh. These are different source snapshots, not uncertainty ranges. This table preserves v4.1.1, even as the live pages change versions.
AA’s September 1 writeup puts max at 66 and xhigh at 65, with xhigh costing $1.04 less per Index task in that launch snapshot. High scores 62, matching Fable 5 at max. AA evaluated Fable 5.1 with default server-side fallback, which routed some requests to Opus 4.8 or Opus 5 and accounted for roughly 4% of output tokens across the Index. Its five effort levels used 13.1M to 143.7M output tokens across the suite, an elevenfold range.
AA’s September 3 coding-agent comparison reports 70 for Fable 5.1 in Claude Code. Astra in Codex scores 67, approximately level with Opus 5, Fable 5, and Muse Spark 1.3. Those are composite Coding Agent Index scores from that dated comparison.
The live leaderboard has since moved to v1.5. On September 10, its independent DeepSWE v1.1 result for Fable 5.1 at max effort is 64.3% in Claude Code. AA also reports $12.39 per task for that configuration across the full three-benchmark Coding Agent Index. Its public per-benchmark data omits a DeepSWE-only cost. That suite-wide bill cannot supply the missing chart coordinate, and Claude Code is a different harness from Datacurve’s mini-swe-agent.
Where Muse Spark 1.3 fits
Meta reports 75.4% on DeepSWE v1.1. Its evaluation methodology specifies Muse Spark 1.3 at max effort with a mini-swe agent. The comparison with Muse Spark 1.2 also changes the effort setting from xhigh to max.
I have kept 1.3 as self-reported and pending independent review because it was absent from the Datacurve leaderboard at this source check. Both hollow diamonds are overlays of the same reported score. They never change the official frontier or receive an official rank.
Both pricing entries are visible by default. Standard (non-contributor) uses an estimated $2.75 per task, and Contributor uses $0.25. Each can be selected or hidden through the model filters. These estimates came from the initial comparison notes. They are not measured DeepSWE bills, and token prices alone cannot establish them.
| Tier | Input | Cached input | Output |
|---|---|---|---|
| Standard | $1.25 | $0.15 | $4.25 |
| Contributor | $0.10 | $0.002 | $0.20 |
Meta’s pricing page says contributor usage is used to improve its products. Standard usage is not. That makes the discounted tier a separate purchasing decision even for the same model.
Meta also reports roughly 25% fewer tokens and 20% fewer tool calls in comparisons by its engineers. Applying those reductions to the 1.2 leaderboard row gives roughly 74k output tokens and 80 steps. Those are projections from general engineering workflows, not measured 1.3 DeepSWE statistics. Tool calls and agent steps may also be counted differently.
The missing inputs are the actual input-token volume, cache mix, and billed usage for that evaluation. A reproducible cost calculation needs uncached input tokens times their rate, cached input tokens times their rate, and output tokens times their rate, summed across the run and divided by the number of tasks. Until those numbers are available, I use the two scenarios to explore where the claim might land.
Check a second kind of work
Terminal-Bench adds a different view of agent work in terminal environments. I keep its version and agent name next to each score because both change what a result means.
| Model / effort | Agent | Resolution rate |
|---|---|---|
| GPT-6 Astra[max] | Codex | 58.2% ±2.8 pp |
| Claude Fable 5.1[max] | Claude Code | 57.9% ±3.8 pp |
| Claude Opus 5[max] | Claude Code | 51.8% ±3.4 pp |
| Gemini 3.8 Flash[high] | mini-SWE-agent | 19.1% ±3.4 pp |
The source labels these as 95% confidence intervals. View the Terminal-Bench leaderboard.
Gemini’s strong DeepSWE result does not carry over to the same position in this Terminal-Bench 4.0 selection. Astra and Opus run through their native coding agents here, while Gemini runs through mini-SWE-agent. Task mix and harness both differ, so this is evidence about those evaluated systems. It cannot isolate how much of the gap comes from the model itself.
Meta’s 88.8% Terminal-Bench claim is for version 2.1, under its native-harness evaluation setup. It does not belong in the 4.0 table. A future Muse 1.3 result needs the same benchmark version before I add it there.
For a purchasing decision, I would add a small set of tasks from my own repositories. I care about regressions, time to a patch I can accept, and the total bill including retries. The public benchmarks help me choose which configurations to try first.
Sources
- Datacurve DeepSWE v1.1 leaderboard
- DeepSWE v1.1 configuration data (before website pricing corrections)
- DeepSWE methodology and evaluation harness
- Datacurve reporting corrections
- Meta Muse Spark 1.3 scores and token pricing
- Meta Muse Spark 1.3 evaluation methodology
- Meta Muse Spark 1.3 announcement and efficiency claims
- Terminal-Bench 4.0 leaderboard
- Anthropic Fable 5.1 system card, §8.3 DeepSWE v1.1 (p. 168)
- Artificial Analysis Fable 5.1 pre-release evaluation, September 1
- Artificial Analysis Fable 5.1, max with fallback
- Artificial Analysis Fable 5.1, xhigh with fallback
- Artificial Analysis Fable 5.1, high with fallback
- Artificial Analysis Fable 5.1, medium with fallback
- Artificial Analysis Fable 5.1, low with fallback
- Artificial Analysis coding-agent comparison, September 3
- Artificial Analysis live Coding Agent Index
- DeepSeek V4.1 Flash model card and technical report
- DeepSeek V4.1 Flash launch and API pricing
Updates
- Added DeepSeek V4.1 Flash at low, high, and max using DeepSeek’s self-reported effort-25, effort-75, and effort-100 DeepSWE v1.1 points and estimated usage and cost coordinates.
- Added DeepSeek’s Terminal-Bench 2.1 results to the analysis without mixing them into the separate Terminal-Bench 4.0 comparison.