Frontier coding models, measured over time

A living comparison of DeepSWE v1.1 scores, cost per task, and complementary coding evals. Official results alongside pending claims, with an interactive Pareto frontier.

Published · Updated

The top DeepSWE scores are close enough that cost changes the decision. GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 sit within half a percentage point in this snapshot. Their average task costs range from $2.36 to $11.84.

I’m keeping this page as a running comparison of frontier coding models. I want to see what a model solves, what the run costs, and whether the result holds up in the agent I would use.

Last source check . Datacurve’s leaderboard is dated . This is a curated snapshot.

DeepSWE at a glance

DeepSWE v1.1 has 113 tasks across 91 repositories and five languages. Datacurve writes original engineering tasks with functional and regression tests. The leaderboard uses mini-swe-agent across models, which makes it useful for comparing models under a shared harness.

Pass@1 measures success on one attempt per task, averaged over the evaluation. Average cost per task includes unsuccessful attempts. A model that spends more time on a failing task can still leave you with a larger bill.

COST / CAPABILITY

DeepSWE v1.1

Official runs · 113 tasks · mini-swe-agent · checked

All 70 published configurations across 28 official models, plus self-reported Muse Spark 1.3, Fable 5.1, and DeepSeek V4.1 Flash results.

32 models · 75 plotted configurations · 1 score only
Pass@1 versus average cost per taskCost decreases from left to right. The upper-right is better. Cost uses a logarithmic scale. Each colored line connects reasoning levels of the same model version. The optional dashed line is the official Pareto frontier. Hollow diamonds are self-reported scores with estimated costs. Results with unknown costs appear in a separate score-only panel. Focus or select a point to read its details. Every configuration is also available in the table.$0.01$0.10$1$100%10%20%30%40%50%60%70%80%DeepSWE scoremost efficient ↗Avg cost per task · log scaleGPT-6 AstraXHIGHGemini 3.8 FlashHIGHGPT-5.6 LunaMAXGLM-5.3 FlashMAXMuse Spark 1.3 (Standard)MAX · EST.Muse Spark 1.3 (Contributor)MAX · EST.DeepSeek V4.1 FlashMAX · EST.
Same model, different reasoning levels● Measured configuration◇ Self-reported / estimated cost

SCORE ONLY · COST NOT REPORTED

Cost per DeepSWE task is unknown. Select an entry for its source and evaluation details. It has no position on the cost axis.

Hover, tap, or tab to a point. On narrow screens, scroll the chart horizontally or inspect a model below. With more than six models visible, labels identify frontier models, self-reported results, and the inspected model. Filter to six or fewer to label every model.

GPT-6 Astra [xhigh]

Official leaderboard result · On this selection’s official frontier

Harness · mini-swe-agent

Pass@1
74.1% ±3 pp
Cost / task
$6.52
Output tokens
30k
Agent steps
29
View score source

Official frontier members in this selection are GPT-5.6 Luna [low], GPT-5.6 Luna [medium], GPT-5.6 Luna [high], GLM-5.3 Flash [max], GPT-5.6 Luna [max], Gemini 3.8 Flash [medium], Gemini 3.8 Flash [high], GPT-6 Astra [xhigh]. Pending results never change that frontier. Solid lines join published reasoning levels of the same model and evidence type. Values between points are not measurements.

Configuration table 76 results
Selected DeepSWE v1.1 configurations · self-reported entries may use estimated costs · unknown costs remain unplotted
Model / effortPass@1USD / taskOut tokensStepsEvidence
GPT-6 Astra[xhigh]74.1%±3 pp$6.5230k29Official
Gemini 3.8 Flash[high]73.8%±1 pp$2.36143k166Official
Claude Opus 5[max]73.6%±4 pp$11.84118k99Official
GPT-6 Astra[high]73.2%±3 pp$5.7227k27Official
GPT-6 Astra[max]73.2%±1 pp$12.3761k28Official
Claude Opus 5[xhigh]73.2%±3 pp$9.0792k89Official
Claude Opus 5[high]72.8%±2 pp$6.0864k73Official
GPT-6 Astra[medium]72.8%±3 pp$4.3820k26Official
GPT-5.6 Sol[max]72.7%±3 pp$6.4660k61Official
Gemini 3.8 Flash[medium]71.0%±2 pp$1.97125k147Official
GPT-5.6 Sol[xhigh]70.7%±1 pp$3.6041k44Official
Claude Fable 5[xhigh]69.9%±3 pp$13.4180k68Official
Claude Fable 5[max]69.7%±4 pp$21.63119k88Official
GPT-5.6 Terra[max]69.6%±3 pp$3.9672k76Official
GPT-5.6 Sol[high]69.4%±1 pp$2.6628k37Official
GLM-5.3[max]69.0%±3 pp$3.9980k124Official
Claude Opus 5[medium]68.9%±1 pp$3.2937k52Official
Claude Fable 5[high]68.6%±1 pp$9.1857k59Official
Kimi K3[max]68.5%±5 pp$4.6581k98Official
Grok 4.6[medium]67.5%±2 pp$3.4550k70Official
GPT-5.6 Luna[max]67.2%±4 pp$0.6173k102Official
GPT-5.5[xhigh]67.0%±6 pp$7.2346k82Official
GPT-6 Astra[low]67.0%±1 pp$2.1911k20Official
Grok 4.6[xhigh]66.7%±2 pp$5.5071k87Official
Gemini 3.7 Flash[medium]65.5%±3 pp$2.0394k117Official
Claude Fable 5[medium]65.4%±4 pp$6.0940k48Official
Gemini 3.7 Flash[high]65.3%±2 pp$2.18107k125Official
Grok 4.6[high]65.2%±2 pp$4.3861k79Official
GPT-5.5[high]64.4%±3 pp$5.1031k62Official
GLM-5.3 Flash[max]63.4%±4 pp$0.2473k123Official
DeepSeek V4 Pro[max]62.8%±6 pp$1.67106k155Official
GPT-5.6 Sol[medium]61.1%±2 pp$1.4218k31Official
GPT-5.6 Terra[xhigh]60.2%±2 pp$1.7040k43Official
Claude Fable 5[low]59.6%±3 pp$3.7625k38Official
Claude Opus 4.8[max]59.0%±2 pp$13.22135k120Official
Claude Opus 5[low]58.1%±2 pp$1.6620k36Official
Qwen3.8 Max[xhigh]57.5%±3 pp$3.7395k111Official
GPT-5.6 Luna[xhigh]56.9%±2 pp$0.3145k71Official
Muse Spark 1.2[xhigh]54.9%±2 pp$3.7099k101Official
Claude Opus 4.8[xhigh]54.4%±4 pp$8.0186k95Official
GPT-5.5[medium]54.0%±3 pp$2.7520k46Official
Claude Sonnet 5[max]53.8%±4 pp$26.40214k268Official
Gemini 3.7 Flash[low]53.8%±3 pp$1.8373k130Official
GPT-5.6 Terra[high]53.8%±4 pp$0.9122k34Official
Grok 4.5[high]53.8%±2 pp$2.4236k61Official
DeepSeek V4 Flash[max]53.3%±4 pp$0.46108k153Official
Muse Spark 1.1[xhigh]53.3%±3 pp$2.3674k96Official
Claude Opus 4.8[high]51.8%±5 pp$4.2850k73Official
GPT-5.4[xhigh]51.8%±2 pp$5.6571k70Official
Claude Sonnet 5[xhigh]49.7%±3 pp$11.89121k186Official
Claude Opus 4.8[medium]48.7%±2 pp$3.4441k66Official
Claude Sonnet 5[high]48.2%±5 pp$7.4387k147Official
Gemini 3.6 Flash[high]46.7%±4 pp$2.2196k117Official
GPT-5.6 Sol[low]45.4%±2 pp$0.8211k23Official
GPT-5.6 Luna[high]44.2%±3 pp$0.1626k49Official
GLM-5.2[max]43.8%±2 pp$3.9278k129Official
Grok 4.6[low]41.6%±2 pp$1.0416k44Official
Claude Opus 4.8[low]40.8%±1 pp$2.2929k54Official
Claude Sonnet 5[medium]39.8%±3 pp$4.0857k108Official
GLM-5.2[high]36.3%±5 pp$2.8454k122Official
Gemini 3.5 Flash[high]36.1%±4 pp$3.4576k105Official
GPT-5.6 Terra[medium]35.1%±3 pp$0.4712k25Official
Kimi K2.7 Code[default]30.5%±1 pp$2.8259k149Official
Claude Sonnet 5[low]30.5%±1 pp$2.1936k77Official
Claude Sonnet 4.6[high]29.9%±4 pp$5.5276k134Official
GPT-5.5[low]27.0%±2 pp$1.209k28Official
GPT-5.6 Terra[low]24.1%±1 pp$0.349k21Official
Gemini 3.1 Pro Preview[high]11.7%±1 pp$2.1428k76Official
GPT-5.6 Luna[medium]11.3%±1 pp$0.048k24Official
GPT-5.6 Luna[low]1.5%±1 pp$0.013k12Official
Muse Spark 1.3 (Standard)[max]75.4%Range not reported~$2.75estimate~74kestimated~80estimatedSelf-reportedPending review
Muse Spark 1.3 (Contributor)[max]75.4%Range not reported~$0.25estimate~74kestimated~80estimatedSelf-reportedPending review
Claude Fable 5.1[max]67.4%Range not reportedNot reportedNot reportedNot reportedSelf-reportedPending review
DeepSeek V4.1 Flash[low]66.0%Range not reported~$0.19estimate~38kestimated~95estimatedSelf-reportedPending review
DeepSeek V4.1 Flash[high]72.0%Range not reported~$0.30estimate~60kestimated~120estimatedSelf-reportedPending review
DeepSeek V4.1 Flash[max]74.2%Range not reported~$0.48estimate~95kestimated~140estimatedSelf-reportedPending review

The comparison includes all 70 configurations across the 28 models in Datacurve’s current v1.1 dataset, plus the Muse Spark 1.3, Fable 5.1, and DeepSeek V4.1 Flash claims. Muse has separate Standard and Contributor entries, each with its own model filter. Use the searchable model dropdown on the chart to choose a comparison, and the Options dropdown beside it for the scale, frontier, uncertainty, and self-reported toggles. Fable 5 and Luna are included at every published reasoning level. DeepSeek V4.1 Flash connects low, high, and max with the standard reasoning-series line; its hollow markers identify the self-reported points. Fable 5.1 appears in the score-only panel and configuration table because its DeepSWE cost is unknown.

The upper-right is the useful direction on this chart, as on Datacurve. Cost decreases toward the right. The default logarithmic cost scale keeps the roughly $0.01 to $26.40 range readable. Switch it off for a linear scale. The axes rescale to the visible models, so a narrow comparison uses the full plot area. Each solid curve connects the published reasoning levels of one model. Hover or select any point for that configuration’s score and cost. Small comparisons label every model. Larger comparisons label the frontier models, self-reported results, and whichever model you inspect.

A configuration is on the Pareto frontier when no other included configuration is at least as good on both cost and score and strictly better on one. Turn on “Show official frontier” to see it as a dashed line. It is computed from the visible official configurations and updates with the model filters. With all models selected, it uses the full 70-configuration snapshot. A configuration below this line could still be useful for a different workload.

What I take from the current results

Gemini 3.8 Flash at high effort is the cheapest of the three leading configurations here. Astra at xhigh buys a small increase in the point estimate for a higher task cost. Opus 5 at max costs more than either and has a slightly lower score in this run.

Luna at max reaches 67.2% for $0.61 per task at Datacurve’s corrected prices. GLM-5.3 Flash reaches 63.4% for $0.24. Both extend the frontier at lower budgets. Fable 5 peaks at 69.9% at xhigh for $13.41, so it sits below the cost/score frontier in this run.

The reported uncertainty ranges overlap. A 0.3-point gap between Astra and Gemini is too small to treat as an established quality advantage from these numbers alone. The chart uses the point estimates to calculate dominance. It does not test statistical significance.

Astra also uses far fewer output tokens and agent steps in this snapshot. Gemini’s lower bill comes with a longer trajectory. Steps are useful context, but wall-clock latency also depends on token throughput, tool execution, and provider delays. I would measure latency separately before choosing an interactive default.

Where DeepSeek V4.1 Flash fits

DeepSeek’s V4.1 Flash model card reports a continuously controllable reasoning effort from 1 to 100. For consistency with the rest of this chart, I label the three supplied evaluation points at 25, 75, and 100 as low, high, and max. This is an editorial chart mapping, not DeepSeek’s API terminology: its technical report maps the public API presets low, high, and max to 50, 75, and 100. The report does not tabulate an exact DeepSWE coordinate for effort 50, so I have not inferred a separate medium point from the figure. The displayed sweep rises from 66.0% at low to 72.0% at high and 74.2% at max. These remain self-reported points outside the official frontier until an independent run supplies comparable uncertainty and billing data.

At max, the reported 74.2% reaches the same band as Claude Opus 5 and Gemini 3.8 Flash. DeepSeek’s own comparison reports Opus 5 at 74.0%; the Datacurve snapshot plotted here has Opus at 73.6% and Gemini at 73.8%. The estimated V4.1 Flash bill is about $0.48 per task, roughly 4% of Opus 5’s measured $11.84 bill in this snapshot.

The low setting is the more consequential result for routine work. Its reported 66.0% already exceeds DeepSeek V4 Pro at max in the Datacurve snapshot, 62.8%, while the estimated 38k output tokens are about 64% below Pro’s measured 106k. The estimate is about $0.19 per task. High provides a middle point at 72.0%, about 60k output tokens, and roughly $0.30.

DeepSeek also reports Terminal-Bench 2.1 scores of 82.4%, 88.0%, and 90.6% at the chart’s low, high, and max points. Those numbers show the same monotonic effort tradeoff, but they do not enter the Terminal-Bench 4.0 table below. Benchmark versions and harnesses are part of the result.

DeepSeek’s launch documentation lists output at $0.60 per million tokens and uncached input at $0.155 per million, with lower cache-hit rates that vary by peak window. The chart’s cost, token, and step values are estimates based on the release report rather than measured Datacurve bills, so the points use hollow markers.

Where Fable 5.1 fits

Anthropic’s system card reports 67.4% on DeepSWE v1.1, averaged over five trials. Section 8.3 is on page 168. The standard configuration is adaptive thinking at max effort. The DeepSWE section does not specify the harness or provide cost, token usage, steps, or uncertainty.

Fable 5.1 is selectable alongside Fable 5. The chart keeps Anthropic’s result as a score-only entry because that run has no published cost. The result is self-reported and remains outside the official frontier.

Artificial Analysis provides a separate five-level effort sweep on its Intelligence Index v4.1.1. That index combines nine evaluations, including Terminal-Bench v2.1, SciCode, HLE, and GPQA Diamond. The scores and weighted task costs below belong to that suite.

Artificial Analysis Intelligence Index v4.1.1 · Fable 5.1 · source check 2026-09-10
EffortIntelligence Index scoreUSD / Intelligence Index task
max66$3.69
xhigh65$2.65
high62$1.43
medium60$1.00
low58$0.77

Pre-release evaluation with default server-side fallback. Model-page costs are shown above. The September 1 launch article recorded $3.76 at max and $2.72 at xhigh. These are different source snapshots, not uncertainty ranges. This table preserves v4.1.1, even as the live pages change versions.

AA’s September 1 writeup puts max at 66 and xhigh at 65, with xhigh costing $1.04 less per Index task in that launch snapshot. High scores 62, matching Fable 5 at max. AA evaluated Fable 5.1 with default server-side fallback, which routed some requests to Opus 4.8 or Opus 5 and accounted for roughly 4% of output tokens across the Index. Its five effort levels used 13.1M to 143.7M output tokens across the suite, an elevenfold range.

AA’s September 3 coding-agent comparison reports 70 for Fable 5.1 in Claude Code. Astra in Codex scores 67, approximately level with Opus 5, Fable 5, and Muse Spark 1.3. Those are composite Coding Agent Index scores from that dated comparison.

The live leaderboard has since moved to v1.5. On September 10, its independent DeepSWE v1.1 result for Fable 5.1 at max effort is 64.3% in Claude Code. AA also reports $12.39 per task for that configuration across the full three-benchmark Coding Agent Index. Its public per-benchmark data omits a DeepSWE-only cost. That suite-wide bill cannot supply the missing chart coordinate, and Claude Code is a different harness from Datacurve’s mini-swe-agent.

Where Muse Spark 1.3 fits

Meta reports 75.4% on DeepSWE v1.1. Its evaluation methodology specifies Muse Spark 1.3 at max effort with a mini-swe agent. The comparison with Muse Spark 1.2 also changes the effort setting from xhigh to max.

I have kept 1.3 as self-reported and pending independent review because it was absent from the Datacurve leaderboard at this source check. Both hollow diamonds are overlays of the same reported score. They never change the official frontier or receive an official rank.

Both pricing entries are visible by default. Standard (non-contributor) uses an estimated $2.75 per task, and Contributor uses $0.25. Each can be selected or hidden through the model filters. These estimates came from the initial comparison notes. They are not measured DeepSWE bills, and token prices alone cannot establish them.

Meta’s listed USD per million tokens · checked 2026-09-11
TierInputCached inputOutput
Standard$1.25$0.15$4.25
Contributor$0.10$0.002$0.20

Meta’s pricing page says contributor usage is used to improve its products. Standard usage is not. That makes the discounted tier a separate purchasing decision even for the same model.

Meta also reports roughly 25% fewer tokens and 20% fewer tool calls in comparisons by its engineers. Applying those reductions to the 1.2 leaderboard row gives roughly 74k output tokens and 80 steps. Those are projections from general engineering workflows, not measured 1.3 DeepSWE statistics. Tool calls and agent steps may also be counted differently.

The missing inputs are the actual input-token volume, cache mix, and billed usage for that evaluation. A reproducible cost calculation needs uncached input tokens times their rate, cached input tokens times their rate, and output tokens times their rate, summed across the run and divided by the number of tasks. Until those numbers are available, I use the two scenarios to explore where the claim might land.

Check a second kind of work

Terminal-Bench adds a different view of agent work in terminal environments. I keep its version and agent name next to each score because both change what a result means.

Terminal-Bench 4.0 · selected official entries · checked 2026-09-10
Model / effortAgentResolution rate
GPT-6 Astra[max]Codex58.2% ±2.8 pp
Claude Fable 5.1[max]Claude Code57.9% ±3.8 pp
Claude Opus 5[max]Claude Code51.8% ±3.4 pp
Gemini 3.8 Flash[high]mini-SWE-agent19.1% ±3.4 pp

The source labels these as 95% confidence intervals. View the Terminal-Bench leaderboard.

Gemini’s strong DeepSWE result does not carry over to the same position in this Terminal-Bench 4.0 selection. Astra and Opus run through their native coding agents here, while Gemini runs through mini-SWE-agent. Task mix and harness both differ, so this is evidence about those evaluated systems. It cannot isolate how much of the gap comes from the model itself.

Meta’s 88.8% Terminal-Bench claim is for version 2.1, under its native-harness evaluation setup. It does not belong in the 4.0 table. A future Muse 1.3 result needs the same benchmark version before I add it there.

For a purchasing decision, I would add a small set of tasks from my own repositories. I care about regressions, time to a patch I can accept, and the total bill including retries. The public benchmarks help me choose which configurations to try first.

Sources

Updates

  • Added DeepSeek V4.1 Flash at low, high, and max using DeepSeek’s self-reported effort-25, effort-75, and effort-100 DeepSWE v1.1 points and estimated usage and cost coordinates.
  • Added DeepSeek’s Terminal-Bench 2.1 results to the analysis without mixing them into the separate Terminal-Bench 4.0 comparison.
All writingEmail me