ParetoBench

Composite

Relative score against relative cost

Chart view
Chart display
Relative composite score
Composite legend view
Reasoning-level relative performance versus relative costEach line connects the reported reasoning levels for one model from minimal through max. Higher scores and lower costs are better. Cost decreases from left to right. The x-axis is linear and the y-axis is linear.02040608010080644832160Relative scoreRelative costlowmediumhighxhighmaxOpus 5.5lowmediumhighxhighmaxGPT 6 AstralowmediumhighxhighmaxSonnet 5.5lowmediumhighxhighmaxFable 5.1lowmediumhighxhighmaxGPT 6.1 SolmediumhighmaxSWE 2maxKimi K3lowmediumhighxhighmaxGPT 6 LunalowhighmaxGLM 5.3maxDeepseek V4 FlashlowmediumhighxhighGrok 4.7lowhighmaxGLM 5.3 FlashminimallowmediumhighxhighmaxMuse Spark 1.3mediumhighGemini 3.8 FlashmaxDeepseek V4 ProhighDeepseek V4 Pro 0813reportedComposer 2.5highDeepseek V4 Flash 0731highGemini 3.1 Pro Preview

Methodology

How the composite is calculated

01 · Relative composite

Make unlike benchmarks comparable

The composite uses DeepSWE v1.1, FrontierCode 1.1, CursorBench 4.0, and Terminal-Bench 4.0. CursorBench scores aren’t directly comparable with version 3.2 because its task set changed.

  1. Score and cost are independently min-max scaled within each benchmark. The observed minimum becomes 0 and the maximum becomes 100; a range with no spread becomes 50.
  2. For each model and reasoning level, the normalized values are averaged equally across every benchmark reporting that configuration. If any contributing benchmark lacks a cost, the composite cost is unavailable; the score still appears.
  3. The result is a relative score and relative cost. The Log control changes only the x-axis spacing, never the underlying value.
relative = 100 × (value − min) / (max − min)

02 · Pareto fit

Estimate the composite cost-performance frontier

Pareto mode shows the reported reasoning level highest above the shared fitted contour for each model. Its Pareto score is normalized performance minus the contour height at that configuration’s log cost. Ties favor lower cost. Chart and legend use the same fit from all configurations of selected models, and Log never changes the selection. Individual benchmark views use their own fit; when a curve cannot be fitted, selection falls back to highest score.

  1. Pareto guides appear only in Composite. Only configurations belonging to selected models are considered; at an equal cost, only the highest score survives.
  2. Points are sorted from lowest to highest cost. A point joins the frontier only when it improves upon every cheaper score, removing all dominated configurations.
  3. Frontier coordinates use zero-safe log cost and linear score. We fit a least-squares quadratic constrained to rise with log cost and bend toward diminishing returns, then cap the fitted score at the chart’s score ceiling. Switching to Log transforms the display of those same boundaries without fitting a different curve.
  4. The valid unconstrained quadratic, saturating quadratic, and linear boundary are compared; the lowest-error fit wins. The six solid curves are evenly spaced score intervals around that fitted frontier, including one interval above it. Each band spans one quarter of the visible score range (25 relative-score points when the axis runs from 0 to 100). These are visual guides, not measured model results.
score = a·log-cost² + b·log-cost + clog-cost = log(1 + cost) / log(1 + maximum cost)a ≤ 0 · slope ≥ 0 across the chart

Chart legend

Choose legend models

19 in legend