01 · Relative composite
Make unlike benchmarks comparable
The composite uses DeepSWE v1.1, FrontierCode 1.1, CursorBench 4.0, and Terminal-Bench 4.0. CursorBench scores aren’t directly comparable with version 3.2 because its task set changed.
- Score and cost are independently min-max scaled within each benchmark. The observed minimum becomes 0 and the maximum becomes 100; a range with no spread becomes 50.
- For each model and reasoning level, the normalized values are averaged equally across every benchmark reporting that configuration. If any contributing benchmark lacks a cost, the composite cost is unavailable; the score still appears.
- The result is a relative score and relative cost. The Log control changes only the x-axis spacing, never the underlying value.
relative = 100 × (value − min) / (max − min)
02 · Pareto fit
Estimate the composite cost-performance frontier
Pareto mode shows the reported reasoning level highest above the shared fitted contour for each model. Its Pareto score is normalized performance minus the contour height at that configuration’s log cost. Ties favor lower cost. Chart and legend use the same fit from all configurations of selected models, and Log never changes the selection. Individual benchmark views use their own fit; when a curve cannot be fitted, selection falls back to highest score.
- Pareto guides appear only in Composite. Only configurations belonging to selected models are considered; at an equal cost, only the highest score survives.
- Points are sorted from lowest to highest cost. A point joins the frontier only when it improves upon every cheaper score, removing all dominated configurations.
- Frontier coordinates use zero-safe log cost and linear score. We fit a least-squares quadratic constrained to rise with log cost and bend toward diminishing returns, then cap the fitted score at the chart’s score ceiling. Switching to Log transforms the display of those same boundaries without fitting a different curve.
- The valid unconstrained quadratic, saturating quadratic, and linear boundary are compared; the lowest-error fit wins. The six solid curves are evenly spaced score intervals around that fitted frontier, including one interval above it. Each band spans one quarter of the visible score range (25 relative-score points when the axis runs from 0 to 100). These are visual guides, not measured model results.
score = a·log-cost² + b·log-cost + clog-cost = log(1 + cost) / log(1 + maximum cost)a ≤ 0 · slope ≥ 0 across the chart