How we measure intelligence per dollar
The exact formula behind the AI Intelligence Calculator's headline metric — how tokens per dollar and the intelligence index combine into a single value score, and its limits.
How we measure intelligence per dollar
Every number on this site funnels into one question: how much AI do you actually get for a dollar? The answer is a single metric, and this post shows the exact math behind it so you can judge it for yourself.
The formula
intelligence per dollar = tokens per dollar × (intelligence index ÷ 100)
Tokens per dollar is objective and mechanical:
tokens per dollar = 1,000,000 ÷ blended cost per 1M tokens
blended cost = (input price + 3 × output price) ÷ 4
We blend input and output at a 1:3 ratio because, in typical use, output tokens cost roughly three times as much as input tokens. A model that’s cheap on input but expensive on output gets penalized accordingly.
The intelligence index is the quality weight. It comes from Artificial Analysis’s Intelligence Index, a 0–100 benchmark of general capability. A model with an index of 50 is treated as delivering twice the “intelligence” of a model at 25, per token.
Why the index matters
Cheap tokens only matter if they come with quality. A $0.08 model that scores 10 on the index delivers far less intelligence per dollar than a $0.44 model that scores 53 — even though the first is technically cheaper. The index is what stops the metric from simply crowning the cheapest model.
Models without a benchmarked index fall back to a tier-based quality factor (frontier 1.0, balanced ~0.7–0.8, fast 0.5) so they still rank sensibly rather than scoring zero. You’ll see those as having no index pill but a normal value score.
What the metric is — and isn’t
- It’s a ranking lens. It answers “which model gives me the most capability per dollar on average.”
- It’s not a quality ranking. A high intelligence-per-dollar score can come from very cheap tokens as easily as from high quality. Read it alongside the raw index, not instead of it.
- It’s not workload-specific. Coding, writing and long-document work weight input and output differently, which is why the calculator lets you pick a use case and re-ranks everything.
The full treatment, including data sources and update cadence, lives on the methodology page. And if you’d rather just see the result, the calculator does all of this math live for every model we track.