There is no shortage of AI model rankings. Most try to answer some version of the same question: Which model is best?

We wanted to answer a different one: Which models are best for different kinds of work?

That distinction matters if you are building an AI application. The model that performs well on a difficult coding problem may not be the best choice for extraction, writing, or quantitative work. And even when, two models produce similar-quality answers, the difference in cost or failure behavior can be meaningful.

So we built our own evaluations around the kinds of requests Elektric actually needs to route.

Across 12 category-and-difficulty pilots, we evaluated a common cohort of 14 models from OpenAI, Anthropic, and Google across six types of work, with separate Easy and Hard pilots for each. Altogether, the selected pilots contained 130 prompts, 1,865 candidate answers, and 3,730 model-judge evaluations.

The results did not point to one universal winner. They showed why model selection becomes more useful when you look at the work being done, the difficulty of that work, the cost of generating the answer, and the reliability of the model together.

Six types of work

We divided requests into six broad categories:

Code: covers programming, SQL, debugging, software architecture, and technical implementation.

Quantitative: covers mathematics, statistics, calculations, valuation, proofs, and numerical analysis.

Knowledge and risk: covers work where specialized knowledge and consequences matter, including legal, regulatory, compliance, security-policy, and other expertise-heavy analysis.

Reasoning and decisions: covers planning, comparisons, recommendations, causal diagnosis, strategy, and tradeoffs.

Writing and transformation: covers drafting, rewriting, summarizing, translating, editing, and changing tone or format.

Extraction and classification: covers pulling specified information from supplied material, assigning labels, and mapping information into structured schemas.

These categories describe the underlying work, not superficial features of the prompt. Asking for JSON does not make a task coding. Writing SQL does. Calculating a financial value is quantitative, while making a nonnumeric business recommendation is reasoning and decisions. Summarizing a document is writing and transformation, while copying specified fields from it is extraction.

We also separated each category into Easy and Hard tasks. Easy generally means routine work with limited ambiguity. Hard tasks involve things like dependent reasoning, difficult constraints, advanced derivation, system-level debugging, conflicting evidence, or deeper expertise. Prompt length alone does not determine difficulty.

No single model led every type of work

The easiest way to see the result is to look at the leaders across the 12 pilots.

Category

Easy leader

Hard leader

Code

GPT-5.6 Sol

GPT-5.6 Sol

Quantitative

GPT-5.6 Sol

GPT-5.6 Sol

Knowledge and risk

GPT-5.6 Terra

GPT-5.6 Terra

Reasoning and decisions

GPT-5.6 Luna /    GPT-5.6 Terra

GPT-5.6 Luna /      GPT-5.6 Terra

Writing and transformation

GPT-5.6 Luna

GPT-5.6 Sol

Extraction and classification

Anthropic Sonnet 5

GPT-5.6 Luna

These are average model-judge score leaders within each pilot, not production routing assignments.

That alone tells us something useful: the model that led one type of work did not automatically lead another.

Extraction is a particularly clear example. Anthropic Sonnet 5 led Easy Extraction with an average judged score of 98.35. On Hard Extraction, GPT-5.6 Luna led with 77.8, while Sonnet 5 scored 60.95.

These were different prompt sets, so that difference should not be read as a controlled measurement of difficulty alone. But it shows why a single overall model ranking is not enough for the routing decisions we care about.

Cost and quality did not move together

The relationship between generation cost and judged quality varied substantially across tasks.

On Easy Quantitative work, GPT-5.6 Sol received an average judged score of 100 at an average candidate-generation cost of $0.029052 per prompt. GPT-5.6 Luna scored 99.7 at $0.005378 per prompt. In this pilot, Luna produced a very similar average judged score at a fraction of the candidate-generation cost.

Hard Writing and Transformation showed a different pattern:

Model

Average judged score

Avg. candidate cost per prompt

Severe-failure prompts

openai/gpt-5.6-sol

97.3

$0.077953

0 of 10

openai/gpt-5.6-luna

91.8

$0.008211

1 of 10

anthropic/claude-fable-5

91.0

$0.119196

1 of 10

Here, the highest-scoring model was not the most expensive. Claude Fable 5 cost more per candidate answer than GPT-5.6 Sol while receiving a lower average judged score. GPT-5.6 Luna was much less expensive than both, but it also received a lower average score than Sol and had one prompt flagged for severe failure.

That is why we look at quality, cost, and failure patterns together rather than optimizing around one number.

The benchmark has to match the task

One of the strongest lessons from building these pilots was that “quality” has to be defined differently depending on the work.

For quantitative tasks, judges looked at dimensions such as numerical accuracy, method, completeness, and reproducibility.

For coding, the rubrics included correctness, robustness, instruction following, and test quality.

Writing and Transformation emphasized source fidelity, audience and tone fit, organization, instruction following, and whether the model preserved qualifications and uncertainty.

Extraction and Classification focused on correctness, schema compliance, normalization, completeness, and ambiguity handling.

The prompts reflected those differences. A Code task might ask a model to implement bounded asynchronous concurrency or design payment idempotency across active regions. A Hard Reasoning task might ask it to allocate scarce company resources under financing and dependency constraints. A Hard Writing task might require reconciling conflicting incident records into an executive brief without erasing unresolved uncertainty.

Those are different jobs. Evaluating them with one generic leaderboard score would hide much of the information we actually need.

How we judged the answers

Judge setup varied across the pilots.

Code used combinations of OpenAI, Anthropic, and Google judges. Candidate identity was hidden in the judging prompt, and judges from the candidate’s own provider were excluded. Some Code answers ended up with one valid judgment because other judge responses were invalid.

Several other pilots used combinations of OpenAI and Google judges.

Throughout this article, scores are model-judge evaluations on a 0–100 scale, not accuracy percentages. The Code pilots also used model judgments rather than a separate deterministic execution harness, so their scores should not be read as unit-test pass rates.

How the results inform routing

The benchmark work exists to inform model-selection policy.

When Elektric receives an ordinary text request, it identifies the kind of work and its expected difficulty. That request then maps to a routing policy. The policy may contain one model or a pool of models, and candidates are filtered for production eligibility and the capabilities required by the request.

Benchmark results help determine which models are reasonable candidates for those policies, but benchmark leaders are not automatically production routing choices. The current routing policy differs from the score leader in several cells, while other cells include the benchmark leader within a broader candidate pool.

That separation lets us consider more than one dimension when making a routing decision.

A model can have the highest average judged score but cost substantially more. Another can be cheaper but have more severe failures. A third may perform well but lack a required capability.

The benchmark is evidence. Routing is the operational decision made from that evidence.

What happens when new models arrive

When a new model launches, we can evaluate it against the models already being considered for the same kinds of work.

The useful comparison is not simply whether the new model scores well in isolation. We want to know whether it offers a better tradeoff than the models already underneath the application.

Can it deliver higher judged quality at a similar cost? Similar quality at a lower cost? Does it complete more tasks? Does it support capabilities the existing choices do not?

If the evidence supports a better tradeoff, the routing policy can be updated without requiring developers to change their integration.

That is the practical reason we maintain these evaluations continuously: the models underneath an application can change even when the application itself does not.

What we learned

For the tasks we tested, no single model led every category and difficulty level. The useful question was which model offered the right tradeoff for a particular request, and how confidently our evaluations supported that choice.