Skip to content
J JevAPI.dev

Model routing and cascades

Jev API vs RouteLLM and Not Diamond: The Router or the Decision Behind It

RouteLLM and Not Diamond pick which model answers a request. Jev API can be the thing that decides. How learned routers and typed decisions differ — and fit.

Independent reference · not affiliated with TypeSafe AI

This is the one comparison in the set where the products are not really competing. RouteLLM and Not Diamond answer “which model should handle this request?” Jev answers questions. One of the questions it can answer is “which model should handle this request” — which makes it a candidate to sit inside a router rather than replace one.

Getting the layers straight is most of the work here, and it is where most of the confusion in this space comes from.

The short answer

RouteLLM is an open-source framework with a learned selector; Not Diamond is a managed selector API. Both exist to route traffic between a cheap model and an expensive one. Jev is a model that returns typed answers with calibrated probabilities, and routing is one of the several things you can point it at.

If you already run a router, Jev does not automatically make it redundant. What it offers is a different selector — instruction-driven rather than trained on preference data, with calibrated confidence attached to the decision. Whether that is better depends on whether you have the labelled traffic that makes a learned router worth training.

Three layers, not two products

Most “Jev vs router” confusion comes from collapsing these:

Layer What it does Examples
Gateway Executes the call, handles auth, retries, fallbacks, provider formats LiteLLM, Portkey, OpenRouter, Kong
Selector Decides which model the request goes to RouteLLM’s learned router, Not Diamond, a Jev question
Model Produces the answer GPT, Claude, Jev

Not Diamond is explicit that it is a recommender, not a gateway — it tells you where to send a request and your application executes it. RouteLLM ships a selector and is described across the literature as a reproducible research baseline rather than a production control plane, needing a gateway around it.

Jev is not a router at all. It is a model that can be asked routing questions. Putting it in the selector slot is a design choice, not a product category.

RouteLLM: a learned binary decision

RouteLLM, from LMSYS, makes one call: route this query to the weak model or the strong one. Its approaches include matrix factorisation, weighted ranking, a BERT classifier and causal-LLM routing, trained on preference data. Published results claim up to 85% cost reduction while retaining 95% of GPT-4 performance on MT-Bench, with smaller savings on MMLU and GSM8K.

The constraints are operational as much as technical:

  • It is a framework, not a service. You own evaluation, model updates, threshold calibration, hosting and retraining. One comparison notes the repository has not kept pace with commercial tooling.
  • Its text encoder caps input at 8,192 tokens, which excludes long-context traffic from routing.
  • It depends on the OpenAI embedding API, which shows up as added latency — RouteLLM ranked eighth of twelve on latency in the RouterArena benchmark.
  • The decision is binary by construction. Weak or strong. Adding a third tier means changing the framework.

That last point is where a typed decision differs structurally, and it is worth separating from any claim about which is more accurate.

Not Diamond: a managed selector

Not Diamond sells the selector as a service: $0.05 per million tokens routed, 10–100 ms recommendation latency, a pre-trained router or a custom one trained on your prompts and evaluation criteria, with SOC 2 and ISO 27001 and ZDR options for enterprise. It powered OpenRouter’s original openrouter/auto, which OpenRouter has since replaced with its own rankings.

Its market position is the middle ground — more automatic than a plain aggregator, less operational burden than self-hosting RouteLLM. The independent check on its claims is less flattering: RouterArena placed it twelfth by average cost and eighth on arena rank, largely because it frequently selects expensive models, though its accuracy on long-context tasks was high.

What the benchmark evidence actually supports

RouterArena’s finding is worth stating plainly because it cuts against the premise of the whole category: no router leads across all metrics, and every one tested falls short of an oracle, over-relying on strong models and missing chances to defer to cheap ones. Systems built for the task — vLLM-SR and CARROT — reached roughly 35% lower cost with under 2% accuracy loss, which is a reminder that a learned binary router carries a sharp accuracy-cost trade rather than a free lunch.

Two practical conclusions follow. First, “which router is best” has no answer; the products optimise different workloads. Second, the more reliable pattern in 2026 is a gateway underneath and a selector on top, applied only once static rules stop being enough.

Where Jev fits, and where it does not

Jev can occupy the selector slot, and in some ways it fits the shape better than a learned binary classifier:

  • N-way rather than binary. A choice question selects among up to 255 declared options, so a three-tier or five-tier cascade is one question rather than a framework change.
  • Confidence included. Every answer carries a probability distribution and a confidence, so the router can threshold on its own certainty — escalating when the routing decision itself is uncertain.
  • No embedding round-trip. Questions are answered in 70–500 ms without depending on a separate embedding service.
  • Multiple questions at once. Routing tier, required capability and estimated risk can be asked in a single parallel request for near-zero added latency.
  • Replaceable by instruction. Changing the routing policy means editing the criteria you declare, not retraining a model.

And where a learned router is genuinely better:

  • It learns from your traffic. RouteLLM fits preference data; Not Diamond will train on your prompts and evaluation criteria. Jev is zero-shot and never improves from your logs. If the boundary between “cheap model is fine” and “needs the strong one” is subtle and you have examples, a learned selector is the right instrument — and this is the single strongest argument for that approach.
  • It is a system, not a decision. Thresholds, fallbacks, provider formats, retries and observability all come with the product. Asking Jev which model to use still leaves you to build the part that uses the answer.
  • It is priced per token routed on a specific job. Not Diamond’s $0.05 per million tokens routed and Jev’s $0.042 per million input tokens sit in the same range, but one bills for a routing recommendation and the other for an arbitrary typed question — comparable numbers attached to different work.

How they fit together

Three patterns are worth considering, in ascending order of how much you want to build:

Jev as a gate in front of the gateway. The simplest, and what most teams actually ship. Jev answers “does this need a frontier model at all?” and the gateway handles whichever answer comes back. No router involved — the Claude comparison describes this one from the other side, as an escalation gate inside an agent loop.

Jev as the selector inside an existing router. RouteLLM’s selector is replaceable and the framework is Apache 2.0, so a typed decision can sit where its BERT classifier does. You take on the integration, and you give up the trained weights.

Learned router on top, Jev for the adjacent decisions. The router handles tier selection; Jev handles the questions a router does not answer — is this output safe to return, does this claim survive its evidence, which tool applies here. This keeps each tool on the job it was actually built for, and it is the pattern that matches how most of these systems are deployed in practice.

How to decide

  • You have labelled traffic and a subtle routing boundary → a learned selector, RouteLLM or Not Diamond. Jev cannot learn what you already know.
  • Your routing policy is expressible as rules you can write down → a typed decision is cheaper to operate and easier to change.
  • You need routing plus the operational surface → buy the system, and consider Jev for the decisions around it.
  • You want confidence carried with the routing decision → that is the specific thing a learned binary router does not give you, and a calibrated probability does.
  • You are still writing static if-statements → neither is the next step. Reach for a learned selector only when static rules show diminishing returns; that is where the published guidance converges.

Where to next

FAQ

Is Jev an alternative to RouteLLM or Not Diamond?

Only partly. RouteLLM and Not Diamond are routing systems — they ship a learned selector, the serving code around it, and in Not Diamond’s case a managed API. Jev is a model that answers a bounded question. It can occupy the selector slot, but it does not replace the gateway, the fallback logic or the operational surface those products provide around it.

How does Jev compare on price with Not Diamond?

Similarly, which is a coincidence worth noting. Not Diamond charges $0.05 per million tokens routed; Jev bills $0.042 per million input tokens with output free. Both charge per token rather than per request or per seat. The real difference is what each one does for the money: Not Diamond recommends a model from your catalogue, while Jev returns a typed answer to whatever question you ask it.

Which is more accurate, a learned router or Jev?

Unknown, and the published evidence does not settle it. The RouterArena benchmark found no router leading across all metrics, with every entrant falling short of an oracle; Not Diamond ranked last on average cost because it favours expensive models, while RouteLLM ranked poorly on latency. Learned routers are trained on your traffic, which is a real advantage Jev does not have. Neither family has published a head-to-head against the other.

Can I use Jev as the routing decision inside RouteLLM?

In principle yes — RouteLLM is an Apache-2.0 framework whose selector is replaceable, and a typed decision is exactly what a selector produces. In practice you would be taking on the integration work yourself, and RouteLLM’s own routing logic is trained on preference data rather than written as instructions. The simpler architecture is often Jev as a gate in front of a gateway, leaving the router out.

What does a learned router do that Jev cannot?

Learn from your traffic. RouteLLM’s approaches are trained on preference data, and Not Diamond will train a custom router on your own prompts and evaluation criteria. Jev is a zero-shot model driven by instructions you write — it never improves from your production logs. If your routing boundary is subtle and you have labelled examples to fit it, that is a genuine reason to prefer a learned approach.

More comparisons

All comparisons →