Structured-output frameworks
Jev API vs Instructor and Outlines: Constrain the Output, or the Decision?
Instructor and Outlines make a model return valid structured data. Jev removes the model. How the three differ in cost and failure mode, and when each wins.
Instructor and Outlines are the two well-known answers to “how do I stop my model returning invalid JSON.” Jev is an answer to a different question — “why is a language model involved in this decision at all?” — and the difference in mechanism is what makes the comparison interesting rather than competitive.
If you already use either library, you are further along than most teams evaluating Jev. You have already decided that classifications should be typed. The only open question is whether they should also be generated.
The short answer
Instructor and Outlines both keep the language model and constrain what it emits: Instructor by validating the output and retrying, Outlines by preventing invalid tokens from being sampled in the first place. Both work. Both still pay for a generation.
Jev removes the generation. A decision that selects from a fixed set of options does not need a model that can write, and once you stop generating text you stop paying for it and stop waiting for it. If your classifications are high-volume, that is the whole argument. If they are not, the libraries you already have are fine.
Three mechanisms, not three products
| Instructor | Outlines | Jev API | |
|---|---|---|---|
| Approach | Validate output, retry on failure | Mask invalid tokens during generation | Model returns only declared types — no generation |
| Still runs an LLM | Yes | Yes | No |
| Reported success rate | 95–98% (with retry) | 99.9%+ | 100% well-formed by construction |
| Works with API providers | Yes, 15+ | No — self-hosted only | It is an API |
| Cost basis | Model’s generation cost + retries | Self-hosted inference | $0.042 / MTok input, output free |
| Typical latency | Model latency + retry latency | Model latency, slightly reduced | 70–500 ms |
| Can it return prose or a rationale? | Yes | Yes | No |
| License | MIT | Apache 2.0 | Commercial API |
Read the “still runs an LLM” row as the load-bearing one. Everything else is a consequence.
Instructor: validate and retry
Instructor decorates a provider client, takes a Pydantic model as response_model, and guarantees the returned object matches it — by attempting the parse, and on failure feeding the validation error back to the model and asking again. It supports 15+ providers, offers partial streaming via Partial[Model], and is the default choice for Python teams not hosting their own models. Reported recovery rates exceed 95% for schemas under about 15 fields.
Two properties matter when comparing against Jev:
The retry is real cost and real latency. A 95% first-attempt success rate means one call in twenty is paid for twice, and the second call is issued only after the first has finished. At scale that is not a rounding error; it is a budget line and a p95 tail.
Its parser is strict by design. Instructor expects clean JSON. Markdown-fenced output, or a model that reasons before answering, is exactly the case it handles worst — the gap that libraries with more forgiving parsers were built to fill.
None of this is a criticism. Instructor is reliable at the job it was given. The job is just more expensive than it looks, because the model has to produce tokens before anyone can check them.
Outlines: constrain generation
Outlines compiles a schema into a finite-state machine and masks every token that would make the output invalid, so the model physically cannot emit malformed data. That is why its reported success rate is 99.9%+ rather than 95% — there is no retry, because there is nothing to retry.
The trade is access. Constrained generation needs logits, so Outlines runs against self-hosted inference — transformers, vLLM, llama.cpp — and does not work with API-only providers. Its own documentation and third-party comparisons also place it behind XGrammar for high-throughput production serving, positioning it for experimentation and custom grammars.
For teams already serving their own models, Outlines is the strongest guarantee available: the output is valid because invalid output is unreachable. That guarantee covers syntax. It does not cover whether the label is correct, and it does not make inference cheap.
Where the cost actually goes
The three approaches differ most in what happens when something goes wrong, and that is the comparison tables tend to omit.
- Instructor: failure costs a second generation, plus the latency of discovering the failure.
- Outlines: failure is structurally impossible, but the whole inference bill remains.
- Jev: failure is a wrong-but-valid option, priced the same as a right one.
Notice that none of the three eliminates the possibility of the answer being wrong. Instructor’s retry loop catches type errors, not mistakes. Outlines guarantees well-formedness, not correctness. Jev can be confidently wrong, as its CEO conceded on the Hacker News launch thread. Anyone who has shipped a classifier knows this is the ordinary problem of classification rather than a product defect — but it is worth being clear that schema validity was never the hard part.
The hard part, in all three cases, is knowing whether 0.9 means 0.9. Jev is the only one of the three where that is a training objective rather than a downstream calibration exercise; the Claude comparison has the measurements, and the classifier comparison covers what it costs to get there the hard way.
Choosing
Keep Instructor if you are calling an API provider, your volume is modest, the call sometimes needs to return a rationale alongside the label, or you want one interface across many providers. It is the lowest-friction option and it is already working. For how a provider’s own structured-output mode compares with either library, see Jev API vs GPT.
Keep Outlines if you serve your own models and need a hard structural guarantee. Nothing else gives you that.
Consider Jev if the call is a bounded decision — a label, a route, a score, a yes/no — running at volume, and the generation is overhead you are paying for out of habit. The migration is small: declare the options, send the state, read the typed answer — the three primitives are the entire interface.
Do not migrate if the answer must be explained. Neither Instructor nor Outlines is required for that, but the language model underneath them is, and Jev has nothing to offer there.
One structural point that the libraries make easy to miss: because Instructor and Outlines treat the model as untrusted, everything downstream is written defensively. Jev’s answers arrive constrained, so the defensive code has nothing to catch. That is a smaller codebase and one fewer class of production incident — though not, as the failure discussion above makes clear, a smaller set of things that can go wrong.
Where to next
- Jev API vs GPT — the same argument against a frontier model rather than a framework.
- Jev API vs Claude — including the calibration measurements, if confidence thresholds are what you are after.
- Jev API vs traditional ML classifiers — if the honest alternative is not a library but a model you train.
FAQ
Do Instructor or Outlines replace Jev?
No, and they do not compete directly. Instructor and Outlines exist to make a generative model produce valid structured output — Instructor by validating and retrying, Outlines by masking invalid tokens during generation. Both still run a language model. Jev has no generation step to constrain, so it produces typed answers without one. They solve the same symptom from opposite directions.
Are Instructor and Outlines reliable enough that I do not need Jev?
They are reliable at producing well-formed output and unreliable at producing fast, cheap decisions. Instructor reports 95–98% success rates with retries, and Outlines claims 99.9%+ through constrained generation. Both numbers describe whether you get a parseable object, not whether the answer is right, and neither addresses latency or per-call cost. If your pipeline is correct and affordable, there is nothing to fix.
Can I use Jev together with Instructor or Outlines?
You can, but there is rarely a reason to. Instructor’s value is validation and retry around an unreliable generator; Jev’s answers are already constrained to your declared types, so wrapping it adds a dependency without removing a failure mode. The one case that makes sense is a mixed pipeline where some calls still go to a language model and you want one interface across both.
Does Outlines work with API providers like OpenAI or Anthropic?
No. Outlines works by masking invalid tokens during generation, which requires access to the model’s logits — so it runs against self-hosted inference through transformers, vLLM or llama.cpp. It cannot constrain an API-only provider. Instructor is the opposite: it wraps 15+ provider clients and works wherever you can call an API, which is why it is the more common choice for teams not hosting their own models.
What does Jev cost compared with running Instructor over an LLM?
Jev bills $0.042 per million input tokens with output free. Instructor over a frontier model bills the model’s full generation cost — including the retries it triggers when validation fails — so the comparison is Jev against a metered generation, plus whatever retry rate your schema produces. At low volume the difference is noise; at millions of classifications it is the reason teams switch.
More comparisons
- Frontier language models Jev API vs GPT-5.6 Terra Not rival products — Jev replaces the classification call, GPT keeps the reasoning call. The stack usually wants both.
- Frontier language models Jev API vs Claude Sonnet 5 Claude reasons about a decision and can explain it. Jev makes the decision and cannot. Most agent stacks need Claude doing the first job more than they need it doing the second.
- Traditional ML classifiers Jev API vs Traditional ML classifiers A trained classifier is cheaper per call, explainable and yours. Jev is faster to start, survives label changes, and needs no data. The deciding question is whether your labels are stable.