Classification Never Left. It Just Got a New Home in LLMs.
Jev and Laya are new models built only for fast, structured decisions: choice, score, or true and false, not text generation. They read input once and return a label.
Join the DZone community and get the full member experience.
Join For FreeBefore large language models became the face of AI, most of machine learning was about one thing: sorting stuff into buckets. Is this email spam or not spam? Is this transaction fraud or normal? Which support team should get this ticket? That is classification, one of the oldest and most useful jobs in supervised learning.
Then LLMs arrived, and everyone started asking chat models to do this same sorting work. It works, but it is a bit like hiring a novelist to fill out a form. The novelist can do it, but you are paying for a full essay when you only needed a checkbox ticked.
Two new tools, Jev and Laya, are trying to bring classification back to its roots, built the way it should be for software pipelines rather than conversations. This piece goes deeper than a quick overview: what these tools actually are, how to call them, what is running under the hood, and where they sit next to the rest of the LLM stack.
A Quick Refresher: What Classification Actually Is
In supervised learning, you show a model many examples where you already know the right answer. Emails labeled spam or not spam. Tickets labeled billing, technical, or sales. The model studies the patterns and learns to predict the label for new, unseen examples. The output is not a paragraph. It is a label, sometimes with a probability attached, like "billing: 92 percent confident."
| Classic ML term | Plain meaning |
|---|---|
| Training data | Past examples with known correct answers |
| Feature | A signal the model uses to decide, like words in an email |
| Label | The bucket the example belongs to |
| Confidence score | How sure the model is about its answer |
| Model | The thing that learned the pattern from the data |
Where LLMs Entered the Picture, and Where They Struggle
When ChatGPT and similar models showed up, people realized you could just describe the categories in plain English and ask the model to pick one. No training data needed, no feature engineering, just a good prompt.
That is genuinely useful. But it comes with a cost. A chat-style LLM answers by writing one word at a time, checking each word against everything before it, then writing the next word. Even if the final answer is a single word like "billing," the model still goes through this slow, token-by-token process to get there. For a single question, that is fine. For a pipeline that needs to sort a million support tickets a day, it adds up in both time and money.

Notice the extra step at the end too. The output is text, so your software has to read that text and turn it back into a structured decision. Sometimes the model rambles, adds a caveat, or phrases things slightly differently each time, and now your parsing code breaks. Tools like Mellea from IBM Research take a swing at the same problem, wrapping LLM calls with strict types and retry rules so output is always one of a fixed set of values, a pattern covered in Open-Source LLM Tools Worth Your Time. Jev and Laya push the same idea a step further by removing text generation from the equation entirely.
Meet Jev
Jev comes from TypeSafe AI. It does not generate text at all. You send it a piece of text and a set of typed questions, and it returns a choice, a score, or a probability for each one. TypeSafe calls this a "System One Model," a name borrowed from Daniel Kahneman's split between System 1, the fast instinctive mind, and System 2, the slow deliberate one that a chat model resembles, thinking token by token to produce an answer. Jev launched into early access in September 2026. The team includes Diogo Almeida, a former OpenAI researcher who co-authored the InstructGPT paper behind RLHF training, and the company raised a $40 million seed round led by DCVC.
Jev is not a package you pip install and run locally. It is a hosted API, and it has already been wired into several developer platforms.
Calling Jev directly, through LiteLLM's pass-through endpoint:
curl -X POST 'http://0.0.0.0:4000/typesafe/v1/systemone' \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"
}
}
}
}'
The response comes back as structured JSON, not a sentence:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"probabilities": {"billing": 0.08, "technical": 0.85, "sales": 0.07},
"confidence": 0.82
}
},
"usage": {"input_tokens": 312, "output_tokens": 48}
}
Calling Jev through the Vercel AI Gateway in Python:
import os
import requests
response = requests.post(
"https://ai-gateway.vercel.sh/typesafe/v1/systemone",
headers={
"Authorization": f"Bearer {os.environ['AI_GATEWAY_API_KEY']}",
"Content-Type": "application/json",
},
json={
"model": "typesafe-ai/jev",
"state": "I was charged twice for my subscription.",
"questions": {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for money back?",
}
},
},
)
print(response.json()["answers"])
Jev is also reachable through Opper and OpenRouter, both of which keep TypeSafe's original request and response shapes so existing client code barely changes when you switch provider. One neat downstream use: LiteLLM uses Jev internally to decide whether an old tool result in a long agent conversation is still relevant, and drops it if Jev scores it below a threshold, trimming context without involving the main model at all.
Meet Laya
Laya comes from Convai Innovations, and unlike Jev, it is fully open-weight. You can run it yourself, on a GPU, on a Mac with Apple Silicon, or through the Hugging Face Hub.
Laya ships three checkpoints, and the choice between them is mostly about language coverage and speed:
| Checkpoint | Encoder | Parameters | Context | Best for |
|---|---|---|---|---|
| laya | ModernBERT-large | 421M | 512 tokens | English |
| laya-multilingual | mmBERT-base | 322M | 1024 tokens | 100+ languages, roughly 2x faster |
| laya-typed-decisions | ModernBERT-large | 421M | 1024 tokens | longer typed-decision workflows |
Installing and running Laya:
pip install laya
import laya
agent = laya.load("convaiinnovations/laya")
result = agent.predict(
{
"subject": "Duplicate charge on invoice 4411",
"body": "We were billed twice for March. Please refund the duplicate.",
},
{
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "invoices, payments, refunds",
"technical": "bugs and outages",
"sales": "pricing",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["not urgent", "soon", "blocking"],
},
"churn_risk": {
"type": "noul",
"instructions": "Does the user threaten to cancel?",
},
},
)
print(result["answers"]["department"]["choice"])
If your traffic mixes languages, Laya includes a small router that picks the right checkpoint per request instead of you hardcoding it:
from laya import Router
router = Router() # lazy-loads only what a request needs
router.predict({"body": "I was charged twice"}, questions) # -> laya
router.predict({"body": "मुझसे दो बार शुल्क लिया गया"}, questions) # -> laya-multilingual
On a Mac with an M-series chip, there is also a native MLX build that skips PyTorch entirely and runs fully offline once the weights are downloaded once:
pip install laya-mlx
import laya_mlx as laya
agent = laya.load("aac6fef/laya-mlx")
result = agent.predict(
"I was billed twice. Please refund the duplicate today.",
{
"department": {
"type": "choice",
"instructions": "Which department should handle this request?",
"criteria": ["billing", "technical", "sales"],
}
},
)
Laya is trained with reinforcement learning against strictly proper scoring rules, a training method where the model can only earn the best reward by reporting its true confidence rather than an inflated one. It never generates text, so there is nothing for your code to parse and nothing for it to hallucinate.
What Is Actually Running Under the Hood
It helps to open the hood a little, because the underpinnings of Jev and Laya are not some brand new invention. They are a clever remix of parts that have existed for years.
The encoder, not the decoder. A chat-style LLM like GPT or Claude is built mostly from decoder blocks, the part of the transformer designed to keep generating the next word based on everything written so far. Jev and Laya lean on the other half of the original transformer design, the encoder. An encoder reads the whole input once and builds a rich understanding of it in a single pass, without needing to predict what comes next. Laya's checkpoints sit on ModernBERT and mmBERT, both encoder-only models, which is exactly why they answer in milliseconds instead of seconds.
Small and dense, not huge and sparse. These models stay small on purpose, in the hundreds of millions of parameters rather than the hundreds of billions you see in frontier chat models. A smaller model with a narrow job — answering a typed question about a piece of text — needs far less computing muscle than a model that has to be ready to write a poem, debug code, and hold a conversation in the same session.
Trained differently, and judged for honesty, not eloquence. Jev is trained mostly on synthetic examples built for the choice, score, and true-or-false questions it needs to answer, which is closer to how a classic supervised learning classifier trained on a labeled dataset than to how a general chat model learns from scraped internet text. Laya goes further on the confidence side: its reinforcement learning setup only rewards the model for reporting honest probabilities, the same idea behind Brier scores and log loss, tools statisticians have used for decades to judge whether a forecaster is well calibrated, not just whether it is right.

Put together, you get a model that behaves less like a chatbot, and more like a fast, well-calibrated cousin of the classifiers data scientists have been building since long before anyone said: "prompt." A lot of what looks new in 2026 is really an old, well-tested idea wearing a transformer-shaped coat, a theme that shows up again in Shingling in the Generative AI Era, where a decades-old text similarity technique turns out to still be quietly useful for spotting near-duplicate content in the generative AI age.
The Three Answer Types, in Detail
Both tools boil every question down to one of three shapes.
| Question type | What it returns | Classic ML equivalent | Example use |
|---|---|---|---|
| Choice | Selected label, a probability per option, overall confidence | Multi class classification | Ticket routing, intent detection |
| Score | Expected level on an ordinal scale you define, plus a distribution | Ordinal regression | Urgency, frustration level, severity |
| Noul (true or false) | A single calibrated probability from 0.0 to 1.0 | Binary classification | Phishing detection, churn risk, spam filtering |
If you have ever trained a classifier with scikit-learn or built a logistic regression model, this table should feel familiar. What changed is how the model gets to the answer and how easily it understands raw, unstructured text without you having to hand-engineer features first.
Why the Confidence Number Matters So Much
A model can be right 90 percent of the time and still be badly calibrated if it says "99 percent confident" every single time. In production, that difference decides whether you trust an automated decision or send it to a human. If a ticket comes back "billing, 55 percent confident," a good system routes that one to a person for a second look instead of trusting it blindly, the same way a bank flags a borderline fraud score for manual review rather than auto-approving or auto-rejecting it. This is the same "trust but verify" thinking behind guardrail and safety layers in the wider agent stack, covered in the security section of Open-Source LLM Tools Worth Your Time, where a model-level judge checks risky outputs before they reach a user.
Speed and Cost, With Real Numbers
Convai Innovations published a head-to-head benchmark of Laya against Jev's own published figures. Worth reading with the usual caution that one side ran its own comparison, but the gap is large enough to be worth noting.
| Metric | Jev (published) | Laya (fine-tuned checkpoint) |
|---|---|---|
| Single question, typical latency | ~400 ms average (range 70 to 500 ms) | 38.4 ms (p95: 42.1 ms) |
| 10 questions, batched | ~1,500 ms serial | 156.0 ms |
| 50 questions, batched | multi-second, rate limited | 721.4 ms |
TypeSafe's own figures put Jev's accuracy at around 68 percent on its workflow evaluations, priced at $0.042 per million input tokens with no charge for output tokens, since it never generates any. That still lands it well to the left of general-purpose LLMs on a cost chart, just not as far left as a small open-weight encoder running on your own GPU. Treat both sets of numbers as a starting point for your own testing rather than a final verdict, since neither has a large body of independent benchmarks yet.
A Simple Decision Guide
| If your task is... | Reach for... |
|---|---|
| Writing an email, summarizing a document, holding a conversation | A regular chat style LLM |
| Sorting tickets, tagging content, scoring risk, routing requests at high volume | A classification model like Jev or Laya |
| You need it hosted, with no infrastructure to manage | Jev, or Laya through a hosted endpoint |
| You need it self-hosted, open weight, or fully offline | Laya, including the MLX build for Apple Silicon |
| A mix, reading a document then deciding who handles it | Both together, LLM for reading, classifier for the decision |
| Deciding how several agents or tools fit into one system | Worth reading up on agent framework design first |
That last row is the most common setup as teams move from single model calls to full pipelines. A Field Guide to AI Agent Frameworks and Loop Engineering: The Layer After Prompt, Context, and Harness Engineering go into how these pieces, agents, tools, and fast little classifiers like Jev and Laya, get wired together into something that runs reliably in production rather than just in a demo.
Where These Tools Still Fall Short
Worth being honest about the rough edges before you commit to either one.
- Neither tool has a large body of independent, third-party benchmarks yet. Most published numbers come from the vendors themselves.
- Jev is closed and hosted only. If you need full data control or offline operation, Laya is currently the only option of the two.
- Both are built for short, well-defined questions. Neither replaces an LLM for open-ended reasoning, multi-step tasks, or free-form writing.
- Calibration claims are strongest on the benchmarks each vendor chose to publish. Test on your own data before trusting the confidence numbers in a high-stakes decision.
The Bigger Picture
None of this is really new math. Classification with confidence scores is one of the oldest ideas in machine learning. What Jev and Laya are doing is repackaging that old, reliable idea with a modern transformer brain underneath, one that can read messy real-world text the way an LLM can, but answer the way a classifier always has: fast, structured, and with a number attached that tells you how much to trust it.
As more teams build pipelines with LLMs doing the heavy thinking and lightweight classifiers doing the quick sorting, expect more tools like this to show up. The chat model got all the attention for the last few years. The classifier, quietly, is coming back for its turn.
More from me on the pieces this connects to:
Opinions expressed by DZone contributors are their own.
Comments