Stay connected with KayaToday, follow us on Instagram and Facebook for the latest news and reviews delivered straight to you.
There is a quiet inefficiency running through most enterprise AI deployments today, and it has nothing to do with hallucinations or model safety. It is something more mundane: companies are using large language models to generate paragraphs of text, then throwing almost all of that text away, just to extract a single label. Is this support ticket urgent? Does this document belong in this context? Should this transaction be flagged? The answer is always one of a handful of options, yet the system burns through tokens producing prose to arrive at it.
TypeSafe AI’s Jev, released in early access on September 15, is the most prominent recent challenge to that pattern. Jev does not generate text at all. It takes application state and a set of typed questions, then returns choices, scores, and yes/no answers with calibrated probabilities attached. Response times run between 70 and 500 milliseconds, and pricing sits at $0.042 per million input tokens, according to TypeSafe. Output tokens are free. Within days of launch, Jev had integrations with agent frameworks including Pydantic AI and LangChain, TypeSafe had to pause its signup queue, and a wave of open-source alternatives began appearing on GitHub and Hugging Face.
The Hidden Cost of Asking a Writer to Be a Judge
To understand why this matters, it helps to understand what language models are actually optimised for. Pretraining teaches them to predict the next token in a sequence. Post-training with reinforcement learning from human feedback teaches them to produce responses that people prefer. Both stages produce strings of text. Enterprise automation, by contrast, needs typed values, predictable latency, stable cost per API call, and outputs that can be scored against ground truth in a repeatable way.
When a generative model is used to make a decision, even when it gets the decision right, the application pays for token-by-token decoding. It then has to parse and validate the output, and handle the cases where the model drifts from the expected format, adds unsolicited commentary, or refuses to answer. There is also a calibration problem that is easy to underestimate. Generative models tend to be poorly calibrated when asked how confident they are in a decision. A system that cannot reliably estimate its own uncertainty cannot decide when to escalate a borderline case to a human reviewer, which is precisely the kind of judgment that matters most in production.
TypeSafe’s launch materials make this argument directly, contrasting sequential token generation with Jev’s parallel sampler and contrasting self-reported confidence with calibrated probabilities. The company says its headline figures of roughly 194 times faster and 445 times cheaper come from its own workflow evaluations and likely sit at the high end of real-world gains, which is an honest caveat worth noting.
Why Classification Is Suddenly Viable Again
Classification is not a new idea. By 2019, researchers had already demonstrated zero-shot classification using natural language inference models. The Yin et al. paper repurposed those models to test candidate labels against an input, and facebook/bart-large-mnli became a common default for classifying text against labels supplied at request time. The problem was brittleness. Those models struggled outside their training domains, so many real deployments still required labeled data and task-specific training runs, which added cost and time.
What has changed is the quality of pretraining. BERT, the foundational model from 2018, was trained on roughly 3.3 billion words. ModernBERT, released in December 2024, was trained on 2 trillion tokens with an 8,192-token context window. Stronger pretrained models can now handle a much wider range of classification questions defined at request time, reducing the need to train a separate model for every new task.
The open-source reproductions of Jev make this visible because their internals are public. SemIf trains no model at all. It reads option probabilities directly from a frozen Qwen3.5-4B, and its developers measured about one second for 21 questions against more than five seconds for generated JSON, with the two methods agreeing on 18 of the 21 choices. Laya builds a decision model on a 421-million-parameter ModernBERT-large backbone and answers in roughly 33 to 40 milliseconds per question on an Nvidia T4. Other projects apply LoRA adapters to Qwen3.5 or run on DiffusionGemma. The implementations differ, but they share the same core idea: using stronger pretrained models to make decisions without generating text.
The reproductions also surface the limits honestly. Laya’s developers report that its base checkpoints score near chance on their typed-decisions benchmark, with substantially better results only after fine-tuning. TypeSafe has disclosed little about Jev’s internals beyond describing a new architecture, a parallel sampler, and a post-training method it calls Reinforcement Learning for Calibrated Decisions, or RLCD, which optimises the model to return probabilities that match how often it is actually correct. Making calibrated uncertainty its primary training objective is TypeSafe’s central technical bet, and whether it holds up at scale remains to be seen.
The Playbook That Never Left Large ML Teams
Jev is the most visible example of a broader pattern: builders taking frontier-quality pretraining and applying pre-ChatGPT deep learning tactics on top of it. Distillation and small specialist models never really disappeared from companies with mature machine learning teams. They just became less visible once the public conversation shifted to large generative models.
LinkedIn’s approach illustrates how far this can go. Erran Berger, who led the work as VP of product engineering before becoming CTO, described on VentureBeat’s Beyond the Pilot podcast how the company fine-tuned a 7-billion-parameter teacher model on a detailed product policy document, then distilled it through a 1.7-billion-parameter intermediate model down to a 0.6-billion-parameter student that serves production traffic. “There was just no way we were gonna be able to do that through prompting,” Berger said. The process became a repeatable recipe reused across LinkedIn’s AI products. Shopify has taken a similar approach with an internal platform that lets R&D teams distill a frontier model into a fine-tuned open model for a single subtask in roughly a day, with evaluations built in. Farhan Thawar, Shopify’s VP and head of engineering, said the distilled models run anywhere from 2 to 30 times cheaper and faster than the frontier models they learned from, and in some narrow tasks outperform them.
The common thread is that generation gets reserved for tasks that genuinely require it: drafting, summarising, writing code. Routing, tagging, triage, and scoring go to cheaper, faster models that meet the application’s accuracy and reliability requirements without the overhead of text generation.
What Teams Should Actually Do With This
The practical starting point is an audit of existing LLM traffic, sorting API calls into two buckets: generation tasks and decision tasks. Developer Flavio Copes has published a worked example of the math. A company spending $10,000 a month on AI, with $6,000 of that going toward decisions, would save $5,700 a month if those decisions moved to a path costing 5% as much. Actual savings depend on how much traffic can realistically move and what the replacement infrastructure costs to run.
The caveats are real and worth taking seriously. Jev is in early access, accepts text only, and currently serves from the West Coast of the United States. TypeSafe has said it will take time to establish whether its pricing is sustainable. Pydantic’s documentation notes that text in the input written to steer Jev’s answers can succeed, meaning Jev-based guardrails should sit alongside deterministic checks rather than replace them. The open-source alternatives remove vendor dependency but hand teams the work of hosting, fine-tuning, calibration, and ongoing monitoring.
For engineering teams in Malaysia and Singapore building on top of AI APIs, the latency and cost implications are particularly concrete. API calls routed through US-based endpoints already carry geographic overhead, and paying generative model rates for decisions that need only a label compounds that inefficiency. The broader lesson from Jev’s reception is that the field is correcting a category error that crept in during the generative AI boom: not every AI task is a generation task, and the engineers who understood that distinction before 2022 are finding their skills relevant again. The teams best positioned for the next phase of enterprise AI are those that can pair that older expertise with the stronger pretrained models now available, building automation that is cheaper, faster, and far easier to audit than the text-generating alternatives it replaces.
Read More: Microsoft’s Copilot Overhaul Bets on Persistent AI Agents and App Hosting, But Pricing Stays Murky