Rethinking AI: From Language Models to Calibrated Decision Systems
ChatGPT broke Diogo Almeida’s heart.
Almeida, a former OpenAI researcher instrumental in developing the chatbot and pioneering reinforcement learning from human feedback (RLHF)—a technique largely responsible for today’s AI advancements—found himself disappointed despite the model’s impressive capabilities.
“We have lightning in a bottle, and yet it is not useful,” Almeida told TechCrunch. “I’ve been battling that problem since then. It took me a while to come to the conclusion: The problem is we are optimizing for human language … We have been super good at human language for four years, but it’s not useful for automation because computers speak a different language.”
From OpenAI to TypeSafe AI: Pursuing a New AI Paradigm
Two years ago, Almeida left OpenAI to found TypeSafe AI, a startup dedicated to addressing this fundamental disconnect. This week, the company unveiled a new transformer-based model called Jev. Unlike traditional large language models (LLMs), Jev does not produce text outputs. Instead, it generates probabilities, or what TypeSafe AI terms “calibrated decisions.”
By moving away from language, Jev achieves remarkable benefits: it is exceptionally fast and cost-effective, and its outputs cannot hallucinate because users define the outputs upfront. Input tokens are billed by the billion rather than the million, and output tokens are free, making it a highly economical choice for developers.
Image Credits:TypeSafe AI
Developer Enthusiasm and Real-World Applications
Developers have shown strong interest in Jev, with demand so high that the company briefly could not serve new API users. Jev’s primary appeal lies in software automation, offering developers a cheaper, faster, and more robust way to integrate intelligence into their workflows.
For instance, Pranit Sharma, a software engineer at Vercel—a company focused on agentic infrastructure—shared that switching from OpenAI’s ChatGPT Luna 5.6 to Jev for command safety classification resulted in performance improvements of five to 18 times faster and more accurate outcomes.
Similarly, Bryo AI CTO Nikhil Mudholkar tested Jev against Gemini for classifying business emails. While Gemini edged out slightly in accuracy, it was 10 to 20 times more expensive. Mudholkar emphasized Jev’s unique advantage: “it is the only one that hands back a real probability which makes it ideal for automating workflows!!”
Augmenting and Overcoming Limitations of LLMs
Beyond replacing LLMs in some contexts, Jev can also augment them, serving as a smart monitor against hallucinations and misbehavior. Using costly agents to check other agents can quickly become prohibitive, but Jev’s affordability and speed make it an efficient choice for tracking LLM agent traces and preventing jailbreaks.
Armin Ronacher, CTO of Earendil—which develops the open-source model harness Pi—explained the practical benefit: “At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it.”
Ronacher also highlighted Jev’s potential in model routing, where its low cost and speed enable real-time selection of the most appropriate model for a given workload—a task that would be prohibitively expensive if handled by an LLM.
The Vision Behind Jev and Future Prospects
Named after the 19th-century economist William Stanley Jevons, whose paradox observes that falling commodity prices can lead to increased consumption, the model embodies Almeida’s vision that cheaper intelligence will drive widespread, emergent, and distributed smart software—much like the early internet rather than centralized mega-apps.
“We think that there’s just going to be smart software all over the place in a way that’s emergent and distributed … much more like the early internet than you know like the mega apps that people are trying to build right now,” Almeida said.
While Almeida remains discreet about Jev’s precise architecture—speculated to build upon an open-weight LLM—the company positions Jev as a “System One model,” focusing on intuition rather than complex reasoning, and trained exclusively on synthetic data via a new method called “reinforcement learning from calibrated decisions.”
“We made an early bet that we will be making all of our data, and that has been one of the best bets I’ve ever made in my life—better than our launch, in my opinion, better than RLHF,” Almeida told TechCrunch. “Half of [our company] is a lab that basically owns this entire subfield of statistically well-understood synthetic data, and that is now my life joy.”
Although Jev currently stands alone in its niche, experts like Ronacher anticipate competitors will soon emerge, given the clear utility of this novel approach. “We should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don’t have to be creative yet,” he noted.
Looking ahead, TypeSafe AI plans to develop additional versions of Jev across new modalities. When asked about the company’s identity, Almeida emphasized a grounded focus: “The main product of frontier labs is fear or hype. I would like our main product to be intelligence…[but we are] not a lab in the sense of, you know, like bet on infinite wealth, or a religion, or building God in a data center, or whatever is the thing of today.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Read more Here.
