In 2011, psychologist Daniel Kahneman popularized the ideas surrounding dual process theory in his book “Thinking, Fast and Slow”. Its main thesis is a differentiation between two modes of thought: “System 1” is fast and instinctive, while “System 2” is slower, more deliberative, and more logical. The book summarizes decades of research to suggest that we have too much confidence in human judgment and proposes methods to make better decisions based on probabilities rather than heuristics.
According to Kahneman, we humans struggle to reason probabilistically. We tend to associate new information with existing patterns or thoughts, rather than creating new patterns for each new experience. The easier it is to recall the consequences of something, the greater we perceive those consequences to be. We fail to account for complexity and the fact that our understanding of the world consists of an insufficient set of observations. We underestimate the role of chance and therefore falsely assume that the future will be like the past.
One of the biggest challenges for software developers incorporating LLMs into system architectures is that for the past eighty years, computer science has involved cold, hard, deterministic logic. Compiled source code produces instruction sets that flow through chains of AND, OR, NOT, NAND, NOR, XOR, and XNOR gates made of transistors. The same input always results in the same output. This speed and predictability are what have made computers so reliable for high-stakes systems. LLMs, on the other hand, are highly probabilistic systems that take text input and produce a variant output each time. This output may also be hallucinated, non-factual, and convincing. They are not appropriate for high-stakes systems without other systems or people to verify and correct them.
The feats of frontier LLMs border on the miraculous. However, as of this writing, frontier models still have startlingly high hallucination and inaccuracy rates. Why are the rates so high? It’s a side effect of how they work and how they’re trained. Reinforcement Learning from Human Feedback (RLHF) optimizes models for human preference rather than factual correctness. RLHF trains models to fabricate plausible-sounding explanations rather than signal a lack of knowledge, inadvertently encouraging hallucinations through reward hacking.
The question is not “to AI or not to AI”, but what AI architecture should we choose and where and when is it appropriate to apply it. General purpose transformer LLMs like ChatGPT were introduced to the public as chat bots because they are trained on natural language, are suited to that type of low stakes assistant interaction, and the labs needed to collect real human training data to improve them. When software serves as an assistant with disclaimers that it “can make mistakes”, there’s an implied assumption that the user remains in control and is responsible for it. It was only in the past year that open source experiments like OpenClaw pushed the promise of LLMs into autonomous use cases with tool calling abilities and outright control of computer systems. The industry, public and stock market has since proceeded to expect that use case is viable even though the underlying architectural paradigm remains unreliable.
Many argue that general-purpose transformer large language models are not the future of the frontier; its just the one that the brute-force approach of scaling compute demonstrated first and attracted venture capital. A recent study by MIT notes scaling laws have reached diminishing returns from sheer parameter and data expansion. The AI industry cannot rely on raw compute scaling alone to advance model capabilities. Progress will require shifting investments toward architectural innovations, algorithmic efficiency, and post-training methods like reasoning or search. Furthermore, the rising cost of training frontier models against shrinking operational lifespans strains the economic viability of traditional compute scaling for enterprise deployment. That isn’t great for frontier labs. In the meantime, practical application of AI in enterprise systems needs more reliable and cost effective options.
There are three problems with all RLFH models in applied systems engineering contexts:
- They may have gaps in training data that don’t allow for sufficient “known knowns” to accurately make a decision or provide an answer.
- Its internal calibration of confidence is obfuscated to the end user.
- RLFH methods reward a model for providing an answer over admitting uncertainty.
Diogo Almeida, an OpenAI veteran, co-inventor of reinforcement learning from human feedback (RLHF), and contributor to InstructGPT and GPT-4, has noted that this glaring reliability problem requires entirely different approaches.
His argument is that products like ChatGPT and Claude Code represent increasingly capable iterations of a single paradigm: AI assisting a human who remains in the loop. If every automated process requires human supervision to guard against inherent unreliability it imposes a ceiling on what automation can achieve.
During a July presentation at the AI Engineer World’s Fair, Almeida asked a critical question: why does artificial intelligence seem simultaneously amazing and disappointing? He highlighted two contradictory views of the current technology landscape.
On one side, progress appears extraordinarily rapid:
- Models routinely surpass human performance on benchmarks.
- Capabilities continue to expand.
- Autonomous operating windows are lengthening.
- Systems accomplish complex intellectual tasks.
On the other side:
- Systems remain unreliable.
- Software engineering has not fundamentally transformed.
- Businesses cannot trust models with high-stakes decisions.
- Humans remain necessary for basic operational tasks.
- Most commercial products remain limited to chat interfaces and coding assistants.
Why can models solve advanced mathematics and protein folding problems yet fail at basic operational decisions? Because current systems were designed using RLHF and human-in-the-loop training rather than for direct task automation. This does not fix the failures related to how a model stores and infers “facts” or how it internally measures its own uncertainty.
Preference optimization has a distinct consequence, according to Almeida: models excel at appearing correct. He uses a humorous example. Give ChatGPT an audio file of farts and ask it to evaluate the “music.” Rather than identifying the recording, the model produces a thoughtful artistic interpretation.
Almeida’s new AI model, Jev, uses a different paradigm. It makes fast, inexpensive judgments without hallucinations instead of generating text. Rather than returning natural language for text prompts, Jev returns typed probabilistic decisions for other software or AI models to use. In term of Kahneman’s framework, Jev is a “system 1” solution.
Jev wasn’t trained on human preferences. It classifies input and provides clear direct confidence scores as probabilities. Jev uses a training method called Reinforcement Learning for Calibrated Decisions (RLCD) that optimizes the model so that its output probabilities reliably match actual accuracy. If Jev assigns a 90% probability to a structured choice, it is trained to be correct approximately 90% of the time, making its uncertainty metrics mathematically honest.

An AI’s probability score acts as the mathematical version of Daniel Kahneman’s WYSIATI (What You See Is All There Is) rule, representing confidence based strictly on its internal “known knowns”—its training data and prompt. By making probability explicit and mathematically calibrated, RLCD turns an AI’s confidence score into something actionable. Software developers can set strict thresholds—instructing the system to automatically execute a decision if the calibrated score is above 95%, but escalate to a human the moment the score drops, explicitly acknowledging when the AI has hit a blind spot.
To be clear, classifiers are nothing new and single pass classifier models date back to 2018 and zero-shot label pipelines to 2019. JEV offers something that older classifiers don’t, which is dynamic instruction following. Unlike static classifiers that require retraining when labels or policies change, JEV can adapt to new semantic instructions at runtime. Jev will likely be the first of many lightweight fast “universal classifiers” and may influence the training designs of future models that provide greater transparency.
Back to our three problems and how RLCD helps solve them:
- Gaps in “known knowns” – RLCD relies on proper scoring rules (such as log-loss) where predicting a 50% probability on an uncertain token yields higher expected reward than declaring a 99% probability and getting it wrong. When RLCD models reach knowledge boundaries created by training gaps, they actively output low confidence scores.
- Obfuscated Internal Calibration – RLCD explicitly targets calibration as the primary loss metric. The model is forced to output explicit probability tokens alongside its categorical choices. Training guarantees that across N outputs tagged with 80% confidence, roughly 80% are empirically correct.
- Incentivizing False Certainty – RLCD sets a mathematically balanced betting game. The reward function scales non-linearly: high-confidence wrong predictions incur severe negative penalties. This makes “expressing doubt” the optimal strategy for the model whenever its internal representations show high variance.
This approach is invaluable when AI interactions require constrained answers for deterministic handling.
With Jev, developers provide a state value, such as a JSON object or a string like “My card was charged twice.” The model evaluates this state using question primitives—Choice, Score, and Noul—which return typed, probabilistic responses.
For example, a routing query asking which department should handle a customer service ticket returns structured data: {“billing”: 0.08, “technical”: 0.85, “sales”: 0.07}, accompanied by a confidence score of 0.82.
While this format offers little value to end users seeking conversational answers, it provides software developers with the structured data required for reliable classification and decision-making.
This principle mirrors general software engineering: apply appropriate abstractions instead of using one general-purpose tool for every problem. More importantly, it marks a broader evolution in artificial intelligence where models transition from systems designed to communicate with humans toward computational components embedded within autonomous software architectures.
The most important implication of Jev is not revolutionary intelligence potential, but its speed, cost efficiency and practical fit in autonomous systems design.
The performance claims are substantial. Jev is extremely fast. TypeSafe reports that Jev is 193.6 times faster and 444.6 times cheaper than frontier models in peak in-house testing, with a cost per decision of approximately $0.0004 and pricing at $0.042 per million input tokens, with output tokens priced at zero. Independent testing from Every corroborated these directional claims, finding Jev roughly 25 times faster and 580 times cheaper than Claude Fable 5.1 on extraction tasks at 0.35 seconds versus 8.83 seconds per passage.
The economics change dramatically when intelligence becomes this cheap. For automated software systems making thousands or millions of small decisions, repeated language generation is prohibitively expensive. An autonomous agent may perform dozens of decisions during a single task. If each decision introduces substantial latency, cumulative delay becomes noticeable. Faster bounded models could operate as an intermediate control layer, making rapid decisions while larger models handle more complex stages.
This design inherently shifts more of the burden to smart systems engineering. While the underlying mechanics remain probabilistic, this reality is made explicit through its typesafe output and “multiple choice” I/O, which lends itself to more thoughtful application and thus greater reliability. Developers must intentionally design their agent pipelines to accommodate clear probabilistic boundaries in business logic rather than relying on generalized generative models to paper over structural errors.
This division of cognitive labor carries profound implications for industry economics as low-cost automation solidifies its position as a primary revenue driver. When the marginal cost of routine machine decisions approaches zero, enterprises can scale autonomous agents across complex multi-step workflows with lower infrastructure bills and lower user-facing delays. Industries dependent on massive transaction volumes, real-time data extraction, and continuous operational triage will likely shift toward modular, tiered intelligence stacks. Companies that successfully integrate these layered cognitive models will capture a distinct competitive advantage, transforming high-frequency automation from an expensive luxury into a high-margin foundation for enterprise growth.