I have integrated Jev into many of my projects because I find the idea of putting a probabilistic assessment into software as naturally as calling a function genuinely ingenious. Choices run through a program all the time: who should handle a request, which process to start, when to wait for more information. These are often tiny steps, but the behavior of the whole system depends on how they fit together.

Usually, when we talk about artificial intelligence, we look at what the model produces in front of us. An answer, an image, a piece of code. Jev, the model from TypeSafe AI, draws attention to what happens a moment earlier: the point at which a program needs to understand enough about a situation to choose what to do next.

Imagine a request arriving at an IT service: “Since this morning, the export keeps getting stuck, but only with the new client's files.” Before it even answers, the system needs to work out what kind of job it is dealing with. A known procedure might be enough, or someone may need to examine the logs; perhaps the message still contains too little information to choose. Handing everything to a generative model is one option, but each of these steps has different requirements and could be handled by a different component.

It is a hypothetical example, yet it captures what interests me about Jev. An assessment becomes a component with which to build other things. From there, it is fairly natural to imagine systems that distribute work among conventional code, small models and more powerful services, using the resources each step actually needs.

This way of designing software meets another shift: capabilities we associated with a remote service a few years ago are arriving in models we can run on our own machines. Spark-X2.5 and MiniCPM5-2B are part of this movement; Apple sells computers while explicitly promoting local AI execution as one of their uses. That raises an economic question: if a useful decision becomes much cheaper, how much does that change the expectations of those investing in selling it to us?

An Answer the Program Can Use

Jev receives a state, called `state`, and a set of questions. The state is the available context: for example, the text of the request, information about the service and a description of the resources the program can use. Each question has a response type defined in advance. The code receives structured values and probability distributions, without having to search a paragraph for the decision it needs.

TypeSafe calls this family System One, drawing on the distinction between fast thinking and deliberate reasoning popularized by Daniel Kahneman. The company uses the analogy to describe an approach to bounded tasks: assess sufficiently precise questions separately and leave the program to combine the answers.

A `Choice` question asks the model to choose among options described by the developer. In our example, it could distinguish a request for instructions from a problem report that needs investigation. The response includes the preferred option and the probabilities assigned to the alternatives. Receiving the distribution as well lets us see whether one category clearly dominates or the model is divided. The options must, of course, cover the possible situations: without a category for unrecognized cases, even an unrelated request will end up looking like something on the list.

With `Score`, the developer defines a scale through levels described in words. We might assess how much a reported problem prevents someone from working, distinguishing an inconvenience they can work around from a blocker for which no temporary solution exists. The result can fall between two levels and comes with the probabilities assigned to each. Describing the conditions matters: “serious” and “very serious” say little on their own to anyone, including a model.

`Noul` answers a yes-or-no question by returning a probability between zero and one. “Does the message mention a workaround?” is a question of this kind. A value of 0.8 expresses the estimate that the statement is true; it does not mean that the workaround works 80% of the time or that the problem is almost solved. These are different interpretations, and confusing them would change the program's behavior.

All three forms can be used in the same call. According to the documented API contract, the questions are assessed in parallel and separately against the supplied state. The system can therefore obtain several signals without turning each one into another round of conversation. Adding questions still means more text to process and costs to consider.

LLMs can already return JSON and responses constrained by a schema. Anyone who uses them to classify documents will know this possibility well. TypeSafe says it trains Jev to produce calibrated assessments, and it builds the interface around them. What interests me is the whole path from the question to a result the code can use, and the efficiency that path can achieve on the chosen task.

The format solves part of the problem. A valid category can be wrong, just as an integer variable can contain the wrong number. TypeSafe also publishes the model's limitations, including difficulties with calculations, ambiguous contexts and irrelevant information. In our export example, then, a well-formed response tells us that the program can read it; checking its meaning requires real examples from our service.

Once the Number Arrives, a Decision Still Has to Be Made

Suppose the system considers it likely that a request can be resolved using the documentation. Sending it straight to a local assistant may make sense if a mistake can be corrected with a second question. The same probability could be wholly inadequate for changing a production configuration. The context is similar, but the consequences of the two actions are very different.

The decision has to account for this gap. The model offers an assessment; the people building the software establish the cost of being wrong and which steps are allowed. Waiting, asking for a detail or involving a person can be a normal part of the process, without forcing every request toward an automated answer.

Choice and Score include a `confidence` field that summarizes how concentrated the distribution is. A distribution clustered around one option looks more confident than one spread across many. That number describes the model's response; it does not certify that the next action will succeed. To understand its usefulness, we need to compare it with outcomes on our own material, preferably in the language and conditions in which the system will operate.

Calibration also concerns this comparison. If, across a set of cases, a model assigns an 80% probability to events that occur roughly eight times out of ten, its estimates are well calibrated in that range. This property helps us choose sensible thresholds, but it cannot tell us in advance which two cases will be wrong. A change of domain or model version therefore calls for another performance check, rather than treating the threshold as a universal constant.

In our service, we could use the assessments to separate information retrieval from operations that change something. Documentation can be consulted automatically; a change to the system requires further checks. Spending limits and permissions remain rules in the program. This design distinction makes it possible to use the probabilistic component in many places without handing it every responsibility at once.

This is also why I find the idea so interesting for infrastructure. A complex system constantly makes decisions about where to run a job and with what priority. Sometimes a measurement is enough: if a node lacks sufficient memory, it is ruled out. At other times, the system needs to interpret a loosely structured request or work out which procedure matches the situation described. A semantic evaluator could help with these steps, while measurements and constraints continue to be handled by code.

We can imagine a control layer that classifies incoming work and an allocation mechanism that chooses among the resources actually available. A request that only needs a text reorganized might stay on a local machine; a difficult analysis would take another route. The policy can change with the load, the budget or the time available. For this to work, though, the classification must predict well enough which route will be able to complete the task: calling a request “simple” does not make it so.

TypeSafe already documents application-level routing among procedures, models and people. Extending this to infrastructure management is a design possibility to assess case by case. In ordinary high-speed packet forwarding, waiting for a remote call for every packet would be incompatible with the timing requirements. An evaluator of this kind would more plausibly belong in the layer that interprets a situation and updates a policy, leaving network mechanisms to apply it.

When combining many results, we also have to consider the relationships between them. Two questions assessed separately may concern dependent events: if they share a cause, or the outcome of one affects the other, simply multiplying the probabilities can produce a wrong estimate. Whoever understands the system needs to represent these dependencies in the design, along with the consequences of each choice.

Before the Chat Window, the Email Filter

We have become very good at debating which LLM gives better answers, and that can make us forget how broad the history of artificial intelligence is. LLMs are part of machine learning, which in turn occupies part of AI. A small model capable of classifying a message belongs to that history too, even if it has nothing to tell us in a chat window.

In 1998, Mehran Sahami, Susan Dumais, David Heckerman and Eric Horvitz published a paper on using a Bayesian approach to filter junk email. The problem was an everyday one: estimate which category a message belongs to and decide what to do with it. The different costs of errors were already part of the picture. Letting an annoying advertisement through and hiding a legitimate message do not cause the same harm.

The filter has a bounded task, receives clues and contributes to a decision. Its success is measured by how well it separates the messages we want to read from those that waste our time. Recognizing this continuity helps us see Jev in terms of the assessments it makes easier to insert into a program, within a history of probabilistic decisions that long predates the product.

In infrastructure, machine learning has been working on equally concrete problems for some time. Resource Central, described by Microsoft Research, uses predictions about virtual machine behavior to help manage resources in Azure. Knowing something about the expected workload can improve placement decisions and server utilization. A predictive capability becomes useful because it enters a system that knows how to use it.

These examples reveal a limitation in how we talk about AI: the visibility of a result tends to determine the importance we assign to it. An astonishing piece of writing travels easily; better resource allocation is noticed mostly when it is missing. Yet the latter can affect how a service works on every request.

Models that can be queried through instructions make some of these possibilities accessible to people who do not want to build a classifier from scratch for every new need. We can try a question, compare it with known cases and make it part of a procedure. For a stable, frequently repeated task, a purpose-built classifier or a rule may still be the best choice. The advantage is having more tools and being able to test them, instead of turning every problem into a conversation with the largest model available.

The Capability Comes to Your Computer

Meanwhile, the range of things we can run directly on our own machines is growing. Spark-X2.5 is distributed in versions with 1.7 billion and 4 billion parameters, with documentation for local use, including through MLX on Apple silicon. OpenBMB's MiniCPM5-2B offers another example of a compact model: its name refers to roughly two billion parameters excluding embeddings, while the total is around 2.5 billion. The available weights allow it to run with compatible local tools.

Parameters are values learned during training. Their number gives an indication of size, but does not by itself measure how capable a model is or how much memory an application will need. The way the weights are represented and the context the system must retain during processing also matter. A compact model can require much more memory when we ask it to work on very long texts.

Quantization reduces the numerical precision used to store the weights, often allowing lower memory use. The trade-off needs to be tested on the task: two files carrying the same model name can behave differently if this representation changes. The time devoted to reasoning also affects the result. A ranking without the test conditions therefore tells only part of the story.

The impression that capabilities we paid for not long ago are becoming accessible today does have a measurable precedent. In the AI Index 2025, Stanford records a more than 280-fold drop in the inference price needed to reach GPT-3.5-equivalent performance on the MMLU benchmark between November 2022 and October 2024. This is a threshold on a test, not the entire experience of using an assistant. Still, it shows how quickly the price of a defined capability can change.

The evaluations published for Spark-X2.5 and MiniCPM5-2B mainly compare them with other open models. They are not enough to claim that a small model today is equivalent in every respect to any commercial service from two years ago. Building an application calls for a more specific answer anyway: can this model do our job well enough, within the time and memory available?

If the answer is yes, we can choose where to run it. Being small and being local remain different properties: a small model can be offered through the cloud, while a machine with plenty of memory can host a much larger one. Specialization is a separate choice too. A compact general-purpose model and a dedicated evaluator can fit into the same system without doing the same job.

In a 2025 position paper, NVIDIA researchers argue for using small models for many recurring agent tasks and heterogeneous systems when different capabilities are needed. Their proposal turns the discussion toward how to distribute the work: reserve the most capable component for the steps that require it, and establish which activities can be entrusted to the others. The authors make the case for the advantages; whether they pay off still has to be tested in the actual application.

Apple shows that this possibility also interests hardware sellers. In presenting the Mac Studio with M5 Max and M5 Ultra, announced in August 2026 and available from September, the company explicitly promotes on-device AI alongside other professional uses. Unified memory and software for running models become part of its case for buying the computer. The most capable workstations address workloads different from those of a model with a few billion parameters, which does not inherently require a top-spec machine.

Some computing capability can thus be purchased with the device, rather than consumed exclusively through paid API calls. For a developer, that means being able to experiment with resources they control; for an organization, it means assessing whether some workflows can remain on its own infrastructure. When the data and models really do stay there, the need to send material to an external service also decreases.

The connection with Jev needs to be precise: in the form described here, Jev is a remote API service. Pairing it with a local model produces a hybrid system, not a fully offline application. If the router needs to read the content to choose a route, that content is sent to the router. We therefore need to design which information the decision requires, rather than merely moving the last step onto our own computer.

The Price of a Call and the Cost of the Work

We can return to the request about the stalled export. An initial check rules out unavailable resources. An assessment of the message helps choose the procedure; a local model can formulate a clarifying question using the relevant documentation. If the case calls for a more difficult analysis, the program can involve a more powerful remote model or a person. This is one possible architecture, whose value depends on how many cases each route can resolve.

The router has its own cost and adds a wait. If it often chooses the wrong route or sends almost everything to the most expensive service, the extra step could make the result worse. To assess it, we also need to count repeated attempts and the work needed to correct mistakes. Measuring the cost per successfully completed task is more useful than comparing only two prices per million tokens.

The Jev price list consulted on September 28, 2026, gives a price of $0.042 per million input tokens, with free output. One million requests of a thousand billed tokens each would therefore amount to $42 in API charges alone. That is a calculation based on the price list, not the full cost of a service: there is the surrounding software, the other processing and the time spent preparing and maintaining the system. The actual size of each request also depends on how much context and how many questions we send.

A price like that makes it worth trying assessments at points where the cost per call might previously have discouraged an experiment. It does not tell us, however, what they cost the provider to produce. In introducing Jev, TypeSafe acknowledges that demonstrating the sustainability of the price will take time and that it cannot yet prove the service is not subsidized. For anyone discussing the economics of AI, this distinction matters as much as the price list.

Locally, the calculation takes a different form. Hardware is purchased and depreciates; it consumes energy, requires maintenance and can handle only a limited amount of work at once. A machine used intensively for suitable tasks may make economic sense, while buying capacity that sits idle for most of the day can cost more than a remote service. The cloud retains an advantage when we need to absorb a spike or occasionally use resources that would not make sense to own.

There is also the cost of learning to use the tools well. Preparing a representative set of examples, observing errors and checking an update again takes time even when inference is cheap. As the execution price falls, this work can become a larger share of the total expense. That is a good reason to keep the parts of the system that already work simple, and concentrate new components where they add something measurable.

AI That Works Can Reshape Its Bubble

The debate about the AI bubble often mixes two questions: how useful the technology is and how much money the companies offering it can make. The spread of small models and specialized components makes the distance between the two clearer. We can use much more AI and, over the same period, be willing to pay less for each operation.

If a function that required an expensive service yesterday runs on the customer's computer today, the work still gets done. Who gets paid changes. Value can move to whoever builds the application, integrates the system, sells the hardware or provides support. Some of the benefit can remain directly with the people using the software. For the provider that expected to sell that operation thousands of times, however, the change can narrow the commercial opportunity.

Even without moving to local execution, a cheap component that selects which requests deserve an expensive model can reduce the number of calls needed to complete a workflow. The provider of the most powerful model then has to demonstrate its advantage on the difficult steps. The fact that a system contains AI no longer tells us how often, or at what price, it needs to use a particular provider.

This leads to the possibility I find interesting: some investments could come under pressure precisely as AI becomes more useful. If a business plan assumes high revenues from tasks that become commonplace and cheap, the technology's success can change the conditions on which that plan was built. Growing the user base is not enough on its own; revenue per use, the costs of serving those users and the capital already committed all matter.

There is movement in the other direction too. Lower costs make previously rejected uses worthwhile, and those new uses can generate much greater demand. The name Jev contains this very bet: TypeSafe derives it from William Stanley Jevons and invokes the relationship between efficiency and rising consumption. The founders expect cheaper decisions to expand the uses of artificial intelligence enormously.

That is an economic expectation to be tested. Savings on each task and growth in the number of tasks can happen together, without the second always offsetting the first to the same extent. Moreover, an expansion in overall demand does not guarantee that the gains will go to the companies that invested on the assumption that the market would be divided in a particular way.

The energy question also requires us to look at the whole picture. In its 2026 report Key Questions on Energy and AI, the IEA describes efficiency gains and growth in data center electricity demand. Less demanding tasks can coexist with systems used more often and for longer jobs. The price of a Jev call or a model's parameter count cannot, on their own, tell us what the final energy consumption will be.

Distributing some inference across devices also leaves the question of training open. A model we comfortably run locally was built through a different process, with different computing requirements. And applications will continue to exist that need more capable models, large amounts of memory or centralized workload management. Their relative importance is what we need to watch, rather than assuming that every local advance makes a data center disappear.

To assess market expectations, then, we will need to observe which activities continue to require a remote service and how much customers are willing to pay for its advantage. Jev and local models add alternatives against which to make that comparison. Someone building a system with different resources can change providers for part of the work without having to abandon the entire application: that possibility also affects the bargaining power of those selling models.

Back to Designing Choices

My enthusiasm for Jev comes from this too. I want to be able to build a program in which AI plays a part at specific steps, and know why each intervention was put there. I can use a probability to route a request and keep the limits within which the work is performed in the code. When requirements change, I know which part to revisit and which results to compare it against.

The export request, in the end, has to reach someone or something capable of resolving it. The user wants to get back to work, not find out how many models we involved. For the designer, though, that difference matters: using fewer resources to obtain a reliable result leaves room for other functions, other projects and people who previously could not afford them.

I would like to judge AI by the room it opens up as well. If a capability becomes cheap enough to use every day on our own hardware, or simple enough to insert into a clearly bounded part of the code, we have gained a concrete possibility. What it will be worth on the stock market is another calculation. In the program, meanwhile, we can start deciding where we need it.

Bibliography and Documentation

  • TypeSafe AI. Introduction; Choice; Score; Noul; Confidence; AI primer; Intent routing. Technical documentation, accessed September 28, 2026. Question contracts, distributions, calibration and software composition.
  • TypeSafe AI. Models; Jev 1.13 jaggedness. Current documentation, accessed September 28, 2026. Version, pricing and stated limitations.
  • Almeida, Diogo. Introducing System One Models & Jev. TypeSafe AI, September 15, 2026. Product introduction, origin of the name and qualifications concerning results and economic sustainability.
  • Sahami, Mehran; Dumais, Susan; Heckerman, David; Horvitz, Eric. A Bayesian Approach to Filtering Junk E-Mail. AAAI Workshop on Learning for Text Categorization, July 1998. Email classification and the different costs of errors.
  • Microsoft Research. Machine Learning for Systems and Tiered AIOps, Resource Central section. Project page, accessed October 9, 2026. Machine learning applied to cloud resource management.
  • XHToken/SparkLLM. Spark-X2.5. Official repository and documentation, accessed September 28, 2026. Model versions, published evaluations and local execution.
  • OpenBMB. MiniCPM5-2B. Official model card, accessed September 28, 2026. Size, release of model weights and evaluations.
  • Apple. Apple introduces new Mac Studio with M5 Max and M5 Ultra, August 25, 2026; The new Mac mini and Mac Studio are available today, September 22, 2026. Product positioning that includes on-device AI.
  • Stanford Institute for Human-Centered Artificial Intelligence. AI Index 2025: State of AI in 10 Charts. April 7, 2025. Historical changes in the inference price at a defined performance threshold.
  • Belcak, Peter, and colleagues. Small Language Models are the Future of Agentic AI. arXiv 2506.02153, June 2, 2025; version 3, September 22, 2026. Position paper on small models and heterogeneous architectures.
  • International Energy Agency. Key Questions on Energy and AI. April 16, 2026. Efficiency per task, aggregate demand and the economic conditions of data center expansion.