From the course: AI Model Trends: The Latest You Need to Know
Jev: Make fast, cheap decisions at scale with this new model from TypeSafe AI
From the course: AI Model Trends: The Latest You Need to Know
Jev: Make fast, cheap decisions at scale with this new model from TypeSafe AI
Ask a chatbot one simple question, and it suddenly thinks it's breaking into Hollywood or writing the next great novel. But there's a new model focused on decision making that's changing that. So you can stop paying for the director's cut. Jev is TypeSafe AI's model for those kind of decisions. Think of it as a new kind of primitive that doesn't replace ChatGPT or Claude, but sits next to a chat model. It's built for helping your apps make decisions that your code can trust with just three simple primitives that you can use whenever you need decisions instead of a conversation. You can ask it to pick from a menu by using choice. You select from some options you predefined so that the code branches reliably without any text scraping. The other option is to ask it to score something on a scale. You rate against the rubric that you wrote to get clean numeric values for sorting. And finally, you can ask for something called a null. This is a yes or no answer as a number. It returns a probability from zero to one. And then you establish the cutoffs. So you can specify that you'll take something where it's 80% sure or above as a yes and maybe treat decisions that score less than 20% as a no. It's really convenient if you're using something like OpenRouter, where you can choose from different models. Or if you're making thousands of decisions where token cost and wait time is going to compound, this necessitates that you combine different models because Jev isn't a replacement for your writing model, but it's the cheap, fast first pass that takes care of decisions for you. How you combine this with other models does make a difference. You can run Jev first on questions, and if it's very sure about the answer, then keep the answer and move on. If you're not getting a clear signal from Jev, you may want to go to a bigger chat model or perhaps a human. Out of Jev's three available primitives, like choices and scores, the one that's a little bit hard to understand is what we call a null. It means how yes is this yes. By setting up these different thresholds, Jev can take care of massive amounts of calls at speed. This board holds candidate ideas for me to talk about during the day, and then it provides me with some potential posts. All this information comes from newsletters, the Web, and other subscriptions that I have. Each one of these items gets a specific score. This sort of classification is the perfect job for something like Jev. The question about every entry is pretty simple. How good is this idea for a post? Let's look at how Jev would handle something like this. You don't really chat with Jev. You send it the card as an object and then you ask a single type question, title, URL, and the short description. There's not a long prompt like you do with an LLM. Jev has a question type called score. You write a rubric with five different levels from worst to best. And then Jev is going to pick which one of these options fits the card and return that as a number, noting how confident it is. Level zero would be for what I would consider junk engagement bait or anything that doesn't make the cut. Level four would be pretty much ready to draft. And levels in between are sort of a middle ground. Since I already have an OpenRouter API key, I can use that for chat models. But with Jev, you would hit a different endpoint, the decisions API. You send along the card's details as well as the single fit score. You get a type score and a rank for your board. To test this out, I ran Jev on three different types of scores to make sure that I proved that there is a benefit to this before I implement it. Now, these are the results of my tests. Jev was about 19 times faster and about four times cheaper than Luna on the same rubric. And the model was just as accurate as other models. If you map that out to a thousand triage decisions, you can see that Jev is cheaper than GPT Luna, but significantly cheaper than Claude Opus. OpenRouter ran these same kind of tests on 3,000 banking customer support issues through both Jev and Claude Opus 5. And although it was about three points behind Opus on accuracy, it ended up being about 13 times faster and cost about 1/22 of what Opus did. If you're building product at scale, you're probably making thousands of decisions. And so far you've been using the wrong model. For classification, Jev is a no-brainer. It's going to save you a ton of money. And it's not that hard to implement into your current harness.
Contents
-
-
-
Jev: Make fast, cheap decisions at scale with this new model from TypeSafe AI4m 43s
-
Fable 5.1: A cache-busting top model7m 11s
-
Kimi K3: Designing a prototype with the leading value open model8m 1s
-
Opus 5: How it compares to state-of-the-art (SOTA) models9m 20s
-
Gemini 3.6 Flash and 3.5 Flash-Lite: Full-featured apps with the Interactions API6m 14s
-
-
-
-
-