Radar

radar · IA & agentes

OpenRouter State Of Models: Jev, Open-Weights, and Tokenomics

¶1What's up engineers? Andy Debb Dan here. With the release of the new Jev system 1 model class and all the overnight clones, [music] this is a great time for us to break down the state of AI models as we enter the last quarter of 2026. [music] Why is that? Because it's important to understand the intelligence available to you and the trends around them so you can do [music] your best work at the right performance, speed, cost trade-off.

¶2I have five big trends we'll break down in this video. I'll throw them up on the screen now. We'll [music] use open router to explore these trends. If these ideas interest you, stick around and let's scale our compute so we can scale [music] our impact. Trend one tokens go up.

¶3Quick announcement. Early sign up for the phase 3 [music] Aentic engineering product is officially available. Sign up to get on a private email list to understand what advantage is coming next [music] to you and every engineer who's on the edge. I can't wait to share this with you. This will be by far one of the most valuable things I've ever created.

¶4Look at Open Router's weekly token chart. Usage keeps going up, but you have to look at the numbers to understand how astronomically high token usage is now. Roughly one year ago today, there was 4.5 trillion tokens for the entire week. Guess what that number is today? It's not 10 trillion, not 20 trillion, not 50 trillion, not 100 trillion.

¶5Last week, there was a total of 146 trillion tokens used. It's not exponential growth, but it's close. Some people keep calling AI a bubble, but the growing usage numbers we can concretely see on Open Router shows demand for these tools that's nearly exponential. And you've likely seen this in your own work. The more tokens you can gain access to that's intelligent and cheap, the more use cases you keep finding.

¶6This is true for you, it's true for me, it's true for the entire engineering ecosystem. And we're all slowly learning how to make our tokens more economically valuable. This is the art and science of tokconomics. Now, you'll notice something really interesting here. There are barely any anthropic models in the top LLM leaderboard.

¶7That's because OpenR does not count tokens from other language model providers. Anthropic is targeting a massive $2 trillion IPO despite losing almost $50 billion. And this makes us ask an uncomfortable hard question with all the great growth Anthropic is seeing with all that token spend coming into Anthropic. Is this $2 trillion IPO value justified? And the real question is, can their explosive growth continue?

¶8Will we keep seeing the token spend that we see on Open Router come through to the model providers? If Anthropic has any chart that looks just like this and their usage, the answer will be yes. They will make $200 billion in annualized recurring revenue by the end of 2028, which is their projection. That's how they're able to make this insane IPO claim. So, we have a couple really, really interesting narratives going on here, and it all depends on language model usage continuing to go up nearly exponentially.

¶9If this trend continues, we can put our fist down and say there is no bubble. But if this trend does not continue, I think the bare case for Anthropics IPO is going to be true and then it'll likely follow with OpenAI and then the House of Cards starts falling down. I'm more in the optimistic camp, but pretty soon we're going to see which one of these two cases is true. Regardless, what's clear here is token usage is going up. What does this trend mean for us?

¶10I think the key question every engineer needs to ask is, is your token spend going up because you're getting more useful work done or is it going up because you're trying to use more tokens and you can't find a way to use them efficiently? Are we getting more productive use out of our token spend? This is tokconomics. It's not enough to just use more tokens. I always say on the channel, scale your compute to scale your impact.

¶11But it really matters what you're doing with these tokens. A lot of engineers are opting for cheap models and they sacrifice their own time and their own business's performance while other engineers are able to leverage these models properly to get the job done at cheaper prices. You want to be the engineers in that second bucket. You want to be utilizing these models without sacrificing your performance. And the Jev model by Typesa is a huge contributor to that.

¶12We'll talk about that more in a second. You can see the astronomical growth here. 257% up week overweek coming in at 2.6 trillion tokens. More on that in just a second. And this is the game we're all ultimately playing.

¶13More tokens does not equal more value. It's highly correlated, which is why I like to say scale your compute to scale your impact. But burning tokens to feel more productive is not progress. If it was, vibe coding would be the ultimate thing to do. It'd be the best place to spend your time.

¶14But it's not. Generating code is the easiest part of software engineering. The hard part is everything else. I like to aim for a steady rise in my compute usage versus an astronomical dump of tokens. Now, this brings us to the interesting question of given the entropic numbers and given the narrative around language models, who's winning this?

¶15What model provider? What model lab is actually winning the AI [music] race? So, who's actually winning the AI race? Just by glancing at the market share on Open Router, you see Open Weights and in the top three spot, we see DeepSeek, Google, and OpenAI. Now, the more fascinating part here for me is this.

¶16If we hover over, you know, a weekly breakdown, notice who you'll see in the top three spots. Deepseek, OpenAI, Google, Deepseek, OpenAI, Google, Deepseek, OpenAI, Google, Deepseek, OpenAI, Google. Almost the entire year with others, which is, you know, everyone else that's not in the top 10 popping up sometimes in third place. So, why is this? You know, everyone keeps saying that uh Google has fallen out of the AI race and they're not relevant.

¶17I have no idea what they're talking about. Uh Google with the Gemini Flash series is one of the most effective workhorse models available. It's one of the most useful models available in my model stack. You can see this. It's one of my favorite models in that A tier that can do 90% of everything you're trying to do with your language model.

¶18And inside of my PI coding agent, Gemini 3.8 Flash is my default. And there's a great reason for this. It's because it has the best trade-off between performance, speed, and cost. For everything it does for you, for everything this thing gives you, it's just a fantastic model. Now, there are a bunch of competitors in the A tier right now.

¶19This is definitely the most contentious place to be for a language model right now. But that's a great thing, right, for you and I, the engineer, the intelligence explosion is continuing, right? We just had the new uh cloud opus 5.5, we just had sonnet drop, we had the 6 soul and the six luna model drop. And uh, you know, there's more compute available than ever. We have to give Luna a big shout out here.

¶20This is a really, really powerful model that frankly is keeping OpenAI in the LLM rankings on Open Router, right? A lot of this volume is Luna. These three companies are sitting at that intersection. They're sitting at that sweet spot, the workhorse A tier model level where it's good enough to do useful work. Deepseek, Gemini, and there's Luna and Terara is here as well.

¶21This one is less popular, but they're sitting at that great workhorse level where cost, speed, and performance really intersects. As mentioned, Open Router traffic only tracks models going through Open Router. When I'm looking at the anthropic IPO coming up, I'm asking myself, this is like a fun thought experiment you can run. How many tokens of this 65, let's round down, 60 billion of the rumored current annualized revenue run rate that Anthropic has is coming from raw token spends. Like what's the total token count from Enthropic?

¶22So you can quickly easily do this math. Let's just assume that $60 billion run rate. What if it were all Opus 5.5 tokens at Opus 5.5 pricing? Just to put things into perspective, it looks like open weights models are winning and these cheap fast performant models are completely dominant. But if you just do a little bit of on the napkin math, you'll find something very different here.

¶23If all the 60 billion run rate where Opus 5.5 tokens, that means that Enthropic is processing 65.2 trillion tokens per day. As we just looked at, we know this Open Router is doing 140 trillion per week. That's insane. That means they're doing something like 450 trillion tokens per week. And if you don't like my simplified numbers here, right?

¶24Assuming that this is all Opus tokens, which they're not, cut it in half, right? It doesn't really matter. You can cut it in half again if you want. That means they're doing 32 trillion per day or 228 trillion per week, which is still 1.5x the total volume on Open Router. So, don't fully believe this market share here.

¶25You have to really really understand that the AI labs are getting massive, massive direct traffic. And of course, you know, Anthropic is known for being an enterprise leader. So, they're seeing these enterprise token spending numbers. Might be number seven down here, but don't let this fool you, right? [laughter] These open router numbers are really a representation of individual engineers and small to medium-sized businesses spending tokens.

¶26I'm sure Open Rider has some enterprise, but the real enterprise are fully bought into Anthropic, OpenAI, Google, and maybe a couple others. Who's winning the AI race? Right? You see a bunch of openweight models here on Open Router. Who's really winning?

¶27Everyone is winning. Again, let's scroll back up to the top models. It's just spend across the board is going up. It's going up massively. Something like 50, 60, 70, 80x of increase over the entire year.

¶28I remember looking at token volume a year ago on open router thinking, "Wow, these numbers are insane. They're insane. They're insane." But they just keep growing. So, who's winning? Everyone is winning.

¶29Token spend is going up across the board. Chinese open weights and now some close weights models are completely dominant here in terms of this specific model that's executing. But we have to give a lot of credit to the frontier model labs for paving the way. And then we have to give credit back to open weight models coming out of China for making it affordable. But then we have to give credit back to open AAI and to anthropic because guess how these models are getting trained.

¶30They're getting trained from distillations from the state-of-the-art is a very popular very well-known thing that's happening now. So there's a circular optimization that's happening at least from you and I's perspective right from the engineers perspective cost going down, intelligence going up. It's a great thing. Everybody's winning. There's no one model wins.

¶31There's no one AI lab wins. This is the optimal scenario we're in here. But something super exciting has happened recently. You know about it. You know, it's in all the headlines right now.

¶32Jev, we have a new class of models to work with. What's old is new again. We have zero shot classifier models back on the scene with Jev [music] leading the way. Trend number three is a old new model class led by the dev model. So, first off, 2.68 trillion tokens up 200%.

¶33This is the fastest grower outside of Mimo Flash and the rumored uh Space Bunny alpha model which is rumored to be a new Minimax model. So Jev is seeing explosive growth here. One of the crazy things about Jev and the usage numbers behind Jev on open router specifically is that Jev uses a fraction of the tokens that language models do when they're inside of agents. And it is still 12th place this week. This just came out.

¶34It's brand new technology or it's old technology made new and highly optimized. You've seen all the Jev clones come out. Everyone talking about fine-tuning models. Every other engineer saying that they've done this before. Yeah, this is classifier models.

¶35I myself personally have worked on a lot of these traditional classifier, you know, scikitlearn models. But it's not that simple. And let me show you the chart that proves it. Jev from Typesafe is one of the most scalable models I've ever seen. And it's one of the most reliable models.

¶36This is a really important chart that not many people will pay attention to. Uptime Jev just doesn't go down. Now, this is in contrast to many, many, many, many, many other models with much worse, you know, latency and much worse uptime. Speak of the devil, I'm filming this right now and Claude just had another outage, a blip in the system. But it's insane to me that uh everyone thinks that they can create a Jev clone when they're not really understanding that this is a super generic classifier that's zero shot that can be applied to many many different scenarios.

¶37Now definitely if you have a bunch of training data, you can outperform Jev with a small model. But what you won't outperform is Jev across many many different use cases. Contrary to popular belief, Jev is not easy to replace. Engineers that were with the channel last week know that I am actively embedding Jev into my PI coding agent. I'm already benchmarking this model.

¶38Last week we covered 10 levels of Jev for a Gentic engineers. Uh when you finish this LLM breakdown, check out that video to understand how to use Jev for real production use cases all the way up to powerful agentic use cases. But you can see here just with the Jev model just one additional model and a couple of tool calls I am outperforming on both speed and cost. You can only see cost here. I need to add speed to this PI agent comparison UI.

¶39You can see here, you know, on a relatively simple 400 500K input um prompt. I'm saving 5 cents and specifically I'm saving about 20%. The percentage is the thing to pay attention to as I'm benchmarking using Jev, a classifier model inside of PI. You really want to keep your eye on this. I'm using the new Sonnet 5.5 model here.

¶40By the way, this new classifier model is groundbreaking. Why is it groundbreaking? Why does Jeb matter so much? It's because this is a new species of models that you can use alongside traditional language models and agents. As viewers of the channel know, I think ins ors.

¶41It's not about this model or that model or that provider or that model provider or that one or this tool or that tool. It's about using them together. Think ins. All right. It's Deepseek V4.1 Flash and Jev.

¶42It's Claude Opus 5.5 and GPT6 Astra together. Combine compute. Don't select compute. You want to be using all these models at the right time in the right combination. And Jev is a new species of model that we can now use.

¶43Interestingly, Open Router has realized this. They have this new, if we go to models, they have a new decisions category specifically for models like this, uh, shout out Jev for being one of the first ones. And now everyone else is is coming out of the woodworks to create their own Jev clone or classifier model that does roughly the same thing. I think they're going to find that what Jeb actually gives you in addition to like a great implementation like they've really nailed this null choice score. They have more types on the way.

¶44They have more models on the way which I'm really excited about. What they've really nailed is the speed. So the latency is consistently great and they've nailed the reliability right which a lot of the leading model labs need to take after. I know Jev is much simpler but reliability has been a massive second priority, third priority for these model providers. So really really happy with Jev.

¶45This is an important release. I am right now in a flurry to explore Jev, looking at additional use cases, really figuring out how to use this model at scale broadly. As you can see here, I am actively rolling this into my PI coding agent. I have this in my default model, so my agent can make quick decisions when it needs to. As we covered last week, you can save a bunch of input tokens.

¶46And you can do this at scale with every agent you execute. My estimates are coming in. I'm saving about 20% token spend with Jev than without it. That's going to add up a lot day after day, week after week. So keep your eye peeled on these new classifier models.

¶47Experiment with them. Don't just put them in your code. Put them in your agents. Use tools together. This new model is going to be around for a while.

¶48That leads me to the next really simple idea, really important idea. Uh just to remind engineers, let's talk about quote unquote free tokens. Just want to leave a quick note here on models like Space Bunny Alpha. I'm sure it's some great new incredible model. It's rumored to be the next minimax model.

¶49One of the things I don't like about Open Router is that they do have this free model deal, which I know is like a great thing, especially if you're on a budget. You know, the problem with these models is that if it's free, you're the product. It's that same simple idea that's been true throughout all technology. You can be pretty confident that these model providers are retaining your prompts. They're reading your prompts.

¶50They're using them. They say they don't use them for training. I don't really buy that. That's kind of the trade-off here for models like this. That being said, you know, this model has jumped up into second place.

¶51So, it's clearly very incredible, very cost effective, but also it's just free, which really paints the picture of what the language model ecosystem appreciates, wants, and what we really need for models to be really effective at scale. We need them to be very, very cheap. We know cost is going down. We know intelligence is going up. But as I look at this trend, the usage of free tokens on Open Router, I'm constantly thinking usage is going to keep climbing and the cheaper tokens are while maintaining performance, the more use cases emerge.

¶52And so even though I don't like the free models on Open Router because I fully believe they're taking all your prompts, all your data, and training on them. You know, I'd rather pay than have them take proprietary information. But get it. The bigger trend that this paints for me is that models like this like point toward that future. It's a continuation of the trend.

¶53The cheaper models are, the more cost effective they are, right? Without losing too much performance, the more tokens will be spent. So, we can kind of see that across models throughout the market share. The cheaper models that are still effective tend to rise up, right? That's why once again, Deepc, Google, OpenAI are doing a really, really great job.

¶54This stealth model has exploded into near first place. It's pretty incredible how effective it is to release a lowcost model because it drives a lot of attention, a lot of usage solely based on its price which as I like to explore the open router usage data also shows a crack in this usage right because why is a lot of this usage coming to open router is because the models here are of a wide variety and they're also dirt cheap right some of them are even free because they train on your data so wherever you're getting your information sources from especially on these like large data uh analysis tools, you always have to understand who's providing the data, where do the stats come from, what are the users like. And you know, an open router, as valuable as this is, a lot of the usage comes from users that want price down over everything else. And that's in contrast to engineers using Claude, for instance, or OpenAI who want maximum performance, right? It's insane that Enthropic has these numbers or some numbers like this.

¶55Even if you drop it all the way down to Open Routers weekly numbers of 140 trillion, if OpenAI actually has these numbers, it's insane because it means that their tokens as they price them are much much more expensive than all the models on Open Router. Uh we do have that new Haiku model coming soon. We'll see what they do with that. Based on the Sonnet release, they're going to price it exact same. So, they're going to need to really create a price competitive Haiku model.

¶56I would love to see this come in under a dollar. that would make it potentially competitive with some of our workhorse models like the Gemini 3.8 Flash. But at the same time, you know, I'm not really expecting Anthropic to move on their prices here. So, you know, that's kind of a caveat I like to mention here. I'm always cognizant of user usage data when I'm checking out these language model trends.

¶57As you can see here, just as a general trend, if your tokens, if your tool, if your products are very, very cheap to consume, but still provide value, users will show up. Now this is really hard thing to fully digest because you have to make money somehow off your tokens, right? This is tokconomics. So that's some insight into the market share. One last thing I want to point out and focus on is the top applications flowing through open router.

¶58Now this [music] is really fascinating. So there's only one big trend that I like to look for on top apps and that's where the PI coding agent is in relation to cloud code. For the longest time, Pi has been trailing cloud code in the fifth, sixth, seventh most used application through open router, but um it's been slowly catching up to cloud code usage. To me, this is a sign of a broader understanding of the agentic tools of harness engineering. The more engineers that realize how important it is that they own their agentic coding tool, that they can customize it, that they can control it, the bigger this number is going to get.

¶59And that's the trend I've been seeing very very slowly, very very smoothly. I'll be releasing my yearly predictions video coming up in December here. You know, just to give one of them away through open router usage, I expect Pi to take over cloud code finally for the first time ever in usage again on open router. As I mentioned, there are a bunch of caveats to the open router data as a reliable source because this number coming through cloud code is actually much much higher if you include anthropic memberships and the subscriptions. The numbers are much much much higher.

¶60This is still a very cool thing to see as engineers, you know, use more openweight models as they get hardware that they can deploy. I'm also really excited, you know, I'm going to be picking up one of the M5 Ultra Max Studios coming up here in October and I'll be benchmarking that. I'll be sharing all the stats around what models you can run on a 512 unified memory Mac Studio. So, make sure you like, make sure you subscribe so you don't miss that video. That's going to send my personal token usage even higher and other engineers as well.

¶61again on these great openweight models that have been made available to us. You can see codeex generally lower usage here. Klein being the best I would say followup to pie coding agent. Kilo code not bad either. None of them get close to the simplicity and customization you get out of pi.

¶62So this is a trend I like to watch and keep my eye on. This is a large advantage that you can get as an engineer. If you control your agent coding tool and add more token spend here, you're going to have an advantage that other engineers don't have. Most engineers that watch this channel understand that by now, but if you don't, I'm going to leave a couple videos linked in the description where we utilize and really showcase what you can do with the PI coding agent. And one of them, I added several very, very powerful tools from Jev and coded it into the agent harness, added tools, modified the out of the box experience.

¶63You can see from that, by knowing how to harness engineer, I'm going to save about 20% token spend on every agent I boot up now. So, that's a really solid, very clear, significant advantage you get when you own your harness, not when you're renting them. We're renting enough stuff right now, right? We're renting intelligence. You want to rent as few things as possible, and you want to own as many as you can as it's useful for the time and investment you have to put into it, right?

¶64And I think owning your agent harness is absolutely worth it. So, those are the big trends I'm keeping my eyes on right now. The pie is growing. There's not a single winner. It's not open AI or anthropic.

¶65It's not closed weights versus open weights. It's everyone is winning, right? The pie is growing and the headliner chart is this one. Token spend is up something like 80x year-over-year from 4.5 trillion tokens to 445 trillion tokens. It's not exponential, but it's close.

¶66Every time someone thinks something is slowing down, every time someone thinks there's a wall approaching, there isn't. Growth continues. I think the only question from a financial perspective is how harmful and how much is the US economy really banking on the anthropic IPO and the open AI IPO going well because this could pop the financial AI bubble. As we can see here, there isn't really an AI bubble. This is very raw useful technology.

¶67People point to crypto a lot saying that that technology was preemptive. I think crypto is also useful, but it was much harder to measure its usefulness. Now we know that crypto is also very useful, very relevant. But it's much easier to see here in the open router data, right? And again, this is just the public data.

¶68Enthropic openi, they have their own internal numbers that probably look a lot like this, maybe even steeper. What's old is new. Again, we have the Jev model creating space for new species of models, right? We have these new decision models coming into play here. These are very, very powerful models to help you make decisions in your work and in your software.

¶69Again, check out last week's video where we broke down 10 levels of Jev for Agentic Engineers. The pie is actively getting bigger. There are no winners, but what there is is a bunch of options for you and I to scale their compute, to scale their impact. I think the big question for you and I going forward is constantly thinking, how can I increase my token spend in the most useful way? And so oftent times what I do is I throw state-of-the-art models at the problem I'm trying to solve.

¶70And then if I have to solve the problem repeatedly, I try to scale it down to cheaper, faster, more performant models. That's the big reason I like to maintain this model stack. Soon I'll be releasing in a gentically updated model stack that conglomerates a lot of the best benchmarks I look at for long horizon outloop agentic coding work. So you can, you know, play with that. You can analyze that to see how it compares to your model stack and how you think about the tier list of models.

¶71That's the name of the game. It's understand the technology, understand the intelligence available to you. So you can deploy it at the right performance, speed, cost trade-off for your work, for your use cases. I think when we look at models like Jev, we can expect to see a continuation of new species of models, not just language models to put in agents, but classifiers, decision makers, and other more specialized models that we can't even imagine right now. So big shout out to Jeff for paving the way.

¶72Again, check out last week's video to deeply understand Jeff. I think the most important thing every engineer can do right now is be prepared for the intelligence explosion to continue. That means using models like Jev to route to the right models to the right agents to the right workflows and it means understanding what model you need for the right time at the right price. You want to be managing your tokconomics for your inloop agentic coding work. And more and more the big unlock for engineers really focus on the signal is outloop agentic coding and more specifically outloop agentic engineering.

¶73For every engineer that made it to the end of this video, you know, big shout out to you. Thank you. That pre-signup for the phase 3 agentic engineering product is now available. Sign up is the first link in the description. There's going to be an optional quiz so I can optimize the product for you as soon as you sign up or as soon as you register if you're already an Aentic engineering [music] member.

¶74This will get you on the email list. We're going to be sharing exclusive big ideas and content from that next product that's coming [music] up. The Facebook product is consuming all of my time right now and making it the best thing possible for you and it's really turning out great. You know, it's crazy [music] to say this, but I think this is going to be the most valuable product I've ever created. Look out for an early December, late November launch.

¶75You [music] know where to find me every single Monday. Stay focused and keep building.