¶1a brand new open-source model just dropped that is competitive with the best models on the planet. It's free. You can download it right now. And that's not even the most impressive part. Take a look at this trailer.
¶2It paints such an optimistic and really simple vision of what the future of artificial intelligence can be. A guy going fishing while AI is doing work for him. Here's another one playing tennis while AI is doing scientific research. bouldering while putting together spreadsheets. All of this is such a clear and simple vision of what the future can be.
¶3And I just love it how it's framed, how simple it is. This video really spoke to me, and I think we need more of this, less houses on fire, which is the anthropic route. So, the model is Quen 3.8 Max. This is another in a series of incredible open-source models from China. This is absolutely a frontier grade model coming in at 2.4 trillion parameters.
¶4And why do I even mention parameter count so early in this video? Well, that's going to become more clear later while I do some comparisons to Soul and Fable and OpenAI's next model release, which hopefully will come soon. But the point to remember is Kimmy K3 just came out and it was right around this parameter size, right around 3 trillion parameters. Now, we have the second model out of China, open- source, open weights, that is right around that same size, and it is extremely capable. I just want to get that right out right away.
¶5This model is excellent. Now, of course, let's go over the benchmarks, but take all of these benchmarks with a grain of salt because they can be gamed, they can be benchmaxed, and a model can be really good at these specific benchmarks and not good more generally, which we've seen from a lot of models. Kimmy K3 is incredible, but it certainly is not as good as a Fable. So, here we have Sweep Pro 67.7 as compared to the leader right here at 80. This is Fable.
¶6We have Terminal Bench, one of, if not the most important Deentic coding benchmarks out there, coming in at 86.6 versus 84.6 for Fable and just under GPT 5.6 Soul. And what you're going to see, I'm not going to read all of these, but Quen 3.8 8 Max is absolutely competitive with the best of the best US closed source AI models. Absolutely, without a doubt. At least according to these benchmarks. All right, so Alibaba decided to highlight the fact that this new Quen model is really good at reproducing research papers.
¶7And if you're wondering why they're highlighting it or why that's so important, that is kind of the required skill for automated AI research, right? So basically, Quen 3. Max is given a research paper and then they task it with reproducing the results. That requires the model to read the research paper, understand what it's trying to accomplish, write code, and then run the code and get the same output, the same results as the research paper. And so if a model's really good at understanding a research paper and reproducing the results, that's kind of the step right before it's actually able to make its own discoveries, write its own research papers, prove out its hypothesis using code.
¶8It's kind of the step right before that. And so they even detail it. The catch Quent 3.8 Max started from nothing but the paper and a set of GPUs. No starter code, no ready-made pipeline. And here's the important part.
¶9turning reproduction into self-evolution. It invented and tested 18 improvement ideas of its own across four rounds. This is recursive self-improvement. Now, it's based on an idea that was put forth by a human researcher. But again, if you can expand the scope of that recursive loop, it might actually be able to discover new algorithmic unlocks and new AI research ideas on its own.
¶10And again, I think this is probably pretty intentional by the Alibaba Quinn team, but autonomous chip design and closed loop feedbackdriven optimization. Autonomous chip design. One of the things that China is behind on is chip design and chip manufacturing. But if they had a model that could help them design these chips, that would no longer be their weakness. And that's kind of what they're highlighting here.
¶11Again, Quinn 3.8 8 Max has independently achieved the autonomous execution of the entire silicon design flow which is incredible and we've already heard of the leading chip labs in America using AI to help them design nextG chips. Now at the bottom of the blog post they have this full benchmark table which shows all the scores of each of the leading Frontier models as well as each of the benchmarks it was run across but it's kind of hard visually to see which model is doing better. But I actually had Codex rebuild the chart and here it is. Now we can actually visually see which model's doing best. So for multimodal reasoning you can see Quinn 3.8 Max absolutely dominating.
¶12Here we have visual agent and coding. Here this is Fable 5 dominating across the benchmarks for document and office intelligence. We have Quen as the winner. Real world and spatial understanding. Quen as the winner visual perception, Quen is the winner.
¶13So you can see this is an incredibly capable model. And by the way, if you want to see this table that's easier to visually recognize which model is doing best, I'll put it on forwardfuture.com and I'll link it down below. And the next thing I want to show you is the pricing because what opensource is known for and specifically Chinese open source is that they are incredibly efficient and inexpensive. But it's not as simple as thinking lower price equals better. I'll tell you what it really is right after this.
¶14If you run an e-commerce business, you know how difficult the entire process of finding a supplier, negotiating with them, making sure they're legit. That entire thing is such a headache and so tedious. That's why I'm super excited to tell you about Axio Work today because it essentially solves all of those problems. They've built a team of AI agents that go out and find suppliers that are right for a global audience. And it not only did that, it found reorder stats and other pieces of information that helped me decide if this supplier is trustworthy and I want to actually work with them.
¶15Then I was able to select the manufacturers, select the suppliers, and begin the negotiation process. The Axio sourcing toolkit even goes out, messages the suppliers, and tells them to give me their full pitch. I can watch the conversation happening in real time or after the fact, all from the messaging tab. You can step in and guide the conversation or message the supplier manually at any time. You can even schedule tasks like campaign monitoring and competitive research.
¶16If you're building a lean e-commerce business without a full team, Axio work is your solution. And they are giving my viewers 7 days free to try it out. Go try out Axio Work. I'm going to drop a link down below. Click it.
¶17Let them know I sent you. It helps. Thank you. And thanks again to Axial Work for sponsoring this video. Now, back to the video.
¶18All right, so here's the price. This is from Open Router, and right now it's $2 per million input and $6 per million output. And to put that in perspective, GPT 5.6 Soul is $5 per million input and $30 per million output. And Fable is $10 per million input, $50 per million output. So you can see this is a much much less expensive model than either GPT 5.6 Soul and Fable.
¶19But again, that doesn't tell the whole story because the price per token is only one half of the equation. The other half is how many tokens does it actually take to accomplish the same task. And so if Quen takes three or four times as many tokens, it's essentially the same price as a fable as a soul. Artificial analysis has a really good chart for exactly this. Now, unfortunately, the new Quen model is not yet tested against the artificial analysis benchmark, which means it does not show up in this chart.
¶20But you can see right here, Quen 3.7 Max, $128 per task completed. Whereas, you can see all the most expensive models are all anthropic. And here's Kimmy K3 Max, which comes in at about 30 or 40% less. And here's 5.6 6 soul coming in at $1.23. So really what matters and I've said this a lot, how much does it cost to actually accomplish a given task, not how much the tokens are.
¶21Okay, so now I want to talk about the model size. Right now, the two best models out of China and specifically the Chinese open source labs, Moonshot has Kimmy and Alibaba has Quen. They are kind of in the mid 2 trillion parameter model size. That is the absolute frontier of what these Chinese open source companies are capable of. And that's excellent.
¶22These models are incredibly efficient. Again, we'll have to wait and see what the actual cost per task completed is, but these are really, really good models, and I am very appreciative that these open- source Chinese companies are putting them out for free. We can download them. We can run them ourselves on our own servers. We could put them on our computer if we have a big enough GPU.
¶23But I think the model size is still very telling as to where we are in this kind of closed source US lab versus open-source Chinese lab race. Fable is rumored to be at about 7 plus trillion parameters, which would be about twice the size of what Kimmy and Quen are right now. And OpenAI is gearing up to launch their new training run model, which publicly has been codenamed Astra. And that model is likely to be 7 plus trillion parameters as well. The US closed source AI labs are at twice the size which you need the best GPUs in the world to accomplish.
¶24Plus, you need a lot of them. And again, that's kind of where China's weaknesses. It's been reported that some of these Chinese AI labs are smuggling in some of our top GPUs from Nvidia, but generally they don't have nearly the same compute as the US does. And so it really does seem everything comes down to how much compute you have. And that's going to be important.
¶25Keep that in mind for later in the video when I talk more about the race between Chinese open source AI and American closed source AI. But let me put something plainly. In the short term, right now, these incredible open-source openweights models coming out of China are really good for the ecosystem. They give a lot of great alternatives to paying anthropic and open AI substantially more per token. They give us enterprise a lot of optionality into owning their own models, serving it themselves, fine-tuning it how they want, and ultimately not being beholdened and have this incredible platform risk with an open AI and with an anthropic.
¶26And one of my biggest fears with artificial intelligence in general is the concentration of power. And so I love seeing alternatives to open AI models, to anthropic models. I love seeing these open source models that anybody can get their hands on and play with and optimize for their own use case. Not to mention that it's just fun as a tinkerer to be able to download these models, put them on your own computer, play around with them, run them. It's just cool to do.
¶27But I still have the midterm and long-term fear that the United States and US enterprise becomes overly dependent on Chinese artificial intelligence. these incredible open- source models from China get released. We obviously don't need to use inference provided by China. That means literally sending your data to China. Instead, we can download the models.
¶28US inference companies can run them, can serve them, and everything stays on US soil. Okay, great. So, why is that a problem? Well, US companies have this decision to make. They can use Fable, they can use Soul, and they can pay top tier prices for them, or they can get 95% of the capability and pay a fraction of the price for some of these open-source models.
¶29And what do you think they're going to choose? For the vast majority of use cases, the absolute frontier just isn't necessary and in fact might be overkill. And so, they're going to choose these open source models. And these open source models, the best in the world, are coming from China right now. And so although we're able to download them, eventually what will likely happen is the model will be co-designed with the chip to be the most efficient and the most optimized for that specific chip.
¶30So US inference providers might be able to charge one rate, but if you actually go and pay a Chinese inference company, you might be able to pay a fraction of even that amount. The point is with extreme code design between model and hardware, it might actually put us as we're becoming dependent on Chinese open source, highly dependent on Chinese chips as well. And that causes a lot of geopolitical risk for the United States. And so I'm extremely thankful to companies like DeepSeek, like Moonshot with Kimmy, like Alibaba with Quen. They are putting out incredible models and the world is benefiting from them.
¶31But ultimately there is geopolitical risk there. I just want to point that out. It's not my opinion. It is just a fact. These are adversarial countries we're talking about.
¶32And so here are my conflicting thoughts. And I go back and forth on this all the time. Open source is incredible. The open- source models coming out of China are incredible. And they're putting a lot of competitive pressure on US closed source labs like OpenAI and Enthropic because they are just so much less expensive.
¶33And so does that ultimately commoditize the model layer? Does that commoditize intelligence? And does that drive down the value of companies like Anthropic and Open AI? And then the conflicting thought is maybe it doesn't matter. Anthropic and Open AI have the best models in the world.
¶34They have so much momentum and we talked about automated AI research. This all leads to recursive self-improvement. And if anthropic and open AI are already well on their way to recursive self-improvement, all of the open source models in the world don't matter. Because ultimately, when you have more compute than anybody, the best models in the world, and those models are getting better and more efficient more quickly than anybody else in the world, there is no competition. And we just saw that last week where OpenAI literally said GPT 5.6 Six Soul helped increase the efficiency of their models so much they dropped the price by 80%.
¶35And so we're already seeing that. So I'm not sure. Let me know what you think down below because I'm so conflicted on whether open- source will ultimately drive competitive pressure against the closed source Frontier Labs or ultimately it doesn't matter because RSI is the only goal. Now I do think open source is incredibly important. I made an entire video about open-source.