¶1better than Fable, half the price. This is Opus 5. We have a new model, Claude Opus 5. Now, if you remember, Opus was their kind of top tier in their family model. There was Haiku, Sonnet, and Opus.
¶2Then, of course, Fable came out. Fable is their biggest best model they make. And now, we have a version five of Opus. I tested it again. I didn't have a lot of time.
¶3and they gave it to me very late. But let's look at the benchmarks. Whoa, wait a second. Wait a second. Are are you all seeing this?
¶4Is Opus beating Fable 5 on almost every benchmark? Aentic terminal coding, Frontier Bench, an incredibly important benchmark for coding. 43 as compared to 33 on GDP val which is an open AI created benchmark which tests realworld practical tasks. 100 point improvement from Fable 5. How could this be?
¶5Whoa. Arc AGI 3. Shout out to Greg at Ark Prize. WA 30%. That's crazy.
¶6This is browse comp 90% versus 87%. Humanity's last exam basically the same with tools basically the same. A slight improvement. We have OS World, which is computer use. the model's ability to control your computer, literally figure out where things are on the screen, click buttons, do tasks.
¶7A fourpoint improvement for Deep Su, the best benchmark in my opinion. Deep Sui, I'm so glad they're including this now. It had a slight drop. So, I think if I were to point to any benchmark and say this is probably the vibes you all and myself are going to feel, this is it. So, a very slight drop, basically the same.
¶8Now, I can't wait to find out what the price is of this model. Okay, so Frontier Code, basically the same. Uh, Automation Bench, a very nice improvement. Uh, 9point improvement there. On the legal benchmark, we actually had a decently significant uh decrease from 13.3 to 11.7.
¶9Sorry lawyers. Uh, Healthbench Professional, we did have a decrease as well. Biom Mystery Bench looks to be just about the same. This is crazy. I did not expect this.
¶10Opus 5. Wow. Opus 5 is also highly efficient. It outperforms other models for a similar or lower cost per task. Now, if you've been watching my videos over the past week, you know I've been talking so much about cost per task.
¶11Wow. Wow. Look at this. So what we're seeing here on the yaxis is the total score of OS world and remember that is the computer use benchmark. On the x-axis we're seeing cost per task and what we're seeing is this orange is opus 5 and it is higher than this kind of gold yellow and blue which is fable 5 and opus 4.8 eight, but it's also to the left more, which means it's better and cheaper.
¶12This is a really big deal. Now, here's GPT 5.6 Soul, which is interesting. It has such a wide spectrum of price. Uh, but even at the very top to match performance of Opus 5 at its lowest setting, it is costing over twice as much. Wow.
¶13Crazy. Crazy. Unreal. This is like This is very impressive. This is Automation Bench.
¶14the same. Oh my, look at this. This is so crazy. So, okay, once again, a substantial improvement. And by the way, I want to say I'm very happy that they're putting so much emphasis on the cost per task.
¶15That is what matters. For example, with Kimmy K3, it was half the price of GPT 5.6 Soul, less than half the price of Fable 5, yet it took twice as many tokens to accomplish the same task. So, thus it's a wash. You're paying the same price. It's no longer sufficient to just look at the price per token uh and and and think, okay, well, that's the price I pay.
¶16Well, it's half the price. It's going to be less expensive. No, that's not it. It is the main metric you need to be looking at is cost per task. That is it.
¶17Now, what we're seeing here, this is incredible. So, it's it's not only cost per task, but it's also pass rate per task, which I think is fantastic. So, we have Opus 5 coming in way above everything else, way above and significantly less expensive. I'll also I'd like to say I said this a bunch of times in my videos in the last few weeks. I said, you know, people looked at Fable and they said, "Wow, it's so expensive.
¶18It uses so many tokens." And I said, "Well, that's their first version. They're going to optimize the hell out of it." And that's what we're seeing here with Opus. They basically used Fable probably, this is my assumption, to to really get the most out of Opus. Thanks to the sponsor of this video, Box, they tested Claude Opus 5 quite a bit. So, our benchmark of realistic document grounded tasks across 12 industries benchmarked here against the prior generation Claude Opus 4.8.
¶19The tasks mirror the analytical work that knowledge workers actually do. Reading source documents, reconciling numbers, running due diligence, and reviewing expert output for errors. Here's what we found. So, we do have a pretty sizable bump. Now, this does kind of reflect the bump we saw from Fable.
¶20I I can't believe I'm saying this, from Fable to Opus 5. So, from 63 to 78, great due diligence 65 to 76. Report drafting from data 67 to 69. Expert review just about the same. And then data analysis, a nice big six-point jump on complex enterprise knowledge work.
¶21Claude Opus 5 is a clear step up from the prior generation. It's advantage concentrated on the exhaustive multi-step analysis that drives real decisions. Opus 5 is coming soon to Box AI. Thank you to Box for sponsoring. Shout out to Box.
¶22Go check out Box. We use it at Forward Future. All of my agents use Box. We set up the Box CLI. They just know to put stuff there.
¶23It's really nice. And go build on top of Box if you're a developer. They're awesome. They've been a great partner. So, thank you to them.
¶24So, here's humanity's last exam. Once again, what we're seeing, the pricing actually looks to be about the same. Um, you know, it depends on your thinking level. That's what we're seeing. Each of these dots most likely represents a thinking level.
¶25Uh, we have Fable 5 and Opus 4.8, but Opus is the best. It's crazy to me that Opus is outperforming Fable. Does that mean the government is going to take it offline soon or we're going to have a lot of those kind of weird, hey, we can't let you use Opus 5 on this. You're going to have to fall back to Opus 4.8, which I do get pretty often. Um, but this is a really good score.
¶26Okay. Uh, Frontier Bench, we have here's GBT 5.6 Soul. So, right here is the max score coming in at about 3536% and coming in at about $12. Now, we have essentially a slightly better score for a little bit more of a cost at this thinking level for Opus 5, but for just like another $2 per task, you actually get another 5%. Um, and then interestingly, the score actually went down at the highest thinking level.
¶27Uh, and it was more expensive. So, that's interesting. Keep that in mind as you're thinking about which model to use. Okay, Arc AGI 3. the benchmark is it gives the model a a game to play or a set of games and nothing else.
¶28It doesn't even give it the name of the game. It doesn't tell it how to play. It doesn't tell it how to complete it. All it does is give it a game, right? So, like we'll click start and you have to basically just figure out what to do and you have a certain amount of moves, certain amount of time based on moves to complete it.
¶29And again, you have no idea what it is. And so an AI model's dropped in here and just says go and that's it. The games are solvable by humans. Uh but apparently very difficult for AI. Now look at this.
¶30The best model in the world prior to this was coming in at let's see maybe 8%. Now we had this massive jump up to 30%. This is such a massive jump. I don't think our three prize expected it. Opus 5 is stronger than Opus 4.8 on cyber security tasks, but it remains substantially behind Mythos 5 at developing exploits.
¶31Very interesting. Very interesting. So I think like the Mythos model has no guardrails on cyber. And I think that's the key. And so Opus 5 probably has not only guard rails, but they probably built it in a way in which they tried to remove some of that cyber capability, which is interesting.
¶32And again, this is all speculation. Um, which is interesting. Typically, when you remove capabilities from a model, you hurt the model generally, right? It it performs worse generally. But we're actually seeing the opposite effect here.
¶33Let's check out the blog. Let's see if it has any new information here. Yeah, I mean, look, this is stunning. Opus 5 outperforming Fable 5 at almost everything except for cyber security. And that that was probably their entire intention.
¶34Uh, Claude Opus 5 built a working wind tunnel. Nice. Very cool. I mean, it's not the first time that we've seen fluid simulation, but still impressive nonetheless. Claude Opus 5 is available today on all platforms priced at $5 per million input tokens and $25 per million output tokens.
¶35The same price as Opus 4.8, half the price of Fable, half the price of Fable, and comparable to GPT 5.6 Soul. So, this is Anthropic's answer to 5.6. 6 Soul. Now, also just keep in mind 5.6 Soul is likely the last iteration of the GPT5 series of models before OpenAI actually releases their next GPT6 model. Okay, interesting.
¶36So, automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5. So, they're doing the same thing. If they're safety classifiers, and hopefully they're not overly aggressive, it will fall back automatically. Route to a new model, you do pay the price of the fallback model to be clear, but it's annoying.
¶37I don't like when they do that. Yeah. So, so we're seeing already people saying Opus 5 has fallback to Opus 4.8 due to safety filters. They're probably more lax than Fable, but they still exist. Look at this novel problem solving by cost.
¶38Look at this score. 30% on the ARC AGI prize and it was able to do so less expensive than GPT 5.6 Soul less expensive and more than tripled its score. That is so crazy. Do we think that the government is going to leave it? Do you think they've already shown the model to the US government?
¶39Do you think they've already gotten approval for it? I would be doubtful that they did not share this model with the government that they didn't share the scores. And I think the interesting thing to note, and I'm going to just go back one more time to this. This is probably why we're not going to see a roll back. This is probably why Opus 5 is here to stay and was already reviewed by the government.
¶40But what we're seeing here, so exploitation success, a 13 with Mythos, zero with 4.8, and a four with Opus 5. So they really reduced the performance of the model's ability for cyber capabilities, but also improved the quality. And again, it's kind of crazy to be able to do that. Usually, when you put guard rails on a model, you also just reduce the overall performance of the model, but that's not what we're seeing. Here's Turk.
¶41Uh, Opus 5 rounds out our Cloud 5 family beautifully. Interesting. Does that mean Haiku 5 is not coming? I think it's an incredible daily driver. Pair it with Fable for planning, brainstorming, or fixing the hardest bugs.
¶42Pair it with Fable. Why would he say pair it with Fable? Interesting. There are three models now at the paro frontier which means a combination of quality and cost. Those models are now Opus 5, GPT 5.6 Soul and Kimmy K3, an open-source model.
¶43And I made an entire video about that model. Check it out right here.