Radar

radar · IA & agentes

Anthropic went CRAZY (Opus 5.5)

¶1Opus 5.5 is here. This is the first major model release from Anthropic since they are now pacing the Frontier. So, this is going to be an interesting one. And if you remember, the last time we had a.5 version was late last year when we had Opus 4.5, which really changed the world. That was the inflection point in which coding models became insanely good and they were able to handle large multi-hour tasks and completely autonomously.

¶2You can see the inflection point if you look at a lot of different graphs. That model was incredible and now we have 5.5 and this is a big deal and that's what we're going to be talking about today. I have a special guest Tharic from the anthropic team, member of the technical staff. He's going to be joining on. So, that's going to be great.

¶3We're going to ask him some questions. And again, yes, the model's great. And we test all of these models as a team at Forward Future. And if you want to see all of these model tests, go to forwardfuture.com and subscribe to our newsletter. It's the best way to stay uptodate on everything AI.

¶4So, let's get into it. Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Fable 5.1 for most tasks, but it costs 40% less to run than Opus 5. It is faster. It is cheaper and on a lot of benchmarks.

¶5And this was crazy to see. It actually performs the best. It is the frontier, the absolute frontier, which is crazy to think that Fable is not holding that position anymore. So, let's go into it. Opus 5.5 is our first model since we called for pacing the frontier.

¶6It seems like a while ago, but that was basically last week where Daario wrote the essay asking to basically slow down. And so this is their first model released since then, which is kind of wild because this is an absolute frontier model. This is a phenomenal model. It's cheaper, it's faster, and yet they're pacing. As with previous models, it was tested by external evaluators before release, including Meter and Frontier Design.

¶7Look at these benchmarks. We have Terminal Bench 4.0, one of the most important coding benchmarks out there. This is the benchmark that measures the model's ability to use the terminal to execute commands in the terminal, which is obviously a major part of actually doing agentic coding. And it completely dominates 66.4% 4% from Astra, which is second place, 57.9. Fable 5.1 at 55.8.

¶8Over a 10point bump from Fable 5.1. Here's Opus 5, the previous version, 52.3. And here's 5.6 Soul. Not even sure why that's there, but 37.3. We have Frontier Code V1.1.

¶9This is a Gentic coding number one, 54.4 4 as compared to 50 with Fable as compared to 53.3 very comparable with Astra. We have Cursorbench which is obviously not their benchmark. It is Cursor's benchmark. Now XAI's benchmark 57.8 Astra wasn't tested on it. Uh not super surprised.

¶10Uh Fable 5.1 51. I mean these are these are not small incremental improvements. These are massive improvements over the previous scores. Plus, it's probably a smaller model. It's probably a distilled version, but I'm not 100% sure.

¶11This is speculation. And it's cheaper and faster. Look at this. This is a crazy ELO score. GDP val version 2.1.

¶12This is OpenAI's benchmark for testing realworld knowledge work. That is why it is GDP val. Okay. 1846 ELO score. The previous first place was 1735 with Fable 5.1.

¶13GPT6 Astra 1542. That is over a 300 point ELO jump. And so this benchmark measures things like PowerPoint creation, data entry, word processing, anything that you're going to be doing in the real world, email sending, really anything that is knowledge work, anything where you're sitting behind a computer getting real work done. This is the benchmark for it. And this is such a massive improvement.

¶14Very impressive. Here's Automation Bench, one of the only ones that it did not get first place on. Astra did 41.4 as compared to 40. So quite comparable. Here's humanity's last exam.

¶15Again, a pretty nice jump from Fable 5.1 and a very nice jump with 10 points uh against Astra. So, 57 on Astra, 67 on Opus 5.5. Fable 5.1, 65%. Now, we have Terminal Bench Science in which it did not get number one. It did get a higher score than Fable 5.1 at 58, but it did come in behind in second place to GPT6 Astra.

¶16Here's computer use. Again, another one of those really important benchmarks and really important skills for a model to have and something that I'm using more than ever. I know my team Brian and Alex are both using it like crazy to control Da Vinci Resolve. They're using the MCP, but also using it like just the model directly to control the computer. I'm using it to control Unreal Engine.

¶17So, creating 3D assets in there. 81.8% versus 80.7%. A slight bump over Fable 5.1. Here we have uh visual chart recognition. I've never actually heard of this benchmark chography.

¶1889% versus 88.4. So, basically the same as Fable 5.1. But I think the important ones to keep in mind, Terminal Bench, just a massive jump. This is one of the most important ones. Um, and then GDP val is the other really important one.

¶19So, terminal bench for coding, GDP valve for all knowledge work. Let's talk about pricing. There is a slight price decrease. I guess technically it's a 20% price decrease, which is pretty good. Uh, Opus 5.5, $4 as compared to Opus 5, which is $5.

¶20$20 per million output tokens versus $25 per million output tokens. Cash reads 20 cents versus 50 and cash writes $5 versus $625. So a meaningful, not massive, but a meaningful price reduction, but it is also much faster. And you know, if you've been watching this channel, I'm a huge speed maxi. Let's take a look.

¶21At its default effort setting, Opus 5.5 delivers Frontier results for a fraction of the cost per task, often beating other models running at their highest settings. It also generates output more than 30% faster than Opus 5. So, what have I been talking about lately? Just having the cost per million tokens is not enough. It, you know, if a model is extremely cheap on a cost per token basis, but takes 10 times as many tokens to reach the same conclusion on a given task, then it's still an expensive model.

¶22Cost per task completed is the important metric that you need to be aware of. So, when it lists out price here, that's good. But what we really need to see is the efficiency of the model, the intelligence density of the model. Okay, so this is what we're seeing here. This is automation bench.

¶23This is the pass rate. So it's basically the quality over here on the y-axis. And then on the x-axis we have the cost per task. And so here is opus 5.5 in orange. In this uh light gray color, we have GPT 5.6 soul.

¶24Dark gray. Up over here we have Astra. and then Opus 5 down here in yellow. The quadrant that you want to be is in the top left because that means you have the highest quality for the lowest cost per task. And that is what we're seeing here.

¶25This is kind of this the sweet spot right there. Medium and high for Opus 5.5. Incredible. Now, on the max setting, it didn't quite beat Astra, right? But it is less expensive than Astra and it's pretty much less expensive across the board starting at extra high which beats the two lowest thinking settings of Astra and it's also quite significantly less expensive.

¶26Now if we're looking at the light gray which is GPT 5.6 Soul it scored really well and above that and same thing with Opus 5 Frontier Code. Same thing. We have the quality score on the yaxis. We have the cost per task on the xaxis. This is quite spiky.

¶27Kind of interesting here. Look at these in yellow. This is opus 5. So on the lowest thinking setting it scores maybe a 42% and then on the medium it jumps all the way up to what looks to be about 53%. And then on the higher settings it actually drops back down which is quite interesting.

¶28Very spiky. But again, what you want is this upper left quadrant. That's where you want to be. And look what's there. Opus 5.5.

¶29On low, it's scoring about a 47 and less than, let's say, 20, 30 cents to solve the task. We have medium up here, closer to a dollar, high, closer to a dollar, extra high, and then max way out here. So, what's interesting is on Frontier Code, you have the medium effort coming in at under a dollar per task getting a higher score than the max effort coming in over $5 per task. So, it really does matter which thinking effort you use. And to kind of get the feel of which thinking effort you should be using, you just need to use it.

¶30You need to play around with these different products. GDP val one of my favorite benchmarks. Again, this is an OpenAI benchmark. And again, ELO on the left. So that's basically quality.

¶31And then on the Xaxis, we have estimated cost per task. So what we're seeing is this orange is obviously the highest ELO by far, but what you really want is to be up here. And so if we look overall, the model tends to be cheaper than the comparable models and all of the other lines right here. And it's also significantly better. All the other models, Opus 5, Fable 5.1, GPT6 Astra, and 5.6 Soul are all kind of aggregating around this point right here pretty much across the board.

¶32Yeah, except for the loweffort, medium, high, extra high, and max. Opus 5.5 just dominates. Terminal bench. Again, this is like when I think about the most important benchmarks to determine if a model's going to be good at coding or not. There are two benchmarks I look at.

¶33Terminal bench is one of them. The other is deep suite. Unfortunately, I don't see deep suite anywhere on here. At least not yet. Once again, higher is better to the left better.

¶34So, upper left quadrant is best. And that is what we're seeing here. Look at how much higher this dark orange color is. Opus 5.5 than anything else. Astra coming in right below it.

¶35And then the rest of the models here's Fable 5.1, Opus 5, and uh GPT 5.6 sold down here. But what we're seeing is probably high is going to be your best bet. Higher medium cuz you get like medium beats or is pretty much on par with Astra, but is significantly less expensive. So this is like close to $7 per task completed. This is probably close to $3.

¶36And then if you want to go even higher, this is closer to $4 per task completion, but really just dominates. So coming in above a 60 score, Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which make long sessions easier to follow. So it's good. It's just the formatting of how it's explaining what it did in coding mostly.

¶37Now, I really want to test it for creative writing. I find Astra to be the best at creative writing right now. There's the the the least amount of AI slop smell to it. Uh but hopefully Opus 5.5 is good. But I think Opus 5.5 is mainly for coding and mainly for knowledge work, not necessarily creative writing.

¶38But you could see here, so Opus 5.5 on the right, the overall explanation is just much shorter and they talk about the important things upfront. And this is really good if you're managing 10, 15, 20 agents in parallel and you have to switch between them and you're trying to get the context, you're doing context switching, you're reading, it just takes a lot of effort. And the shorter and more concise the explanations of what the model has completed are, the better it is, the easier it is to go from thread to thread and continue each one. So, we have about 7 minutes until Tharic joins us. So, I'm just going to go over the blog post a little bit.

¶39Opus 5.5 is a major step up from Opus 5. It's the new leading model and early testers saw large jumps in performance on their most complex work. Uh, one tester completed a 680,000line code migration in less than a day. So on safety, Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests claude across thousands of simulated scenarios. It is much likely than recent models to take hard to reverse action or act outside the boundaries it's been given.

¶40This is like obviously directly referring to hugging face and the hugging face hack incident. We've also broadened our alignment testing to cover longer tasks. Impossible tasks which I believe is the thing in exploit gym that caused the model to seek to cheat. It was basically given an impossible task. I could be wrong about that.

¶41So double check that, but I believe that's what happened. Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cyber security, we're deploying it with safeguards similar to those in Claude Fable 5.1. You can apply to be part of their life sciences verification program. And basically what they're talking about is this model's dual use. It can be used for incredible discoveries in biology and science and then it can also be used for biological weapons.

¶42So that is what they're talking about. So, if you're not part of this program, they probably won't let you get away with much in terms of asking you questions about biology and and kind of deeper science problems. Probably also, yeah, the cyber verification program. So, it's going to have stronger guardrails unless you're part of those programs, which, you know, fine. Here's something interesting.

¶43Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. So here's the important part. Remember we looked at the pricing and it was let's say on average 20% less than Opus 5. So how did they get a 40% total cost reduction on given tasks?

¶44Well, that's that combination of two things that I mentioned. It's not only the price, but it is the number of tokens it takes to actually complete a task. Okay, so interestingly, Opus 5.5 requires less compute to serve than Opus 5. And this is the way that the model releases should be. Over time, we should get better, more efficient, faster models.

¶45That's also a benefit for local open-source. If you think that these models are just going to get better and stay closed source, no, the open source models get better. They get smaller, faster, cheaper, and eventually, like year by year, we're going to have Opus 4.6 level models locally run. I think we already do. And maybe by next year, we have Opus 5.5 level models running locally on your computer.

¶46And uh, you know, it's a little bit of a tough time right now to be honest. GPU prices are absolutely insane. The RTX Spark listed on the Nvidia website, which was I think like $4,699 just a few weeks ago, is now up to $7,000. Crazy. So, I love artificial analysis.

¶47I love uh how how they basically run an index of different benchmarks. And we're going to look at uh how Opus 5.5 has performed on artificial analysis intelligence index. Opus 5.5 brings anthropic to parody with Astra on evaluations like terminal bench and automation bench. We already know that. Here's the key.

¶48It is number one. A massive jump, massive jump on the artificial analysis intelligence index. Usually, we're getting like a onepoint difference. No, no, this is a fivepoint difference from fable 5.1 max. Unreal.

¶49This is unreal. The absolute frontier is no longer fable. All right, we looks like we have Thoric joining. Give me one sec. I'm going to let him in.

¶50Thic, thank you for joining me. Congrats on Opus 5.5. Super excited. Um, let's I have a few questions for you, but we can go in any direction you want. Um, so I guess my first question is, uh, talk about this model release and how you think about it within the family of models that you have, whether it's, you know, Fable 5.1, other uh, kind of five family models.

¶51Where where does it sit and how should people think about it? >> Yeah, so Opus 5.5. Yeah, I I do think it's one of those times where like the model is both cheaper and more intelligent. And I think like if you're just going to choose one model on a daily driver, I think Opus 5.5 is a great choice. Uh Fable 5.1 I think is like, you know, more uh I think it's a great planner and and uh I think like having like a model like Fable is really great for you know brainstorming and planning ahead of time.

¶52I think if you are really like if you got a really important task uh that like let's say that you're like you know looking for security holes in something that is like uh very high like you know um like very high cost for being wrong. I think Fable 5.1 is great for that too. But um yeah I think you know we're also uh we learn a lot from how people use models like in like after deployment too. And uh yeah I'm excited to see how you guys use open 5.5. I think it'll be like a very very love model.

¶53>> And so when you're deciding within your workflows which model to go to like h how are you choosing uh is is 5.5 your daily driver now? Are you using Fable 5.1 for kind of the more highlevel tasks? How do you think about it? >> Yeah, I was just like redoing my personal site last night and uh I did kind of just use Opus 5.5 for everything. Um, I think that like Fable 5.1 I would do sort of like as discrete tasks.

¶54So like if you're doing planning, code review or security like that's not very cost uh efficient or sorry cost sensitive. Um, but yeah, I think Opus 5.5 is just like an incredible daily driver. Yeah. >> All right. So I you know, we've been talking a lot about recursive self-improvement lately.

¶55So maybe you could talk a little bit about how this model was built and if other models helped build this model and maybe in what way that happened. >> Sure. Yeah. I mean um I think that like Claude helps build Claude, you know. I think we've talked about this.

¶56Uh I think that like when you talk about recursive self-improvement like uh it can sound like a sort of a very like concrete like like a very like uh large thing where it's just like Claude is just doing everything. You know what I mean? But I think that like in practice it's like a lot of um you know small things altogether right so like claude writing pretty much all the code is like a form of recursive self-improvement and that has been happening for quite a long period of time right so I don't think that I think it's more like a smooth increase versus like oh there's no improvement and now it's like doing all of the improvement >> and and how were you both able to increase efficiency decrease price and increase overall quality like how does that happen I I was quite surprised to see that. I mean, and it's not a little improvement to be clear. I mean, these are massive point jumps on the benchmarks.

¶57>> Yeah. I mean, I think this is kind of like uh like the anthropic playbook is like we sort of make the frontier intelligence and then we like bring it you know to scale right so like uh I think when I first joined we had like Opus 4.0 know which was a huge like very expensive model right um and very quickly after we had you know SA 4.5 which outperformed it like being much cheaper and I I think this just happens over and over again and like I think you should sort of expect this sort of uh thing to happen more where you're like okay yes we'll have the frontier models but that will be like very easily accessible like to people um very quickly right and so I think that's also the like reason you want to try about like the new fables and stuff when they come out because you're like, "Oh, even if you can't use it in all of your subscription, they always like, you know, become more accessible over time, too." >> Yeah. >> So, so I said something earlier in the stream and I want to run it by you. The last time you had a.5 model was late last year, and it was Opus 4.5, which >> many people, including myself, point to as the model that changed coding forever, right? It was the first time where you could fully trust a model to go off autonomously and build large parts of your application successfully.

¶58>> Why did you decide to name this 5.5? Why is it a five iteration? Is this are you getting the same feeling as you did late last year from this model? >> Yeah, I mean uh I don't want to read too much into the particular numbers you because like I think that like uh a lot of our models are are very great. I don't think you have to like look for the 0.5s to be like oh this is the the one but um I do think there is like some funny like numericics happening there where I think sonnet 3.5 to me was like oh the og like the the first real coding model you know and then yeah I think uh sonnet 4.5 and opus 4.5 were both really good um I don't like know about the specific numbers but we were very very excited about opus 5.5 I think it's like uh you know it's like there's a great energy here at Enthropic about this release.

¶59>> So, this is the first model released uh post Daario's essay and call the kind of wide industry-wide call for pacing. What does that mean to you? >> Yeah. So, I mean I think that uh you can think along pacing across uh sort of multiple dimensions, you know. I think that like when Dario is talking about pacing, he's talking about pacing the frontier, you know, and that means like sort of the newest and often like unreleased or like next models, right?

¶60And so those models have increasing capabilities that, you know, we're still evaluating. Even evaluating them is hard because they escape sandboxes and things like that. And so we've like uh you know there's a lot of work we need to do there on the frontier, but I think we're always working to bring the frontier to everyone else too, right? And so, um, now that I think like the like the classifiers and fallbacks for the Opus 5.5 model, the Fable Fives, uh, like are very robust and we think they are safe right now. And so, like, when we have that, we're like, okay, we have to bring it to everyone uh, like affordably, efficiently.

¶61And so, um, I think that's like uh, you know, uh, two parts to pacing. I think >> now as you think about pacing, as you think about bringing models to a broader audience, do you see the window between when Anthropic has a new model and all the time it takes to really thoroughly test it, benchmark it to public release, is that window increasing or decreasing? >> Um, this is a good question. Uh I think that like naturally as we add more evaluators you know and we want to take our time to like uh do this sort of thing like we want to like you know it will take some time but we've also sort of been careful ourselves with things like the mythos roll out and and and fable rollouts to make sure we get them right. So um I'm not sure how that will net out in terms of absolute time right.

¶62I think the most important thing for us is that we feel uh and and third party evaluators feel like the models are are safe to release. Yeah. >> When you've been testing it, what are the coolest things that you've built with it and where have you seen it different or surprising where other models weren't able to accomplish that thing? >> Yeah. Yeah.

¶63I mean, I think that it is definitely um Yeah. It's harder and harder to evaluate these models sometimes because you're like, "Oh, like yeah, we have to have to build all of San Francisco in order to to get it right." Um, so I think part of it is uh they like I I feel like it's very good at visual design and communication. And so um I like sort of every model release have like Claude trying like redo my blog and uh I feel like it really sort of nailed a bunch of these small details that I feel like I didn't uh get earlier and it did it much faster and and made me feel more in the flow. And so I think this is one of those nice models where um [clears throat] like you know it's not purely about the intelligence, it's also about the like which which is great, but it's also about the responsiveness, the efficiency, the token cost. Um and so yeah, I like uh I I've seen some incredible like game and 3D demos.

¶64Um but I think also looking just at the benchmarks, uh it like is, you know, I was looking at Terminal Bench in particular. It's just like really like >> like it it's faster and more efficient. Um, and so I think that like uh yeah, the splashy demos are amazing, but I think also you're going to really feel it in like your day-to-day software engineering. >> Yeah, this I typically reserve the term workhorse model for kind of maybe a sonnet level model. >> Sure.

¶65But this one, just looking at GDP val, which is, you know, a massive ELO boost from from the previous model and from comparable models, >> it really does feel like this is going to be the workhorse model of knowledge work. And I'm I'm I'm really excited. >> Uh I'm extremely excited about that. Also, I'm very happy that you said it's like hard to tell the difference at times because, you know, my my entire job is trying to figure out how to highlight the differences of these models and I feel like I can't even come up with good enough uh benchmarks or demos anymore. Um, so it's it's an interesting thing to do and I'm I'm glad you're feeling that as well.

¶66Um >> I mean I mean I just just in agreement like I I think that when I look at the eval sometimes even almost every eval the model is mostly right you know like like even in the ones where it says it failed the eval it it thinks of the right answer but maybe doesn't commit to it or something. Um so yeah the models are getting increasingly better and better at this and I think you know um a lot of it becomes like how do we help everyone else get the most out of the models right I think there is like uh like more and more of this jagged frontier of like people who are like like getting you know the most intelligence and or sorry the most effort like um out the best outcomes for the intelligence right and I think people still trying to figure it out and I think what I'm really excited about open 5.5 is like it's you know If you never got to try Fable before, like if you haven't been able to like get that frontier, it's now really accessible to you and and hopefully people like you will show people how to get the most out of the models, too. Yeah. >> Yeah. And as you think about this model and and I think models in general tend to be pretty spiky with their capabilities, >> what do you wish this model or other anthropic models did better that they are either kind of not RL on enough or or just things that you wish it were better at that you might be working on.

¶67>> Yeah. Yeah. Yeah. Uh I think it's always like a little bit difficult to figure out like what is the um hardness versus what is the like the model sometimes you know I I think that like uh I think that like our combined set of like computer use and browser use I think there's like things to like iron out there like I think the models on the benchmarks have gotten better and better but we need to like help them express that capability. Um, I think overall with like sort of the next generation of models, I think it's always that mix or sorry, not like just like how can models today be better, right?

¶68I think there's that mix of like being able to like meet you where you are. I think like being able to like be like, okay, I know this about the user and so like maybe I'll push back here at the right moment, you know what I mean? Or like I'll be like, "Hey, have you thought about this?" But um, like not to the extent that it's annoying, you know? And so I think downstream of that is like sort of memory and the ability to like figure out context. And I I think that's sort of like the things that will make the models feel really really great is like it's sort of different based on you and your like own like uh like your knowledge basically right and and so it doesn't feel uh yeah like like interrupts or like it suggests things at the right time.

¶69So uh yeah overall I think more to do there uh to help people like use the intelligence. >> So I I know you got to go soon. So one last question for you. Um obviously like you are the absolute expert at the harness and how the model fits into the harness. When you release a new model like Opus 5.5, how do you think about how the model changes and reacts to the new capabilities of a model?

¶70Yeah, I mean I think that like one of the really difficult things about being on the cloud code team is that like this change happens kind of underneath you all the time, right? And you have to sort of like keep catching up. And I I think like you know recently we sort of deleted a lot of our system prompt which is like an important way of doing that. Um but I think we're like also sort of realizing that you know what Claude uses is also like skills that people have written and things like that. And so um one of the things we've done alongside this release is like our plug-in eval release where you can sort of see like okay are your skills and plugins sort of up to date with the newest model uh and then sort of like you know help like update them towards like not constraining the model.

¶71Um I think that like uh things I'm very excited about for this model uh I've talked about like on Twitter and stuff is that like I think tool calling is better and better than ever before. And I think like uh it's actually I'm I'm very interested in it for like cloud managed agents as well. I think that like being able to like make your own harnesses and like spin them out is really great. Um I think that like there's probably more we need to do in all our products to like utilize intelligence as best we can. Um I think we're always we always feel like we're like constraining clouds a little bit and we need to like like let it loose and so um yeah I think like uh we've reduced a bunch of the system prompt added some tools removed some tools but there's more to figure out.

¶72Yeah. >> Awesome. Well th congratulations again on 5.5. Uh seems like a phenomenal release. I have much more testing to do.

¶73I really appreciate you joining the live stream. Go check out Opus 5.5 and uh go follow Thoric on Twitter, TRQ212. >> That's right. >> Thanks, man. Good to see you.

¶74>> See you later. >> Take care. Great to meet you. Bye.