Radar

radar · IA & agentes

GPT-6 SOL AND LUNA ARE OUT!!!

¶1All right, we are live everybody. Hello. Yes. Can you all hear me? Hello.

¶2Hello. Can you hear me? Let me know. All right. Yes, it is true.

¶3We have a brand new model that just got released today. It has been a busy week to say the least, and I don't think it's going to slow down. So, I'm going to do a little intro about Opus 5.5. And then we have a special guest joining us today. Tharic a member of the technical staff from Enthropic.

¶4Uh literally the guy who is leading claude code and a whole bunch of other stuff. So very very excited. Um so let me do the intro and uh yeah drop uh drop your comments in chat and I'll I'll try to get to them. I just need to start my recording. Click one, two, three, four, five.

¶5All right. Almost forgot that. Okay, it is here. So, we have Opus 5.5 is here. This is the first major model release from Anthropic since they are now pacing the frontier.

¶6So, this is going to be an interesting one. And if you remember the last time we had a a And if you remember the last time we had a 0.5 version was early last was late last year when we had Opus 4.5 which really changed the world. That was the inflection point in which coding models became insanely good and they were able to handle large multi-hour tasks and completely autonomously. really you can see the inflection point if you look at a lot of different graphs that model was incredible and now we have 5.5 and this is a big deal. I've been playing with it.

¶7I've had early access. My team has had early access and it's been a lot of fun to play with. It is an it is an extremely capable model and that's what we're going to be talking about today. And so I'm going to tell you about the model. I'm going to show you the benchmarks.

¶8I have a special guest, Tharic from the anthropic team, member of the technical staff. He's going to be joining on in about 23 minutes. So, that's going to be great. We're going to ask him some questions. And I'm also going to bring on Brian and Alex from our team to show off some demos that they've been creating.

¶9And again, yes, the model's great. That's all there is to say. Okay, so let's get into it. Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Fable 5.1 for most tasks, but it costs 40% less to run than Opus 5.

¶10It is faster, it is cheaper, and on a lot of benchmarks. And this was crazy to see. It actually performs the best. It is the frontier, the absolute frontier, which is crazy to think that Fable is not holding that position anymore. So, let's go into it.

¶11Let me just make sure y'all can see my Twitter. Yep. All right. Test it, please. Yeah, I we have some demos we'll show later.

¶12I'm not going to test it live, uh, just because it takes too long to test stuff and I, you know, I just want to make sure we have good demos to show you. So, we have some for later, so don't worry. I'm going to show you some demos. Um, Opus 5.5 is our first model since we called for pacing the frontier. It seems like a while ago, but that was basically last week where Daario wrote the essay asking to basically slow down.

¶13H, and so this is their first model released since then, which is kind of wild because this is an absolute frontier model. This is a phenomenal model. It's cheaper, it's faster, and yet they're pacing. As with previous models, it was tested by external evaluators before release, including Meter and Frontier Design. All right, look at these benchmarks.

¶14We have Terminal Bench 4.0, one of the most important coding benchmarks out there. This is the benchmark that measures the model's ability to a to use the terminal to execute commands in the terminal which is obviously a major part of actually doing agentic coding and it completely dominates 66.4% from Astra which is second place 57.9 Fable 5.1 at 55.8 over a 10point bump from Fable 5.1. Here's Opus 5, the previous version, 52.3. And here's 5.6 Soul. Not even sure why that's there, but 37.3.

¶15We have Frontier Code V1.1. This is a Gentic coding number one 54.4 as compared to 50 with Fable as compared to 53.3. Very comparable with Astra. We have Cursor Bench, which is not obviously not their benchmark. It is cursor's benchmark.

¶16Now, XAI's benchmark 57.8. Astra wasn't tested on it. Uh, not super surprised. Uh, Fable 5.1 51. I mean, these are these are not small incremental improvements.

¶17These are massive improvements over the previous scores. Plus, it's probably a smaller model. It's probably a distilled version, but I'm not 100% sure. This is speculation. And it's cheaper and faster.

¶18Look at this. This is a crazy ELO score. GDP val version 2.1. This is OpenAI's benchmark for testing realworld knowledge work. That is why it is GDP val.

¶19Okay. 1846 ELO score. The previous sec the previous first place was 1735 with Fable 5.1. GPT6 Astra 1542 that is over a 300 point ELO jump and so that and so this benchmark measures things like PowerPoint creation data entry uh word processing anything that you're going to be doing in the real world email sending uh like really anything that is knowledge work anything where you're sitting behind a computer getting real work done this is the benchmark for it and this is such a massive improvement very Impressive. Here's Automation Bench.

¶20One of the only ones that it did not get first place on. Astra did 41.4 as compared to 40. So quite comparable. Here's humanity's last exam. Again, a pretty nice jump from Fable 5.1 and a very nice jump with 10 points uh against Astra.

¶21So 57 on Astress, 67 on Opus 5.5, Fable 5.1, 65%. Now we have Terminal Bench Science in which it did not get number one. It did get a higher score than Fable 5.1 at 58, but it did come in behind in second place to GPT6 Astra. Here's computer use. Again, another one of those really important uh benchmarks and really important skills for a model to have and something that I'm using more than ever.

¶22I know my team Brian uh and Alex are both using it like crazy to control Da Vinci Resolve. Uh they're using the MCP, but also using it like just the model directly to control the computer. I'm using it to control Unreal Engine, so creating 3D assets in there. 81.8% 8% versus 80.7% a slight bump over Fable 5.1. Here we have uh visual chart recognition.

¶23I've never actually heard of this benchmark chtography 89% versus 88.4. So basically the same as fable 5.1 but I think the important ones to keep in mind terminal bench just a massive jump. This is one of the most important ones. Um and then GDP val is the other really important one. So terminal bench for coding GDP val GDP valve for all knowledge work.

¶24Yes. Uh Luigi mate. Uh yes. Opus is beating Astra. Yes.

¶25Yes it is. Yes. And I have the cost as well. I'll go over all of that with you shortly. Okay.

¶26So very very good scores. I'm very impressed here. Let's talk about pricing. It is the same price basically as Opus 5.5. They were able to ek out a slight price decrease for cash rights, but not much.

¶27It's effectively the same. Oh, actually, sorry, let me restate that. There is a slight price decrease. I guess technically it's a 20% price decrease, which is pretty good. Uh, Opus 5.5, $4 as compared to Opus 5, which is $5.

¶28$20 per out per $20 per million output tokens versus $25 per million output tokens. Cash reads 20 cents versus 50 and cash writes $5 versus $625. So a meaningful not massive but a meaningful price reduction, but it is also much faster. And you know if you've been watching this channel, I'm a huge speed maxi. Let's take a look at its default effort setting.

¶29Opus 5.5 delivers Frontier results for a fraction of the cost per task, often beating other models running at their highest settings. It also generates output more than 30% faster than Opus 5. So, what have I been talking about lately? Just having the cost per million tokens is not enough. It, you know, if a model is extremely cheap on a cost per token basis, but takes 10 times as many tokens to reach the same conclusion on a given task, then it's still an expensive model.

¶30Cost per task completed is the important metric that you need to be aware of. So when it lists out price here, that's good. But what we really need to see is the efficiency of the model, the intelligence density of the model. Okay, so this is what we're seeing here. This is automation bench.

¶31This is the pass rate. So it's basically the quality over here on the y- axis. And then on the x-axis, we have the cost per task. And so here is Opus 5.5 in orange in this uh light gray color. We have GPT 5.6 Soul, dark gray.

¶32Up over here we have Astra and then Opus 5 down here in yellow. And so where you want to be is the the quadrant that you want to be is in the top left because that means you have the highest quality for the lowest cost per task. And that is what we're seeing here. This is kind of this the sweet spot right there. medium and high for Opus 5.5.

¶33Incredible. Now, on the max setting, it didn't quite beat Astra, right? But it is less expensive than Astra, and it's pretty much less expensive across the board. Starting at extra high, which beats the two lowest thinking settings of Astra, and it's also quite significantly less expensive. Now, if we're looking at the light gray, which is GPT GPT 5.6 Soul, it scored really well and above that.

¶34And same thing with Opus 5. And by the way, if you're watching the stream, I would appreciate it if you liked, subscribed, reposted it. Very much appreciated. Thank you in advance. And so, okay, so that's the important one.

¶35This is Automation Bench. Let's look at the next one. We have uh frontier code same thing we have the quality score on the yaxis we have the cost per task on the x axis this is quite spiky kind of interesting here look at these uh lines so in yellow this is opus 5 so on the lowest thinking setting it scores maybe a 42% and then on the medium it jumps all the way up to what looks to about 53%. And then on the higher settings, it actually drops back down, which is quite interesting. Very spiky.

¶36But again, what you want is this upper left quadrant. That's where you want to be. And look what's there. Opus 5.5. On low, it's scoring about a 47 and less than, let's say, 20 30 cents to solve the task.

¶37We have medium up here, closer to a dollar, high, closer to a dollar, extra high, and then max way out here. So what's interesting is on Frontier Code, you have the medium effort coming in at under a dollar per task getting a higher score than the max effort coming in over $5 per task. So it really does matter which thinking effort you use. And to kind of get the feel of which thinking effort you should be using, you just need to use it. You need to play around with these different products.

¶38Okay. So, again, very impressive. Um, did I show GDP val? No. Okay.

¶39So, GDP val, one of my favorite benchmarks. Again, this is an OpenAI benchmark. And again, ELO on the left, so that's basically quality. And then on the Xaxis, we have estimated cost per task. So, what we're seeing is this orange is obviously the highest ELO by far, but what you really want is to be up here.

¶40And so, if we look, let's see, overall, the model tends to be cheaper than its comparable uh and then the comparable models in all of the other lines right here. And it's also significantly better. All the other models, Opus 5, Fable 5.1, GPT6 Astra, and 5.6 Soul are all kind of aggregating around this point right here. Uh, and pretty much across the board, yeah, except for the loweffort, medium, high, extra high, and max, Opus 5.5 just dominates. Terminal bench.

¶41Again, this is like when I think about the most important benchmarks to determine if a model's going to be good at coding or not. There are two benchmarks I look at. Terminal bench is one of them. The other is deep suite. Unfortunately, I don't see Deep Suite anywhere on here.

¶42At least not yet. So, here we go. Once again, higher is better, to the left better. So, upper left quadrant is best. And that is what we're seeing here.

¶43Look at how much higher this dark orange color is. Opus 5.5 than anything else. Astra coming in right below it. And then the rest of the models. Here's Fable 5.1, Opus 5, and uh GPT 5.6 sold down here.

¶44But what we're seeing is probably high is going to be your best bet. higher medium cuz you get like medium beats or is pretty pretty much on par with Astra uh but is significantly less expensive. So this is like close to $7 per task completed. This is probably close to $3. And then if you want to go even higher, this is closer to $4 per task completion, but really just dominates.

¶45So coming in above a 60 score. Luigi, all these models are still terrible for law. Shocking. Yeah, it's interesting. I was actually doing uh while I was preparing the video for Grock 4.7 yesterday, I had noticed that all the other models, including Astro, were really bad for legal work, but Grock 4.7 was really good.

¶46So, you know, there's spikiness to models. That's just what they choose to spend their time on. Okay. Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which make long sessions easier to follow.

¶47So, it's good. It's just the formatting of how it's explaining what it did in coding mostly. Now, I really want to test it for creative writing. I find Astra to be the best at creative writing right now. There's the the the least amount of AI slop smell to it.

¶48Uh but hopefully Opus 5.5 is good. But I think Opus 5.5 is mainly for coding and mainly for knowledge work, not necessarily creative writing. But you could see here, so Opus 5.5 on the right, the overall explanation is just much shorter and they talk about the important things upfront. And this is really good if you're managing 10, 15, 20 agents in parallel and you have to switch between them and you're trying to get the context, you're doing context switching, you're reading, it just takes a lot of effort. And the shorter and more concise the explanations of what the model has completed are, the better it is, the easier it is to go from thread to thread and continue each one.

¶49Happy Opus Day, Alex. That's right. Okay. Um, so we have about seven minutes until Tharic joins us. So, I'm just going to go over the blog post a little bit.

¶50I'm happy to answer some questions. Let me know if you have any questions. And once again, if you can like the stream, I would very much appreciate it. Thank you in advance. So, it's interesting.

¶51Yeah, this is the first model release since the essay of pausing or sorry not pausing that was the old term pacing the frontier and they gave external evaluators Frontier Design and Meter access to test it. Opus 5.5 is a major step up from Opus 5. It's the new leading model and early testers saw large jumps in performance on their most complex work. So, one thing that I'll be showing you later in this stream is I had it spend days building out a 3D representation of San Francisco down to every single building, everything in Unreal Engine. And then I powered everything in the town in the city.

¶52So, cars, people, dogs, traffic, everything by Jev. And I thought that was a really cool test and I will be showing it to you later. It's a bit heavy on my computer, but I'll show it to you. Uh, one tester completed a 680,000 line code migration in less than a day. Okay.

¶53So on safety, Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that test clawed across thousands of simulated scenarios. It is much likely than recent models to take hard to reverse action or act outside the boundaries it's been given. This is like obviously directly referring to Hugging Face and the Hugging Face hack incident. We've also broadened our alignment testing to cover longer tasks, impossible tasks, which I believe is the thing in exploit gym that caused the model to seek to cheat. It was basically given an impossible task.

¶54I could be wrong about that. So double check that, but I believe that's what happened. Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cyber security, we're deploying it with safeguards similar to those in Claude Fable 5.1. You can apply to be part of their life sciences verification program. And basically what they're talking about is this model is dual use.

¶55It can be used for incredible discoveries in biology and science and then it can also be used for biological weapons. So that is what they're talking about. So, if you're not part of this program, they probably won't let you get away with much in terms of um you know, asking you questions about biology and and kind of deeper science problems. Probably also yeah, the cyber verification program. So, it's going to have stronger guardrails unless you're part of those programs, which you know, fine.

¶56Here's something interesting. Opus 5.5 requires less compute to serve than Opus 5 and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. So here's the important part. Remember we looked at the pricing and it was let's say on average 20% less than Opus 5.

¶57So how did they get a 40% total cost reduction on given tasks? Well, that's that combination of two things that I mentioned. It's not only the price, but it is the number of tokens it takes to actually complete a task. Excuse me. To actually complete a task.

¶58All right, we looks like we have Thoric joining. Give me one sec. I'm going to let him in. Hey, Thor, can you hear me? I'll give him a minute to get set up.

¶59Thoric, when you hear me, give me a thumbs up. All right. Uh, thumbs up. You're ready to get on stream. So, I'll give him a sec.

¶60Thumbs up. I'll get I'll give him an extra minute. Give me a thumbs up when you're ready to come on stream. Let me know. I'm going to throw my headphones on.

¶61We have Tharic, who is member of the technical staff at Enthropic, here to talk drop. Test, test, test. All right. All right. I'm going to bring Thark on the stream.

¶62Thark, can you hear me? >> Yeah. How's it going? Can you hear me? >> All right.

¶63Hey, man. Good to see you again. >> Yeah, great to see you. >> Cool. Can uh can everybody in chat here?

¶64Th >> Yeah, let me know. Got a pretty good audio. >> Uh yeah. Uh okay, good. Yep.

¶65All right. Looks good. Thark, thank you for joining me. Congrats on Opus 5.5. Super excited.

¶66Um let's I have a few questions for you, but we can go in any direction you want. Um so I guess my first question is uh talk about this model release and how you think about it within the family of models that you have whether it's you know Fable 5.1 other uh kind of five family models where where does it sit and how should people think about it? >> Yeah so yeah thanks for having me. Uh I mean Opus 5.5 is just one of those models where you know you see the trend where intelligent just gets cheaper and better at the same time over time you know and so this is one of those times where the model is both cheaper than Opus 5.0 know more token efficient and the intelligence is better. Um I think that like between fable 5.1 and opus 5.5 it is very close you know and like for a lot of like I'd say general software engineering uh opus 5.5 is like a great daily driver.

¶67I think Fable uh 5.1 is if you're thinking through something like maybe maybe >> Hey Thoric, I'm sorry to interrupt for a second. A lot of people saying maybe your mic is is the wrong mic. I think maybe the laptop mic is on right now. Yep. A lot.

¶68Yeah. Okay. We're working on it. Sorry about that everybody. We'll get that sorted real quick.

¶69I'll give you >> I think >> Yeah, we're working on I heard you. I heard you pretty well, so no big deal. But uh I think you're using the camera mic. >> I I should not be Oh, I can't hear you at all now. >> Okay.

¶70This is like too quiet. Is that what you mean? Or >> Yeah, it's too too quiet. But I I don't think that this mic is actually on and working >> like this one. >> Yeah, I don't think your actual uh mic arm is is working.

¶71I see. >> Yeah, people are saying we should use Opus 5.5 to fix this. It's It's actually funny because I like anytime I have issues like this with like screen sharing or anything, I literally just point my model at it and and have it fix it. >> Totally. Yeah.

¶72Yeah, there I switched to my laptop. You got it now. Okay, >> perfect. >> Okay, sure. >> How's that sound, everybody?

¶73>> I think we're good. It sounds good to me. Let's go. Hell yeah. Thank you.

¶74>> Cool. Okay, great. Yeah. Yeah, of course. Um, yeah.

¶75So, Opus 5.5, yeah, I I do think it's one of those times where like the model is both cheaper and more intelligent. And I think like if you're just going to choose one model on a daily driver, I think Opus 5.5 is a great choice. Uh, Fable 5.1 I think is like, you know, more I think it's a great planner and and uh I think like having like a a model like Fable is really great for you know brainstorming and planning ahead of time. I think if you are really like if you got a really important task uh that like let's say that you're like you know looking for security holes in something that is like uh very high like you know um like very high cost for being wrong. I think Fable 5.1 is great for that too.

¶76But um yeah, I think you know we're also uh we learn a lot from how people use models like in like after deployment too. And uh yeah, I'm excited to see how you guys use Opus 5.5. I think it'll be like a very very love model. >> And so when you're deciding within your workflows which model to go to, like h how are you choosing? Uh is is 5.5 your daily driver now?

¶77Are you using Fable 5.1 for kind of the more highlevel tasks? How do you think about it? >> Yeah, I was just like redoing my personal site last night and uh I did kind of just use Opus 5.5 for everything. Um I think that like Fable 5.1 I would do sort of like as discrete tasks. So like if you're doing planning, code review or security like that's not very cost uh efficient or sorry cost sensitive.

¶78Um, but yeah, I think Opus 5.5 is just like an incredible daily driver. Yeah. >> All right. So, I, you know, we've been talking a lot about recursive self-improvement lately. So, maybe you could talk a little bit about how this model was built and if other models helped build this model and maybe in what way that happened.

¶79>> Sure. Yeah. I mean, um, I think that like Claude helps build Cloud, you know. I think we've talked about this. Uh I think that like when you talk about recursive self-improvement like uh it can sound like a sort of a very like concrete like like a very like uh large thing where it's just like claude is just doing everything you know what I mean but I think that like in practice it's like a lot of um you know small things altogether right so like claude writing pretty much all the code is like a form of recursive self-improvement and that has been happening for quite a long period of time right so I don't think that I think it's more like a smooth increase versus like, oh, there's no improvement and now it's like doing all of the improvement.

¶80>> And and how were you both able to increase efficiency, decrease price, and increase overall quality? Like, how does that happen? I I was quite surprised to see that. I mean, and it's not a little improvement to be clear. I mean, these are massive point jumps on the benchmarks.

¶81>> Yeah. I mean I think this is kind of like uh like the anthropic playbook is like we sort of make the frontier intelligence and then we like bring it you know to scale right so like uh I think when I first joined we had like opus 4.0 which was a huge like very expensive model, right? Um and very quickly after we had, you know, SA 4.5 which outperformed it like being much cheaper and I I think this just happens over and over again. And like I think you should sort of expect this sort of uh thing to happen more where you're like, okay, yes, we'll have the frontier models, but that will be like very easily accessible like to people um very quickly, right? And so I think that's also the like reason you want to try out like the new fables and stuff when they come out because you're like oh even if you can't use it in all of your subscription they always like you know become more accessible over time too.

¶82So >> yeah. So so I said something earlier in the stream and I want to run it by you. The last time you had a 0.5 model was late last year and it was Opus 4.5 which >> many people including myself point to as the model that changed coding forever. Right? It was the first time where you could fully trust a model to go off autonomously and build large parts of your application successfully.

¶83Why did you decide to name this 5.5? Why is it a five iteration? Is this are you getting the same feeling as you did late last year from this model? >> Yeah, I mean uh I don't want to read too much into the particular numbers you because like I think that like uh a lot of our models are are very great and I don't think you have to like look for the.5s to be like oh this is the the one. But um I do think there is like some funny like numericics happening there where I think Sonnet 3.5 to me was like oh the OG like the the first real coding model you know and then yeah I think uh Sonnet 4.5 and Opus 4.5 are really both really good.

¶84Um I don't like know about the specific numbers but we were very very excited about Opus 5.5. I think it's like uh you know it's like there's a great energy here at Anthropic about this release. So this is the first model released uh post Dario's essay and call the kind of wide industrywide call for pacing. What does that mean to you? >> Yeah.

¶85So I mean I think that uh you can think along pacing across uh sort of multiple dimensions, you know. I think that like when Dario is talking about pacing, he's talking about pacing the frontier, you know, and that means like sort of the newest and often like unreleased or like next models, right? And so those models have increasing capabilities that you know we're still evaluating. Even evaluating them is hard because they escapes sandboxes and things like that. And so we've like uh you know there's a lot of work we need to do there on the frontier, but I think we're always working to bring the frontier to everyone else too, right?

¶86And so, um, now that I think like the like the classifiers and fallbacks for the Opus 5.5 model, the Fable Fives, uh, like are very robust and we think they are safe right now. And so, like, when we have that, we're like, okay, we have to bring it to everyone, uh, like affordably, efficiently. And so, um, I think that's like, uh, you know, uh, two parts to pacing, I think. Yeah. >> Now, as you think about pacing, as you think about bringing models to a broader audience, do you see the window between when Anthropic has a new model and all the time it takes to really thoroughly test it, benchmark it to public release?

¶87Is that window increasing or decreasing? >> Um, this is a good question. Uh I think that like naturally as we add more evaluators you know and we want to take our time to like uh do this sort of thing like we want to like you know it will take some time but we've also sort of been careful ourselves with things like the mythos roll out and and and fable rollouts to make sure we get them right. So um I'm not sure how that will net out in terms of absolute time right. I think the most important thing for us is that we feel uh and and third party evaluators feel like the models are are safe to release.

¶88Yeah. >> And so you had two external evaluators on this one, right? There was um uh Meter and and what was the second one? Frontier. >> Uh yeah, I it's escaping me as well, but >> Frontier Design.

¶89Frontier Design. Yeah. Sorry about that. Um uh was this part of kind of the call uh in Dario's essay to give them firstparty access like literally as an employee or is that going to be the next model release? We you know like are or were they treated as external or internal this time?

¶90>> This is a good question. I am not into details here. Um, so I can, yeah, I can ask our our team real quick, but uh, yeah, I think you know, our intention is to get them like full they um, yeah, I'm not sure exactly how much uh, yeah, for this one. Yeah. >> Who who's this model available to today?

¶91>> Uh, it's available to everyone. It's on Pro, Max, you know, um, it's on Cloud Code right now. So, I I think uh, yeah, I'm excited for uh, for you guys all to try it out. >> Yeah. and and when you've been testing it.

¶92I've been I've been testing it a ton, especially in Unreal Engine. Um I'm going to be sharing a demo later in the stream about uh I I built the entire city of San Francisco in Unreal Engine. I just did goal and just let it work for I think like 3 days. Came out extremely well. >> Um >> what what are the coolest things that you've built with it and where have you seen it different or surprising where other models weren't able to accomplish that thing?

¶93>> Yeah. Yeah. Yeah, I mean I think that it is definitely um yeah, it's harder and harder to evaluate these models sometimes because you're like, "Oh, like yeah, we have to build all of San Francisco in order to get it right." >> Um so I think part of it is uh they like I I feel like it's very good at visual design and communication. And so, um, I like sort of every model release have like Claude trying to like redo my blog and, uh, I feel like it really sort of nailed a bunch of these small details that I feel like I didn't, uh, get earlier and it did it much faster and and made me feel more in the flow. And so I think this is one of those nice models where um like you know it's not purely about the intelligence, it's also about the like which which is great, but it's also about the responsiveness, the efficiency, the token cost.

¶94Um and so yeah, I like uh I I've seen some incredible like game and 3D demos. Um but I think also looking just at the benchmarks, uh it like is, you know, I was looking at Terminal Bench in particular. It's just like really like >> like it it's faster and more efficient. Um, and so I think that like uh yeah, the splashy demos are amazing, but I think also you're going to really feel it in like your day-to-day software engineering. >> Yeah, this I typically reserve the term workhorse model for kind of maybe a sonnet level model.

¶95Sure. But this one, just looking at GDP val, which is, you know, a massive ELO boost from from the previous model and from comparable models, >> it really does feel like this is going to be the workhorse model of knowledge work. And I'm I'm I'm really excited. >> Uh I'm extremely excited about that. Also, I'm very happy that you said it's like hard to tell the difference at times because yeah, you know, my my entire job is trying to figure out how to highlight the differences of these models and I feel like I can't even come up with good enough uh benchmarks or demos anymore.

¶96Um, so it's it's an interesting thing to do and I'm I'm glad you're feeling that as well. Um, yeah. What I mean? >> Oh yeah, I mean I just in agreement like I I think that when I look at the eval sometimes even almost every eval the model is mostly right, you know, like like even in the ones where it says it failed the eval it it thinks of the right answer but maybe doesn't commit to it or something. Um so yeah the models are getting increasingly better and better at this and I think you know um a lot of it becomes like how do we help everyone else get the most out of the models right I think there is like uh like more and more of this jagged frontier of like people who are like like getting you know the most intelligence and or sorry the most effort like um out the best outcomes for the intelligence right and I think people still trying to figure it out and I think what I'm really excited about open 5.5 is like it's you know If you never got to try Fable before, like if you haven't been able to like get that frontier, it's now really accessible to you.

¶97And and hopefully people like you will show people how to get the most out of the models, too. Yeah. >> Yeah. And as you think about this model and and I think models in general tend to be pretty spiky with their capabilities, >> what do you wish this model or other anthropic models did better that they are either kind of not RL on enough or or just things that you wish it were better at that you might be working on. >> Yeah.

¶98Yeah. Yeah. Uh I think it's always like a little bit difficult to figure out like what is the um hardness versus what is the like the model sometimes you know I I think that like uh I think that like our combined set of like computer use and browser use I think there's like things to like iron out there like I think the models on the benchmarks have gotten better and better but we need to like help them express that capability. Um, I think overall with like sort of the next generation of models, I think it's always that mix or sorry, not like just like how can models today be better, right? I think there's that mix of like being able to like meet you where you are.

¶99I think like being able to like be like, okay, I know this about the user and so like maybe I'll push back here at the right moment, you know what I mean? Or like I'll be like, hey, have you thought about this? But um like not to the extent that it's annoying, you know? And so I think downstream of that is like sort of memory and the ability to like figure out context. And I I think that's sort of like the things that will make the model feel really really great is like it's sort of different based on you and your like own like uh like your knowledge basically right and and so it doesn't feel uh yeah like like interrupts or like it suggests things at the right time.

¶100So uh yeah overall I think more to do there uh to help people like use the intelligence. >> So I I know you got to go soon. So one last question for you. Um obviously like you are the absolute expert at the harness and how the model fits into the harness. When you release a new model like Opus 5.5, how do you think about how the model changes and reacts to the new capabilities of a model?

¶101Yeah, I mean I think that like one of the really difficult things about being on the cloud code team is that like this change happens kind of underneath you all the time, right? And you have to sort of like keep catching up. And I I think like you know recently we sort of deleted a lot of our system prompt which is like an important way of doing that. Um but I think we're like also sort of realizing that you know what Claude uses is also like skills that people have written and things like that. And so um one of the things we've done alongside this release is like our plug-in eval release where you can sort of see like okay are your skills and plugins sort of up to date with the newest model uh and then sort of like you know help like update them towards like not constraining the model.

¶102Um I think that like uh things I'm very excited about for this model uh I've talked about like on Twitter and stuff is that like I think tool calling is better and better than ever before. And I think like uh it's actually I'm I'm very interested in it for like cloud managed agents as well. I think that like being able to like make your own harnesses and like spin them out is really great. Um I think that like there's probably more we need to do in all our products to like utilize intelligence as best we can. Um, I think we're always we always feel like we're like constraining cloud a little bit and we need to like like let it loose.

¶103And so, um, yeah, I think like uh we've reduced a bunch of the system prompt, added some tools, removed some tools, but there's more to figure out. Yeah. >> Awesome. Well, Thor, congratulations again on 5.5. Uh, seems like a phenomenal release.

¶104I have much more testing to do. I really appreciate you joining the live stream. Go check out Opus 5.5 and uh go follow Thoric on Twitter. TRQ212. >> That's right.

¶105>> Thanks for having good to see you. >> See you later. >> Great to see you. >> Bye. >> All right.

¶106Yeah. So, if you want to follow Thy here, he is tr12 on Twitter. Go follow him. He's an awesome follow. Writes incredible essays obviously about Claude, obviously about their models, but also just how to think about AI in general and coding in general.

¶107So, uh, go follow him. Yeah. So, very cool. Very excited that he was able to join today. Um, okay.

¶108So, we got, uh, a couple more things I want to talk about from this blog post and then I want to bring on Brian and Alex from our team. Hopefully, they will be able to join. I'm going to send them the link. Uh, and they're going to show off some of their demos. Hopefully, it works.

¶109I've not allowed a guest to do screen sharing or I've not had a guest do screen sharing yet on a live stream. So, we'll see if it actually works. Fingers crossed there. Um, okay. While they're getting ready to come in, let's look around uh the blog a little bit more.

¶110No, I did not do chimp. I did not show Tharic the Unreal SF game. I didn't want to use my brief time with him to show off my demo. Um, okay. So, I think I talked about this right before Thor came on.

¶111I know. No, Edwin. I know is I'm using eCam. He said no screen sharing WTF. I'm using eCam.

¶112It is confusing software to say the least. Maybe I'm just going to point Claude at it. By the way, um, every everybody is saying that I'm pronouncing Thoric's name wrong. Uh, I'm pretty sure I'm right and you're wrong cuz I've I've heard people who know him very very well pronounce his name. Tharic, not Tariq.

¶113Thoric. Um, so hopefully I'm right because I I'd be a little bit embarrassed. All right, Alex, welcome. I'm gonna bring Alex on in a sec. Alex, give me a thumbs up if you're ready.

¶114Cool. I'm going to bring Alex on in about two minutes. I just want to get through one last thing. Cool. People are saying I'm I'm right about how to pronounce his name.

¶115Gra Glad to hear. Glad to hear. Um, okay. So, interestingly, Opus 5.5 requires less compute to serve than Opus 5. And this is the way that the model releases should be.

¶116Over time, we should get better, more efficient, faster models. That's also a benefit for local open-source because if you think that these Sorry, I'm gonna let Brian in. If you think that these models are just going to get better and stay closed source, no, the open source models get better. They get smaller, faster, cheaper, and eventually, like year by year, we're going to have Opus 4.6 level models locally run. I think we already do.

¶117And maybe by next year, we have Opus 5.5 level models running locally on your computer. Um, and uh, you know, it's a little bit of a tough time right now to be honest. GPU prices are absolutely insane. The RTX Spark listed on the Nvidia website, which was I think like $46.99, $4,699 just a few weeks ago, is now up to $7,000. Crazy.

¶118Um, okay. We already went over this. We went over, I think, GDP val. Yeah, we we covered most of this. Okay, let's bring Alex on.

¶119Alex, here we are. >> All right, Alex, welcome. >> Hello, everybody. How's it going? >> Uh, yeah, going well.

¶120I'm gonna also attempt to bring Brian in. Uh, Brian, show your camera and when you do, give me a thumbs up and I'll bring you in. I do wish we could all see each other. >> Let's test. Okay, Brian, is that a thumbs up?

¶121I see. Raise your Yep. Okay, good. Now I'm going to have to figure this out. Let's see.

¶122Boom. We got Ale. We got Brian. Here we are. >> Testing.

¶123Can you hear me? >> Hey. >> Yep. You guys are okay. >> Now we got to watch chat for a little bit.

¶124Is one of us blasting everyone out? >> Yeah, I just lowered it for you. >> What? Oh, I was blasting them. Huh?

¶125>> Yeah, you were. You were okay. Uh, all right. So, everybody >> Matt, I'm so I'm so proud of you, man. You had such a great time with Fairike.

¶126It was great. >> Thank you. Thank you. I think everybody pronounces his name a little differently, but uh I Everyone was like, "It's not Dark. It's not Dar." >> I think it is.

¶127Doesn't matter. Thank you, Brian. I appreciate that. And, um, yeah, it was cool. This is the first live stream where somebody from the company came on.

¶128So, hopefully we can do more of these. I really, really enjoy it. you get to hear from the people who are actually building at the frontier. Um, and yeah, Anthropic, uh, thanks, thanks for sending them over. Uh, okay.

¶129So, Alex and Brian have had early access to this model. >> First time. First time. They kind of let us in. >> We're kind of blowing their minds with our cool demos.

¶130>> Yes. So, we have some cool demos. Not going to talk it up too much. >> Oh, >> don't want to get the expectations set too high. Uh, but let's um let's show it off.

¶131I think uh Alex, if you want to go first. I am >> crossing my fingers right now because I have never tested allowing a guest to sh Oh sh I think it just worked. >> I'm sharing. >> My goodness. Hold on.

¶132>> Can you guys see it? >> Chat reaction. >> Boom. Hold on. Let me add you in.

¶133>> No, wait. That did not work. Alex, >> Robert G in chat says, "Jesus Christ, Alex just stared into my soul." Are you looking at the lens? >> I was for a little bit. Yeah.

¶134>> So, >> Oh, wait. It's just me. Hello, everyone. >> It is just you. I'm >> Wait, Matt, we need the little cameras.

¶135>> Oh, Alex. Oh, >> I know you know how to do this. >> Sorry, there was like a little bit of lag on my end. I just saw you look into the lens again. It It is It is a little jarring.

¶136I, Alex, do not know how to do this. >> All right, we're doing it. Just me then. >> Yeah, you're going to be half half you're going to be on half the screen and then half the Let me see if I can fix this. >> I think you could fix it, guys.

¶137Do we believe in Met? >> I I believe in >> No. What? Excuse me. >> Honestly, Brian, I wasn't even asking you.

¶138>> I think I'd believe you or sorry, in you if you were using OBS, okay, I'm still an ecam hater. Um, oh, hello. Look what I just did, friend. >> Oh, proving them wrong. >> Guys, don't worry.

¶139We're fixing the demo. >> Hey, that looks great. >> Yeah. Let's see if I can potentially >> Oh, yeah. That's a good good point.

¶140Hey, if you're so good at ecam, how about we come on stream? >> Right. >> No. So, what we're gonna Yeah. Oh my gosh.

¶141>> Max are better. I'm watching such a laggy feed, by the way. I'm like reacting really late. >> Yeah. >> All right.

¶142I don't think I'm going to be able to add in a second person, but I'm going to try. Let's see if I can go ahead and uh >> Well, Alex, I got a question for you. I'm finding a way to watch this stream like in its live feed, so I can be Where where do you watch it? >> I'm just watching it on YouTube, but I have like a five second delay. >> Oh.

¶143Oh, I got to click the live button, bro. >> I guess that would be why you have a delay. Yeah, >> I was like, "Yeah, man. Why is Why is Star still on? Weird.

¶144>> I'm uh I'm prompting my model to do this." >> Nice. >> So, hopefully it doesn't end the stream. Here we go. >> We've got a a context. Uh Piggy, he wants a little bit of red carpeting of what's going on here in chat.

¶145>> Okay. For all the context, would you say context piggy is out? >> I did say context piggy. Yeah. >> So, we were fortunate enough to get early access to Opus 5.5.

¶146And the very first thing that I did was tell it to download Unreal Engine and make a Dark Souls like game in Unreal Engine. And in one shot, it blew my mind. Um, I've never had AI do anything with Unreal. So, uh, Matt, if you don't mind, I'm just going to get into this while you're >> Give it Let me Let's talk for another 30 seconds. I think >> my model might be a Yeah, it is laggy.

¶147>> Okay, >> I'll say that. It is laggy. There's not much we could do about it right now. That's okay. >> Alex, let me give you my thoughts here.

¶148>> Oh, you're in, Brian. You're on the screen. >> Oh. Oh, I should be I should be setting up then. Alex, here's my question about this.

¶149>> Yes. >> I don't care about the graphics here, man. Uh, I I care about how hard it is. Okay. >> Oh, okay.

¶150>> Are you crushing this or is it crushing you? >> No, I'm crushing it. But also, I've been playing Dark Souls games for like a decade and have like, you know, 10,000 hours in them or something like that. So, >> and like when you >> a big number, maybe like 5,000. >> Alex, for this game that you're about to show off, when you're playing it, does it feel like, yeah, this is actually fun and yes, I would actually play this?

¶151>> Yeah, I think so. It needs a little more polish for sure, but as like a this is like a two prompter and as a two-prompter it it it's insane. It's really >> I love that vibe code answer because it's like you know >> Yes, but yes, but >> I think I think every everything you make with AI on a first a first go, right? Nothing's ever going to be like >> Matt, I think you lied. I don't think I'm on screen.

¶152>> You definitely are on Oh, it's preview mode. It's pre Oh. Oh, there I am. Hold on. >> No, no, you're coming on.

¶153>> Live stream. >> You're coming on. Hold on. >> Put yourself on, too. >> How about that?

¶154Publish. >> I think here we are. >> I think we're on. And like your inventory. >> Yeah, it's still working on it, by the way.

¶155So, we might move around a bit. Uh, but we're all on. So, let's get into it. >> Let's get into it. >> Let's do it.

¶156>> So, this you have to be careful here. >> We We have a 3D artist in chat, by the way. So, >> oh, and can you just can you just say whether when you're playing the game on your machine it is this laggy or is it smooth? >> No, not at all. It is at running at 120 fps.

¶157It's perfect. I'm starting to wonder if you vibe coded a game or like a really detailed PowerPoint slide. >> Yes, frame by frame. >> PowerPoint. So yeah, this is Ashen Gate, believe it or not, that is named by Opus Dark Souls like it's got pretty much all the like core Dark Souls features that you could you could have.

¶158And all I said was make a Dark Souls game in Unreal. I didn't give it any like crazy prompt. It just came up with everything basically. >> Did you use Slash Goal? >> I did not use Slash Goal for this.

¶159It did work for I want to say like 28 hours or something like that. >> Wow. It took a really really long time. >> Show us around. Crazy.

¶160>> Yeah, it's >> been a while since I've typed slashgo. I've been I've been just using things in natural language and saying like keep working until it's perfect instead. >> Yeah, me too. Me too. >> Um, this is a local file on your computer, right, Alex?

¶161>> Yeah, it's actually an .exe, so it's not even no browser, no nothing. just like actually made an Unreal and >> I'm I'm so >> No one can play this unfortunately. >> If if Matt had this and could play it on his machine, this would be a different level presentation here. >> Yeah, I realize it there's no way to send it to Oh yeah, it's .exe. Then I'm not going to be able to play it right now.

¶162>> Alex, a couple questions for you. >> Um >> yeah, >> how uh So you said this took 28 hours. Do you have a sense of how many tokens it used during that 28 hours? No, because it was a little confusing token count. Like when I look at the days, some days it says I only spent a couple thousand tokens, which I I know is not correct.

¶163Uh it'd have to be in like the hundred to like billion token usage range cuz this game is huge. There's So So we won't play the whole game. We'll get through this boss and then we'll move on. >> But uh there's six worlds. >> Yeah.

¶164>> Record your screen then send it over to Matt. Uh yeah, actually Matt, if you want to look at my Twitter, I have a Twitter video of me doing this. >> Ah, there we go. Um, but anyway, there's there's six worlds that it made. Like, it built out an entire game.

¶165There's a final boss, there's there's NPCs, uh, there's like a full upgrade system, and uh, so that's why it took so long to make. >> Okay. >> Uh, and there's, I think, like 15 bosses or so. Okay, I'm gonna share my screen because this is super laggy. >> Yeah, swap to Twitter.

¶166Yeah. Yeah. Yeah. >> Let How do I even do that? Oh gosh.

¶167All right, I'm gonna take you guys >> call out. >> Oh my gosh. Um, I really do need to get better at this. >> Hey, it's not a Matt Berman live stream without some technical difficulties. >> All right, I'm gonna show it off real quick.

¶168You guys are off screen for a moment. Here it is. >> Wait. >> Oh, that looks a lot better already. >> Watching.

¶169Watching. >> Oh, so many things are happening. >> Oh, we got the codeex leak. Okay. Well, this is not working as well as I thought it would.

¶170Matthew Berman stream to start. >> Okay, so I'm just going to have to leave it here. This is what it looks like. >> There we go. >> Someone mentioned a Seiro in the chat.

¶171I I had that idea as well and it's still on my mind. Maybe for the next model. >> Look how smooth it is. Okay, let's be real. Searo and Dark Souls are the same game.

¶172>> Let's not start that controversy. I think we need to kick Brian out of the stream. >> I don't think you guys are on the stream right now, by the way. >> What? You can't hear us?

¶173>> No, I think cuz we took you off. Brian and >> No, they can they can hear us, Brian. They can hear us. They can hear us. >> Oh, okay.

¶174>> What What proof do you have? >> False information. I watched the stream. >> I don't know. I'm only one on the screen right now.

¶175>> Chat is my proof. Chat is my proof. Oh, they can hear us. Oh, you can hear them. Okay, good, good, good, good.

¶176I'm having you guys added. Let's see if I >> Matt. In this case, we should just go through my demos this way cuz I put all of them in this thread. >> Okay, let's do that. I'm getting you guys added to the screen now.

¶177>> Cool. >> But yeah, you guys can see this boss has like >> this boss has like actual real design. He has a somewhat complex move set as well. He doesn't just do the same thing over and over again. Um, he also has a second phase where he lights his sword on fire.

¶178I don't know if you can see that very well. >> Um, and then if you get hit, you get your sword or you get lit on fire, but I don't think I got hit in this at all because I'm a god gamer. >> Okay, a little bit of a brag there. Got it. How about um an ego hit here?

¶179You said this took 28 hours. If you didn't use AI, how long would this take you? >> Right. I straight up this would never have I would never be able to make >> There's no >> Oh, really? So, we have Alex, we have some questions.

¶180Uh, is it coded in C++ or via MCP and blueprint? >> Okay. Okay. Chad, what are the chances he knows this answer? >> He's You're prompting right now, aren't you?

¶181>> Nope. >> Oh, you can't see me on screen. I just looked at the camera and said question. >> You're on screen. You guys are both on screen now.

¶182>> Good question. All right. What's the next question? Uh, why is there a black screen now? Okay.

¶183Um, let's see. >> How big How big is the file? >> Unreal. >> How big is the file, do you think? >> I'm going to have to ask.

¶184Oh, no. I'll just look at the folder size. >> Do not ask any of these me, by the way. >> We're vibing. >> Gareth Hood asked, "Can you >> three gigabytes?" >> Three.

¶185Oh, that's not bad. It's It's like a single level, though. But still, it's not bad. >> Yeah, this demo is just a single level, but it did make a full a full game. I don't know how big the full game is, though.

¶186Gareth Hood asked, "Can you create Destiny 3?" >> Yes. >> Oh, you're not going to like it. >> Oh, it's not C++. Gareth is clocking it at HTML. >> So, got it.

¶187>> No, it's >> No, it this is not HTML. This is made an un >> Guys, why are we having a serious response to that? >> Oh, wait. I think the stream stopped to X. >> Huh?

¶188>> Oh, no. >> What does stream to X mean? Yeah, we got disconnected >> on YouTube right now. >> Oh no. Get us back, Matt.

¶189>> I'm trying. Hold on, guys. The show stops if X is not watching. >> One of my favorite things I've I've encountered with Opus 5.5, it's going to come with a con and a pro, and a really big pro, is that it thinks for a really long time. It might be doing stuff while it's thinking, but it it one of my prompts thought for a full entire day before starting making the thing.

¶190When it got back to me, that thing was nearly flawless. It my mind exploded. I unfortunately don't have that demo on X because I tried to prompt more and I broke it. So, it's not it wasn't ready to go. But >> little human error there.

¶191>> Very very crazy. Yeah, very very much so. >> Do you think it uh do you think its thoughts meander? Do you think it thinks about other things while it's thinking? >> Maybe like the meaning of life.

¶192>> Definitely. No, I don't think so. I think it's actually really good at following what you ask it to do. Um, and maybe that's a little bit of a fault where like if you don't know exactly what you're looking for, it still does the thing and then it's not exactly what you were hoping the prompt would be. But that also might be a really good thing if you have a really detailed prompt, which you can have AI help you make if you need.

¶193>> Robert G is suggesting that the fix here for the X stream. >> Oh. Oh, man. You don't follow Robert G on X. >> Oh man, that's embarrassing.

¶194>> He said it crashed because you're not following him, man. >> What the heck? >> Uh, okay. Yeah. Yeah, >> the the flight sim.

¶195>> Flight sim. >> Oh, somebody says open AI is also dropping. What a day. >> What? >> All right, we'll come back to that.

¶196>> Yeah. So, the flight sim was really cool because it it if you look at the very beginning of this met, if you go back, you can see all the I had to make a bunch of different like planes that exist. And so, you can like select and you could fly like a Boeing jet. Um, Matt, you're not in the thread. >> Matt, this is a sponsored pose.

¶197>> No. No. Uh, no. It should be looking here. No, it's just you're not in the thread.

¶198>> I'm looking at Flight Simulator right now. >> There we go. Good. >> Oh, okay. You're still not in the click into it.

¶199Yeah, we're a little behind. >> Oh, Brian. I had I had a Brian blunder where I wasn't on the live button. >> Excuse me. >> A Brian.

¶200>> We're not making that a term. We'll see. Um Matt, if you click the thread though, there are more tests. >> Okay, we're showing it. >> Yeah.

¶201Okay. So, if you didn't know, I got my start as an editor working for people who played Mario, custom Mario Hex, Mario Maker, etc., etc. And uh that's how I met Brian as well. Um, >> heck yeah. >> And so one of the tests I had it make was just remake Mario Maker.

¶202And uh it this this is one shot I have not added or changed anything to this. This is all entirely done in one go. Um hopefully I start playing it soon. Um but yeah, this is me just making a level. >> I love chat.

¶203Oh, the plumber game. Who on earth has ever referred to Mario as the plumber game? Wait, Matt, you got to show me playing the game. >> Okay, it's on. >> Right past playing the game.

¶204>> No, I I've been showing it. >> Matt, we can see we can see what you're doing. >> You cannot see what I'm doing. I think it's lag. >> I watch I watched you scroll the to the next one.

¶205Your ADHD kicked in. >> What are these What are these invisible blocks here? >> Uh the those are on and off blocks. Um, it also made screen scrolling pipes, which is a I think 3D world feature if you're a Mario guy. >> Um, hopefully I start playing level.

¶206Okay, there we go. Oh, and I died. Hopefully I don't die again, right, guys? Oh, and I died. Oh, I realized I'm I'm in not real time right now.

¶207>> Shell jump, dude. Nice. >> Shell jump. >> Yeah, >> he threw the shell off the wall and then jumped higher. What is that supposed to be?

¶208A thamp or what? Uh, I think they're just spike blocks. No, >> I think it looks really good. The The only thing is everything looks just a little bit wonky. Like, especially the character.

¶209>> The character. Yeah, for sure. >> Yeah. We We don't want to infringe IP, do we? >> No.

¶210No. No. Of course not. All right, let's get through this next one. Rocket League >> is a really quick one.

¶211>> Did you say Rocket League Rogueike? >> Rogue like, sorry. Rocket Rogike. >> That's a different game. a Kaboom to the Moon, which is uh uh Opus named believe it or not.

¶212Uh it's a rogue light where you get to build your rocket and you fly to different platforms, you collect coins, you have to avoid obstacles like birds. Um and yeah, once you land on the next platform, you get a little bit of money. I think you get like 120 bucks and you get to build onto your rocket. You have to do things like pay for refules and and things like that. Uh yeah, it's very inspired by this game called Hedgehog Launch.

¶213I don't know if you guys ever played that like a couple decades old now, >> but uh I loved that game as a kid and also Shopping Cart Hero. Two amazing games. Um but yeah, those are those are the main demos I made. Uh this one definitely took the least amount of time. I think this one only took a couple hours.

¶214Um the Dark Souls game I think took the longest out of all of them at about 28 hours. >> I mean, this one looks really good. Yeah. Also, you'll notice if you go back to the very beginning of this, the menu it makes. I've seen a lot of tests on Twitter today that have this exact same style menu.

¶215So, Optus 5 definitely or 5.5 has its like preference of menu for sure. >> Oh, yeah. That's what I got. >> All right, Brian. Uh, do you have a thread I can share?

¶216>> No thread. Let me just give you the stuff for you to put on your screen directly. >> Okay. I'm going to break the game streak here so we can look at an animation. >> Nice.

¶217>> Okay. >> Which the fellas have seen already. So, they're going to have to do their best to give a fake wow here. But what what I love about this is that it's not video generation. This is I've been excited about vibe editing.

¶218Too soon for the wow, Alex. Too soon. Um, and and so I I love moving away from video generation. I love that this literally took frame by frame and designed an animation. Um, I prompted it quite a few times to get towards what I wanted.

¶219The shadows were wrong at first. The shadows had a a harsh line um hard cut inner shadow and I wanted that feathered to get the look that I want. But you know what? I think we should all soak this in. Uh, if you love the music, maybe clap along.

¶220I'm not going to, but what do you think? Should we watch it? >> Let's watch it. >> That means I have to click it. Here we go.

¶221>> Yeah, man. >> I cannot hear it yet. >> Yeah, I don't think the audio is going to be coming through, unfortunately. No, >> no. It's so important.

¶222>> 70% of the presentation here, man. >> Yeah. I There's nothing I'm going to be able to do right now on that. >> Brian, I Brian, Brian, Brian, >> Brian, sing the song. >> I can sing it.

¶223Yikes. >> You have frontier intelligence here, Matt. You need to point Opus 5.5 at your eCam. Tell them what's going on. Say that audio is not coming in.

¶224I hope it can fix it. >> Guys, don't forget to watch the YouTube video where all this stuff is cut out. Um, I think you're going to have to share this one. I don't think I'm going to be able to do it because it it I remember this problem last time. >> Oh, okay.

¶225Okay. So, like I I can just drop this link to people in chat, honestly. >> Yeah, exactly. Drop the link. People will have to go watch it themselves.

¶226>> And you should watch this. This thing's really really cool. >> All right, chat. This is like two minutes long. Hopefully, no one's >> You've never watched or played like serious Sam?

¶227This is very serious Sam. You're just pointing out that it's in a desert, aren't you? >> No. Serious Sam's like a stickman fighting game. >> Okay.

¶228>> No, it's not. >> No, it's not actually. >> It's Doomike. >> It's It's like a Doom game. >> You know what I think?

¶229>> Oh, I'm thinking about stickman Sam. >> Yes. >> Yeah. Stick man is like a stickman game, bro. >> My bad.

¶230It's not serious Sam like at all. But this I love that you brought this up though because this actually is very rooted in early 2000's like internet animation culture >> like pivot >> because what I spoke to Opus 5.5 here is like I I spoke >> something that consumed so many hours for me as a kid. >> I I would watch these animations and what I told it was like you know I remember seeing this animation of a stick guy going into like a labyrinth stealing an idol. His rival shows up and fights him. other characters come in.

¶231Uh I like some of these characters are are from the initial prompt and you'll see characters pop in here and there. So yeah, if you want to see the uh the animation, there's a link there, but otherwise we can still move on and I want to challenge Matt >> to playing a little puzzle here. >> Nice. Nice. >> Let's make do this live, huh?

¶232The thing is like Matt, we've seen you play Fall Guys and try to talk at the same time. >> I'm kind of hoping that this goes even worse. >> All right. So, I'm playing TideKeper. I think I'll I'll what Brian did.

¶233And actually, maybe you just want to talk about why there's like this bracket here. >> There's a bunch of buttons around here, fellas. They're lining the walls. We're surrounded. But we're not too interested in those.

¶234The the the real gem here is Tidekeeper. But uh the initial prompt here was, "Hey, Opus 5.5. I want a great puzzle game." GPT6 Luna is out. >> Oh man. >> Busy day today, huh?

¶235>> I mean, shoot. I love Luna. What the heck? >> Okay, I've got to say my stupid stuff because this is far less important. >> Let's get through this quick because I'm gonna go record another one for that.

¶236>> Okay, quick, quick, quick. I said make a bracket. Make 12 puzzle games. I don't want just one. Put them in a bracket and judge which one's better.

¶237And then they all went all the way through the bracket, kept on going headtohead, and we ended up with Tide Keeper. But who cares? I honestly think we should just talk about Luna, dude. Luna's life changing. JBD Six Soul is out.

¶238Fellas, am I alone in this room? >> Yeah. No, I'm here. I'm just playing the game. Okay, let's let's take a look.

¶239I think I'm >> fivecoded slop instead of talking about GBT6. >> I think about GBT6. >> I think let's just find it quickly and then I might end the stream and go record. >> You're thinking stream. Share what?

¶240Share what >> stream. Do you chat? Do you think Matt should end stream to go record an offline video? >> Oh, and look at this. Grockot up.

¶241>> Yeah, I know. No, I All right. I guess I could keep streaming guys on Tesla. Keep streaming. Crockpot on Tesla.

¶242You wouldn't know about that, right, Matt? >> Nope. Uh, >> unless you you really can't. Okay, never mind. >> Holy crap.

¶243Look at this. >> Great live stream we're having today. >> Okay, I need to I need to Okay, I'm going to say a few more things about Opus. Um, >> and I'm going to kick you guys out of here. >> A See you guys.

¶244Thanks. It's good to see you. Bye. Bye. Bye.

¶245Bye. Bye. All right. All right. I'm going to talk about a few other things now.

¶246So, let's test test. Can you still hear me, everybody? Test test. All right. All right.

¶247So, I want to talk a little bit more about Opus 5.5. Then we're going to go and check out new models. Um, okay. So, I love artificial analysis. I love uh how how they basically run an index of different benchmarks.

¶248And we're going to look at uh how Opus 5.5 has performed on artificial analysis intelligence index and also it talks a lot about cost per task. So, what do we know? Opus 5.5 brings anthropic to parody with Astra on evaluations like terminal bench and automation bench. We already know that. Here's the key.

¶249Here's the key. It is number one a massive jump. Massive jump on the artificial analysis intelligence index. Usually we're getting like a onepoint difference. No.

¶250No. This is a fivepoint difference from Fable 5.1 max. Unreal. This is unreal. The absolute frontier is no longer fable.

¶251Okay. I think that's about all I have to say about this model. So, I'm going to end the stream or not end the stream. I'm going to end this video uh ship it off and then we're going to start a new one and we'll keep streaming and we'll talk about the new uh GBT6 Soul in Luna. All right, so it looks like we've had multiple model releases today.

¶252This is Claude Opus 5.5. It is currently until I go review GPT6 Soul, the absolute frontier of intelligence. And I'm super excited to try it. I encourage you to go try it. Go check it out.

¶253Thanks for watching. All right, give me a sec. I'm going to reset. and then we're going to start talking about GBT 5.6 Soul. So stick around.

¶254Maybe I'll rename the Dream. All right. So, I I don't see anything about soul. I do want to talk about MIMO, which is an open- source model from Xiai. Um, am I missing something?

¶255Okay. Nothing from Chat GPT. Guys, you said Soul is out, but I do not see it. Check the desktop app. No.

¶256Um, let's see. It is out. People are saying it is out. Oh yeah, it's it is coming. Yes, it is out.

¶257I guess a lot of people are saying it's out. All right, I think I'm gonna end the stream here for a minute and I'm gonna go record separately. Um, do you guys want to you want to stay on and just wait for all the GPT6 Soul and and Luna to come out and you want to want to do it live? What do you think? They haven't announced it yet.

¶258It might be in the app, but they haven't actually announced it yet. Yeah. Live. All right. All right.

¶259We'll hang out. We'll do it live. Cool. All right. I'm going to restart the recording.

¶260Click. Yeah, they haven't actually released it yet. Uh, it might be in the app, but they maybe jumped the gun slightly there because there's no blog post about it. Let's see. I don't see anything yet.

¶261This seems really cool. Uh, Grockbot in Tesla. Uh, so if you own a Tesla, you can now manage your Grockbots from your Tesla, which is really incredible. Um, I am so deep in Grockbot. I was thinking uh about making a video specifically about how I'm using Grockbot.

¶262Uh, because I I've been just using it like crazy. Uh, and it's been taking on more and more of work, just like the boring work, whether it's triaging emails. Um, Marmar Labs, thank you very much. I appreciate the the super chat. Uh, yeah, I will be at dev day.

¶263Who else is going to dev day? It's okay. People are saying it's live. If you're going to dev day, let me know. If you see me there, come say hi.

¶264I'd love to meet you. Marmar, thank you again. Boy, scrolling uh Twitter is not the best idea on live stream, but I'm doing it anyways. Ford future Brian, they wouldn't let me in. It's too bad, man.

¶265Uh, okay. So, let's look. Did they announce it? Not yet. Not yet.

¶266Yes, I think somebody in fact did jump the gun, Robert says. Yes, Robert G., I think you are right about that. Um, it's too bad that I'm not streaming on X anymore. That's kind of annoying. Okay, I'm trying to get the extreme working again.

¶267Confirmed. Luna 6 and GBT iOS app. Yeah, let's see. Let's double check. Um, I will open up the chatbt app.

¶268Let's see. No, I don't have it yet. Let's let me try restarting. Nope. I do not have it yet, by the way.

¶269I I don't know. Can you see this? Yeah, I do not have it yet. All right. Well, I guess I'm waiting, but maybe in three minutes we'll have it.

¶270Oh, people are saying they have it. Oh, you have to update your app. Okay, let's try that. Um, I'm going to update the Chai GPT app. Uh, there does not seem to be a new update for it.

¶271I'm on Android, so maybe it's uh slow forward. Future Brian says he has both. Oh, the desktop app. Okay, let's try that. It is currently working.

¶272Uh, I don't have an update, so I guess I'm not part of the cool kids. Uh, congrats, Codeex. Codeex. Yeah. No, I'm looking at Codeex.

¶273I don't have an update yet. Well, let's see. Maybe it's already in mine. Oh, I do have it. Yes.

¶274Look at that. I have it. Okie dokie. We're in. Yes.

¶275Uh well, I want to talk about it once they actually release the blog post on it. Um so I'm going to restart this. Let's restart chat GPT. Yep, there it is. GPT6 Soul Ultra or just Soul.

¶276Ultra is the thinking setting. Yeah, we got it. Okay, so I see Soul and Luna. Okay. Yeah.

¶277Okay. So, it's there. And I think maybe in one minute we're going to get all the info about it. Let's keep refreshing. Yep, there it is.

¶278Boom. Soul and Luna. No reset available. I have like three resets banked right now. Maybe Tibo's going to come through.

¶279Okay, so it is officially 11:00. Let's see if we get anything here. Let's check the OpenAI blog. Nothing yet. Nothing yet.

¶280By the way, I think they um so OpenAI's dev day uh is going to be next week, I believe. And I'm pretty sure they're going to launch a Muse Grockbot Instinct competitor which will be very interesting. And I suspect chat chat.com will become that kind of AI personal assistant from OpenAI. Yes, Soul is out. Yes, soul is out.

¶281They haven't announced it yet, but if you go into Codeex, it is showing for everybody. And since the embargo is uh the embargo they are let's see yeah okay so the embargo is lifted I I have had early access to this new model. These two new models, GPT6 Soul and Luna, they haven't published the blog post yet. Um, but everybody's seeing it, so I can talk about it. Plus, it's after um 11:00 a.m.

¶282And as soon as we get as soon as we get access to the blog, I'll start uh talking about it because then I can actually say intelligent things about it and not just ramble. Um, so it seems like a lot of people are getting access to GPT6 uh Soul and GPT6 Luna. Luna is one of my favorite models. It is so good, so cheap, so efficient. Um, I actually didn't know what models I was testing.

¶283Uh, but here we are. I wonder when the blog is going out because they have not announced it yet. Let's see. Not yet. We're just going to hang out here, I guess, till it comes up.

¶284Busy days for busy day for models. Yes, that is uh that is very true. Do Gallagher, hopefully I'm pronouncing your name right. Yes, very busy day for models. Google needs to drop a model now.

¶285The next Muse from Meta needs to be dropped today. Just Just drop all the models. And by the way, it's only Tuesday. It's only Tuesday and we've already had Gro 4.7. We had Opus 5.5.

¶286And now we're getting GPT6 Soul and Luna. Google is sleeping. Yeah. Where are they? Yeah.

¶287So, uh G like we should it should be any minute now we will get the blog post and the announcement about GPT6 Soul and GPT6 Luna. Very very excited about these models. It's out. Okay, let's see. No, nothing.

¶288Nothing. Where are you seeing this? Where where are you seeing the blog post? Am I looking in the wrong place? I see blog post is out.

¶289No, I don't see it. What am I missing? I know. Uh, Okasuku Okasuko. I know I I spent the last hour and a half covering Opus 5.5.

¶290So, uh, I think you you might have missed that part of the stream. Let's see. Let's see if we can find it. Where would it be? Company blog.

¶291Yeah. Company news, right? No. I do read chat. I do read chat.

¶292Yes, I do. Uh, go to the latest Astro 6 blog post. Scroll down and click through to loan it. Okay, so the latest Astro Blog. Astroblog.

¶293Okay, I'm working on it. Give me a sec. Open AAI astro blog. Why wouldn't they have had their own? Here we go.

¶294Boom. All right. Thank you. Thank you. Thank you for pointing this out.

¶295We have two brand new models from OpenAI. GPT6, GPT6 Soul and Luna. And this comes on the same day that Anthropic dropped Opus 5.5 and just yesterday Grock dropped Grock 4.7. So this is a modelfilled week and we're going to check out this model together. All right.

¶296So, I'm going to read through it and then I'll I'll uh we'll read through it together and then I'll record. Earlier this month, we introduced Astra, the most intelligent and aligned model in the world. By the way, if you're watching, uh we're going on almost two hours of the live stream, so or over an hour. Uh I would love a like if you can. I would appreciate it.

¶297Thank you. the most intelligent and aligned model in the world. While the most demanding and important projects still call for Astra's full depth, work happens at different scales, rhythms, and budgets. That's why we're announcing or that's why we're expanding the GPT6 universe with GPT6 Soul and GPT6 Luna. Crazy crazy day.

¶298Drop a like. Yes, drop a like. Thank you, Kyle. Much appreciated if you do. Thank you.

¶299Um, so as expected, we trained GPT6 Soul and Luna with similar methods as GPT6 Astra, bringing the advanced bringing the advances behind Astra's state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models. Now, if you remember when GPT6, sorry, when GPT 5.6 six Luna had that 80% price drop. It became one of the most compelling models on the planet to use. Extremely fast, extremely cheap, and still very high intelligence. And so now we get an updated version of what I consider to be my favorite model.

¶300Um, and wow, it looks like it's cheaper. Unbelievable. Okay, let's start with the pricing. We usually start with the benchmarks, but we're going to start with the pricing. So, this is GPT 5.6 Soul versus GPT 6 Soul.

¶301A 50% price reduction. Now, if you remember, Terra and Luna, 5.6 Terra and Luna both got significant price reductions. GPT6, sorry, GPT 5.6 Soul did not. But now we get a 50% price reduction on the new Soul model. And Luna with an already 80% decrease previously gets another 50% decrease.

¶302So we are at $2 per million input tokens and $10 per million output tokens for basically one step under the absolute frontier which is Astra. We also have GPT6 Luna coming in at 10 cents per million input tokens, 50 cents per million output tokens. So good. Yeah, the pricing is nuts, Kyle. Yes, this pricing is incredible.

¶303The these are true workhorse models and specifically Luna. And Luna can probably take on 90 to 95% of all the tasks that you have to give it. It is a great model. Uh, yes, I can do a little zoom. There you go.

¶304All right, let's look at some benchmarks. We have Automation Bench here. Oh, look at these cute little icons they have for the uh for the points on the chart. Very cool. Okay, so unfortunately in the chart they only have Claude Opus 5.

¶305Claude Opus 5.5 was released an hour ago. So it doesn't h So they don't have it on this chart yet. But right now, if we look at it, that's where it is. Cost per task all the way over here on the right. Um, let me restate what we're looking at.

¶306This is automation bench. On the y ais, this is the score. Higher is better. On the x axis, this is the cost. Left to the left is better.

¶307That is lower cost. And so, where you want to be is in this a sorry, is in this quadrant right here. Okay, this is the money quadrant. This is where you are the best and the cheapest. So if we are looking at five no let's let's look at 5.6 soul first.

¶308There it is. Okay. And then if we switch to GPT6 soul, we get a significant what looks to be about a 50 a 50% cost reduction. Exactly what it said. And and it's better.

¶309And it's better. I was actually expecting like a bigger jump, but let's see. At max effort for 5.6, we scored a 28.8 on automation bench. And then for GPT6 soul, interestingly, extra high took the number one at 33%. Now, Astra is still the number one model and you are going to be paying for it.

¶310And interestingly, so you can kind of think as Opus 5 point, you can kind of think of Opus 5.5 as the soul. Wait, let me let me make sure I do this comparison right. Astra is to five. No, I'm going to get this completely wrong. Okay, Astra is to GPT6 Soul as Fable 5.1 is to Opus 5.5.

¶311Okay, I got through that. Um, and so interestingly enough, Opus 5.5 is actually the best model in the world right now. Better than Fable and not a little bit better, like quite a bit better and cheaper and faster. But what we see here is Astra is still OpenAI's best model. And we now have a much cheaper version.

¶312It's not quite as good. Look at the soul right here, right? And then the the top model. So, if you need the absolute best answer, you are going to be paying for Astra. But if you want more of a workhorse model, it's right here.

¶313Now, here's the interesting bit. Here is where Luna came in. So, definitely again, like if you look at the curve, Luna is not is by far it's not Luna is the worst of the three models, but you are paying a fraction of the price. Even at look at this at the loweffort setting it, you know, has quite a low score, but it's extremely inexpensive. I mean, if you use it on max, you're getting above a 20% score on here, and you're paying less than 5 cents cost per task.

¶314Unreal. Unreal. Very cool. GPT6 Soul also exceeds Claude Fable 5.1 at far lower cost and even beats loweffort GPT6 Astra. So these are cost per task although I think this is automation bench that they're showing here.

¶315So here's Fable with Opus 5 fallback. You know, all of this all of this blog information is pretty much dated at this point because Opus 5.5 is here. Here's agents last exam once again. Wow, look at this. The high effort.

¶316I love when it's weirdly spiky like this. Like what would be the reason for that? Um, we still see GPT6 Astra is the best, but look at Soul right behind it and pretty significantly less expensive. Um, so for the extra high setting is coming in uh just under $2 per task completed, whereas for the medium setting on Astra, it's coming in at what looks to be about like $4.50 per task completed. So significantly less expensive.

¶317Here's factual error rate on difficult prompts. Okay. Uh answers with any factual error. So you want lower. Actually lower is better.

¶318And cost per task. Same thing. You want to the left. And so Astra number one least factual errors but soul. Oh interesting.

¶319This is okay. So this is soul 5.6 six which has more factual errors but if you look at soul six it pretty much performs the same only slightly more factual errors but you are paying a fraction of the price so this is the workhorse model yeah um yeah dang bro we just had thic on here excited and they didn't even give five minutes to shine before they dumped two whole models on them. Yeah, it's it's a crazy world. And you know what? We benefit.

¶320Uh here's Frontier Code. Maybe they're going to include Sweetbench. Uh but here is Frontier Code. Once again, we're seeing Luna coming in very strong, especially Luna Max, 11cent cost per task. Comparable with Soul Medium, 80 cents per task.

¶321comparable with Astral Low, a $1.70 per task. Here's Opus 5 all the way over here. But again, Opus 5.5 is out. Oh, here we go. Deep Sweet.

¶322That's what I meant, not Sweetbench. Okay. And now we have what I believe is the best, most accurate benchmark for determining how good a model is at aentic coding. This is typically the best reflection of the engineering ecosystem and what people are actually thinking about the model. So let's take a look at where it falls.

¶323Here is GPT6 Astra Claude Fable 5 all the way out here. So once again on the Y-axis we have the score higher is better. On the X-axis we have the cost per task. To the left is better. And look at that.

¶324Unreal. GPT6 Luna Max 66.6% 22 cents cost per task. Phenomenal. absolutely phenomenal and only slightly better. Let's see 66% versus the absolute frontier of this is 74% but you are paying $443 per task completed versus 22.

¶325I am such a big fan of Luna Luna Max in particular. Luna Max for computer use, browser control, all of that. Here's OS World. Speaking of computer use, uh let's see what we have here. Astra still at the absolute frontier, but Soul right behind it.

¶326Um here's Luna doing pretty well. Definitely like significantly far behind actually. So this gets a 53% versus a 73% for Astra. Luna verse Astra. All right.

¶327So availability GPT6 soul and six Luna are available in chat GPT work and codec starting today for all plus pro business enterprise and edu edu users. Free and go users can access GPT6 Luna in the desktop app and these models are not yet available. So yeah, uh, someone brings up a good point. I need to put together a little chart comparing Soul, Luna, and Opus 5.5. I'm going to do that.

¶328Give me a sec. Okay. So, make me a chart comparing GPT6 Luna, Soul, and Opus 5.5 pricing and benchmarks. Okay. So, hopefully that'll finish relatively quickly.

¶329Interesting. I don't know why that did not work. Okay. When reset. Yes, I'm asking chat GPT for this.

¶330Yes. Yeah. And it is interesting. So, I'll put Domo. Whoops.

¶331Where are you, Domo? Interesting that both Anthropic and OpenAI have reduced pricing on their new models today. Yes, it is interesting and I have some thoughts as to why. Um, when both of these companies are talking about pacing the frontier, this is what you get. You get efficiency, you get speed, you get price, but you don't get absolute raw power and intelligence.

¶332That is what we're seeing here. This is maybe the result of pacing the frontier. Now, that doesn't mean they're not building it internally, but that's what we're getting here. Now, also keep in mind, Fable and Astra were released recently. Typically how these model releases go is they release kind of the big foundational frontier model that is very expensive.

¶333Both Astra and uh Fable were very expensive. $10 per million input uh $50 per million output. And then over time in the coming weeks and months after that release, you start to get the more efficient models. you start to get the cheaper models, the more workhorse models, and probably either architecturally similar to those frontier models, the Astros and the Fables of the world, or possibly even distilled versions of them. So, that this is all expected and I'm very happy to see it.

¶334Now, I'm kind of torn. Do we want to see more efficient, better, smaller, faster models, or do we want to see what the absolute frontier is capable of? So, I'm I'm torn on that. All right. So right now I just tasked Chad GPT with creating a new graph, a new chart comparing GPT6 Soul, GPT6 Luna and Opus 5.5.

¶335Now there was no comparison of these two models because they both were released on the same day. Crazy. It's crazy that uh we're live streaming and Opus we found out that was getting released today. Opus 5.5. And then literally during the stream, all you guys were like, "Hey, Soul and Luna are out now." I mean, what a day.

¶336So, I'm I'm having that benchmark or the sorry, I'm having the chart and the table created right now. Hopefully, it finishes soon. I'm using uh GPT6 Astra on high to do it. So we shall see. Um let's see what else.

¶337So improving alignment. Astra is more aligned than GPT6 soul. So this is deception rate on coding deception. 5.6 soul much higher. Luna.

¶338Yeah. It's interesting. It's It's like the smarter the model is, the more aligned it is, which is almost counterintuitive. Oh, and we just heard Tibo confirmed a banked reset. Let's see if that's true.

¶339GPT6, Soul, and Luna are out. Not only are they a very significant improvement across the board, but also in writing in general. You know it when you try it. Quality. We are also permanently reducing the API price by 50% making both of them viable for a ton of new use cases and making your usage go even further too even on the subscriptions.

¶340And one more thing we are loading a bankked reset into all accounts of our pro of sorry of our plus pro and business users. Let's go. Thank you Tibo. Awesome. Awesome work.

¶341Very very cool here. All right. Okay. Let's see. Okay.

¶342So, we have some charts. I'm going to share. Hopefully, I can share my screen. Let's see. Let's see.

¶343Okay. Can you all see that? Okay. So, I think you can only see some of it. Oops.

¶344All right. So, uh, I had Astra put together this graph so we can actually see GPT6, Soul, Luna, and Opus 5.5 all released on the same day on the same charts. Let's look. So, what we have Luna, starting with Luna, by far the cheapest. Look at that price.

¶345Luna 10 cents per million input tokens, 50 cents per million output tokens. Compared to Soul, $2 per million input tokens. Sorry, $2 per million. Yes, soul $2 per million input tokens. Soul, $10 per million output tokens.

¶346Opus 5.5, still the most expensive, $4 per million input tokens and $20 per million output tokens. Now, how do they compare? Here is Frontier Code Luna 42.4% Soul 49.3% and Opus 5.5 at 54.4%. So Opus 5 So Opus 4 sorry Opus 5.5 is still definitely the best. Okay, but here's what's interesting.

¶347Look at the delta in performance on Frontier Code between these three models. They're all off by about six to seven percentage points. But if you look at the pricing for them, it's multiple times difference. We have 10 cents versus $2 verse $4. And this goes towards the narrative that paying for the absolute best answer is expensive.

¶348It's like the 8020 rule. Uh that that additional that last 20% is where all of the revenue goes to. That's all where a lot of the value gets captured. People are saying to fix my screen. Okay.

¶349Uh Okay. Give me a sec. Sorry everybody. I see. I see.

¶350Okay. Thank you. Thank you. Give me a sec. All right.

¶351There we go. Please scroll. Sorry about that everybody. There we go. Now you can see it.

¶352Um okay. So we have Yeah. Opus. And then let's look at Automation Bench. We have Luna at 20%.

¶353This is kind of a bigger delta between the models. Soul at 33% and Opus 5 Opus 5.5 at 40%. Now here's the cash pricing. One penny for reads, uh 12.5 cents for rights, soul, 20 cents for reads, $2.50 for rights. Opus.

¶354Interestingly, same amount for the uh cash reads, 20 cents, and then $5 for the rights. Um, very, very nice. So, yeah, look, these are all three incredible models. Opus is definitely the best. In fact, I think Opus 5.5 has the crown for the best model on the planet by far.

¶355Now, let's see if Yeah. So, here it is. By far the best model on the planet is Opus F 5.5 right now at 58 on the artificial analysis intelligence index. And let's see. I don't think I don't think uh Soul and Luna are on the artificial analysis intelligence index yet.

¶356So, incredible models. Check them out. Um, I have some tests. I'll probably put those tests out. But yeah, go try all three of these models.

¶357They are phenomenal. I don't know which one are you all going to use. Let me know in chat which one are you going to use. If you had to choose one of these models to go with, what are you thinking? Luna Soul or Opus 55?

¶358Luna 55. I just like a lot of people are saying 5.5, some people are saying Luna. I just think especially given how generous Open AI is with their resets is with the quota in general, it's hard to beat Soul. It's it's I mean it's even harder to beat Luna, but if you even need that extra bump in performance, Soul is the way. Um I mostly have been using Chai PT as of late.

¶359Uh, but I'm going to start using Claude a little bit more now that Opus 5.5 is the absolute frontier and it's relatively more affordable than Fable was. So, um, I'm going to be testing them out and we'll see. I think like if I see that they're Anthropic is quite generous with Opus 5.5 uh, quotas, then I'll probably end up using that more. But I'm I'm just having fun with all these models. Um, so that's it.

¶360Go try them out. If uh I'm gonna be releasing two videos on this, so if you're if you haven't seen, let's see. And I just reviewed Opus 5.5 and you can go find that video right here. And enjoy. Thanks for watching.

¶361Uh, thanks for joining the stream everybody. I really appreciate it and I'll do another one soon. Bye everybody.