¶1Google is back. They just dropped Gemini 3.8 Flash and it's actually good. And that's not the most impressive thing about it. So, I'm going to show you the benchmarks in a second, but I just want to give you a little bit of context as to where Google is currently in the competitive landscape. So, a few years ago, Chat GPT came out and Google was nowhere to be found.
¶2Chat GPT took off. Open AAI was doing so well. And of course, Google had to drop everything and try to catch up. And then at a certain point about a year and a half ago, they released Gemini 2.5 Pro, which was at the time a phenomenal model, the best model on the planet. The first model to create a full Rubik's Cube simulation that actually worked.
¶3But it was a fleeting moment for Google. Suddenly, Anthropic and Open AI were just far ahead. Fast forward to the last 6 months, Google has released a number of models in the Gemini family that, to put it plainly, have fallen a bit short. The benchmarks have been pretty good, but the actual usage of the models has definitely left something to be desired. But now, Gemini 3.8, 8.
¶4Fresh off the release of Gemini 3.7, which really was just a few weeks ago. We have a brand new Gemini model, and it's actually quite good. Now, it's really good in some benchmarks, kind of okay in others, but the cost is what makes it special. All right, so here are the benchmarks. Number one, the most important benchmark in my opinion today, Deep Sui.
¶5And this tests long horizon software engineering tasks. It seems to be the most accurate reflection of how actual engineers using these models feel about the models. For example, Fable on many benchmarks is way ahead of GPT 5.6 Soul, but on Deep Suite, it's really close. And that's my experience with these models. Soul is fantastic, Fable is fantastic, but they're not that far apart.
¶6And now check this out. Deep SUI V1.1 Gemini 3.8 Flash coming in at 73.7%. Which is effectively even with Claude Opus 5 which was just released a few weeks ago. Here is GPT 5.6 Soul coming in at 72.7% under Gemini 3.8 Flash. So this looks to be a really good model on the benchmark that I respect most of all.
¶7Then we have GDP valve. This is a benchmark from the OpenAI team testing the models on real world knowledge work. PDF extraction, data analysis, presentation creation, all of that kind of stuff. And it's coming in at 1545, which is not that good. So, Claude Opus 5 absolutely dominating the competition, coming in at 1824, second place 1710 for GPT 5.6 Soul, and then kind of a much less good score of 1545 for Gemini 3.8 Flash.
¶8So on GDP val knowledge work, real world knowledge work, it's just okay. Basically equivalent with GPT 5.6 Terra. However, if you look at the pricing, that's the model size that this new 3.8 Flash is trying to compete against. So you really shouldn't compare it to Soul and to Fable. It doesn't really make sense to compare it to that because it is a fraction of the price.
¶9And speaking of price, let me just show you the price before I show you the rest of the benchmarks. So, input price, 75 cents per million input. Output price, $3.75. Really, really inexpensive. Look how it compares to Opus 5.
¶10It is a fraction of the price, $525. Here's GPT 5.6 Soul, $420. GPT 5.6 Terra, $212. So again like 20 30% of the price of even Terra. Now the 75 $3.75 prices is if you look at the fine print they're introductory prices which expire at the end of this year.
¶11I don't love that they put this in bold right here. So it's 75 and $3.75 but in parentheses in small text under it is the actual price that we should expect after this year. Now, of course, if everybody is clamoring and loving this model, they may just extend the price indefinitely. But if you already plan on raising the price in the future, you should make that price the more prominent price. But even at $1.50 and 750, uh, it's still cheaper than a comparable Terra model.
¶12We have Harvey's legal benchmark absolutely dominating number one, 61.4% 4% and number two was Gemini 3.7 Flash. So for legal work, this is the best model on the planet. And that's where I want to spend a minute. This model and seemingly a lot of these newer models are really good at certain things and not so good at others. So if you're an enterprise company and you're thinking about which model to adopt for your business, it's not as simple as just saying, "Okay, what's the latest from Anthropic?
¶13What's the latest from OpenAI?" You should actually look at the benchmarks and more importantly test them yourselves. Have your own internal benchmarks that you create, which it's not very simple to do, but it's very doable and you can test all of these new models against your own benchmarks and choose the right model for the job. So, if you're a law firm, you should definitely look at this model in particular because it scores so highly on the Harvey legal agent benchmark and also is very cheap. Then we have terminal bench 2.1. It got the number one score at 89.4.
¶14This is a gentic terminal coding. One of the most important benchmarks for knowing how well a model can code. And then we have terminal bench 4.0, the newer version. And look at this. It actually got quite a low score.
¶15So I don't know exactly what changed between 2.1 and 4.0. Okay. So I did a bit of research and terminal bench 2.1 was all but saturated as you can see here. And so Terminal Bench 4.0 is really a brand new benchmark under that family of benchmark, a much more difficult benchmark and that's why we're seeing lower scores. However, look at this claopus 51% absolutely dominating the field.
¶16And then we have of course Gemini 3.8 coming in at 19.1% which is just okay. It beats Sonnet 5 and comes in well under Opus 5 and Soul. Here is humanity's last exam. And surprisingly, Gemini 3.8 Flash comes in at number one at 55.9. So, this is why it's a little bit weird.
¶17Some of these benchmarks, Gemini 3.8 Flash just dominates at and is so cheap, and then on others, it's just okay at best. Here's OSWorld, which is agentic computer use, the ability for the model to control your computer, to control your browser, coming in at 59%, a respectable score, but Opus 5 just dominating at 75. But this is the chart that really matters. This is Deep SWE again, but importantly, it also maps the average cost per task. This is very important because the average cost per task not only takes into account the actual cost per million input and output tokens, but it also takes into account how many tokens are used to solve the task.
¶18Because if one model is half the price of another model, but takes twice as many tokens to complete the same task, it is effectively the same price. So here's what we see. This is Gemini 3.8 a flash and really where you want to be on this chart is as high up and to the right as possible and that's what we're seeing. It's a good model and I do very much trust Deep Sweet so I'm looking forward to using it. Now they released a second model in the 3.8 family.
¶19It is Gemini Flash 3.8 cyber and it is specifically exactly what it sounds like, built for cyber capabilities. And so it is only available to what they call trusted defenders via the Fair Wind program, but it's basically the model without some of the cyber guard rails that come in the kind of normal Gemini 3.8 Flash. And what we're seeing here, this is the Cyber Gem benchmark, and it performs really well. Here's Mythos 5, although Mythos 5.1 is now out. Here's GPT 5.6 Soul, 83%, 83%, even GPT 5.5 Cyber, a model specifically made to be really good at cyber attacks and defense, 85.6 and Gemini 3.8 Flash Cyber 86.2.
¶20Now, unless you're part of that program, you're not going to get to test this model, unfortunately, but you can apply to the program and see if you get in. Now, they did something interesting, which I haven't seen other model companies do. Listen to this. to better capture real world defensive needs which are not limited to just C C++ code bases like in CyberJim. Okay, so Cyberjim is the benchmark and they only have tests built in C and C++.
¶21We also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex code bases spanning 20 programming languages. And it does really well. Here's Gemini 3.8 Flash, Gemini 3.7 Flash, and 3.5 Flash. Now, they did not test other companies models on this internal benchmark, but on 3.8 Flash, it is a massive jump versus the previous iteration version at 3.7 Flash. All right, let me show you those tests now.
¶22So, over the last week and a half, I've tested GLM 5.3, Fable 5.1, and GPT 5.6 Soul against the same set of demos or tests. And this is the first one. This is creating seven individual 3D lowpoly biomes. And Alex, put the other ones on the screen, please, so we can actually compare them. Here is Gemini 3.8 Flash.
¶23And I'm quite impressed. The beach looks really good. The farm looks really good. I'd say the overall detail is definitely less than GPT 5.6 Soul, which so far has been my favorite. Fable 5.1 is probably second place and then GLM 5.3 and Gemini 3.8 probably right around the same judgment.
¶24Now, very similar to Fable, we have this weird kind of glitchy thing right here. And then we have a bunch of little fish in the ocean under the singular boat. I don't know why that happens, but we've seen that before. Here's a little mistake. Look how the water under this ice lake is kind of protruding out from the biome right here.
¶25So yeah, overall it's pretty good. It's definitely not the top of what I've tested over the last week and a half, though. All right. Next, I asked it to create simple web pages for different products. An Apple, a DJX Spark, a rubber duck company, the Galaxy ZFold, and a Tesla Model Y.
¶26And again, please Alex, throw the others up on the screen. Here's the Apple website. So, it looks okay. Very simplistic. No Apple above the fold.
¶27Yeah, this is definitely one of the more simple websites that I've seen. This is kind of weird. It's like a checkout, but it gives me no options to actually select items to put in my cart, and it pre-selected six slots. This is definitely the worst website I've seen. It's not terrible, but compared to the others, this is not as good.
¶28Here's one for the DJX Spark. This one actually looks quite good. a little clipping on this G right here, but otherwise the colors look good. It's kind of interesting how all of these models get the colors right because these are the colors of Nvidia and specifically the DJX Spark. So, a bunch of good stats.
¶29Let's see if we can change it. Yep, we can change it. We can see all the other numbers updating. We have a little terminal right here. I think this is pretty good comparable to all the others.
¶30Here we have a rubber duck company. This is uh definitely not the best. I really like the more kind of childlike look of it where we actually see a rubber duck. This is more of just like a realistic duck emoji. If we scroll down, yeah, I don't know.
¶31Astro duck Voyager. Yeah, none of this really makes a lot of sense. I'm going to put this at the bottom. Also, this is definitely not that good. Here's one for the Galaxy Zfold.
¶32I guess that's a foldable. Am I supposed to click this here? Let's see if I can. Okay. Oh, that's actually kind of cool.
¶33So, I can drag this and actually open and close the foldable. Although, you know, it doesn't look anything like a phone. And in fact, if I close it, you can see you can see the back of what's supposed to be on the screen on the inside of the phone. Yeah, this is this is not good. And then last, the Tesla Model Y.
¶34Once again, a model tries to recreate the Tesla Model Y, basically looking like it's using Microsoft Paint. And it is really bad. But you know what? The other ones were really bad, too. So, I can't really blame it.
¶35Uh, we are able to design our Model Y. We have the performance all-wheel drive. We have the long range all-wheel drive. We can change the colors. Let's see if it changes the colors up here.
¶36It does. There's our little preview of our car. We can say how many miles we're going to drive per day and our average estimated 5-year gas savings. This is actually pretty nice. Oh, this is super nice.
¶37So, vision only neural autopilot. This is a pretty clean animation to show that off. Obviously, like very simplistic, but definitely uh better than what I've seen elsewhere. So, I'd say this is probably comparable to the other ones. Overall, I'd say this probably falls right about or under where GLM 5.3 was.
¶38All right. So, for the next test, I asked Gemini 3.8 ate flash to create a PowerPoint deck all about data centers and this is what we have now the nice thing is it knew to use my forward future branding these are the forward future colors forward future.com again so that looks really good this is our typography uh how the physical backbone of the internet the cloud and AI actually works okay first page looks good uh second yeah it's it's okay it doesn't have strong design chops but like the structure is there the information looks correct, but certainly if I'm going to create a PowerPoint and especially if I'm focused very much on design, I'm probably going to a Claude model. All right, Alex from my team just sent me this link. He ran a test with Gemini 3.8 and the purpose is to create a 3D topographic map of Mount Everest. And this is phenomenal.
¶39I I am super impressed. So, I can drag it around. I can zoom in. We have different points. I have different sliders.
¶40Oh, nice. So, I can get a cut angle straight through Mount Everest. Custal depth offset. Oh, that's really cool. Solar azimuth.
¶41I hope I'm pronouncing that right. I've never even heard that word before. Yeah. Very cool. Vertical exaggeration.
¶42Ooh. Yeah. Yep. Yep. This is really cool.
¶43This is a very, very good topographic map. Here's a 2D projection. Yeah. Wow. This is really cool.
¶44I'm very impressed. Oh, and look at this. We actually have a bunch of different points on Mount Everest. Here's camp 2, for example. And so, if we switch back to the 3D terrain, there it is.
¶45Yeah. So, this is really cool. This is an absolute winner. I have nothing to compare it against, but in my mind, this is great. All right.
¶46And Alex just sent me another game that I was completely just playing in a distracted way, but I'm going to show it to you now. This is a very simplistic Doom recreation. This is basically just one prompt. Uh, let's see if I can find a baddie to kill here. Here we go.
¶47Yeah. So, you know, you can see it. It's very simple, but with a few more prompts, I think this is actually decently fun to play. Yeah. I, you know, pretty darn good.
¶48And real quick, I just want to tell you to go visit forwardfuture.com. That's our newsletter. And if you want to stay uptodate on all the latest AI trends, new model drops, go there. forwardfuture.com. Check it out.
¶49It's free. It's awesome. So, congratulations to Google. This is definitely a good model. It's not quite the Frontier, but for the price, extremely competitive.
¶50And you can actually see how it compares to Fable 5.1, which I covered yesterday. Hey, go check out that video right