¶1What's up engineers? Indie Deb Dan here. Kimmy K3 was released mid July. Deepseek V4 Flash late July early August. Quinn 3.8 early August.
¶2Muse Glimmer. Neotron 3.5. Grock 4.6. Deepseek V4 Pro. Gemini 3.7 Flash.
¶3Quinn 3.8 27 billion and GLM 5.3. We have just experienced an intelligence explosion. We have more than five model releases within 5 days across all tiers of models. The big AI labs are clearly panicked. Fable is a permanent part of the clawed subscriptions now and OpenAI recently slashed prices for Terra Luna and now they're testing 50% cuts on GPT 5.6 sold on Open Router.
¶4The LLM pricing wars are in full effect. The intelligence explosion emphasizes two [music] key questions for us engineers. What's the best way to use these models together to outperform them individually? and how can we build for an age of rapid change so we can leverage the models with the best performance, speed, [music] cost trade-offs for our valuable engineering work. I've been engineering for over 15 years and in technological revolutions where rapid change is the norm.
¶5[music] One engineering principle stands far above the rest. The most flexible system wins. [music] Today I want to share my thoughts on the intelligence explosion and share a V2 of one of my favorite custom [music] PI coding agents, the fusion harness. We're going to talk about why the fusion harness is so valuable and then upscale all of it to talk about inloop agent coding [music] and outloop agentic coding. If you want to leverage these ideas for your engineering work, stick around [music] and let's capitalize on this massive intelligence explosion.
¶6There are three commands I want to share with you to give you concrete ideas on how to use your flexible agent harness more effectively. Now, I recommend the piecing agent, but feel free to use or build any tool that gives you the capabilities you're going to see next. Before I showcase these, we need to address the intelligence explosion because there's a lot here and there's never been more opportunity for us engineers thanks to this wide selection of compute. I maintain a running model stack of the best models. Here's all the new models that were just released over the past month.
¶7And then, of course, I have my favorites. Let me give you my thoughts on the current model landscape. Starting at the floor here with the lightweight models that can run on your machine. Muse Glimmer was released. Meta is still in the race.
¶8We'll see how they perform in the upcoming model releases. Iron 3.5 Lightning, really, really great release coming out of Nvidia. I want to see where this goes. Then we had the incredible Quinn 3.827 billion. I'm really really happy about this model release, but I'm also looking for that mixture of experts release.
¶9I'm curious if they're going to release something like this or if we have to wait for another version. Then we pass the B tier and move right into our daily workh horses. This is where a lot of the leverage is being created right now for engineers deploying language models into their engineering agents and their product agents at scale. We have the new GLM 5.3. I am picky and stingy with my model providers, so I'm waiting for this to come to a US-basedbacked company.
¶10We of course had the incredible Deepseek V4 Flash release. This model is absolutely cracked. The pricing is cracked. Even after they increase it, it's still going to be an incredible model to work with. And it's in a tier with the Deep Seek V4 Pro model that was just released.
¶11We're going to look at this one today and compare it side by side. The incredible Gemini 3.7 Flash. Out of all the models that were released and the intelligence explosion we're experiencing right now, this is my favorite model. Again, you'll see why in this video. Then we had Quinn 3.8 and Grock 4.6 in addition to Kimmy K3.
¶12These are two models entering the S tier that are open weights. Okay, so obviously Grock, not open weights, but Kimmy 3 and Quinn, incredible models. Of course, it's basically impossible to run one of these unless you have tons of capital. But through model providers like my favorite Fireworks, through Open Router, these models are now available and they're fractions of the price of the current state-of-the-art. The Fable 5, the Opus 5, the GPT 5.6 Soul.
¶13So, a lot of great intelligence to work with. Let's jump into the Fusion harness and let me really highlight two of my favorite models in the Intelligence Explosion and we'll work through them to see how they compare directly with Fable 5. Open up the terminal. I've got Fable 5 running right next to GBT3.7 Flash, right next to DeepSeek V4 Pro. Let's see how this agent [music] team performs side by side on three specific tasks.
¶14So, Duck DB, one of my favorite in-memory in process databases, has just released a preview of their V2. When this happens, I sit down with agents to understand it, to gro it, and to really see what became available, what's new. This is a perfect use case for a fusion harness. We have multiple sources of compute, multiple sources of intelligence we can tap into to understand new tools, to understand new technologies, to help us harness them in the best way. So, let's take the V2 of the Fusion harness we've discussed on the channel before, and push it to the next level/ FH opinion.
¶15So, I want the opinion from all of my models, from all of my compute on one issue. Here's the prompt we're going to run here. I'm going to go fire this off. We want to understand what the first most important release coming out of DuckDB is. What's the most important thing we should focus on for a developer building a local first analytics application should test first.
¶16So, you can see here our agents are getting to work. Gemini Flash is already completed. Deepseek V4 Pro getting to work here. 33 tokens per second so far. And then, of course, the Behemoth, the Monster, the best model available.
¶17Claude Fable 5 completed. It's worked there already. You can see the massive difference in token costs from just one simple prompt. And here you can see our models completed. So, a couple things to note right away.
¶18When you're using models, there are three things you should be paying the most attention to. Performance against whatever you're trying to get them to do. Speed and then cost. We have that all right here. Performance, speed, and cost.
¶19Context windows. model aliases so that they don't start competing against each other by understanding the model name. This is a very very subtle thing you'll notice when you're using these models together a lot. You can never reveal the name of the model to the other model. Otherwise, they'll start emitting weird behavior.
¶20We'll talk about that more in a moment. Gemini 3.7 Flash is moving like the Flash. Insanely quick model. Deepc4 Pro 80ps. Claude Fable 5 the slowest but also the most powerful but also vastly more expensive.
¶21in order of magnitude more expensive. In fact, this trend is going to continue throughout. But you can see here the fusion harness is presenting us with the results of all of our models. What does this mean? Why is this valuable?
¶22We have three unique perspectives on the results. So, what is the most important feature we should pay attention to from the release of this new tool based on our suite of intelligence? The most important thing we should focus on is the new variant. We can now use variant to gain automatic structural decomposition at right time from our in-memory duck databases. This is the choice from Fable 5.
¶23This is a choice from Gemini Flash. And you can see here Deepseek V4 Pro actually drifting a little bit. Interesting. Wonder why that is. They all wrote out a little quick experiment on how to actually execute on the task that they recommend.
¶24So this is a very very simple prompt and you've seen this before in many different ways and configurations. Instead of just running one model to give you feedback on something, run n models and then push it a little bit further. Understand what it costs to deliver that single unit of result because guess what? That's going to scale as you scale your efforts with your [music] agents. Let's move to a much more powerful prompt that gets out of what everyone else is doing slashfusion harness debate.
¶25Inside of the duct DB announcement post, they said something really interesting here. You can now use duct DB as a server with quack and connect. They broke down exactly how you can do this very cleanly and concisely. This is really interesting because now it questions how we should best use duct DB. Is this a simple in-memory analytics processing tool that optimizes on the columns or is this more like a server that we should connect multiple instances to?
¶26We don't know. Let's ask our team of intelligence. What do they think about this? That's exactly what this prompt was. Debate this claim using the duck DB v2 preview announcement.
¶27And then here's the claim as a single quote. And now they're going to start debating. So this is a different agentic workflow built into the fabric of my PI coding agent as an extension. This is behavior you're not going to find in any out- of- the-box agentic coding tool. We're having multiple models debate in several rounds.
¶28So let's walk through this. And it's not just one model. It's three models. And in fact, this system goes up to five models. But you can see round one just completed.
¶29They all gave their concrete opinion on the question we presented. Debate this claim. You can see that Fable says the claim is wrong. You can see from Gemini 3.7 Flash team should continue treating DuckDB as a primary embedded analytical engine and Deepseek V4 Pro's position is that we should use this as a pilot as a test. We should not move this to a production default generalpurpose multi-tenant server.
¶30Every single model is converging on the same answer here. This is important. This is valuable information. Now, they all spent different amounts of time and tokens on this. As we can see here, of course, our big offender being Fable.
¶31That one response cost me 15 cents. Definitely drop a like. Definitely drop a subscribe so uh YouTube can pay for this. But if we scroll down here, you can see that the next round has started. So, my agents aren't just giving me a single opinion.
¶32They're not just saying something. They're debating. Okay? So, let's scroll down and let me show you exactly what this looks like. Right?
¶33This is round two updated side. So, what happened there? every agent shared its response with every other agent in the system. And so we get this type of structure if we look top down, right? Everyone shared their opinion and now they're going to the next round of debate.
¶34You can see how this can be very valuable. We're having unique perspectives and diverse opinions coming out of our language models. Compete, compare, contrast around a central idea, a central topic of debate. So you can see here, no one is budging, unchanged. It makes it a credible network analytic endpoint worth piloting.
¶35Okay, opinion map. There is no pro side position here on change. Position unchanged from Gemini 3.7 Flash entrance deepseek V4 Pro0813. And then here's the cool part. You can see the agents taking in the opinions of the others.
¶36So we have our rune agent, flux agent, drift agent hiding the model name so that these agents don't get weird and start competing and start sabotaging each other. Something that just emerges naturally. I'm not sure what that's about, if that's part of the training data or how that emerges, but it's something important to note. If we keep scrolling down here, you know, you can see reputations, agreements, what changed my mind, nothing did. You can see that they're basically just reinforcing each other's decision now.
¶37And then we have our final closing statements. So, the debate is ending. Everyone gives a final answer. Reject the claim. Reject the claim.
¶38Reject the claim. And of course, we've prompted the system so that they have a nice structured debate and it all makes logical sense. And that ends our debate. There's the total cost for the entire session. This is what we're spending so far.
¶39Quick call outs here. When you're comparing these tools side by side, you can very, very quickly see how fast are these models and who is performing the best. So far, performance is dead even. And you can see, of course, the cost. Part of what's going to happen here.
¶40As you already know, we're just going to keep emphasizing how expensive these state-of-the-art models are versus our A tier compute. The intelligence explosion is really giving us great A tier workhorse compute that is competitive for most tasks with the state-of-the-art. That's a really really big driver that's happening right now. We're getting cost down. I care less about cost down and I care more about effective agent hour per token cost.
¶41That's what really matters here. And in this A tier, we're getting a lot of heavy hitters specifically in this top row. The benchmarks are really really spelling success for this top row. The gap here is really large on a cost basis. If we look at Gemini 3.7 Flash, less than a dollar in, four bucks out.
¶42Deepc V4 Pro going even further beyond that. And then we look at these state-of-the-art model prices, it gets really, really hard to keep using these unless you're inside some subscription plan that is kind of keeping you in a nice bubble defended from the true costs of these models. A lot of people don't know that there's some hidden text within a lot of these models. For instance, GPT 5.6 series. Once you pass that 280K mark, the price doubles for your input and 1.5x for your output.
¶43So, this model is not as cheap as it looks. If you compare Opus 5 and GPT 5.6 Soul, Opus 5 is better and it's cheaper. Now, just like in last week's video, there are some problems with Opus. We employed some serious system prompt engineering to fix a lot of the ticks to address how verbose, how many output tokens, and how loadbearing Opus 5 is. Check out that video.
¶44I'll link that in the description for you if you want to master system prompt engineering and get more leverage and stop hyperfixating on your skills. But you can see here that on many dimensions, not all models are created equally. And that's what the model stack is about. We're combining compute. We're not selecting it.
¶45That's great. So that's debate. We're putting models against each other against a specific thesis that we want to flesh out more. This is great not just for learning new tools or technology or giving you direction. This is really, really great for decision making.
¶46Strategic decision-m is one of the highest leveraged things an engineer, a PM, pretty much anyone can do. If you take an action that gives you 10x, 100x leverage, that can change everything. Some decisions you and I commit to, they're month-long decisions. They're year-long decisions. So, if you're going to put that much time in, you better be sitting down with some of the best intelligence available to you to help you make the decision or at least flesh out the decision you're planning on making or giving you counter points.
¶47So that's the debate command. This is easily my favorite command of the three. But the next one I'm going to share with you is the [music] most powerful/ FH collaborate. Now our agents have context. We've been having these discussions.
¶48They're debating. Now I'm going to say this. Build a working demo of the top three features from the DuckDB release executable via Astral UV single file scripts. Fire it off. And now they're going to start a collaboration agentic workflow.
¶49Again, this is built into the tool. When I boot this agent up, I have a unique workflow that basically no one else has that gives me outsiz leverage on my agents and therefore more ROI on my engineering work. Now, of course, you're here. You're with me. Big shout out to you.
¶50I'm going to link this tool in the description for you. You're going to have the V2 version with all these tools available to you. Definitely drop a like, subscribe, all that good stuff. if this is valuable information and if this is helping you understand the landscape even just a little bit better. We're starting out with the same workflow.
¶51Our agents are looking at what's going on here. They're understanding the current project structure. Now, something really interesting is going to happen. As soon as Deep Seek V4 Pro finishes up here, all of our agents are going to present a plan. Once again, we're scaling our compute to scaler impact.
¶52We're not just getting one plan from one model. We're having every model in our system give us a concrete plan. And then, as you can imagine, we have our two builders here. and we have one architect. The fusion harness always has one architect.
¶53This is where you're going to want to put your most powerful model that you're willing to spend on because the architect is what's going to take all of the opinions, all the plans of our agents, and it's going to put them together into something incredible. So, let's look at the plans, right? Here's a proposed end state from every single model. We're getting multiple perspectives from our agents. Cloud Fable 5, Gemini 3.7 Flash, Deep Seek V4 Pro, and they're basically putting together a plan.
¶54Now, you can see here, very similar. This is a relatively simple thing to do. We want three Astral UV scripts to showcase the three top features we should pay attention to from the DuckDB v2 preview release. They're breaking down their plan piece by piece task dependencies. And then they're doing something interesting.
¶55They're all explicitly talking about their team, who should build what, and then we have a collisions and safety concerns. We've basically done a little bit of templating with a simple system prompt to help guide our agents create their own plan. So, if you watched last week's system prompt engineering video, you might notice something really awesome here. We're utilizing that inside of the system prompt for all of our agents. You can see here we have our reference points T1, T2, T3 that we can quickly reference between any one of our agents or all of our agents if we want to.
¶56You can see here T1, T2, 33, T3, different tasks. There's a risk analysis. We have our Rs, so on and so forth. And we have clear language. We're clearly communicating with our agent.
¶57There's no loadbearing garbage in here. And again, I'll link last week's video in the description for you if you want to understand how we fixed smartass opus 5 by recentering on the most important skill any engineer can learn, prompt engineering. The simplest things often have the greatest value if you do them right. Prompt engineering is one of them. So, our architect agent is creating a simple task list where you build and there's a dependency on previous tasks.
¶58You're very familiar with systems like this. The great part about this is that this is getting built on the fly. And our architect agent assigns an owner and a mode for each one of these tasks. And we have task dependencies. Classic stuff.
¶59If we scroll down here, you can see our agents are just getting to work. And we've built in our till done lists that we've talked about in the past. It's a to-do list for our agents except we are explicitly assigning tasks and the dependency workflow for each task. So this lets our models collaborate on a solution and deliver a unique solution using many many models. I'm using three models here mostly for simplicity.
¶60We can see these three models here side by side by side. But there are versions you can run here. If we open up a new terminal here I have J fusion 5 where we have five models put together [laughter] and we could do the same thing with other models. You can scale this up or down as much as you like. The whole idea here is you want to combine compute.
¶61You don't want to select compute. You want to use the best model for the job and you want to use a tool like the pi coding agent that's extensible that lets you customize the very experience of your agentic engineering. And so as I'm working through testing, validating, building with models for inloop agentic coding and outloop agentic coding. I'm constantly comparing these models to truly understand what does a token here give me? How fast does that token come out?
¶62And how valuable is that token versus the cost of the token. Once again, same theme, nothing new. Fable 65 cents for that run. Gemini 3.7 Flash just 7 cents and Deepseek V4 Pro just 5 cents for that run. So, really important stuff.
¶63You can also see here something really fascinating. Gemini 3.7 Flash is likely the fastest, most useful model in existence right now. It is just insanely fast. You know, a lot of people are saying Gemini Google, they're falling off the AI race. They can't produce a pro model.
¶64They're not competing anymore. When I look at Gemini 3.7 Flash, I see a strategic decision to optimize for speed, cost, and intelligence and to say we don't need to be at the frontier to be relevant here. And I don't know about you, but when I look at this and I use this model, they're 100% right and they're doing it. I of course have private benchmarks for these models and I sit down like this and run these models together side by side by side to really understand what's happening. So you can see tasks getting completed here one by one.
¶65Fable's done. Flash is done. Deepseek V4 Pro getting to work here. One thing I have noticed about Deep Seek V4 Pro, it thinks a lot. This is one of the heavy deep thinking models.
¶66Very similar to Kimmy K3. Some of the other models like the new local model uh Quinn 3.827B. Great model. Also thinks a lot and is thinking a lot without mixture of experts and it really really hurts the performance and the time to get those tokens back on your local machine. Obviously throw this on a GPU node somewhere, a cluster somewhere and it runs a lot faster.
¶67But when we're thinking about lightweight models, models we can own running locally. This is where the Quinn 3.8 model has been a little disappointing for me. It's performing really, really well. If we open up Artificial Analysis, it's an insane model that's getting above that 50 mark. But when I sit down to use this thing, the context window becomes a serious problem.
¶68So, all that being said, it's an incredible model. I'm using this model locally, but we're absolutely going to need some performance speedups on this model, which I am very, very confident they're coming. It's pretty crazy that at least on this benchmark, you can't take benchmarks as law, as truth, for everything. Really, you need to deploy these models against your specific use case to really know if they're valuable, but just as input to your decision, we have an open weights model that you can put on your device that's running at the performance of GBT 5.6 Luna Max. So, very, very powerful.
¶69It's also apparently one intelligence point away from GLM 5.2 Max. This is where things kind of fall apart. Is it really that good? It doesn't really feel that good. Always take the intelligence benchmarks with a grain of salt and build great extensible flexible tools where you can compare them yourself.
¶70You can see here our agents are working through that checklist. What is this? This is multi- aent orchestration. What I've just showcased here are three different multi- aent orchestration patterns you can deploy into the fabric of your Asian harness. Big fan of cloud code.
¶71Still using cloud code. The subscription is one of the big pieces keeping me continuously using cloud code. But there are things that you just can't do when you're locked into someone else's agent harness. When you're locked into someone else's product, that's by design, but it's also going to hold you back, especially if you're an engineer that is pushing into the new role where we are still software engineering, but with autonomous technology that can take actions on your behalf. If you're really pushing into that, the agent harness is the thing to own.
¶72And if the tool you're using, closed source or open source, isn't customizable enough, isn't flexible enough, doesn't let you change as models continue to change, you're going to run into some roadblocks that you don't need to. And other engineers like myself and like many other engineers that follow the channel, these are just not problems we're going to run into because we have both a closed source option and an open- source super flexible option. I totally get cost optimizing. I totally get using out of the box models, especially for enterprises, like you really need a lot of extra stuff. But for experimenting, for really understanding the landscape, you really, really, really want to own your agent harness.
¶73And then the next step is to own your model. More on that later. Obviously, that's a much harder problem to solve. We talked about that a lot in the isthropic stealing your data while you pay for it video. I'll link that in the description for you as well.
¶74But you can see here our agents are just working through the checklist. It looks like right now Flux, which is our Gemini 3.7 Flash, our lightning flash fast model, is appending the demo receipt to our just file so that we can quickly run this demo of all these features. Last but not least, our architect is responsible for the final integration. It's doing the validation. If you're in the AI industry, this is stuff you've seen before, but we're combining compute to do the job of what we would normally deploy one model for.
¶75We can deploy many, understand the model, and get work done and have our compute check itself. are adding more compute, adding more scale so that we can have more impact. On the model level, I highly recommend you check out Gemini 3.7 Flash. It's incredibly cheap. Big fan of Deep Seek V4 Pro.
¶76This is going to be the open weights model if you have the hardware, if you have the technology to host this and also fine-tune this. This is going to be a fantastic model much like Deepseek V4 Flash. Even as I look at all these models quote unquote catching up when I sit down to do hardcore advanced agentic engineering work, I'm still orchestrating with the clawed fable 5 model. I'm not orchestrating with Opus 5. I prefer Fable 5.
¶77Opus 5 is fine. It has a lot of problems. It's too hungry. It's looking for problems where they don't exist. And again, last week we covered how you can try to attack this in the system prompt.
¶78That works decently well. But you know, for a lot of the work I'm doing in loop, I am opening up cloud code, throwing Fable at it, and then I'm using my many, many custom variants of my PI coding agent to get the job done. And this is one of the variants that I'm putting more time into because I'm getting a lot of nice results out of this model out of multiple model perspectives. Again, as there is an intelligence explosion, it's going to become more and more important for you to understand each model. And this is a great way to understand multiple models at the same time.
¶79Here's our final result. Here's our final cost analysis. As you can see here, it's about an order of magnitude cheaper to use 3.7 Flash or V4 Pro over a state-of-the-art model like Fable. I know a lot of engineers watching the channel that like to complain in the comment section while you're really using the API while you're really spending those tokens. Yes, I am.
¶80If you want outsized returns, you have to do things others aren't doing. This is one of them. It's one of the easy ways to spend capital to get an advantage. Of course, use your subscription. Of course, look for the good deal.
¶81Again, we talked about this. When you're using the API, you have higher levels of protection for your IP. There's real ZDR built in. So anyway, but here is our fusion harness result completed. Our architect says collaboration is closed.
¶82Here is the result. You can see there's all the deliverables. If we open up VS Code, we can see exactly what was built in this codebase. Here is our demos. You can see here I also threw the system prompt from our fixing Opus 5 video from last week right in here.
¶83Gave it to all my models. We can look at the configuration. Super super simple. Here's the one I was using for this. Append system prompt.
¶84You can append multiple system prompts. Pretty nice. But then just model name thinking, keep it simple, don't over complicate things. And here is our demos that our agents built showcasing the new duct DB work. Good stuff here.
¶85I'm not going to run this. Obviously, this isn't the point of the video. This is more for my personal exploration of DuckDB. I'll go ahead and leave the demos in this codebase. If you're interested, I'll port it over to the full fusion harness codebase, which is available to you.
¶86Link in the description for the full breakdown for the full fusion harness. This is the V2 version. We did a video on this a few weeks ago. I really just want to stress that idea. You want to think in ands or it's not GPT 5.6 soul versus fable.
¶87You want to use them together. You want to stack up compute. Combine compute. Don't select compute. That is the catchphrase for this tool right at the top.
¶88Combine your compute. There are many other commands I did not run here. I'll leave this as expatory work you can do if you're interested in the fusion harness. Auto validate is a big one. The fusion command is another big one.
¶89So, link in the description for you. Think about the leverage that's stacking up versus an engineer using a tool like this versus using some out of the box tool. They run three models. They get three concrete opinions. They get three different builds.
¶90They get three different plans. Then they combine it or they select the best. There's so many ways to build flexible pi coding agents. Build flexible tools and then use these models together to outperform their normal distribution. More and more and more engineers are using out of the box tools.
¶91Once you understand that you can harness engineer your own specialized units surrounding your agents, then comes the next level that we've been talking about on the channel week over week, the software factory. This is where you combine agents plus code to outperform either alone. This is where the next generation leverage is. Very few engineers are really pushing toward this edge. This is where the real alpha is now.
¶92It's not about you sitting in a terminal babysitting your agent, prompting back and forth over and over for every problem you have. This is about you using outloop agenta coding. It's about stepping out the loop where you build software factories that work on your behalf without you. So that's the next step of leverage. If you don't know what to focus on, if you feel like you're stalled and you understand prompt engineering, context engineering, harness engineering, and if you understand that it's not loop engineering you're after, it's building and managing different variants of the software developer life cycle.
¶93If you get that, the next step is the software factory. You have to get out the loop. The technology is there. The models as you can see very very clearly are there. Now it's about you taking that next step.
¶94Stop blaming the model. Stop blaming the tools. Everything is in your control. If you do this right, you can change everything. But you have to put your foot forward.
¶95You have to take responsibility for [music] the results you're seeing. The next level is the software factory. We're going to keep seeing more of these intelligence explosions. I'm going to keep tracking all these models. My goal here on the channel is to be a trusted source in input to your decision-making [music] process to help you push your agentic engineering far beyond the rest.
¶96Huge thanks from me to [music] you for watching to the end. Drop a like if you can see the opportunity the intelligence explosion is unlocking for all of us. You know [music] where to find me every single Monday. Stay focused and keep building.