¶1What's up engineers? Indy Devdan here. It's time to focus. This is going to be a big video and a lot of engineers are going to miss the leverage I'm going to share in this one. I don't want you to be one of them.
¶2Your software factory can unlock an unprecedented level of engineering results. But there's a massive roadblock every agentic engineer runs into at some point. A roadblock you might have. Where should your software factory run? Where do your agents plus code actually run?
¶3Most engineers allocate a tiny corner of their own computer for their agents, or they lean too heavily on CI/CD or some container. The best engineers aren't doing this. Engineers at big AI labs running the most insane workloads you can imagine do something else because they've realized one simple fact. The more autonomous and secure your agents are, the more they can do for you riskfree. On this channel, across hundreds of agentic engineering videos, you've heard me say scale your compute to scale your impact.
¶4Compute applies equally to GPU compute and CPU compute. Old school bare metal. An agent sandbox is the best space to place your agents. [music] Why is that? Why can't we just use a container and call it solved?
¶5It's because agent sandboxes give you three key advantages: true isolation, insane scale, and agency. Here's the raw reality of the state of agentic engineering. If you are inside the loop, you are the bottleneck. In this video, I'm going to break down exactly how you can step outside the loop by putting your software factory inside an agent sandbox. Let's put your [music] factory in a box.
¶6Here's the tech stack we're going to be using to put our factory inside a box. Don't focus on the tool. Focus on the purpose and choose the best tool you need. Here's the architecture of the system as well. We'll throw this up on the screen.
¶7And why did I put the second orchestrator agent inside the sandbox? It'll all make sense as we progress. Let's start with a powerful workflow you can't do on containers on your device. Best of N. We'll boot up Herder.
¶8This is my terminal multiplexer of choice for watching agents, kicking off jobs, running arbitrary commands, application directory. And we're going to use claw code fable as our tople orchestrator inside of our software factory. We're going to use pi as our agent SDK. Don't just limit yourself to one agentic coding tool. The tools you use directly limit what you believe is possible.
¶9Let's go ahead and kick off the skill that activates our agents understanding. This is going to be SSF, agent, sandbox, orchestrator. I'll break this all down in a second. It understands the skill. It's got all the context primed, ready to go.
¶10And it can see that we have one VM running. We're not going to touch this one. We're going to open up brand new VMs. We're going to run the best of end pattern. We're putting our software factories into an agent sandbox.
¶11Let me show you exactly what that looks like. I'm going to open up VS Code. I'll grab my demo prompt here. Copy this. Feel free to pause.
¶12We're going to break this down in a second. I'm going to paste this in and fire it off. This top level orchestrator agent that operates outside of all of our sandboxes is going to start orchestrating this workflow. If you were here for last week's video where we talked about the super simple software factory, you're going to love what you're going to see next. So, we're running five different agent configurations.
¶13Default, Frontier, Deepest, Seek, Open Weights, and Top Speed, each with its own sandbox. And they're all going to solve the same problem. But we're not just launching agents. We're running an AI developer workflow. We're running the full software developer life cycle against our simple application which you're going to see here in a second.
¶14But you can see here our orchestrator is kicking off not just agents inside sandboxes. Our agent is building our software factory and placing the entire factory inside our agent sandbox. And this unlocks a lot of really really powerful things. When you use agent sandboxes, you're getting isolation, you're getting scale, and you're getting autonomy. Every single agent owns the computer.
¶15They're not just running on a little corner of yours interfering with your work, you interfering with their work. They have everything they need to own the outcome. We are using exe.dev as my primary sandbox tool. One of the best, simplest ways to get started with agent sandboxes. I'll of course link all the resources we're going to talk about throughout this video in the description for you, but their tagline keeps it nice and simple for you.
¶16Computers for developers and agents. Durable sandboxes, fast, secure, and sharable. The agent sandbox is a developer device for your agents. Let's just like really atomize everything. Let's say exactly what things are.
¶17There's a lot of hype and noise in the AI industry. I try to dehype things as much as possible as complex as something like this is. We're not just throwing a expensive claw opus cloud fable Asian at this. We are understanding what model needs to run where with what code surrounding it. Right?
¶18That's the software factory. That's the next step. So this is a lot more powerful than just a skill and an agent and spin up this web app inside of this service. Orchestrating other agents, orchestrating sandboxes does require state-of-the-art intelligence. I always track these models inside of my model stack here.
¶19You've probably seen this if you've been with the channel. I delineate between three tiers of models. We have state-of-the-art, we have workhorse, and we have our lightweight models, which are basically the floor. These lightweight models are defined by their ability to be run by previous generation GPU node or directly on your Mac device if you have the unified memory for it. And then we have our workh horses.
¶20These are all of the powerful models that aren't quite state-of-the-art. The newest entry which you're going to see here is the brand new Deep Seek V4 Flash0731. This model is absurd for its price. You're going to see a model configuration that we're going to kick off that I call Deepest Seek, which is just V4 Flash models. We're going to see how they compare to a bunch of other model plus harness configurations inside of our software factory.
¶21If you're still fixated on models, you are behind. That's not the name of the game anymore. We're entering an age of abundance of compute. And now it's about focusing on what model plus what code do you need to do the job. And the software factory is that system.
¶22So our agent is finished. You can see we have the application layer and the agentic layer. Let's go ahead and jump into them. Open these all in Chrome. Let's see what's going on in our system.
¶23We're going to open them all up in Chrome. And as you can see, we're working on Inkwell. This is a simple mock application. This is not the important piece. We're using this as a demo.
¶24What's important here is our software factories. So, what you're going to see here is 1 2 3 4 five unique software factories running in their own agent sandboxes. So, inside of that prompt, this user interface is like not clear. There's too many buttons. Things are too out in the open.
¶25And we want a more unified experience for our Inkwell writing application. Instead of just running this once on one agent or two agents, I'm running them in software factories. We're running the simple software developer life cycle. Again, we covered this last week. The super simple software factory is a key piece of this, but it's not our key focus.
¶26We're using this as our observability system into our sandbox. We have discrete URLs for every sandbox. This is our default Asian config. This is our front tier Asian config. This is our deepest seek Asian config.
¶27We then have our open weights configuration and then our top speed. And as you can see here, we've opened up two applications for each set. The software factory and the actual application. So you can see we have five unique versions. These are each inside an agent sandbox.
¶28If we hop back over to exe.dev, do a refresh, you can actually see all of our sandboxes up plus the one I had up earlier. Okay, so we have five new redesigned agent sandboxes. So again, like really important concept. I'm not giving my agents a corner of my computer. I'm giving them an entire computer to do all the work they need to with isolation, with scale, and autonomy.
¶29Let's click into our software factories. You can kind of see our model breakdown here on the left. And we can do this for every one of our versions, right? Let's go ahead and click into everything. Now, we can get an understanding of the models that are running, the agent harnesses that each model is running, the tools available.
¶30I want to keep uping the discussion here. It's not just about a single agent anymore. One agent is not enough. just like one prompt was not enough. We then started running sub agents.
¶31We then started chaining agents. Maybe you put together your own custom agent harness thanks to great tools like the PI agent harness. Fantastic. At some level you realize that agentic engineering is software engineering. We have a new primitive which is the agent autonomous software and we need to put them inside of our existing developer workflows that we ourselves would run.
¶32Hence the term AI developer workflow. We're adding AI to the work you and I used to do as engineers. every single day. So, we have this new tool. We need a new frame of mind to understand how we can really use this in the best way.
¶33And spoilers, you know, the whole point of me creating this video and sharing it with you here is that when your agents have their own developer device just like you do, guess what they can do? They can perform and act and even outperform you. That's the whole goal here. A lot of engineers are afraid, you know, of these great models progressing further and further and further. I am not.
¶34This is a new tool in your toolbox. Don't run away from the fire. Run into the fire. These agents can do incredible work on your behalf for you, for your business, for your career, for your users, but you have to leverage it. So, inside of every single sandbox, we have placed a set of agent configurations and we built an AI developer workflow that walks through the software developer life cycle and our agents are shipping work on our behalf.
¶35Here's a key word without us. The key idea that I really want to communicate here is that if you are in the loop, you are always the bottleneck. Now, there are exceptions to this. If you're doing hands-on agentic coding work that requires you, that requires your expertise and you know a great exception to this is like if you are building the system that builds the system, then of course get in the loop, do the work, go hands-on, prompt back and forth, babysit your agent, you need to do that at some level. But there's a lot of work you might be doing.
¶36I know for a fact teams are doing and god forbid we talk about enterprises like there's a ton of work all these sets of engineers are doing that they don't need to be. Why not? because they can have a team of agents do it for them. And not just a team of agents, a team of agents plus code plus engineers. This is our default agent set.
¶37This is our frontier agent set. This is our deepest seek agent set. Spoiler, everything's deepseek v4 flash. We're saving an incredible amount of cache of compute. Then we have our open weight.
¶38So pure openweight implementation, no closed labs here. At the end, by the way, we're going to compare all the changes made to Inkwell. in each one of these versions. Back to the source. That's why this is called best of end.
¶39Say you have a great idea. Say you have multiple directions you want to go or even one single concrete direction. In our example here, we just pass in this one singular prompt. And this is actually running our redesign quiet room. We want a quiet room as our design principle for this application.
¶40The current version is cluttered. We want it to be simple and clear. So this is the prompt we pass into every agent sandbox and then into the software factory inside each agent sandbox. As you can see, I'm not just running one agent. I'm running a pipeline of agents plus code.
¶41And of course, inside the incoming task, we can dial into all the results here. You know, this is a classic agent observability system. We've talked about these on the channel in the past. You know, you can see the request here. Redesign Inkwell.
¶42But then, you know, we can look at the plan workflow. Of course, we can see all the thinking, all the agent calls, the gate passes. We have deterministic gate checks on our nondeterministic systems. JSON error, it fixes JSON error, so on and so forth. Again, we talked about this last week, so check out that video if you want to dig in depth into our super simple software factory that I have fully open sourced.
¶43Oh, nice. GLM 5.2 review has kicked off here. Love to see that. Pure open weights. In our super simple software factory, we are not just doing lightweight prompts.
¶44We are fully in control of the system. What do I mean by that? We are selecting a specific coding agent. Of course, the model, of course, the thinking level, of course, the tools. Oh, what are these tools?
¶45These tools came from our custom agent harness. Yes, we are harness engineering inside of our software factory. If you don't have custom harnesses, you can customize and control. You don't have a software factory. The software factory is your full control over the agents and your code.
¶46We have a nice simple purpose breaking down what this phase is, which is just the plan phase of this workflow. And then we have a nice description. We have our prompt engineering. Okay, so what is prompt engineering? It's both your system prompt and it's your user prompt.
¶47Okay, so on and so forth. I'm not going to download this again. Check out last week's video where we really dive in to the super simple software factory. If you've been with the channel, you're going to notice we're going to start compounding some really heavy-hitting ideas. We're going to lose a lot of engineers.
¶48A lot of engineers are going to look at this and just think, I don't need this. I don't need to build this. Insert whatever excuse or lack of focus or whatever. I don't really know. I don't want to judge anyone, but engineers that can focus right now.
¶49This is the time to focus because the next leap is available. It's your software factory. And then you can really leverage and scale your software factory by putting your factory in a box. So that's all we're doing here. At a high level architecturally, it's relatively simple.
¶50The implementation details do make things a little more complex, but you can see here our default run is running, you know, gem flash, deep flash, GLM52, and GBT Luna. Just a lot of great compute. If you're constantly going back and forth on what model you're using, you're missing the point. You need a model stack, not a single model. Combine compute.
¶51Don't select compute. Let's go and look. Of course, our Opus 5 going crazy. 13 minutes on the plan phase to be expected. I'm sure this is in high thinking.
¶52If we select this, look at our Asian config. Yeah, of course we're in high thinking. Maybe should have dropped that down for this demo. Who cares? Our deepest config.
¶53[laughter] All pure DeepSk implementation, saving a ton of cash, using a ton of tokens. Love to see this. Again, we're entering an age of abundant compute. And you know, by the way, we're going to dial into the system a little bit more here. You can see here our top level orchestrator agent has stopped.
¶54It's not doing anything. All it did was kick off the orchestrators inside our sandbox. So again, we have this three- tier architecture and there's a reason for that. I'll explain that in a moment. We have this three-tier architecture where we have an outloop orchestrator.
¶55We have an in sandbox orchestrator and then we have our software factory. Looks like we have a full fail here. One of our jobs, open weights, completely failed. ChemK3 bombed. You can see here we have a nice detail of why this happened.
¶56That's okay. This is another reason why best event is so important. It looks like it couldn't get the JSON format out. We could dial into the results. I'm not going to waste your time with this.
¶57This is why we use best event. This is why we scaler compute to scaler impact. And what do we have here? This is our top speed. Um, of course, top speed is already done.
¶5814 minutes for this full run. 2 million toes. And look at all of our fast models. If you just care about speed, you want to be looking at Gemini 3.6 Flash, DS4 Vlash, 0731, and of course the incredible GBT 5.6 Luna model. We can start looking at our Inkwell applications.
¶59Then we'll dive into more of the like heavy technical side of it. Okay, I just want to like really connect with you on the core idea here. Five instances. Looks like one of them failed. Let's go ahead and just like start refreshing.
¶60So here's the v1. Let's refresh. There is our updated version. So this is just one sandbox. We have our entire app and our Asian factory and all of our Asian configs and an orchestrator inside this one box.
¶61Our agent has a computer. Our agent can do more. It can solve more problems. It is fully autonomous. And guess what?
¶62The blast radius is zero. The blast radius is the box. Very, very important stuff. And API keys are completely solved here. You know, for all of our PI coding agents running on the software factory level, they're using an open router provisioned key.
¶63Really important concept. It's all going to be here linked in the description for you. But let's keep looking at our versions. So, our default agents put this together. Looks really, really good.
¶64We have the title. We can start writing. We feel a lot more focused than our V1. So, here's that V1. We have hotkeys all over the place.
¶65It's not clear. I don't feel focused. We have the sidebar open. I'm distracted. But now in this new clear version, I am focused at the same time.
¶66We didn't lose our sidebar. We can still do all of our editing and we can do our delete publish. So our actions are there. Total words. Let's go over to our frontier version and let's actually check on this because it might not even be done.
¶67Okay. Yeah. So Cammy K3 and our Frontier version is getting to work. This is not complete yet. So refreshing this won't do anything.
¶68So let's move on. Deepest seek. [laughter] Let's see how our deepest seek is doing. Very very nice. Uh let's see if we have our sidebar.
¶69Good. We don't want to lose anything. Menu at the top right. Very very good. Love the way that looks.
¶70We can scale up, scale down. Three words saved. We are just writing we feel focused. Shout out to the Deep Seek engineers. What an incredible model and what insane prices.
¶71When we look at artificial analysis, we go to flash. Guess who's at the very bottom of the cost per task benchmark. And guess who's at a decent amount of speed, right? 110. Okay, very good.
¶72And guess who's in that workhorse tier? It's Deepseek V4 Flash. Now, I'm super stoked. I'm waiting for two new sets of models to come out. I'm waiting for aside from all the state-of-the-art stuff.
¶73You know, I always want Anthropic OpenAI to really push out the state of the art, but aside from that, I'm really paying attention to two models. And comment down below if you are as well. I'm looking for that Deep Seek V4 Pro. Super curious where that's going to land, if we're going to get into that state-of-the-art territory or if that model is just going to be a super economical uh Pro that has improved benchmarks. And I'm also waiting for Quinn 3.830 billion or 50 billion or 12 billion.
¶74That like 30 billion parameter model that they've been putting out has been incredible. Really really excited for that to run on device but also put that on a B300 and that thing is going to run ultra ultra fast or some other sets of nodes. My point is there is that I'm super super excited for that compute to be available because again I'm looking at all these levels. We're constantly trading off three things. Performance, cost, speed.
¶75Keep your eye on the ball, but also always move where the ball is going, move where the models are going, and have a plethora of options you can deploy into your software factory. So, back to our factory. This looks good. Deepest seek, incredible model, incredible economics. Like, frankly, stock market crashing economics.
¶76If you really start using and really start understanding this model, it's insane what you can do with this. This is a great builder. It's a great worker model. Again, that's why I named these workhorse. It's for your daily 90% of your work.
¶77Previous Deep Seek Flash was in that B tier. This is clearly in the A tier. And A tier is like roughly just to have something to benchmark against. The A tier kind of is in this zone here, right? It's in that 50 plus level in the artificial analysis intelligence index.
¶78And then of course, your state-of-the-arts are like up here. It's more like this, right? This cut off here is what I consider state-of-the-art. It's like that 56 plus level. But again, that's just rough benchmarking.
¶79This is an index and this thing changes. So, you have to be dynamic with comparing your model stack with public benchmarks. Anyway, let's see what else we have here. deepest seek has fully completed. And by the way, once again, this is not just an agent.
¶80This is not sub agent delegation. This is not random cloud code/workflows execution where you have really no idea what workflow it shows or why. We have built a software developer life cycle that we have fine-tuned. We've templated our engineering into the software development life cycle. We've added AI to it and we wrapped it in a software factory.
¶81And this gives us full control over our engineering, over our products. Be careful how much of your engineering work you're handing off to these AI labs. You're handing off to third-party resources. You're handing off to these all auxiliary tools. Don't outsource your thinking.
¶82Stay close to the engineering. This is aic engineering, not vibe coding. Okay? And I say it all the time. I'm going to keep saying it.
¶83Vibe coding is not knowing what your system does and not looking. Aentic engineering is knowing what your system does so well you don't have to look. Big difference. like that. It's like a really fine line.
¶84If you look at a vibe coder and an agentic engineer, it might look like they're doing the same thing. If you spend a little more time, if you really think about what's happening when an agentic engineer sits down to do work, they're thinking in systems. They're thinking about the hundth and thousandth run of the system. They're thinking in observability, reusability, isolation, scale. So, I hope that makes sense.
¶85Again, a lot of engineers, I'm going to lose a lot of my audience here. I know it for a fact. We're at the point in the age of agents where you have to start working for it. So, here we go. Let's see what top speed did.
¶86Open weights bombed. Okay, we just go ahead and close that. And then we had our top speed. How does top speed look? Pretty good.
¶87Still a little bit noise, although you know, nice hot keys in the footer. We have shortcuts at question mark. Okay, we like to see that. And then we have our nice writing view. And that looks good.
¶88We do like that confirmation of saving. And we can scale this up and down. We have some controls here. Overall, really, really good. These workhorse models, Gemini 3.6 Flash, V4, Luna.
¶89These fast, relatively cheap models. They can do a lot of the work you're probably doing right now. That is our changes to our live application here. Of course, this is just a mock application. I can't share production code with you, but um you know, I can share open source stuff.
¶90I can create uh demos, mocks for you. And you know, the the super simple software factory is the same thing. It is a template for how I'm pulling it in to my products, to my client work, to my production level workloads, then doing a lot of fine-tuning and tweaking. Remember, specialization is your product. The advantage you're going to have is not an out-of-the-box agent.
¶91It's a specialized system that does something really, really well at the best economics, right? The best tokconomics. And again, your software factory inside your agent sandbox lets you scale that. It lets you move faster, lets you get out the loop. Best event, fuse the results, move on, hand the work off to your factory.
¶92We really dive into the details of the super simple software factory in last week's video. Check that one out. I haven't looked to see how much engineers like that one yet, but that's there. A lot of value in that one and a lot more value in this one. It's hard to follow, but let me break down some of the implementation details for you so that it's easier to follow.
¶93I'm still observing this stuff. At any point in time, I can enter these sandboxes and I can enter the orchestration. So, let me show you what I mean. In my demo prompt here, we can get shell access. So, let's get shell access.
¶94I'm going to fire this off. And I've got a couple really, really important skills that my orchestrator agent is working off. If you open up cloud skills, you can see here I have four essential skills. Herder sandbox exe.dev super simple software factory and then our super simple software factory orchestrator agent. So this is a composite skill.
¶95I don't like to compose my skills. It makes it really hard to change things. You create these nasty dependency graphs. But when things are unified in a single monor repo, you can make an exception to the rule. Our agent is going to use herder, my terminal multiplexer here to kick off this sandbox redesign fleet.
¶96Check this out. We have five panes. And now my agent is going to boot up SSH access into every single one of them. There we go. A key idea I've been talking about on the channel over and over is a gentic access.
¶97If you want to move at light speed, if you want to move the speed of agents, you need an agent layer wrapping whatever you're trying to do, right? It's that meta layer. It's that building the system that builds the system. You can see here we're in all of our sandboxes, right? Here's our redesign default.
¶98Here's the frontier, open weights, top speed, deepest seek, right? It's all here. And then we can Oh, yeah. I don't have any of my hot keys. We can get status.
¶99We can get log to see the exact change. Document the equal quiet room redesign. Redesign document blah blah blah. And we can go from there. So this is useful.
¶100It's always that same idea in here. I have to do everything manually. So we can go one step further. If we open up code and we go to our third agent access. Let me just show you what this does.
¶101I want the exact same thing except I want to talk to the orchestrator inside each running sandbox. So what's that? It's an interactive cloud code on top of the boxes. This is workflow B. Again, I've taught my agent exactly how to do this.
¶102So, let's see how it does it. We have a new workspace inside of Herder. Let's go ahead and crack it open. We have a nice grid here. There's the boot up command.
¶103We have to trust the system all the time. I'm sure there's some way to bypass this. This is one of the annoying things about cloud code. Yes, I booted it. Just open it for me.
¶104Our agent will likely take care of this for us though in each one of the sandboxes. Let's go ahead and see how it does here. And there we go. Great. We have a policy and we're getting that dangerous mode warning.
¶105We are operating inside of a sandbox. So, we're doing exactly what the anthropic engineers uh want us to do. And there we go. So, notice here we've connected to an agent that's already ran. This is the three- tier orchestration system.
¶106We can run whatever we want to, right? Summarize changes made bullets. Copy that. Fire. Move.
¶107Fire. Move. Fire. Move. Fire.
¶108And again, I'm doing this by hand. You can always do this with an agent. My whole point here is if you don't measure it, you can't improve. But if you can't touch your boxes, if you can't jump in, you can't actually get anything done. Our agents are kicking off.
¶109They're reading the ADW state, the actual agent work that was completed. Again, check out last week's video. There's a whole system around this. Our agents are summarizing what exactly was done in each one of these sandboxes. Let me show you what's going on underneath the hood.
¶110This is going to be fully available to you. Link in the description. Although I do a lot of agentic coding for this stuff, it still takes me a lot of time to really build this out, right? and to really make it reusable for you. If you made it this far in the video, if you're getting value out of this, you know, like, subscribe, comment, let the algorithm know you're interested and let me know you're interested so I can keep delivering this type of content.
¶111You know, if this video doesn't perform well cuz the ideas are too advanced, it's fine. Like, I'll just dial it back and go to safer, more hype, you know, based stuff. But I think you and other top engineers are interested in this. You want to be on the frontier. You want to do as much as you possibly can do with agents.
¶112And also like, you know, I might sound like a super red pill engineer like just trying to maximize productivity, but I think a lot of the value in this technology is that you can just work faster, get everything you want done, and then chill, right? Go back to your life, go back to real life, do your thing. I should probably make that clear on the channel. It's not super important to me what you think of my my philosophy around work life balance, but I just want to mention that a lot of this technology you can use to speed up your work and then chill or speed up your work and get a massive edge in front of your competition. So anyway, factory in a box.
¶113Okay, the whole idea here is that you and I are the bottleneck. We want to be building systems that prevent that now. So what we can do, we can take two great pieces of technology. These are compositions that are emerging. We can take sandboxes, throwaway sandboxes that are also permanent if you want to keep them up.
¶114This is why exe.dev is so important. A lot of other sandbox providers give you 24 hours. They give you a couple days. Is a VM you own, and you can spin them up and down as much as you want. It's nice and secure.
¶115We have SSH access, but then inside of it, we have this factory. We have our Asian factory. Again, check out last week's video, but a lot of that again is going to be super familiar for tactical agent coding members. We have this ADW's directory. It's not just about agents.
¶116It's about agents plus code. And then of course we have the actual application. So here we have what are effectively two agentic layers. Really it's just one. And then we have our application layer.
¶117The whole idea is that you want to get yourself out the loop. If you are letting your agents run on your computer, even inside of a container, that's only getting you isolation. It's not giving you scale and autonomy. In a sandbox, you get to leave the loop. You get to show up when you want to.
¶118As you can see here in Herder, I can still show up into the box. I can still jump in, get things done, understand what's been built inside of every agent box. But the whole idea here is I am out this loop. Great agentic engineering is about showing up at the beginning and the end, the planning, the reviewing, the prompting, and the validating. Uh the more you're in the loop, the more you are the bottleneck.
¶119Okay, so that's what this system allows us to do. So we have a couple interesting things happening here. The orchestrator agent, as you can see here, is inside the sandbox. I put an agent in the sandbox. So we have one operating on all the sandboxes.
¶120On the outside, we have an agent operating inside the sandbox on the inside. And this allows them to run our super simple software factory system, right? And then we have the actual agents running from our pi.dev instance inside of the individual software factories. And we have a great visual here. Jesus, you can see that our state-of-the-art still churning here.
¶121Looks like Kimmy 3 is having a hard time. But anyway, our software factory runs any agent configuration. We're controlling the core 4 on an individual agent basis. We are scaling. We are composing.
¶122We are creating compositions. This is what the architecture looks like. Again, this is the like highle architecture, right? We have our out sandbox orchestrator. We have our in sandbox orchestrator.
¶123And then we have our ADW agents. The ADWs. Once you have two or more, what you really have is a software factory. Outbox orchestrator runs on my machine. Inbox orchestrator runs on the sandbox virtual machine.
¶124And then we have the factory. Again, still in the sandbox virtual machine. Couple key commands are agent ran. I'm not going to go through all the details. We'll be here for a while.
¶125But you can imagine there's a setup flow to this. We need to mount, create the sandbox, set it up. We can tear it down whenever we need to. Inside the orchestrator runs whatever a developer workflow we need. Usually, it's going to be the full software developer life cycle or we're fixing a bug or we're doing a hot fix.
¶126Whatever it is, that's up for you to decide what's the most valuable chain of agents plus code your developer workflow must run. And then we have our actual ADW agents, which of course there are atoms, there are scouters or planners or builders or testers. This is what most engineers are hyperfixated on and really it's a key piece but a small piece of a bigger picture getting actual engineering work done and being able to step away from the loop. Here's the actual setup process our orchestrator did end to end. So we host and then inside the box we set up an execute where we kick off the agent and then of course we can observe from outside the loop thanks to our observability tool and the actual application hosting from exe.dev my sandboxing tool.
¶127And then of course at the end here which you'll see in a moment we're just going to tear it all down. These are ephemeral sandboxes. We spin them up. We do work. We spin them down.
¶128Environment variable keys is a big deal right? So in the system this example codebase system. What I'm doing here is showcasing that. What I've done here is used open router's provision key system. I've limited to 50 bucks.
¶129And then when I tear down all these I'm going to kill the key. So really powerful way to spin up intelligence. Open router is a great tool for this. And it all creates a result that looks something like this. again with the best event system.
¶130Our single prompt fired off in our orchestrator. Orchestrator spun up in sandbox orchestrators to run our ADW system. Whatever model configuration we wanted plus whatever code it needs and then we can pull the best result as you saw here. Not all of our systems finished, but you can imagine your application built in multiple different ways. And this is just one pattern, right?
¶131Best event is a great sandboxing pattern for you to explore potential futures you want in your product, in your system, in your application. And best of n gives you the best of some set of results. And of course the great part about that is you can configure your model selection. Every one of these runs is running a different set a different team of agents. A side effect here is as an agentic engineer as you're running these different variants of your software factory.
¶132You're just picking up on data. How good is this configuration? Is it getting me the results I need in a cheaper, more efficient way? Do I need a state-of-the-art system? By the way, this is still running.
¶133This total workflow has taken almost 45 minutes. I mean, I've been filming for 50 minutes now. This workflow has been running basically the entire time. There's clearly a bug here with Kimmy. So, you know, we'll dismiss this runtime, but you can see here Opus 5 churning up tokens, churning up spend, right?
¶134Do you need this? Interesting question. And then, of course, we have our deepest seek doing things the cheapest and one of the fastest. Definitely churning up some tokens though. Here's read, here's right, but getting the job done right in the deepest seek version here, you pay attention to the title, Deepest Seek, and we go over here.
¶135This version was good. like I feel a lot more focused. It's a quiet room for me to write in. That's just a nice side effect of running these multiple variants of your software factory. It's not just about understanding the difference of the model.
¶136It's about understanding the stack of models plus code and the results and the cost of the results. And ultimately the architecture looks like this. We have the application. We have the agent view outbox orchestrator inbox orchestrator and then our software factory. If you want you absolutely can take your outbox orchestrator, go right to the ADW system, your software factory system.
¶137that's fine. Uh you get a little bit less control there because the inbox orchestrator shouldn't really be doing anything but running and kicking off the workflows. The key here is just keeping it simple. Complexity is always going to come as you add another thing. But in reality, you know, it looks like this.
¶138We have our software factory inside the sandbox, public port for the application, and then a private port. You have to log into exe.dev. But basically, you get the software factory view. This factory in a box system is going to be available to you. Link in the description.
¶139It's going to build on a super simple software factory. I'm going to link that video as well. I skipped over some of those key deeper ideas, but as you can see on the channel, I'm going to be stacking up all of these big ideas so that you can scale your computer, to scale your impact. I'm building and using these systems because I need them to get to the next level. And part of my journey here, part of my goal is to share them with you.
¶140Let me be super clear about this. I come here every single week as a secondary, almost a tertiary goal. All of it comes from my goal to build living software that works for me while I sleep. And then my second and third goal is like democratize this because as we continue to become more capable agentic engineers as you can see and feel this is scary powerful technology especially when you start using it properly especially when you start embedding your engineering into it it's very very powerful there's going to be a lot of economic value created a lot of careers changed and I want that to be dispersed I want that to be democratized equally or as equal as we can technology should not be limited to the lucky few that's why I come here and I share these ideas every single week with you. Let's go full circle.
¶141Why are software factories important? Why is this valuable for your engineering? Software factories let you scale your compute to scale your impact by combining non-deterministic agents with deterministic code. You want the best of both worlds. There will be a point in time when agentic engineering is just called software engineering because agents will be part of the fabric of how we build software as they are becoming already today.
¶142But right now we need to clearly delineate we are agentic engineering. This is not really traditional software engineering yet. And we're absolutely not vibe coding. I know what the system is doing. I could find the exact file doing the exact thing.
¶143Big difference there. That's why software factories are important. They let you compose agents plus code into the workflows that you used to run yourself as an engineer to get results done. We're talking plan, build, test, review, document. We're talking hot fix.
¶144A production issue comes in. Can you solve it in record time with your ADW? We're talking about support requests. We're talking about a bunch of And your software factory is the mechanism. It's the system that lets you do that.
¶145That's why that's important. Why are agent sandboxes important? They're important for three reasons. Isolation, scale, autonomy. Isolate your agents so they don't blow up your production system.
¶146Right? Strip them of their ability to access AWS and GCP and insert the next really important tool. They give you scale. I use one orchestrator agent to spin up five sandboxes to spin up factories of agentic work. A bunch of stuff happened there.
¶147Plan build test review document. and it happened end times. Best event is just one pattern. It's going to be a really popular important pattern because it lets you see multiple futures. Scale is very important.
¶148Many of us, just as we have been for a long time, underestimate the importance of scale. What happens if you could just do this 10 times? What about a 100 times? What could you see? What could you find?
¶149What could you build? What could you learn if you just threw more scale at your problem? Sandboxes are scale for your agents plus code. Sandboxes give you scale. And of course, they give you autonomy.
¶150Don't just limit your agents to a little corner on your computer. You can and should do that when you're working with your agents on net new work, specifically building systems that build systems. I'll just say this, you know, upfront, right in your face. If you're using an agent to directly modify the application layer code, you are wasting time. There are very few exceptions to that.
¶151I hope that makes sense to you. I'm not trying to be a douchebag or talk like I'm above you or anything like that. I'm saying this so you understand where the leverage is. I don't want you to fall behind. So that is autonomy.
¶152Sandboxes give you isolation, scale, and autonomy. When you put these two things together, you get something extraordinary. You get absurd scale inside your sandboxes. When you put a factory in a box, you get absurd scale, absurd leverage. You get really great isolation because you're going to want it [laughter] as you build more and more agents on top of your system.
¶153And of course, you can just scale and solve many, many, many problems at the same time while not limiting the agency of your agents. Okay? Ultimately, you and I, what are we doing here? We're preparing for an age of abundant compute. You see how the landscape is shifting.
¶154You can see these new great cheap workhorse models. You can see the intelligence getting passed down the tier levels. It doesn't matter if you believe it's right or wrong, it's happening. And so, what do we do as engineers? We act like a function.
¶155We take the inputs, we execute, we synthesize, and we deliver outputs. That's all we're doing here. taking the information, make the best decision, act. This implementation of the factory in a box is going to be available to you. Link in the description.
¶156I also highly recommend you check out last week's video. Super simple software factory. It's going to be a banger one. You'll notice once again I didn't mention loop engineering once. I believe this is the wrong term, the wrong phrase.
¶157What we really should be focusing on is the software developer life cycle. If you haven't seen that, I'll link that in the description as well. Take the pieces of everything I'm sharing with you, merge it with your own. I don't have all the answers here, but what I do have is a great direction. Try to get the leverage of the next phase of agentic engineering.
¶158All right, these are relatively advanced concepts, but I don't want you to miss out on the opportunity to scale up. If you made it to the end, you know, hats off to you. Things are going to get more complex from here because all the low-level leverage has been sucked up. Everyone's prompting inside Cloud Code now. Everyone has a terminal agent.
¶159Everyone's using a co-work tool. The next level of leverage is going to take work. It's going to take upfront investment. But as software engineers, there's a lot you can do that you don't realize yet. And that's what I'm going to be here week after week to help you unlock.
¶160You know where to find me [music] every single week. Stay focused and keep building.