Radar

radar · IA & agentes

FIXING Opus 5: PROOF that Prompt Engineering IS NOT DEAD

¶1What's up, engineers? Andy Devdan here. Like me, you've probably gotten sick and tired of Opus 5's insanely verbose responses and its overuse of phrases like loadbearing, worth stating plainly, here's the honest truth, and a bunch of others. Or maybe you're tired of the Anthropic team trying to take credit in your Git commit messages for intelligence you paid for. Or maybe you notice Opus 5 is burning your cash with way more output tokens than any model before it.

¶2You're not alone. Myself and many engineers feel the exact same way. Obus 5 is one of the best state-of-the-art ultra smart models and one of the worst state-of-the-art models ever released because it talks like a complete smartass. In this video, we turn smartass Opus 5 into a precise senior engineer that's enjoyable to work with. How are we going to do that?

¶3We're going to use one of the most important skills any engineer using agents can learn. You know what it is? It's prompt engineering. The skill that was once a complete joke is now the most important skill for [music] any engineer to master to scale their impact with agents. If you hear someone say prompt engineering is dead, completely ignore them.

¶4They have no idea what they're talking about, I've been engineering for 15 years now, and one thing that all the best engineers I've worked with know how to do extraordinarily well is this. Above all, they know how [music] to communicate with their technology. And guess how you communicate with agents? Yes, prompts. There are two ways [music] to prompt engineer your agents.

¶5Most engineers only use one when the second is most powerful by far. By the end of [music] this video, you'll have the prompt engineering expertise to make your AI agent communicate like a precise cracked senior engineer. No matter the model that's running today, it's Opus 5. Tomorrow [music] it'll be something else. Let's fix Opus 5 with system prompt engineering.

¶6What are the two ways to prompt your agents? There's, of course, the user prompt, which you're very familiar with, the single task at hand. But then there's the system prompt, which is the law for every task you hand your agent. Most engineers fixate on skills, skills, skills, skills slash this slash that slash plan/grill/re. Don't get me wrong, skills are powerful.

¶7The user prompt is powerful, but the system prompt is vastly more useful and most engineers never even touch it. Because the system prompt is where every single word you write is multiplied over every single user prompt. The system prompt is essential for setting up great communication patterns with your agents and for reducing those expensive Opus 5 output token costs dramatically. The system prompt is where the real leverage is. Let me show you exactly what I mean.

¶8in VS Code. Here we have the simple three file structure. I've clone down Zuck's the future is for everyone. We're going to do what all of us degenerates do nowadays. We just summarize a long article like this.

¶9We're not going to read this whole thing. We don't have the time. We're too busy with 10 million agents open in our terminal window. To do that, we're going to use a standard powerful clawed opus 5 agent. As we work through this, we're going to constantly improve the system prompt we're passing in to clawed code.

¶10I have this simple command here called compare. So, let's go ahead and fire it off inside of a terminal. We're running herder. This is my terminal multiplexer of choice. You'll see exactly why in a second.

¶11If I type J, you can see all the commands in this directory. And what I'll do is just type J compare. And we'll start with zero. I now have a brand new workspace open. Let's go and hop into it and see what's going on.

¶12You can see here I'm running two cloud code instances side by side. On the left we have smartass opus 5 and on the right we have senior opus 5. Right now these are both in smartass mode. This is just the default system prompt and the default clawed opus 5 model with no tweaks, no changes, no prompt engineering to adjust and improve the system. As you can see all of the ticks and problems and verbosity of this model are coming alive right now.

¶13Zuck's blog is very very long, but you can see here these models spin for a long time. The one on the left is really really going off here with a lot of information. Okay, so 53 seconds, 35 seconds. Most of that was just output text. Look at how many output tokens my agent just torched responding to me.

¶14Most of us truly were not reading this whole thing. Just like we're not reading Zuck's blog, we need concise, compressed information so we can move on or decide to invest more. A lot of engineering with agents is about figuring out where to spend your time. So let's improve the system by prompt engineering the system prompt. So in VS code I have senior opus 5 system prompt and this is getting passed directly into that herder pane.

¶15You can see here we have one version that's just claude code opus out of the box and then we have the other where we're pending the system prompt file using this argument flag. So that's what we're going to do here. And so if we open up this file right now we have nothing in there. So we just got the default response out of it. Let's improve it.

¶16Let's write a great system prompt to steer the behavior of every single run, every single input and every single output from our model. What we're building here is a clear concise communication document. I am typing this by hand. When I'm working with pieces of autonomous software that's going to be multiplied over many, many, many runs, I do it by hand. If you don't understand your technology, if you're just vibes slopping everything, you're not going to get the results who someone paying attention is going to get.

¶17We're going to start really simple here. We're appending this to the system prompt. Clear, concise, actionable communication. And then we're going to talk through the purpose. First things first, you and I maintain a no BS, clear, concise, actionable relationship.

¶18Notice how I'm talking to my agent. I'm not giving it a role. I'm just talking to my agent as almost if I'm talking to like another engineer, okay? And and kind of setting the standards. We maintain a no BS, clear, concise, actionable relationship.

¶19Every word we say together reinforces our clear, concise, actionable communication. I'll jump through some of this. We're here to solve problems and create value. And our communication reflects that. And then I'm going to start the next section, instructions.

¶20And this is going to be where we dial into several ideas. We'll have four sections and then one kind of recap section. Um, this will be the examples and then we'll work through these sections one by one. But I'm writing this out to be able to reference these sections so that I can clearly communicate inside the system prompt what I want my agent to focus on. Okay?

¶21So, pay close attention to the details throughout instructions to maintain our great communication patterns. And then I'm going to do something here that I think is really important when you're prompt engineering. I'm going to explain why. Why? So we can deliver the best possible results for our team, business, and customers.

¶22Just walking through things here really, really cleanly, really concisely with my agent. Let's pause for a second here. Let me just actually delete this and just kick this off again. So let me save that text there. Let's go back into our terminal here.

¶23Let's close this down. Let's fire off our J compare uh with purpose. So it's going to just rerun that exact same thing. Already we have one kind of clear change in our system. And you can see here at the top we do have that reference append system prompt.

¶24We're attaching that. But already we have made a small change in our purpose. This is a good start. As you can see here, it's not going to change a lot of the ticks and the problems inside the model. As you can see, there's loadbearing right there.

¶25That's one of the classic ones with Opus 5. It's using a lot of dashes, but it is communicating a little bit better. You can see that one came in at 35 seconds. Not a huge difference at all from previous runs, but you can see here if we looked through these, we would probably see some glimpses of clear communication, but we need more than this. This is all very high level.

¶26Let's get more midlevel, more lowlevel with our system prompt. So, I'm going to copy this in. And now we have instructions. Okay, so purpose, instructions. Let's start giving some real [music] clear communication patterns.

¶27Now we get into like concrete prompt engineering patterns that are repeatable across all of your work. For instance, here's one clear pattern. Positive patterns and negative patterns. And so what does this section do? Replicate the and I'm going to just write out a H4 here.

¶28Positive patterns. And of course we'll have that and then we'll have negative patterns. And so we're just going to walk through this. And you know, as you can imagine, I'm going to clearly reference this section as behavioral references. So, I'm saying replicate the positive patterns.

¶29And then I'm going to, of course, say avoid the negative patterns. So, we want to replicate these and we want to ignore and not do these. Okay? So, we're doing both positive, do this, and negative. Don't do this.

¶30And now, we can just start writing out these bullets. What do you want your agent to do? What do you want them to not do? I'm just going to start from the top. And this is one of the important things that is missed a lot.

¶31So, I'll say I always see the last thing you write first. Place the most important information there. As you're working with these agents, you have likely noticed this, too. You see the last thing. So, it's important that the last thing you see is the most important.

¶32So, you can, if you want to, skip all the other stuff. What else do we want the agent to keep doing? Positive patterns. Use plain specific language. State each fact once.

¶33There's no need to repeat anything. Match the level of detail to the level of task and request. Okay, so this keeps the agent aligned with the amount of detail and effort you're putting in. Challenge incorrect assumptions directly and explain why. We do not want any sickle fancy here in our responses.

¶34We want a useful engineering partner. Optimize for clarity and engineering value, not quotability. I don't want to have a good conversation. I want to deliver engineering value with clear, concise communication patterns. Okay.

¶35One more note here on domain terminology. Use the simplest domain terminology that compresses information. So just some positive patterns here. You want to write what you want to see. And then of course the negative patterns.

¶36You know exactly what we want to add to this. Avoid words and phrases in this list. Okay. And then here we can just go crazy with all the obnoxious things that this model writes. We can do things like loadbearing.

¶37That's going to be at the top here. We can do things like worth stating plainly. Here's the honest truth, the [snorts] real tension. [laughter] And then there's carry the argument. There are a bunch of these.

¶38Put whichever ones here you want. A couple things I don't like. Personally, I don't like avoid analogies. Discuss what's right in front of us. Not a huge fan of analogies when working.

¶39I just want to, you know, focus on the thing right in our face. Of course, do not overuse em dashes or dash chaining. I think some of these are fine. My writing style personally has a decent amount of EM dashes. I have had to personally dial these back because I don't want people to think I'm just spamming out using models when I'm replying.

¶40I think it's really rude if it looks like you're just responding with AI with zero thought in specific circumstances. Do not flatter, praise, validate, or agree without reason. Again, no sick fancy. Just kind of re-emphasizing the idea. Uh, do not use decorative headings, emoji, or motivative language.

¶41Again, I just want this model to talk like a cracked senior engineer. Let's just solve the problem at hand. Avoid semicolons, fragments, non-standard punctuation. I just want a normal conversation. This is one really important section in your prompt engineering.

¶42You can use this when you're building plans. For the system prompt specifically, this has a lot of weight to it. We're explicitly talking with our model saying, "Replicate these things. Avoid these patterns." So, how else can we prompt engineer our system prompt for consistent results across every single execution? Before we do that, let's actually run this.

¶43Let's see how this is developing. Okay, so I'm going to go ahead and just flip this, place this here, go ahead and just rerun, trash this session, and let's run it again. So, now we have pause negative. All right, so let's rerun a new pair of cloud code again. On the left we have our smartass opus.

¶44No changes. And on the right we're going to have our smart senior concise engineering opus. Okay. So let's see how it does. Now right away you can see something very very cool.

¶45Very few dashes here. We do have a couple dashes showing up. I don't see loadbearing anywhere [laughter] which is good. I'll do a quick search on that in a second. But there's a bunch of these obnoxious keywords that are not getting referenced.

¶4631 second execution. So we are speeding things up a bit. And by speed I mean it's outputting fewer tokens. So, we still have some dashes. We didn't say never use them.

¶47We said use fewer. The language is looking a bit clearer. We have clear headers. You know, no analogies. We do have a dash chain here, private by default.

¶48Maybe you're okay with that. Maybe you don't want to see that. That's up to you. You can prompt engineer that however you like. You can see here we're making some improvement.

¶49And we're starting to save on the output tokens. Even in your subscription, your output tokens still churn up the most of your usage. So, it's important to keep that in mind. We can do this at the system prompt level. Why?

¶50because we want to affect every single input and output from the agent. Your prompts in and its responses out. If you have something that you want to apply globally, that is the time to reach for the system prompt. So, let's keep moving. How else can we improve our communication patterns with our agents?

¶51So, number two, this is a really powerful one that I'm a huge fan of and I think you will be too. Check this out. We can use reference points. What is this? we use.

¶52And again, I'm talking to my agent, right? I'm talking directly to them. I'm not giving them a weird funky role. I'm just talking to them. These models are getting smarter.

¶53And even the lower class workhorse models, the A and B tier models that aren't your S tier state-of-the-art models, they respond really, really well to this additional direction. We use reference points to communicate quickly with each other. And now, what do I mean by that? What is a reference point? Use numbered lists and markdown headings when they improve navigation when presenting three or more findings.

¶54And this is stuff like uh decisions, options, risks, questions, or actions. Assign everyone a short code. Okay, so what does that mean? So use D1 for decisions. And what I'm really looking for here is D1, D2, and then DN.

¶55You can imagine this going on, right? So I'll do this and do a dot dot dot. And I'll skip over this in the video so you don't have to wait for me to type all this out. And then we do dot dot dot for continuation. When we don't have something here, I want it to invent new references for sections we don't have.

¶56Preserve same codes throughout the conversation. Do not create codes for short, simple answers. This is really cool. So you can see this in action right away, right? Let's go ahead and fire off our system and see how the system prompt materially impacts every single prompt.

¶57Because again, that's the scale of what we're dealing with here. The system prompt is the law for your agents. It affects every single task. Let's delete the previous session. Bam.

¶58And let's rerun. This is going to be our ref points. So now we have that new workspace. We're comparing side by side. And of course, smartass on the left and we have our senior opus slowly improving with each change we make.

¶59So you can see here we have the P's, we have the Rs coming up. And so we have risks there. And now our agent is even referencing one of the Rs R six. Okay, so now we can just jump to the reference. Our agent isn't repeating itself.

¶60It's not wasting time in tokens and it's repeating the risks here. And then if we scroll back up to the PS, I assume these are the promises that Meta says it's going to actually make. What is Meta saying it's going to ship? It doesn't really matter what the P stands for. The fact is that we can reference them now instantly.

¶61Okay, so for instance, like I can continue the prompt here, right? So, if I said talk more about R6, and of course, my agent knows exactly what I'm talking about when I say R six because we've created this quick language together. This is risk number six in this section, existential RSI, self-improvement risk that Zuckerberg is talking about here. And our agent is just going to kind of continue breaking things down here, right? What R six actually says, the section claims and sequence.

¶62There it goes. And it's going to break down that section more. We can jump into it. And you can see here we have another reference here. We have FS.

¶63These are findings. That's exactly what we've encoded here. Fs for findings. And so our agent is just using our reference system. And again, like the key I really want to communicate here with you for your engineering work is like um great communication is great engineering and vice versa.

¶64Knowing how to properly communicate with your technology now more than ever is a massive advantage you can use. It's so underutilized and it's so powerful. You can see here we're starting to improve things. We're adding layer by layer. Let's add some more really important [music] layers.

¶65What is the third improvement to our system? Let's go ahead and dial into this. Hard operational boundaries. Okay, one of the very annoying things with some of these state-of-the-art models is that they will do things that they have not been requested. This is part of their reinforcement learning.

¶66This is part of the loops that they're going through when these models are being trained where they're just taught to find the answer at all costs, no matter what it takes. Opus 5 is like really guilty of this. Comment down below. Let me know if you've experienced this. Opus 5 will just find and call out and reference problems you did not even ask for remotely and it's trying to pull everything together.

¶67It's trying to do as much as possible. But in that um it actually loses focus pretty quickly. So hard operational boundaries solve that problem. Let's write that out. So in addition to clearly communicating, it's important that we clearly communicate our work operational boundaries.

¶68What do I mean by that? So deliver only what was requested at the intended scope. That single sentence really is this right? But we can add more detail. Do not widen work into cleanup, refactoring, documentation or any adjacent features.

¶69Just stay focused. That's what I'm saying here. Stay focused. Do not speculate on abstractions for future requirements. Do not claim completion without evidence.

¶70Very, very important. These models are really good at that. Now, this is not a great line to add, but going to add it anyway. And then here's another one. So annoying when these models do this.

¶71Never add a co-author to a commit message. Like nice marketing trick for Enthropic, but they got to stop doing that for completed work. Concisely restate, but do not overload with response detail. There's a better way to say that, but you kind of get what I'm saying. Opus 5 specifically will respond with everything it's done and kind of like do a big recap.

¶72Don't really want that. And let's go ahead and stack another one on here. This is a really powerful one that I think is very useful and it really starts veering into the fact that the system prompt can help you control your inputs and your outputs. This one is aliases. So, let me break this down.

¶73Aliases are reminders of great communication and patterns we want to uphold. When you see these exact aliases, expand them and act as if their expansions were given to you directly. And then we want to clear this up to make sure we're being very clear. If these are referenced in a longer string, they are not aliases. Do not expand.

¶74Okay, so these are exactly what you think they are. They're short codes. They're bash aliases. They're expansions. They're commands inside your system prompt.

¶75And if you want to, these can reference your commands, your skills, and other things you have. But you can also just write them in line like this. STR equals simplify, compress, and repeat your response. Let's write a couple more and then we'll demo these. Okay, so Eli, you already know what this is going to be.

¶76Explain this like I'm We don't want to do five. Explain this like I'm 18. Simplify your language. Shorten your response. What else?

¶77Uh, one of my favorites. Focus on what matters most here. What's the true signal? What's the true value? Boil your response down into the most important thing we need to focus on.

¶78And then ref. Let's just do one more here to reference our previous work. Cuz once you start building up your system prompt, you can reference other sections. So I'm going to say I'll rewrite your responses with reference points, which is exactly what we dialed in above. So now let's try our improved enhanced system prompt.

¶79Same deal. Open the terminal. kill this session with our aliases. So, we're going to compare, boot these up, and now our agent has a little bit more guidance into exactly what it's going to do. Okay, so you can see there we have C123.

¶80That's our core thesis. We have our risks addressed. This is a much better generation really using all of our reference points. Very nice. And done in 30 seconds, done in 30 seconds on the left.

¶81So, you know, sometimes the generations are just fast. These are nondeterministic systems, so they're not always going to be generated faster. And in fact, we can improve that, which we'll do in a second here. But remember, we just added those aliases. So, let's run scr simplify, compress, and repeat.

¶82This is effectively a micro scale you're giving your agent. So, there we go. It's compressing that. It's simplifying it, and it's repeating it. So much more simplified.

¶83I can act on this information a lot faster. Even if our agent responds with a lot, which again we'll tweak in a second, we can improve on that. So, let's run another right focus. This is focus on what matters the most here. here.

¶84So our agent sees that and it's going to do it. Now this should be a lot more concise. There we go. The signal alignment is being redefined from the model holds the lab's value to the agent holds yours. Everything else is downstream of that.

¶85This is a speaking pattern from Opus that you might want to get rid of. Totally up to you. We have why it matters and then pulling some memory stuff here. And then we have our final paragraph there. Right.

¶86A couple things I want to do here. And once you like set up this foundation inside of your system prompt, you can really start to control it and fine-tune it. These patterns and these sections are really important because we can now collapse and we can jump into anyone we need to very very quickly by hand. We are controlling the experience of our agent. Again, I really want to stress this point.

¶87The more multiplicative the thing you're working on is going to be for the rest of your work, the more you should stop, slow down. Don't just throw a prompt at this. Even if it's a state-of-the-art model, as you can see, the state-of-the-art models have problems. They might accomplish your goal, but they'll add or do a bunch of other We're being really, really clear here. We're doing this by hand.

¶88We're slowing down. I know a lot of people are using Whisper Flow. They're just talking in to their device. There is going to be a level of detail that is missed, and you're going to be very exposed and very leaning on the normal distribution of what the models can do, and you're going to be inside that zone. If you're just viro prompting things, there's a time and place for that, but the high leverage places is when you want to slow down and step in to the loop.

¶89And I mean old school, hands-on. So anyway, let's look at this. So I want to do some positive. Where do I want to put this? This is definitely going to be positive.

¶90And I want to say if you can communicate the idea in one paragraph instead of two, do so. And I want to say without losing valuable information, do so. Same idea for one sentence versus two sentences. So I'm just saying be a little bit more compact, be a little bit more concise. This is important.

¶91Do not repeat yourself. State every idea once. Only repeat if it's relevant to subsequent queries. Yeah, something like that. Because we really don't want repetition.

¶92And then I want to add another positive pattern here. Don't use overloaded terms that could mean more than one thing. Use the simplest word that satisfies the idea you're trying to communicate or I should say simplest words and maybe PN that. You might be like, "Wow, you're getting really detailed here." And yeah, that's exactly right. I am getting really detailed here.

¶93And we can see the results of that. So, let's boot this up again. And this is going to be uh fine-tuned. And now let's go ahead and see what this execution looks like. And to be clear here, we are using append system prompt, not overwrite system prompt.

¶94But let's see how uh this execution runs now. Fewer dashes, no loadbearing. We have our reference points policy ask again with reference points. And our sentences are simpler. They're more concise.

¶95They're not getting overloaded where the argument is the weakest. Still outputting a lot here. This is a pretty big one. 35 seconds. So, not great, but it is going to be better than this generation on the left.

¶96Again, these are nondeterministic systems, so it's not going to be consistent. 43 versus 35. Not bad. We could probably do better in our compression. If this happened though, in this specific case, we can just throw an STR at this to get a simplify, compress, repeat, or we can do an ELI.

¶97Throw an Eli at it. explain this like I'm 18 and there you can see the language is really going to get boiled down and sometimes this is just useful right it doesn't mean you're an idiot if you want simpler concise language it means you're effective but you can see here even still the agent is saying a bunch of stuff so I might run a str after this simplify compress and repeat so as we start stacking up the user prompts in between every execution of a user prompt the system prompt is still going to be ever present so you can see argument really broken down here AI is safest when everyone has it not when a Few model labs control it. Other labs say AI is dangerous. Lock it down. He says when just a few people have it with enormous power, that's never gone well.

¶98One safe AI fails because people want different things. Any single AI has to pick whose values win. He's got a lot of really great arguments in here. Highly recommend you read and then compress if you need to uh the kind of core ideas coming out of Zuck's argument here. With these like uh CEO posts, it's always like understand the incentives.

¶99when you're not leading the AI race, of course, you would have a narrative something like this that you would want to share. I agree with him, frankly, on a lot of this stuff, but it is very, let's say, positional of him to put out something like this. That's not the focus of this video, though. I'll link that in the description for you if you haven't seen that yet. One more thing I want to show you that can really, really dial in the performance of your models, and it's a prompt engineering technique as old as time.

¶100Again, if someone tells you prompt engineering is dead, don't listen to them. The PI coding agent has a small system prompt. Cloud Code just got rid of a lot of their system prompt, that is not the signal to not use a system prompt. That is the wrong signal. They're doing that because the models can just do a bunch of stuff without needing it.

¶101If you want the model to do specific things, to perform in specific ways, to not torch your output tokens, to have cool unique capability like aliases or reference points or hard operational boundaries or patterns that you want to see and words and concepts you don't want to see, you must system prompt engineer. The system prompt is the most effective place to prompt engineer because it's applied across everyone of your prompts going into your agent and the responses coming out of your agents. Okay, so one more section here for you. Check this out. Examples to kind of put it all together.

¶102Here are concrete examples of how we do and do not communicate together. Replicate how we do communicate and avoid how we do not communicate. a lot like our positive and negative prompting section except the main differentiator here is we're giving real examples. So, it's just like training data. User is legacy JSON still referenced.

¶103To-do, no, the only match is the file itself. There are no imports, blah blah blah blah blah or documentation links. Very good. Um, we could also, you know, make this tighter. The only match of the file is itself.

¶104We could make this the to-do. And then here's what not to do. Great question. I will research the repository and blah blah blah blah blah blah. Just do the work.

¶105Respond concisely. Right. And then here's another one. Engineering recommendation. And it looks like I'm missing my user prompt here.

¶106So, I'm just going to add that here. User. Should we add reddus to this system? Write something just like this to-do. Do not add reddus here.

¶107This process has one writer, stores from SQL, and has no cross host coordination requirement. Reddus adds a failure domain without solving the correct constraint, the current constraint. What not to do? You are absolutely right. Blah blah blah blah blah.

¶108Again, just a really simple prompt engineering technique that is older than time at this point. Um, I was writing these examples back when GPT 3.5 was the model, GPT4, when there was really no anthropic. This is still relevant today. And that's how you know something is a valuable skill to learn when years ago it was still relevant. These concepts, these patterns, these principles, great engineering is great communicating.

¶109These things just kind of don't go away. That's how you know you have something valuable to hold on to. Now, we can continue doing this. And kind of one big idea, one Easter egg for whoever's still watching, engineers that are really trying to absorb this value. One thing I want to mention here is that you can boot up a model.

¶110Let's say you want to boot up a cloud fable model and you want to run that same prompt that we're running. So let me just go ahead and pull that explain this. Copy paste. And so Fable 5 I think is a great model. I think it responds well.

¶111It doesn't have as many ticks and problems and repeated patterns and verbosity as Opus 5 has. So what I like to do is in context distillation. In context distillation is exactly what you think. That's what examples are, right? It's in context distillation.

¶112Do this, don't do that. And you can do this by pulling out specific context, specific responses from models that you like. For instance, like this oneline core thesis here. Maybe we like this. We just copy this out and we create another one here.

¶113Summarizing a blog. And then user is summarize the blog. and then XYZ. We just template this out. And then we have our to-do text.

¶114And then we can paste the response from the model that we like. And we can go even further. We can tweak the response from the model that we like. Get rid of a couple of lines. Make sure it's formatted properly.

¶115So we can do a dash search to get rid of all of our dashes. All of our framing. Maybe we like that one. We keep that one in. Super intelligence.

¶116Super intelligence. Super intelligence. This should be more concise and compact like this. So maybe we like this version. Then we say not to do text.

¶117Text. We hop back over to the previous version, right? We go to our smartass opus and we copy all this. This is not what we want to see. So, just a bunch of extra information and we want to see a concise, simple response.

¶118Really just get to it. And really, I would want to see something like this, a little bit more broken up, easier to read. Even Fable has some long kind of run-on sentences here. Might break this up a little bit more into a couple section. And then we would fire this off.

¶119We have a new version here. Let's go ahead and run this. It's super correct senior opus. And then we can run our comparison here. So now our agent has examples of how we like to communicate.

¶120It has reference points. We have cleaned up the language. We taught it to respond concisely with fewer em dashes. It's not loadbearing anymore. We have reference points.

¶121This one does not have reference points. So thankfully we've encoded a way to add those. So we could just type ref. And now our short aliases are going to expand the response. There you go.

¶122You can see we now have references. You'll notice this response from Smartass took 41 seconds. Our original response over here took just 22 seconds. So, we are saving our output tokens. And again, if you need to spend output tokens, spend them.

¶123But in this case, you know, when we're uh summarizing things where we're just getting concise responses, we don't want all of that. You know, you can see here we have a nice response. We use our ref alias, which we have encoded into the system prompt. If we just search for ref equals, you can see exactly what this expands into. You could point this to a scale.

¶124You could point this to a command. Whatever you want to do here, you can do. My whole point here is to communicate with you that if you really want to control your agent and get the best results, you must know how to use your tool. A lot of engineers are overleveraged on the user prompt and you get literally less leverage. You get less output by fixating on the user prompt.

¶125Opus 5 and all models coming out, they are going to be smart asses. Quite literally smart. What we want to do is keep the smart but drop the ass. [laughter] With system prompts, you can do that across every user prompt you write. You can use these exact prompt engineering patterns to guide Opus 5 and whichever state of the art models come next.

¶126If we don't like this response, we can shoot up another alias at it. Eli, explain it like I'm 18 or whatever age you want to set your uh response reasoning level to. You can add an Eli 5, you can add an Eli 10, whatever you want to do. Great engineering is about great communicating, right? great communication with your team of engineers, with your agents, and most importantly with yourself.

¶127And this is how you can tap into that. You've noticed this if you're using agents on a daily basis for many tasks. You and I, the developer, are the bottleneck. It's not the model. It's not the tools.

¶128It's not anything else. The hard part now is communicating quickly, concisely, and in the most value accreative way. I'll leave this cracked senior engineer prompt linked in the description for you. It's super super simple. You can do the comparison side by side and see yourself improving every time you add something by running the just compare command.

¶129I'll also add my classic readme and just install and all the quick start so that you don't have to waste your time doing any setup and configuration. I want this to be as simple and quick and easy for you to use. On the channel, we've been focused on much higher levels of abstraction on top of agents like software factories, agent sandboxes. Uh check out last week's video if you want to dive into some really advanced agentic engineering concepts. I'll link that in the description as well.

¶130But every once in a while, it's going to be really important for us to step back into the fundamentals. You don't want to miss the atomic units, the primitives that make up the larger powerful composition pieces because if you stack a bunch of agents together and you burn a bunch of tokens doing in half the time by just putting together the right system prompt, you want to be doing that instead. Don't waste time using really, really clever solutions when there's a much simpler solution. And engineering is about getting the job done, not looking cool, not doing what everyone else is [music] doing, not listening to the idiots that say prompt engineering is dead. That returns false.

¶131Your agent is just another tool. If you understand your tools, [music] you'll understand the results you can get from your tools. Use your system prompt as the law for every [music] task you send to your agents and use your user prompt for individual tasks. If you got value out of this video and you made it to the end, like, subscribe, drop a comment, and let me [music] know how you're leveraging your system prompt to get outsized results and to clean up some of the insane verbosity from models like Opus 5. You know where to find me every single Monday.

¶132Stay focused and keep building.