Ermir Beqiraj
Backend architect. Systems, agents, infrastructure — from inside the work.
all writing

I keep seeing setups where agents brute-force their way through a problem for days, a dozen of them in parallel with an orchestrator on top. It reminds me of people with the crazy desktop setups, who spend more time tuning the shadows than actually using the desktop. I’ve watched some of the same people run agents around the clock for a week and get something that would have come out with higher quality if done linearly in five hours, one person sitting, thinking, and using one agent.

So whenever I try to figure out what actually changes the quality of the outcome, I end up with the same short list of boring things, and most of them have nothing to do with parallel agents, loops or graphs. This is that list, the things I’d set up in any repo before reaching for anything fancier.

Make it feel at home

Think about a new colleague on their first day. They don’t know which tools are installed, which projects exist or how they hang together, what the architecture looks like or where it’s heading. That’s your agent, every single time you start it up.

The first fix is a global instructions file, the one your tool loads on every session no matter which repo you’re in (for Claude Code it’s ~/.claude/CLAUDE.md). It gets loaded every time, so keep it short and plain, and let it point at things instead of explaining them:

All available CLI tools are listed in ~/tools/ENVIRONMENT.md

That list is the first real jump, because agents are really good with a CLI. My favorites are git, gh, az, aws, terraform, docker, sops, fzf, etc. I use twenty-something of them every day, and they’re what makes or breaks real agentic work. A small PowerShell script checks each one is installed and regenerates the file with the latest probe. Once the agent knows your tools, it can reason about your environment by itself.

House rules

Agents don’t want things. They don’t care whether the product succeeds, about your promotion, or how this feature relates to something you built last year. The agent got a task and a list of what’s available, and it’ll run in loops until the conditions are met to reach the exit door. To it, “database already exists, delete it and re-run the migration” carries the same semantic alert as “unused import, delete it”, they both prevent the agent from exiting the loop and calling it complete. To you those are very different messages, so it’s on you to make sure the agent can move around without breaking what can’t be fixed.

Giving it CLI access is the moment to decide what runs unattended and what doesn’t. So git commit could go unattended, git push asks first. terraform plan can run on its own, and a soft rule in the instructions can teach it to always save the plan and apply exactly that one, but if it shouldn’t apply at all, that takes a hard rule, a permission or a hook, something it can’t reason its way past. And if what you’re doing has no real consequences, you could let it apply in a test profile and keep production out of reach.

“With great power comes…” you know the rest.

Now that it has the keys, think about credentials, and here too simple things go a long way.

  • With sops around there’s no reason to keep plaintext secrets anywhere, because sooner or later they end up in a chat log, or a poor model pastes them somewhere.
  • Split configurations by environment. A key that works in dev should not work in test.
  • Wherever the CLI lets you, remove the defaults, so the aws and az CLIs have no default profile and every command has to name one, which pushes the agent to stop and reason about which.
  • Name profiles for what they do, production and not something cryptic, and a good model will flag the guard by itself.
  • Keep production behind a login with a short-lived token, so even if the agent picks that profile it can’t get in until you log in, and you’re smarter than that.

Repository anatomy

This is the next big jump, and unlike the tools it takes real effort from your side, plain unglamorous engineering.

First, the repo’s own agents.md, claude.md, or whatever this week’s convention is, matters less than you’d think. Like the global one, it works best as a handful of pointers and the few rules you’d otherwise keep repeating. Explaining the project is a job for something else.

Think about how the agent works. It has instructions and tools, which we covered above. It gets a task, which is outside the scope of this article. Then it has to find where in the code that task lives and solve it, and that part is decided by your repo.

If your project is worth anything, it won’t fit in a markdown file anyway. Any solution that involves the agent reading thousands of lines of text or code is a poor solution. It might work for some cases, but the time and output are suboptimal and not worth the effort.

What I do is prepare a high-level direction, some of it in the form of documents, some in the form of code, and finally some mix of the two. It’s important to have content the agent can find, and it’s more important to not confuse the agent with what it just found.

High-level documentation

Vision. What this repo is for and where you see it going, short and plain. This is where you draw the boundaries, what belongs here and what stays out. The agent reads the code to learn what it does, but the code won’t tell it that availability has a strict SLA, or that it doesn’t. I’ve seen agents build complicated always-on solutions for tools that operations use from 9 to 5. Most products serve operations, and just knowing who your consumer is lets you simplify a great deal.

Architecture. How the repo is built, in diagrams and words: the APIs and how they’re used, the application, the abstractions and the domains. Whatever pattern you’re using, if it has a name (and I hope it does), name it. The model was trained on these patterns and knows them by heart, so the name alone connects your codebase to everything it already knows, and it makes better decisions.

Whatever breaks for your case. I left this one as a blank cheque, because no two projects are the same. If you keep seeing the agent fall short or waste cycles on the same thing, that’s the document to write.

Keep these documents short and to the point. Their job is to be grepped and understood before the agent starts implementing.

Give it a quick way to prove small units of work

The pitfall I see most often is made in the name of simplicity. It’s a small project, so you keep it light. Then small requirements creep in, you put out the little fires one by one, and one day you look up and you’ve built something ugly, or a poor version of a pattern that already exists.

My advice is to stick with what your brain already knows, because most repo architectures are just separation of concerns. Plain old software engineering, nothing fancy, but it’s what lets the agent understand where something goes, and once it knows where, it knows how to prove it. A domain change gets a unit test. Anything that talks to the outside world, a database, a queue, another service, sits behind an abstraction, so in tests it gets swapped for a fake and the logic around it is tested on its own.

The agent can’t make that call for you, remember it thinks about the task, not your project. Once you’ve made it though, every change comes with an obvious way to prove it. The agent knows a domain change travels with its unit test, so it starts from a green suite and its job is to leave it green, which turns “does this work” into something it can check by itself instead of something it has to tell you.

Give it an easy way to verify bigger chunks of work

Unit tests prove the pieces, but plenty breaks between them, and the agent needs a way to run the whole thing without you in the middle.

So first, take away the blockers, anything that needs a human to operate. Authentication is the usual one, because the agent can’t log in for you. Bypass it locally, and keep the bypass in local config only, which is where splitting configurations by environment pays off.

Then prepare a docker compose for the whole stack. This is the huge advantage, because agents often get caught on the little things between the pieces. The tests pass, but was the service actually registered? The domain logic is right, but a config key is missing, or the response serializes into a shape the consumer doesn’t expect. None of that shows up in a unit test, and all of it shows up the moment a real request goes through.

And give it a way to send that request. Leave this out and the agent improvises, throwing together curl calls and one-off scripts just to get out of the loop. What I found helpful is a small client app in a scriptable language, usually a Python uv directory, that calls the endpoints the same way the consumer would.

And last but not least, speed. Left alone, the agent rebuilds the container image after every change, and that eats a huge amount of time. Instead, mount the build output into the containers, so applying a change is build and restart, and the time from fix to test comes down to whatever the build takes. Write that loop into the instructions file, or the agent goes back to rebuilding.

When to call it done

Pick a small, real task and watch where the agent stops. It asks how to run the service, improvises a curl script, or says it works without running anything. Each stop is a missing piece of the setup, so fix it and try the next task.

It’s done when the agent carries a task to the end and comes back with the proof, with nobody in the middle. From there skills, subagents and the rest start to make sense, because you can finally see where they’d help.

Ermir Beqiraj is a backend architect building AI-integrated infrastructure. This is his personal writing.