
Slop with Green Tests
There is a great deal of anxiety spreading through the tech industry. AI models are improving far too quickly, tasks that looked difficult six months ago are now trivial, and nobody can say with much confidence what the role of a software developer will look like three or five years from now.
My perspective comes from a slightly unusual position because I hire developers, but I still generate code myself. When I think about the kind of developer I would hire today, I do not first think about someone who can produce a certain amount of code per day. I think about someone who can take ownership of a mess, understand the problem, make good decisions and make sure what eventually reaches the customer is within an acceptable standard.
Code got cheap. Attention did not.
In the pre-AI world, generating and fixing code were fairly finite variables. When a feature took three weeks to build, it made sense to invest a lot of time in planning, patterns and abstractions to avoid having to build it twice. Bad code could easily mean days of human rework or a 3 a.m. incident consuming even more of that scarce resource.
That equation has changed brutally. We have never generated, tested, refactored and discarded this much code, and tokens that felt expensive not long ago are getting cheaper by the day. The “clankers”, as they are affectionately called, can produce in minutes an amount of software that once required weeks of human work.
Our capacity to pay attention has not followed the same curve. We still have eight working hours, the same brain switching between product, architecture, users and problems, whilst every agent can come back at any moment asking for yet another decision. If there is one resource that has become relatively more expensive amid all this abundance, it is human attention.
Chemical X
That is where my difficulty with buying into dark factories as some inevitable destination for software development comes from. And it is worth being clear about what I am criticising here: it is not aggressive automation. The more mechanical work I can get off my plate, the better. The problem starts when automating execution also becomes an attempt to remove human technical judgement from the process.
The proposition is seductive: agents write code, other agents review it, tests validate the result, and yet another layer of agents coordinates the whole thing until the human can finally leave the factory. The problem is that we are taking the optimisation of a stage that has already stopped being the hard part to its logical extreme.
It is curious how many of the most impressive examples of dark factories are tools for creating more agents, coordinating agents or making the production process itself more efficient. We have built extraordinary machines to produce other machines that produce software, which is technically interesting but says little about whether we are getting better at solving real problems. Even when an agent reaches a functional solution, it can still take an unnecessarily long route to get there.
It might create five abstractions where two would do, choose an architecture because it is popular, consume far more infrastructure than necessary, or leave behind a codebase that requires even more context for the next agent. Everything works, the tests pass and the dashboard is green, whilst someone experienced looks at the whole thing and wonders why on earth half of it exists. “It works” is a surprisingly small property of good software.

Chemical X is human judgement applied in the right places. I do not want someone manually inspecting every line an agent writes, because that would simply recreate the bottleneck we just removed. I want someone deciding whether the thing should exist at all, whether there is a simpler route, which decisions will be painful to undo and whether what we are building actually makes someone’s life better.
That is why an agent in the hands of an amateur starts to resemble a Labrador holding a slop grenade launcher. When the result is simply lobbed at the next person without anybody really understanding it, the apparent productivity of whoever produced it only existed because the cost of validation was pushed onto somebody else.

Delegating is not dumping
I have seen quite a few people call something delegation when it is much closer to dumping the work and walking away. Delegation requires context, an objective, constraints and some criterion by which you would recognise a good result, even if another entity performs all the mechanical work. Dumping it means throwing the task at the agent and turning up again when it says it is finished.
There is another, perhaps worse, variation: the infamous meat proxy. The human is technically still in the loop but has basically become an intermediary: copy the error into the agent, execute whatever command it asks for, paste the result back, accept the next change and repeat until something turns green. If you can no longer explain why a particular change was made, your presence in the process did not add the human judgement that justified having a person there in the first place.
You do not need to stare at the terminal whilst the model works or obsessively inspect every function, but you do need to understand the route it chose, know where the important decisions are and have enough of a mental model to notice when the agent has taken a bad turn.

That control matters even more as we give agents more autonomy. Otherwise we outsource decisions without even noticing that decisions were being made and end up as an unusually expensive USB cable between the agent and production.
The same confusion appears when someone opens fifteen agents and concludes they are now fifteen times more productive because fifteen terminals are busy. The interesting number is how many of those jobs received enough context, decision-making and validation to arrive at a good result, and that is where we hit another limit that did not scale with token counts.
If AI writes the code, why bother learning programming?
Once a large part of execution can be delegated, it is natural to ask why anyone should continue studying programming in depth. If I can ask for an entire migration, an algorithm or a complete API and receive something functional a few minutes later, spending hours learning how any of it works can start to look like a peculiar form of nostalgia.
The problem is that programming should never have been just about memorising syntax. You learn concurrency so you can reason about what happens when two things execute at once, study databases so you understand consistency and modelling, and learn networking so you can follow what is actually happening between two machines. That knowledge builds the mental model you use to question decisions that may no longer have been made by you.

The recent discussion involving David Heinemeier Hansson (Omarchy) and Arian van Putten (NixOS) is a good example. Even an experienced person can look at a result and attribute the improvement simply to the fact that the new implementation uses assembly; a more careful reading of what actually changed finds data organisation, fewer system calls and other design decisions along the way. The number is the same, but understanding where it came from is what turns the result into reusable knowledge.
There is a huge difference between “the agent made this faster” and “I understand why this became faster”.
This is also why I like the comparison with traditional engineering. Calculators did not make engineers stop learning mathematics, just as CAD did not make the fundamentals of design irrelevant. Those tools gradually removed mechanical work from the process and allowed professionals to operate at higher levels of abstraction, whilst responsibility for the result remained with the person signing off on it.

The standard proposed by Mitchell Hashimoto, co-creator of Terraform, sounds reasonable to me for this new world. You might not be able to reconstruct every line of a system from memory, but you should be able to explain how it works, why the major decisions were made and where you expect it to break. The agent can write the for; you still need to know why there is a for there.
And as you get better at delegating that execution, another problem appears: your ability to produce starts growing much faster than your ability to keep track of everything being produced.
Agents increase output, not your cognitive capacity
Having several agents working at once creates a curious feeling of unlimited capacity. One finishes a task whilst another asks for a decision, the third finds an ambiguity and a fourth has just completed a change that now needs validation. A few minutes of that and you have already switched between several complex problems whilst making dozens of small decisions that would previously have been spread across an entire day.
The most immediate effect of that pace is cognitive overload and decision fatigue. Keep it up for long enough and you are also creating rather fertile ground for burnout, because the amount of mechanical work falls whilst the density of decisions rises. You can finish the day having typed almost nothing and still feel mentally wrecked.
My personal rule has been to limit myself to three projects requiring heavy decisions at the same time. When an agent is working, I try to use that space to review the plan, rethink pending tasks, look at the user experience or simply let my brain recalibrate instead of opening another seven fronts because a terminal became available. Pomodoro and deliberate breaks have become even more useful to me in this setup.
From I-shaped to T- and M-shaped
There is a much more interesting upside to all this extra capacity: it has become far cheaper to be nosy. For decades, career discussions moved from the I-shaped professional, deeply specialised in one area, to the T-shaped one, who keeps that depth whilst becoming capable across adjacent disciplines. AI has dramatically lowered the cost of exploring those additional arms.
A backend developer does not need to become a designer to develop some decent product sense, just as someone focused on interfaces can learn enough infrastructure to participate in an architectural discussion without changing careers. The goal is not to turn everybody into a shallow generalist either. You can remain deep in what you are good at whilst becoming reasonably capable in areas that previously sat completely outside your reach.
I have always been the “walking job-description violation”, so I have been taking advantage of this to poke around networking, CAD, electronics and 3D printing whenever I have some spare time. Not because I plan to become an expert in all of them, but because the cost of crossing the first barrier into a new field has dropped dramatically.

That also changes the vantage point from which a developer sees the product. Once you are no longer confined to implementation, you start moving through the places where technical and product decisions meet, which in turn means learning to review code in a way that is compatible with the sheer volume we can now produce.
Review by risk, not by line count
Reviewing every line with the same level of scrutiny simply does not scale when an agent can produce thousands of them in an afternoon. The attention you spend should reflect how much damage that part of the system can cause and how expensive a bad decision would be to undo later.
I like to start with three layers:
- Contract: what goes in and what comes out. If this is wrong, you have probably built the wrong thing.
- Flow: how information moves through the system. If this is wrong, several components that are individually correct can combine into an incorrect system.
- Implementation: how each part concretely carries out what the contract and flow require.
If you cannot draw the main flow on a whiteboard and explain how information gets from one end to the other, you have not finished the review.
After that first pass, I think about the codebase as a tree:
- Trunk: schemas, business rules, value objects, authentication, authorisation, public contracts and other decisions that hold a great deal together and are expensive to replace.
- Branches: trivial integrations, glue code, helper functions and pieces an agent could simply rebuild tomorrow without putting the whole project at risk.
My current rule of thumb is roughly 50% of my attention on contracts and interfaces, 40% on the core and decisions that are expensive to reverse, and 10% on low-risk commodity code. There is no science behind those numbers; it is simply a way to stop myself spending twenty minutes arguing about a helper function whilst a bad database-structure decision strolls through the front door.
I have also been running into the opposite problem: agents are absolute masters at finding absurd edge cases. More than once I have caught myself spending hours dealing with a chain of theoretically possible situations that would almost certainly never affect a real user, whilst something that would deliver actual value remained stuck. I am not suggesting that we ignore edge cases involving data loss, security failures, financial damage or state corruption; the point is that even robustness needs to be prioritised by risk, otherwise an agent can find an infinite amount of work for you to do.
The same logic applies to the permissions we give agents. If an agent can delete production data, alter infrastructure, access credentials or run destructive commands, the risk is not only in the code it writes but also in the actions it can take. Autonomy should grow alongside our ability to reverse whatever goes wrong.
Automate what does not deserve your attention
Anything a machine can reasonably verify should ideally be filtered before it reaches a human review. Linters, static analysis and tools such as CodeRabbit are useful precisely because they save brainpower on things that do not improve simply because a person spends a few minutes staring at them.
I apply the same idea to AGENTS.md, personas and skills — reusable instructions given to agents. If an architectural preference appears in every project, I do not want to spend part of every session reminding the agent about it. I put that knowledge in the environment and try to reserve my interventions for decisions that still need to be made.
I have been putting some of these workflows into ntorga/agent-starter-kit, including instructions I use to reduce this kind of repetition. I also plan to publish a guide to optimising AGENTS.md based on actual measurements, because we already have quite enough guesswork in this area.

One of the skills I use most is also one of the most irritating: Grill, a 20-line skill by Matt Pocock. Before allowing an important implementation to move forward, it starts an interview and digs into ambiguities that would otherwise be discovered after hundreds of lines had already been written. Who can perform this action? What happens to existing data? How does the system recover from a failure halfway through? After a few rounds, the urge to say “mate, just do it” appears, which is usually a good sign that we have reached the difficult bit.
The cheaper it becomes to implement a decision, the greater the relative cost of making the wrong one. A fast agent can turn a bad assumption into thousands of coherent lines before you finish your coffee, so I would rather make it stop at those points and spend attention before the assumption grows an architecture around itself.
Tests are not a golden parachute
Once we have decided what “correct” means, we need to turn at least part of it into feedback the agent can consume on its own. If a human still has to manually discover after every change that login broke, an API stopped persisting state or an operation became three times slower, we are still spending human attention on the wrong part of the process.
A good test suite acts as the agent’s feedback system. It needs to cover the happy path without stopping there, including expected failures and relevant edge cases, whilst checking observable behaviour and the final state left behind by an operation. Depending on the project’s risk, that might include API and interface tests, authorisation, concurrency, fuzzing or property-based testing; I see little value in chasing a particular number of tests if they cannot tell us when something important has broken.
Performance belongs in the same category. If latency, processing capacity or resource consumption matter to the product, I like to keep reference metrics so that a regression does not quietly become the new normal. I also split the suite into fast, standard and exhaustive levels, giving the agent quick feedback during implementation without making it wait for the entire battery after every tiny change.
Ship small, ship often
Something else that has become even more important with agents is avoiding projects that sit marinating for weeks before anything reaches the real world. If AI lets us build faster, it makes little sense to spend that speed accumulating one enormous change that touches twenty parts of the system and only then discover whether it works in production or whether anyone wanted it in the first place.
I have been favouring micro-releases that ideally take no more than about three days. Again, this is not science, just a rule of thumb: if something is taking much longer than that to produce a usable version, it is probably worth breaking the problem apart. Smaller releases reduce the blast radius, make reversions simpler and, most importantly, shorten the time between a decision and feedback from the real world. For that to work, of course, we need to observe what we just shipped: errors, metrics, user behaviour and any other signal that our hypothesis was wrong.
This is also a defence against one of the more dangerous temptations of AI. Because we can now produce a huge amount of code before getting tired, it is remarkably easy to let the scope grow with it. I would much rather discover in three days that we picked the wrong direction than discover it three weeks and fifty thousand lines later. The speed of agents should shorten our feedback loop, not increase the size of our bets.
Cheap code is not free code
Implementation may have become cheap, but every line still adds future maintenance, attack surface and context that somebody — human or agent — will need to carry when they come back six months later.
That has influenced my technical preferences as well. With an almost unlimited ability to generate implementation, I have increasingly little patience for accidental complexity that could be removed by the compiler, a decent standard library or a smaller dependency tree. On the server side and in the core of the business, that often pushes me towards Go or Rust; on the interface side, TypeScript remains a pragmatic choice where its ecosystem genuinely provides value, although I tend to prefer the lightness of Alpine.js over the lorry that the React ecosystem has become.
The same applies to problems that appear across several projects. If five services need to do exactly the same thing, I do not want five agents producing five slightly different versions of it. I would rather solve it once, validate it and turn it into a reusable piece, which is the idea behind projects such as goinfinite/tk and goinfinite/ui. Besides reducing maintenance and attack surface, it also reduces the amount of context we will need to feed agents later.
Code became cheap to create; it is still expensive to own.
A slightly funny consequence is that knowing when not to build something has become even more important. In the past, plenty of bad ideas died naturally when somebody estimated three weeks of work; today the same idea may survive because it costs an afternoon, even if nobody asked for it and somebody will have to maintain it for the next five years.
User experience is part of engineering
When practically any idea can get an implementation quickly, asking whether something is possible becomes less useful. The question that remains more often is whether it should exist in that form, which pulls user experience much closer to engineering work. That’s right, welcome to 2026, even if you hate working on interfaces.

You do not have to enjoy CSS, but you do need to recognise when a product is forcing unnecessary decisions onto the user, when it has chosen a poor default or when a five-screen flow could have been a single action. Aesthetics matter because people form an impression of quality from what they see, but a beautiful screen that forces the user to fight the product is still a bad interface.
And it does not stop in the browser. An API can be unpleasant to use, a command-line tool can demand options nobody should need to know about, and an error message can send somebody to Google just to work out what happened. A good product engineer tries to shorten the distance between what somebody wants to do and what the software requires them to do in order to get there.
This is territory we used to push towards “product”, “design” or “business”, perhaps because there was enough technical work waiting on our side of the fence. Now that a substantial amount of that work can be delegated, there is more room to cross it.
The product engineer
There is a fair amount of fear that business people armed with AI will start occupying territory that used to belong to developers. They probably already are. What interests me more is that the border has become porous in both directions, and technical people now have much better tools for moving into product, operations, user experience, technical communication and other fields that once charged a much steeper entry fee.
The stereotype of the introverted programmer who receives a specification, disappears for two weeks and comes back with code was already dying before LLMs. It is now even harder to justify the professional who only wants to know about the task dropped into their lap, particularly when talking to users, exploring an idea and building a prototype cost a fraction of what they used to. We already have the tools to cross those boundaries; often what is missing is simply the willingness.
For me, the product engineer is the person who can follow the entire loop:
problem → decision → implementation → delivery → observation → next decision
They do not need to personally execute every stage. They can delegate implementation to an agent, design to a specialist or infrastructure to another team. What does not disappear along the way is their responsibility to understand how each part affects the final result and to keep that loop moving.
Sometimes that requires a sophisticated implementation; sometimes it means deleting half of a feature that grew too large along the way; and every now and then the best technical decision is realising that the code did not need to exist at all.
For a long time we used lines of code, completed tasks and shipped features as reasonable proxies for productivity because producing software was expensive. A machine can now outperform any human by a mile on those indicators without necessarily producing a better product, so perhaps it is time to stop treating them as if they were the final result.
Your output is not code. It is an outcome.
If you do not already do this, it is worth ending the day with a few questions:
- Was the problem actually solved? Could the user do what they wanted?
- Did I keep my manager up to date on progress and provide clear estimates?
- Did the operation become cheaper or the product simpler?
- Do we have fewer incidents, better conversion, better retention or less manual work?
- Where did I mess up, and what could I have done better or faster?
- What technical or cognitive debt did I knowingly accept today and leave documented?
If the future of software development involves less programming, I do not see much tragedy in that. It definitely does not involve less engineering.
A note of caution: know your supplier
Everything above is my opinion about how I intend to work and about the kind of software I am comfortable putting into the world. Plenty of people disagree, and we will probably see more and more products launched where no technical professional truly understood the whole system — or even gave the agents’ output a proper once-over. Some of those products may work perfectly well. Perhaps that model will even become normal.
Personally, I would be rather wary of putting important data, credentials or sensitive information into software that is a complete black box even to the people who built it. The fact that a product works, has a polished interface or even has open source code does not automatically mean that somebody competent understood every decision inside it. Open source helps enormously with auditability, but millions of lines sitting in a repository that nobody responsible has reviewed are still millions of lines you are choosing to trust.
Know your supplier. Know who is behind the software, how they approach security, how they respond to incidents and, above all, whether somebody is taking real responsibility for what was put into production. In a world where manufacturing software has become absurdly easy, trust may be another one of those resources that has become relatively more expensive.
Comments