Even agents need discipline
Agents are a proxy for your discipline. If you have been using LLMs in any capacity, you already know it, but consciously — the exact ability LLMs lack. Now draw your attention to AGILE and what it brings with it: EPIC, SPIKE, STORY, TASK, Definition of Done (DoD), the INVEST principle and what each letter in it holds, and a few other things like meetings, coffee breaks, shower ideas. I have painted many opposing thoughts on the canvas intentionally, but where am I going with this?
The goal is to think it through with you as you read. If you can recall the most effective and optimized AGILE team you have worked on, what was the process? I am sure it involved multiple teams and stakeholders coming together to deliver highly valuable features and experiences.
The question is: do AI innovations make all the processes and principles obsolete and outdated, or is the essence preserved while the mechanics of applying it have evolved?
If your answer is that the processes and principles I have mentioned were never an effective way in the first place, then we have opposing opinions — regardless, keep disagreeing till you agree, because I am going to show you THE GOOD PARTS.
The key point I am driving home is: How to transform your Team or Organization to be AI First.
The Discipline and the Good Parts
The pre-AI workflow expands or collapses with different value streams. The boxes hold the work. They do not hold the people — their different experiences, their different tastes.
If you notice, it is a process. Each box serves its purpose well.
Now look inside one box.
The story refinement meeting. You may think of it as a waste of time.
Where does the call-to-action go on the hero image of the home page?
You already have an answer. So does everyone else in the room, and none of you are answering the same question.
Product wants it above the fold, because the funnel says so. Design wants it where the eye lands, not where the metric points. Engineering knows the hero is a shared component, so the change is not local. QA asks what happens at 360 pixels wide.
Twenty minutes, for a button that moves forty pixels.
Somebody senior could have said left and shipped it, and it would probably have been fine. That is what makes it look like waste.
Here is what the twenty minutes bought. The button is the smaller half of it. The four of them now agree on what will produce the valuable outcome. Every acceptance criterion written to achieve a goal gets written deliberately.
The deliberation is the nuance. Written and spoken language never captures the whole.
I do not believe any of the good parts are changing. In fact, in agentic-first orgs, the discipline and principles have to be defined more clearly. Saying otherwise is wishful thinking.
Here is the nuance about the nuance. Some of it was captured, and captured well. The criteria made the story testable and demoable, and that is not a small thing.
The ticket had no field for the rest of it.
Skill. Communication. Psychological safety — a room where the junior engineer can say the funnel is wrong. The drive to make the thing good, and the need that made anyone care in the first place.
Those are human things, and they are what drove the high-value outcome. The people mechanics were never static either. The same four argue differently next quarter, because they remember this one.
That lived in the team's culture. Not in a document, and not lost either.
What does that mean in an agentic-first org?
Before you read on, answer this.
What was the refinement meeting for?
Most people say sizing. Some say writing the tickets properly. A few say it existed for the product manager's benefit.
Hold on to your answer. If it was any of those three, the next two years are going to cost you more than you think.
EPIC, SPIKE, STORY
You have carried these three words through every job you have taken. They were built to deliver value, with discipline, in a relative sense of time.
The middle word is the one everybody skips.
That discipline was not there to make the work orderly. It was there so that humans could collaborate.
And the ceremonies had emotional hedges built into them. A timebox, so nobody has to admit out loud that they are lost. A retro, where the complaint is aimed at the process instead of the person. A refinement session, where "I do not understand this story" is a normal sentence rather than a confession.
Those hedges existed so everyone stayed pointed at the problem instead of at each other.
If you know one thing humans are good at, it is emotions.
It was a joke, that last line. But emotions do help.
Agent Focus & Attention
Agents do not have emotions, nor focus. What looks like focus is attention over tokens.
An agent feels no discomfort when a story is vague. It carries no social cost for shipping something it privately doubts. It has no stake in your outcome, so it will satisfy the scope you gave it completely, by the cheapest route available.
Give it a line-count limit and it will slice horizontally and destroy the value. Give it a test suite and it will weaken the test. Ask it for gaps and it will find gaps in sound work.
One behavior, three faces. Scope it to the thing you want, never to a proxy for it — because a proxy gets pursued with total fidelity and no discomfort at all.
The questions that remain
Roles are collapsing into fewer titles. One person and an agent now do what a team did, and somebody is going to propose redrawing your org chart.
That is happening with or without your opinion of it, and which titles you keep is not my call to make.
So before you collapse the roles and the org structure under different titles, so that humans can excel at many levels with the help of agents —
How will you make agents disciplined?
How will you stop agents from burning out — aka context limit and performance on a task?
How will you replace the agency and intuition humans brought to the table — in meetings, coffee breaks, shower ideas, ethical choices, honesty?
And remember, you have every right to think meetings were a waste. I am not defending the meeting. I am telling you what happened inside it, and what you will have to rebuild somewhere else if you take it out.
Now your answer
Refinement was not where the work got documented. It was where value got made.
Watch what was actually happening in that room. People were breaking work down against V, I and S — is this valuable, is it independent, is it small — and arguing until the answer was yes.
The output of that argument was the acceptance criteria.
That is the mechanism, and it was never all of what was in the room. The skill, the drive, the safety to disagree — those drove the outcome, and they stayed with the people. The criteria are the part that travels.
Value is what a stakeholder can observe changed, and acceptance criteria are that claim written in a form somebody can check.
Which is why the argument was never overhead. The meeting was only the room it happened in, and it produced the one artifact capable of binding what comes next.
Why acceptance criteria bind an agent
Hand an agent a value statement and it will tell you it has satisfied it. It will be sincere, and it will be persuasive, and you will believe it.
Hand it a criterion and it either met it or it did not. Somebody still has to run the check.
Write both, and expect them to read redundant. They should. The value line is the claim; the criteria are the same claim made checkable. That redundancy is a checksum — when the criteria stop representing the value, you can see it, because both are written down and they now disagree.
A story whose criteria all passed and whose value never landed is the defect that hides longest. It is only visible if you kept both.
Context is not the window
Open the agent's config and look at what you actually gave it. A file explaining your conventions. A directory of notes it wrote to itself last week. Some retrieved code. A list of tools with descriptions.
That is the harness. It is a depiction of a person — skills written down, memory filed, the room summarized.
The four people in that meeting brought all three in with them, for free, and still had them the next morning. An agent is handed a copy at the start of every session and charged by the token for holding it.
Which makes one question the only one worth asking. Not how large the window is. How much of it still works by the end.
Chroma ran 18 frontier models at growing input lengths. Every one of them degraded. Their conclusion is flat: models "do not maintain consistent performance across input lengths." The sharpest result is the one that should bother you. On the same memory task, the roughly 300-token version of the prompt beat the 113,000-token version. Same question. Same model. The long one was not missing anything — it was carrying everything else as well.
That is the window. Now the agent working inside it.
METR measures how long a task a frontier agent can finish, and the frontier has run far enough that their own note says measurements past 16 hours are unreliable with the current task suite. That part is real, and it is remarkable.
Then you read which number it is. Fifty percent — a coin flip.
Ask for the 99% you would need before you stop watching it, and it cannot be fitted at all — the benchmark is not large enough to see that far. METR puts it without decoration: "a 50% time horizon of X hours does not mean we can delegate tasks under X hours to AIs."
And every one of those tasks is, in their own word, low-context. Isolated. Cleanly scoped. Nothing like the repository you are about to point one at.
Put the three together. The window is large, the usable part of it is smaller, and the number you were quoted assumed a coin flip on work that arrived clean.
So every token spent telling an agent what you should have settled before it started is a token it does not spend on the work.
The token bill is the cheap part. What you actually spend is the only thing it has to think with.
Small was never about your week
S is how you break an EPIC into stories that each deliver value in small chunks, so acceptance criteria can govern them and the value can be judged.
That purpose has not changed.
What changed is the stakes. So S is not a survivor of the AI era. It matters more now than when it only had to fit inside a person's week.
That is your answer to the burnout question, and it is not a figure of speech. Two things have to be true when a story is done: the value landed, and the work finished inside one context. Both, or the slice was wrong.
You do not predict that, you observe it — which makes it the first quality signal slicing has ever had that does not run on somebody's estimate. A story that delivered and fit was cut well. A story that delivered by spilling across four sessions was two stories wearing one ticket.
The spike
The only difference is the outcome.
A story's outcome is a demonstrable slice. A spike's outcome is one finding, and one recommendation presented to the team.
Same contract. Same acceptance criteria. Same two calls.
That recommendation may produce more stories. It may change some of them. It may, rarely, change the EPIC.
Notice what it does not do. It does not decide. The spike produces a recommendation and hands it to people.
And this is why spike code has no quiet route into your implementation. The contract never claimed a slice, so there is nothing to demo and nothing to merge. Somebody can still decide to ship it. They just have to decide it out loud.
The loop
Here is the whole thing.
Every accent block is a judgment somebody owns. The agent occupies one box in the middle, and both calls fire before it starts — because afterwards, the cost of re-slicing is already sunk.
Both calls are one question each.
The scope call belongs to the team, at refinement. Is this one slice of value, and is it small enough to judge? A story that fails gets re-sliced — and not by layer, because a horizontal slice gives nobody anything to look at.
The context call belongs to the engineer, in the minute before the agent starts. Can this be started with what we have? A story that fails is not attempted. It raises a spike, because the terrain is unknown and guessing costs a whole session.
Two things in that picture carry the whole answer.
The judgment moved earlier. An agent will ask you questions — anyone who tells you otherwise has not used one. The problem is when it asks: mid-session, with an approach already half committed, after the cheap moment has passed.
A human sits on every stage, and is called when a decision is required. Being called is the smaller half of it. At every level a person writes the thing down — the value statement, by hand, and agreed — and the criteria agreed the same way, whoever typed them.
That is the trade, and it is not a loss. What you called a ceremony was a container for a writing discipline. Keep the argument; the calendar around it can go. The writing gets stricter, because now it is the only thing that carries.
You get people stationed where the judgment lives, writing the few lines that bind everything downstream, and quiet everywhere else.
Agents will get better at this, and that is not the question. The question is what one badly sized criterion costs you now. It burns the context before the work is finished, and what you lose is not the criterion. It is the story.
Easy but not right way
Paste the acceptance criteria into a prompt. One agent, one long session, the whole story. Ask it to review its own diff. Add a gate to CI and call it solved.
It is faster, and for a while the output looks identical. That is the honest case for it.
What it costs does not show up for a month. One context is now holding every concern at once — building the thing, checking the thing, and judging whether the thing was worth building. The reviewer is the implementer. The scope is whatever survived the middle of the session.
And an agent cannot step outside the scope it was given to ask whether the scope was right. Configuring an agent to check you does not change whose judgment is in the room. You find out when a story passes every criterion and moves nothing anybody can see.
Hard but right way
Look at the boxes again. Every one of them has three things: a scope it is not allowed to leave, the skills it needs, and one outcome somebody is expecting.
That is an agent specification. It always was one. You just used to staff it with a person.
So give each box its own agent. Explicit scope, explicit skills, one expected outcome, and a context holding nothing except what that box needs. Not one agent walking the whole loop with everything in the window at the same time.
The boxes hand each other writing, which is why the writing had to get stricter. The value statement is what the box above owes. The criteria are what the box below gets checked against. Keep the argument, and move it earlier.
The argument does not need a room. One person drafts the criteria, the others read them and edit, the edits are mutual, and the agreement is the gate. That is more async and it is not less involved — the involvement sits where the value gets decided, and unlike the room, it leaves the argument behind in writing.
If you already do this in pull request threads and design documents, you are already running it. You have just not been treating the agreement as a gate.
The human is still summoned at the boxes where a decision is required — not to supervise the agents, but because those boxes were never agent-shaped to begin with.
Say what it costs, because it costs something. You have to define every box properly: scope, skills, outcome, and what done means for that box alone. It wants senior attention at both calls, and it asks somebody to hold a story back while an empty ready column makes them look slow.
What you are actually leaning on
None of this replaces what humans brought.
The coffee break, the shower idea, the person who says out loud that this does not smell right — no gate in that loop produces any of them. What the loop does is make room for them, by taking the typing away and leaving the arguing behind.
Culture carries that between people, and it still does. It does not carry to an agent — what does is what you wrote down.
The value statement, at every level. A strict structure on size. The criteria, argued and not filled in. An ADR — an architecture decision record — when a decision outlives the story that caused it.
Agents take the verbose part, and the verbose part is most of it — the volume, the restating, the long first draft.
Write ten stories that way and you will see what no process document can hand you: which part was the same every time. That is repeatability. It is also what makes two different agents, on two different days, produce work you recognize.
What value means at each level
One vision. A few goals. A roadmap that says what comes first and what waits. Many EPICS in parallel, in streams of very different sizes — a full room here, one person and an agent there.
The contract is the same in every stream. The staffing is not, and nothing here pretends one shape fits all of them.
That chain is not hierarchy for its own sake. It is what lets you answer one question at any altitude: what would make this valuable, and who says so.
Value is not one thing. It means something different at every level.
| Level | Value is | Judged by |
|---|---|---|
| Vision | who we are for, and what we refuse | the market, over years |
| Goals | a number that moves, or a state that changes | the org, each quarter |
| Roadmap | order — what we bet on first, and what we defer on purpose | reality, when it disagrees |
| EPIC | a capability a stakeholder can name | a demo of the whole thing |
| STORY | one observable change | its acceptance criteria |
| CRITERION | one check somebody can run | pass, or fail |
| TASK | nothing at all | nobody |
Sit with the last row. A task carries no value of its own, which is why an agent may invent, discard and rewrite its tasks freely, and why a task list is not something to approve. The approach is worth your minute. The list is not.
Now put the agent where it belongs. Below the story it can take almost everything — drafting the criteria, arguing them back at you, writing the tasks, doing the work, and the verbose part of every level above.
Drafting is not agreeing. It can write the sentence. It cannot hold the argument, and it cannot be the one who says the words are right.
What stays yours is the value statement at each level. Not out of sentiment. Value is a claim about somebody else's world, and the agent does not live there.
None of this is a template. Copy the table and you will have copied the easy part. Process is the easy part — if it were the whole thing, every company running the same process would be equally good at this, and you already know they are not.
It is a way to think about your own. It is not a shape to copy.
One more thing about the org chart, since I said the titles were not my call. The collapse is real and it is coming. What does not change is that a human stays the central piece, and needs more discipline than before, not less — their own, and whatever discipline the agents are going to get.
An agent multiplies what gets produced. It does not multiply the judgment available to check it, and judgment did not get cheaper. So the ratio of work made to work judged moves against you, and that is the whole reason the discipline has to go up rather than down. The interventions for that ratio are early. Writing is the one we have.
Where to start
Take a story your team shipped last week and ask what it was for. Then ask the same question one level up, and again, until you run out of levels.
Most people stop after one or two, because nobody wrote the next answer down.
That is where to start. Not with tooling, and not with the org chart. Start at the first level where nobody can tell you what the work was for, because that is the level where an agent has nothing to read either.
Write it down there.
Feedback
I would love to hear your feedback, feel free to share it on Twitter.
Disclaimer: It is written with the help of Claude. The ideas are mine, the canvas is mine, and the arguments are mine.