A team I spoke to recently built an analyst agent and gave it the whole job: bring the data in, analyse it, report on it. Its operator barraged it with questions. Its brain was full. The context window filled, the tokens were used up, and it lost track of what it was for.
Anthropic's guide to context engineering describes a related effect: "as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases." Chroma tested 18 models and found performance "grows increasingly unreliable as input length grows."
I told the team to think of agents as dumb little robots all the way down. One agent pulls data from a domain, one normalises it, one analyses it, one formats the data, one formats the PowerPoint. An agent might take two of those jobs, but you don't give it everything in one go.
I covered decomposition in The Art of Breaking Things Down and it needs repeating, because enterprise teams are still being sold AI as the number one problem solver you can throw everything at.
Five small agents means four handoffs, and someone has to check them. Each agent needs a person who owns it, understands it and takes responsibility for how it performs. Any system a team builds should have that accountability anyway.
These things are only as good as we program them to be. An agent that gets confused may simply have been given too many jobs.