
I Think Subagents Need Better Management, Not More Agents
A practical argument for using fewer subagents, clearer ownership, tighter context, and stronger handoffs instead of maximizing agent count.
There’s been a lot of talk lately about subagents wasting tokens, agents messaging each other too much, spawning five agents where one would have been fine, and honestly I think a lot of that criticism is probably right. I’ve watched some of these workflows turn into a complete mess where every agent needs the same repo context, everybody independently explores the same files, they summarize their findings back and forth, and before anything useful gets written you’ve already burned through a ridiculous amount of usage.
More subagents does not automatically mean more intelligence, and in a lot of cases it probably means the exact opposite because now you have coordination overhead on top of the original problem.
But I’ve been messing around with a slightly different version of this through a CLI/MCP setup, and so far it has been working better than I expected. Instead of letting a bunch of agents loosely talk to each other, I can call something like Fable 5.1, give it the exact workspace and files it needs to understand, let it figure out that part of the problem, and then have it hand very specific execution jobs down to cheaper models underneath it.
That difference sounds small but I think it changes the whole thing.
The point isn’t really to have more agents. If anything, I’m starting to think the goal should be to have as few agents as possible, but make the handoff between them much better.
What I’m actually testing
The main coding session still owns the overall job. If I’m already working inside Codex or Claude Code for three hours, that session has context that is hard to recreate somewhere else. It knows what I’m building, what we already tried, what I rejected, what turned out to be weird in the repo and why certain decisions were made.
I don’t want to throw all of that away every time another agent gets involved.
So the main session stays at the top and only delegates something when there is actually a reason to.
For a bigger piece of work it might decide that one part of the problem deserves its own investigation. Instead of spawning four random subagents and telling them all to “look into auth,” it can call one strong model, say Fable 5.1, and give it the actual workspace plus the specific files or area of the repo that matter.
That model gets a real mission.
Not:
Take a look at auth and tell me what you think.
More like:
We need to replace this provider while preserving our internal user IDs and current session behavior. These are the relevant parts of the repo. Figure out exactly what has to change, what assumptions we need to verify, and come back if anything you find changes the larger architecture.
Now Fable can actually understand that part of the system without having to reconstruct the entire project from scratch.
Once it understands the mission, it doesn’t necessarily need another Fable to do every piece of execution underneath it.
That’s where the worker models come in.
Luna MAX, Opus 4.8, Grok 4.6, and potentially Kimi or GLM when the assignment is scoped well enough. If something still requires much more judgment, obviously use something stronger. I’m not trying to force a cheaper model into a job it can’t handle just to save a few cents.
The interesting part is seeing how much work becomes cheap-model work after somebody smarter has already cleaned up the problem.
The handoff matters more than I thought
I think this is where a lot of multi-agent setups go wrong.
The parent agent gives another agent a vague problem, the second agent has to explore half the repo, it gives a giant summary back, the first agent realizes something is missing, then another agent gets launched to investigate that, and pretty soon all you’ve really done is reproduce one coding session across five separate context windows.
That obviously wastes tokens.
The setup I’m testing is closer to giving the stronger agent enough authority to own its part of the job, then letting it create very narrow jobs underneath itself.
For example, after looking through the auth system, Fable might realize there are three things it needs done. One worker needs to change the provider adapter behind an existing interface. Another needs to write regression tests for current session behavior. Another needs to trace one remaining dependency before anything else changes.
Those are now much smaller problems.
The worker doesn’t need to understand the product vision, the entire migration, the billing system, the conversation I had with the main thread two hours ago, or why we decided not to touch some unrelated abstraction.
It needs the files, the relevant facts, the constraints and a clear idea of what “done” means.
That seems obvious when you say it out loud, but I think it makes a huge difference.
The stronger agent shouldn’t keep running upstairs either
There’s another version of this that would be just as bad, where every worker reports to the manager, the manager reports everything back to the main session, the main session makes another decision, then sends it back down again.
Now you have enterprise bureaucracy made out of tokens.
I don’t want that either.
Once the main session gives the stronger subagent a mission, that agent should own it within clear boundaries.
If Fable is responsible for the auth migration and discovers that the solution stays completely inside auth, let it finish the job. It can inspect the code, dispatch a couple workers, review what they give back, retry something with a better model if needed, and produce one coherent result.
The main session does not need a play-by-play.
If Fable discovers something that affects the rest of the system, then it goes back up.
Say the original plan was to preserve the existing user ID contract, but after exploring the actual code it turns out billing directly depends on an ID from the old provider and there is no clean way around changing it.
Now we have a global problem again.
That goes back to the main session because the decision is no longer just about auth.
I think that boundary is important. The strong subagent should be autonomous inside its mission, otherwise you lose most of the reason for having it.
Exploration first, execution after
This is also something that has been working better for me than immediately telling every agent to start writing code.
A lot of coding tasks are vague because we don’t actually know enough about the code yet.
There’s nothing wrong with that.
The mistake is pretending we do and handing an implementation job to a worker before anybody has figured out what the implementation should be.
So the stronger model can spend some time exploring first.
Maybe it reads ten files itself. Maybe it sends one Luna MAX scout to find every reference to a certain interface. Maybe another cheap worker traces how a value moves through the system. Maybe none of that requires another agent because the Capo can figure it out faster by reading the code directly.
Then, once it actually understands what is going on, it creates the execution jobs.
That separation between exploration and execution seems important to me.
The expensive model is especially valuable while the problem is still ambiguous.
Once the ambiguity is gone, the model requirements can change dramatically.
This is where I think cheaper models get underrated
A cheaper coding model gets judged a lot of the time by giving it the same vague assignment we would give the best model available.
Of course the stronger model usually looks better.
But that isn’t really the experiment I’m interested in.
I want to know what happens after Fable has already looked through the repo and tells Luna MAX:
Change these two functions. Keep this interface exactly the same. Handle these three cases. Do not touch these files. Run these tests and stop if assumption X turns out to be false.
Now Luna isn’t being asked to architect the system.
It’s being asked to execute a good plan.
Same thing with Opus 4.8, Grok 4.6, Kimi, GLM or whatever else ends up fitting into that pool.
Maybe some of them still suck at certain jobs. Fine.
Maybe one turns out to be unbelievably good at clean TypeScript implementation but bad at debugging. Maybe another is better when the assignment still has a little uncertainty left in it. Maybe Kimi or GLM gets good enough on properly scoped coding tasks that it becomes stupid to spend frontier money on those jobs.
I don’t know yet.
That’s why I’d rather measure it than make the architecture depend on whatever model is popular this month.
Context is also more nuanced now
I originally thought a big part of the savings would simply come from smaller context windows. Instead of handing every agent 100,000 tokens of conversation and repo history, give the worker 5,000 tokens containing exactly what it needs.
I still think that matters, but the providers themselves are getting a lot better at solving long-context problems.
Claude already does compaction, where older parts of a long session get compressed instead of being carried forever in raw form. Claude Code also isolates subagent context, which already helps avoid dumping every exploration step into the main session.
OpenAI is taking this further with the newer Astra work in Codex, where they’ve been experimenting with keeping notes across context windows while still being able to retrieve information from older windows later. That is a pretty meaningful improvement over continually summarizing a conversation and hoping every important detail survives.
So I don’t think the argument is that these systems are bad at context.
They’re not.
The distinction I care about is that compaction and retrieval are trying to help one agent maintain continuity over time, while this setup is also trying to decide who needs which context at all.
Even if my main session can perfectly remember everything we’ve discussed for the last six hours, that does not mean a worker fixing one React state bug needs access to those six hours.
Give it the state bug.
Give it the relevant files.
Give it the requirements.
If it finds something that breaks the assumptions in its assignment, let it ask for more.
That feels cleaner to me than giving every agent everything just because the context system technically allows it.
Fewer subagents is probably better
This is maybe the biggest thing I’ve changed my mind on.
When agent orchestration first started getting interesting, there was something intuitively appealing about running a bunch of agents in parallel. You look at a giant problem and think, why not have ten agents attack ten parts of it?
Because sometimes those ten parts aren’t actually independent, and now ten agents are reading overlapping code, making conflicting assumptions and generating ten separate chunks of work that somebody has to understand and combine.
At some point the coordination cost becomes worse than the original problem.
So I don’t think the right goal is maximum parallelism.
It’s probably minimum useful parallelism.
If one strong model can explore something faster than it can explain it to three workers, just let it explore.
If there are two clean questions that can actually be answered independently, send two workers.
If one implementation job is already perfectly scoped for Luna MAX, give it to Luna and move on.
There should be a reason every agent exists.
Otherwise we’re just burning usage because watching agents work looks productive.
So the flow I’m using is pretty simple
The main session owns the overall job and decides if something is worth breaking out.
If it is, a strong model gets a mission plus the exact workspace and context it needs.
That model explores the mission until it actually understands it.
It uses a small number of workers when there is something useful to investigate or execute in parallel.
The workers report back to that model.
That model reviews the work, retries or changes assignments when necessary, and finishes its mission without constantly involving the main session.
If it discovers something that affects the larger project, it escalates that decision back up.
Once the mission is complete, the main session gets the result instead of 50 pages of conversation between every agent involved.
And all of the mechanical shit underneath that, Git commands, worktrees, commits, queues, test execution, budgets, should just be software.
There’s no reason to pay a model to be Git.
So far, it’s working
This is still early and I’m not trying to declare victory after messing with a CLI for a few days, but so far the setup has actually been pretty good.
The usage seems to last longer, the smaller workers aren’t being asked to solve problems they were never good enough to solve in the first place, and the stronger model spends more of its time understanding the problem and giving useful direction instead of doing every little implementation step itself.
That doesn’t mean it will win on every task.
There are definitely going to be jobs where one Astra MAX, Fable 5.1 or Sol MAX session is simply better and faster than trying to split anything up.
That is fine.
The system should be allowed to decide that the best number of subagents for a job is zero.
The part I want to measure
The only way this becomes more than another agent architecture diagram is if it actually saves money or produces better work.
So I want to measure the whole attempt.
Not just how cheap the Luna call was.
How much did the strong model spend exploring?
How many workers got launched?
How many came back wrong?
How many retries?
How much did the entire mission cost?
How long did it take?
Did the tests pass?
Did I keep the implementation?
If Luna MAX is cheap but three of its patches get thrown away, that matters.
If Fable 5.1 costs more upfront but gives three workers such good assignments that every one of them succeeds on the first try, that matters too.
The number I care about is basically the cost of getting to something I actually keep, not how cheap any individual model looks on a pricing page.
Over time you should be able to learn which combinations work.
Maybe Luna MAX becomes the obvious worker for normal implementation.
Maybe Opus 4.8 is worth spending more on when you need a little extra reasoning.
Maybe Grok 4.6 turns out to be great for bounded debugging.
Maybe Kimi or GLM becomes the default for certain jobs because the assignment is clean enough that the quality difference barely matters anymore.
And maybe the data says half the time we should not have delegated at all.
That’s probably just as valuable to know.
Where I’m at with it
I don’t think subagents are bad and I don’t think throwing more subagents at a problem is the answer either.
Both takes are probably too simple.
What seems more interesting is using a small number of them with much clearer boundaries.
Let the main session keep the broad context.
Give one strong model real ownership of a mission when the problem is large enough to justify it.
Let that model explore before it starts dispatching work.
Then use cheaper models for the parts that have become specific enough that they no longer need the same level of intelligence.
And when the main model can simply do something faster itself, let it.
That’s basically what I’m testing right now, and so far it seems a lot more sensible than spawning an army of agents and hoping intelligence somehow emerges from all the messages going back and forth.
The thing I’m trying to optimize isn’t how many agents I can run.
It’s how little expensive intelligence I can use without making the end result worse.
