↓ Skip to main content
  1. Blog/

Stop making agents watch each other

Author
Abhishek Walia
Eternal nerd dealing with complex systems, practical tooling, and the urge to share what I’ve learned.

Recently I learned about Completion Notifications for agentic work within Codex. I told my Codex agents to stop checking on each other and use completion notifications instead. This stopped the orchestrated set of agents from checking on each other with repeated model turns just to report that the work was still running. This started saving my weekly quota to be used on better things than sleep checks every 10 seconds for a 30 minute pytest run.

I added the following three rules to my Codex (global) AGENTS.md so that it applies to every project that uses subagents:

  • Use completion notifications for delegated work. The owning agent must report completion, results, and blockers; do not repeatedly query its status or spawn another agent to monitor it.
  • For long-running commands, keep necessary monitoring with one owner. Prefer completion events over repeated status checks.
  • Report meaningful changes, failures, decisions, or completion. Do not spend tokens merely to demonstrate activity.

That reduced the back-and-forth in my setup. The agent doing the work was responsible for following it through and returning a useful result. For tests, that meant reporting which commands ran, what passed or failed, and anything that needed a decision.

I still want the orchestrator involved when a test contradicts the plan or an implementation choice changes the design. Those are the reasons to read the worker’s findings and steer it. These instructions reduced the repeated status checks in my setup, but context rot is still a problem. Having subagents return summaries helps keep the orchestrator’s context comparitively cleaner.

The visible improvement for me was visibly less noise in the orchestrator thread. Delegated work still consumes additional tokens due to subagent usage, so I wouldn’t use this observation alone to claim that my whole multi-agent setup is cheaper. But for now, I have some peace knowing I don’t need to scroll too much to find what is going on.

Related

I Used One Codex Chat to Orchestrate Another

I just learned that Codex can use one chat to supervise another. These are not subagents, but real threads with one monitoring the other. I am working on a MetaSkill project, and I wanted to try out GPT-6 Astra’s new capabilities to see whether it would work better as an orchestrator or an implementer. The orchestrator experiment is ongoing, and I will tackle the implementer part next. I asked Astra to create a separate chat for the work, assign it the goal, track it, and steer it toward the outcome I wanted. Astra started a Sol chat at Extra High reasoning (exactly as I asked it to), then created a scheduled task in its own thread that checked the worker every ten minutes.

The Design Doc Is the Prompt

AI-assisted development is a multi-pronged thought process. It is a prompting optimization game as well as a model personality/tendency and a big rework problem. The model based on its tendency could decide if it wants to take the prompt too literally or maybe just an outline of what you want and still ignore some pieces that it doesn’t want to care about. GPT 5.5 and Opus 4.7 are prime examples for the divergence in behaviour. Opus is more exploratory in nature while GPT 5.5 takes its instructions pretty seriously (as of today). Once you understand how different models behave, you may unlock the power to work with them as a partner by using the subtle persuasion techniques. I have been dealing with these subtleties of different models for the past 8-10 months of my journey evolving my workflow with them (not the exact same versions, but you catch the drift).