Recently I learned about Completion Notifications for agentic work within Codex. I told my Codex agents to stop checking on each other and use completion notifications instead. This stopped the orchestrated set of agents from checking on each other with repeated model turns just to report that the work was still running. This started saving my weekly quota to be used on better things than sleep checks every 10 seconds for a 30 minute pytest run.
AI-assisted development is a multi-pronged thought process. It is a prompting optimization game as well as a model personality/tendency and a big rework problem.
The model based on its tendency could decide if it wants to take the prompt too literally or maybe just an outline of what you want and still ignore some pieces that it doesn’t want to care about. GPT 5.5 and Opus 4.7 are prime examples for the divergence in behaviour. Opus is more exploratory in nature while GPT 5.5 takes its instructions pretty seriously (as of today). Once you understand how different models behave, you may unlock the power to work with them as a partner by using the subtle persuasion techniques. I have been dealing with these subtleties of different models for the past 8-10 months of my journey evolving my workflow with them (not the exact same versions, but you catch the drift).