↓ Skip to main content
  1. Blog/

Agent Skills Don't Have to Be Flat

Author
Abhishek Walia
Eternal nerd dealing with complex systems, practical tooling, and the urge to share what I’ve learned.

I used to write Agent Skills as flat instruction files. This basically meant that a set of repeatable instructions I had would be converted into one skill, and before too long, I would have around 50 skills per repo. Each skill described one workflow, and they never called each other. That left a lot of repeated text across them.

As I grew tired of it, I experimented with a small orchestration idea. Can I call a skill from within another skill, create semantic orchestration layers, allow multiple skills to interleave with each other, and create a repeatable “pattern” instead of a workflow? The answer I got was a resounding yes. Skills can actually describe a hierarchy and work within it really well.

Think:

main skill/workflow → specialist sub-skill → worker agents → reviewer agents → feedback loop → accepted result → next specialist sub-skill → execution → end

For example:

Task skill
├── grilling skill to finalize what I need from this task
├── Task design skill
│   ├── designer agent
│   ├── reviewer agents
│   └── review loop
└── Task implementation skill
    ├── implementer agent
    ├── reviewer agents
    ├── review loop
    └── validation scripts and hooks

This basically meant that I got somewhat close to progressive disclosure (not exactly the same thing, but close) while still orchestrating things as needed, with more precision, by stitching smaller skills together. It felt like magic.

The hierarchy above depicts a simple model: the main agent follows skill instructions to coordinate the workflow. Subagents are the leaves that do judgment-heavy work. In my workflow, scripts and hooks handle deterministic validation, then feed failures back into the appropriate loop.

I’ve validated this pattern in Codex CLI and Claude Code. The workflow pattern is portable; its execution is not. Each runtime handles skill activation, delegation, hooks, and context a little differently.

By the way, this is not a token-saving trick. Spawning more agents can use more tokens. The real gain is specialization, context isolation, and a better cost per successful outcome.

Related

TMUX and agents: Match made in heaven

Claude Code runs in my terminal all day. Although I have started dabbling with Codex, Opencode and others but Claude Code is still the primary thing I’m looking at for hours at a stretch for now. All of those harnesses run on a dedicated VM, not my main machine. Daily backups, isolated environment, nothing else lives there. At some point I realized my terminal config mattered a lot more than it used to.

The Design Doc Is the Prompt

AI-assisted development is a multi-pronged thought process. It is a prompting optimization game as well as a model personality/tendency and a big rework problem. The model based on its tendency could decide if it wants to take the prompt too literally or maybe just an outline of what you want and still ignore some pieces that it doesn’t want to care about. GPT 5.5 and Opus 4.7 are prime examples for the divergence in behaviour. Opus is more exploratory in nature while GPT 5.5 takes its instructions pretty seriously (as of today). Once you understand how different models behave, you may unlock the power to work with them as a partner by using the subtle persuasion techniques. I have been dealing with these subtleties of different models for the past 8-10 months of my journey evolving my workflow with them (not the exact same versions, but you catch the drift).

Sizing a Kafka Cluster from First Principles

Most Kafka sizing advice is hand-wavy. “Start with 3 brokers.” “Scale when you hit a bottleneck.” “Talk to your vendor.” I have spent years on both ends of that conversation, as the engineer asking and the person being asked, and the answer is almost always some version of “it depends, let’s just see how it goes.” That isn’t sizing. Its the easiest way to say “I don’t know.”.