Building Domain Agents That Know Your Stack
A subagent is a job description, not a personality. What goes in one, why my code reviewer cannot write files, and why most of them run on a cheaper model.
A subagent is a Markdown file with frontmatter. The frontmatter is where most of the value is, and it is the part people skim.
Here is the header of my Nuxt code reviewer:
---
name: nuxt-code-reviewer
description: MUST BE USED to run a rigorous, security-aware review of all
Vue/Nuxt code (pages, components, composables, plugins, server routes,
middleware) after every feature, bug-fix, or pull-request. Use PROACTIVELY
before merging to main.
tools: LS, Read, Grep, Glob, Bash
---
Four decisions in five lines. Each one is load-bearing.
The description is a routing rule
The description is not a summary for humans. It is what the orchestrator matches against when deciding whether to delegate — so it should be written as a trigger condition, not a job title.
"Reviews Vue and Nuxt code" is a job title. It tells the router nothing about when to fire.
The version above lists the artefacts (pages, components, composables, plugins, server routes, middleware) and the moments (after every feature, bug-fix or pull request; before merging to main). It is trying to be matched.
My Vue architect goes further and includes worked examples in the description:
<example>
Context: User needs to optimize data fetching
user: "The monitor detail page is loading slowly"
assistant: "Let me use the nuxt-vue-architect agent to analyze the data
fetching pattern and implement optimized SSR data loading..."
</example>
That looks like padding. It is not. Examples show the shape of a triggering request, which catches cases a list of nouns misses — "the page is loading slowly" contains none of the keywords you would have thought to list.
The words MUST BE USED and PROACTIVELY are doing work too. Without them a reviewer only runs when explicitly asked, which means it runs when you remember, which means rarely. That is the opposite of a quality gate.
Restricting tools is a safety feature
tools: LS, Read, Grep, Glob, Bash
No Edit. No Write. My code reviewer physically cannot change code.
This is the single best decision in my agent setup, and it is about incentives rather than trust.
An agent that can both find and fix problems will fix them. That sounds efficient and is not, for two reasons. You lose the review — findings you never saw, now buried in a diff. And it conflates two jobs with opposite dispositions: reviewing wants suspicion and thoroughness, fixing wants minimal intervention. An agent doing both does neither well.
Taking away the write tools makes the output a report. I read it, I decide what is real, I fix what matters. The reviewer's only job is to be right about problems.
The same logic applies elsewhere. A research agent does not need Bash. A doc generator does not need network access. Default to the smallest tool set that can do the job — not because the agent is malicious, but because capability it does not need is capability that can be misapplied.
Omit tools entirely and the agent inherits everything. Convenient, and almost never what you want.
Model tier is a cost decision
Most of my agents pin a model:
model: sonnet
The reasoning is boring and financial. A WordPress specialist answering "how do I add a field to WooCommerce checkout" is doing recall against well-documented territory. That does not need the most capable model available; it needs a competent one that responds quickly.
Where I leave the model unpinned — inheriting the session's — is anything requiring judgement across a large surface: architectural review, hard debugging, "is this finding real or a false positive".
The heuristic: pin down for recall, leave open for judgement.
Watch for it drifting. An agent that started as a lookup helper and grew into something that makes real decisions should have its tier revisited. Mine have not been reviewed in a while, which is a note to self as much as advice.
Name your agent, not your file
Small thing that cost me twenty minutes: my reviewer lives in code-reviewer.md but declares name: nuxt-code-reviewer. The name field wins. The filename is just a filename.
If invoking an agent silently does nothing, check that you are using the declared name.
What goes in the body
Below the frontmatter is the system prompt. Two things separate a useful one from a generic one.
Be specific about the stack, not the domain. "You are an expert in Vue" produces textbook Vue. Naming your actual patterns — SSR data fetching with useAsyncData, hydration mismatches, composables over stores, whatever your codebase does — produces advice that fits the code in front of it.
Say what the output should look like. My reviewer specifies a severity-tagged report. Without that, review output drifts between a bulleted list, an essay, and a diff, and you cannot skim it consistently. Specifying the output format is cheap and pays every single run.
What to leave out: generic best practice. "Write clean code", "consider edge cases", "follow SOLID". The model already knows. Those words spend tokens to add nothing, and they dilute the specific instructions that do matter.
The honest limitation
An agent cannot see your conversation. It knows its system prompt and the prompt it was handed.
That is the whole point — it is why the isolation is valuable — and it is also the main failure mode. If you have spent twenty minutes establishing that a bug only reproduces under a specific Cloudflare setting, and you then delegate to an agent, that context does not travel. You get a competent investigation of the wrong thing.
So brief them like contractors, not colleagues. Everything needed goes in the prompt. When I catch an agent doing confidently irrelevant work, the cause is almost always my briefing rather than the agent.
Build one, not eleven
I have eleven and use maybe five regularly. The SEO ones in particular fragmented — seo-technical, seo-content, seo-schema, seo-sitemap, seo-performance, seo-visual — because splitting felt tidy at the time. In practice an SEO question rarely respects those boundaries, and I now have to pick between six overlapping agents.
Two or three broad specialists would have served better than six narrow ones. Split when a single agent's instructions genuinely conflict, not when the topics feel conceptually distinct.
Next: why I turned off auto-compact, and what I do instead.
Need a developer who ships fastwithout shipping mess?
I build and maintain WordPress, Laravel and Nuxt applications for businesses that care about performance and maintainability.