The Publisher Agent Runs Five Jobs and Never Edits a File
In August I shipped the tag/pillar system. The registry holds fourteen tags; eleven of them survived the tag audit. Every published post now declares a pillar, and a build-time validator fails if the pillar doesn't match a tag's pillar. The rules and the registry are just files in the repo. A TypeScript file holds the tags. A config file holds the schema. A design doc holds the rules.
What I didn't have was anything that ran those contracts against the actual blog corpus before I merged. So I'd merge a post, the build would pass because the post passed, and I'd later notice what the build can't see. Root-relative links that 404. Descriptions that miss the post's searchable terms. Posts that carry no subject tags at all. The build can only pass or fail the build. It can't catch the things that are technically valid but wrong.
The spec had been sitting there for a week. "Pre-publication validator." Five jobs, no agent.
This week I built the agent.
What I actually wanted
A report. Not an editor. Not a thing that would commit a post on my behalf. A thing that would walk the corpus, run my existing contracts, and tell me what was wrong.
Some of these contracts already fail the build, but only after the merge. The agent runs the same checks as a report before the merge, plus the sidecar-generation job nothing was checking.
The shape I landed on was five jobs, but only four run on a plain report run:
- Job 3 — Tag/pillar validation. Reuses the registry file's own helpers,
resolveTagandvalidatePillarMatchesTags. Flags unregistered tags, pillar mismatches, missing tags. Proposes tags for posts that don't carry any subject ones. The proposals are recommendations for me to consider. Adding tags to the registry stays a PR by a human, per the taxonomy doc's rule that the human is the curator. The agent never writes to the registry. - Job 4 — Internal link sanity. Walks every post's markdown links. Flags root-relative
/<slug>(which 404s — blog posts live at/blog/<slug>, not/<slug>). Top-level pages (/about,/projects,/now) are valid at root-relative. The agent verifies against frontmatter slugs, not filenames. That bit me once already. - Job 5 — SEO checklist. Consumes the SEO practices doc via a loaded skill. Doesn't re-invent the checklist. Drops the
keywordscheck (Google ignores it). Respects thelastModifieddiscipline. No cosmetic date bumps just to look fresh. - Job 6 — Pillar/content advisory. The newest one. Reads the post body, identifies the primary subject, and reports when the subject seems to disagree with the taxonomy doc's decision rule. Advisory only. The agent never re-pillars a post. I told it to stay quiet on ambiguous cases. Only clear mismatches get flagged.
Then there's Job 2, which runs when the run is authorised to write. It generates the sidecar a post's hero needs: a small JSON file that renders the terminal-style window above the post. Velite-typed, terminal-flavoured. The agent writes the file. I review the diff before anything goes in.
The numbering is the spec's. Job 1 was per-post SVG covers, and it died in the card-design reversal (no decorative covers, square corners everywhere) before anyone wrote code. The survivors kept their spec numbers so the spec, the plan and the agent prompt all say the same thing. There is no Job 1.
So that's five jobs total. Four run on a plain report run. All five run when the run is authorised to write.
What I didn't want
I didn't want an editor.
I'd already built a multi-agent review system last winter. Four specialized agents that would debate each draft through two or three rounds. It worked. It also weighed a lot: ~15,750 tokens of prompt overhead across the agents, and 50+ messages burned per review, all against the Claude Code session limits I kept hitting. I tried trimming the prompts to stretch those limits. It was overkill for the actual problem I had.
The actual problem I had was: is this post ready to ship, or did I forget something? That's a yes-or-no question, not a debate.
The publisher agent is the slimmer version. One agent, one pass. It reports. I act.
This matters. If the agent edited the registry file for me, I'd lose the audit trail. The agent runs against the contracts, not the other way around.
How it's wired
The wiring is three files, and the pattern is the part worth writing down: an agent definition, a slash command, and a context skill. Same three-slot pattern the rest of my slash command setup uses.
The skill is in the repo because it references project-specific contracts (the taxonomy doc, the registry, the velite schema), so it's not portable.
The command parses flags before spawning the agent, the same mode-detection approach from that post. One flag flips the mode. A bare slug narrows the target. No args means all posts. Pre-flight checks come before the spawn. The registry exports what it should. The heroes schema is intact. The agent file exists. The posts directory isn't empty. Then it spawns the agent with the mode and the target.
After the agent returns, the command prepends a short summary header and displays the report. The command never commits. The agent never commits. I commit.
The verification command was wrong
The first version of the command file said to run npm run lint to confirm the generated sidecars matched the velite schema. That's wrong. npm run lint is eslint only. It doesn't validate JSON schemas. I caught it before the agent PR merged.
The right command is npx velite (schema-only, no Next build) or npm run build (velite plus the Next build, so heavy). I picked npx velite.
But why was this the command's job at all? The agent already runs a pre-write mental check against the schema before it writes a sidecar: title length, line count, line types, text length per line. That's a necessary check. It's not sufficient. The agent's mental model can drift from the schema. The command's role is to confirm the on-disk artifact actually builds clean.
So the verification step is npx velite. If that throws, the sidecars don't match the schema and I review the diff.
This is the kind of mistake that's easy to make because the doc phrasing invites it. "Run lint to confirm" sounds right because lint is a thing you run. The trap is that "lint" means different things in different projects. Here it means eslint, and eslint doesn't see JSON schemas.
What the agent actually reports
When I run the publish check, the report is structured Markdown. Summary first, then each job's findings, then next steps.
Each job produces errors, warnings, or informational notes. Errors are real build-breakers: a tag not in the registry, a pillar mismatched to its tags, a link to a post that doesn't exist. Warnings are things that are technically valid but carry no information. A post whose only tag is its own pillar tag passes the validator and says nothing about the post. Informational notes are proposals: here's a tag I'd add to this untagged post.
The agent doesn't tell me what to do. It tells me what it found.
For Job 6 specifically, the agent defaults to NOT flagging ambiguous cases. Take the post about a homelab server that also touches on AI tooling. The agent reports the declared pillar and notes that another pillar could equally apply. I decide. If the agent flagged every ambiguous post, half the corpus would be flagged.
The defaults moved while this sat
Drafts wait. The repo moves. One thing changed under this post while it sat: the defaults moved. When I wrote this, hero generation only ran when I asked for it with a flag. Now an all-posts run stays report-only by default, and a single-post run generates the post's missing sidecar in the same run. One flag suppresses that. Another still triggers the batch backfill. The commit that made the change settles the logic: the quality gate was never the flag. It's the never-commit rule plus the velite validation, and both apply the same at one post as at forty.
So Job 2 now runs more than I planned when I wrote this. That's the correct amount. A post shipping without its hero was the gap the opt-in default kept recreating.
The tooling aged too. When I wrote this I ran one tool for the agent work and a second for the orchestration around it. I've settled on a single harness since. The three-slot pattern above is the part that transfers. The tool name isn't.
What's left
The backfill beat Job 2 to it. The day after I wrote this, the sidecar backfill closed by hand. The first two were an experiment. I liked what I saw, so the bulk backfill followed. Thirty-five sidecars across six batches, all in one day. Thirty-nine of the forty-one published posts carry a sidecar now; the gaps are posts from after the backfill. So Job 2's first real test is a comparison run. Generate a sidecar for a post that already has one, and see whether the generated file holds up next to the original. The three hand-authored sidecars from August are still the tone bar.
What I'm watching for:
- The agent's tag proposals. No post has an empty tags array right now, but eleven of the forty-one published posts carry only the pillar tag. That is exactly what Job 3 flags. I'll review each proposal, decide which ones earn their place in the registry, and add them as separate PRs. The registry rule stands: the human is the curator.
- The agent's link audit. I'll fix the root-relative links it flags. The link convention lives in the project's agent notes. The audit makes sure I followed my own rule.
- The agent's SEO checklist output. Some of it will be obvious (description length). Some of it will catch things I'd miss (missing alt text, weird heading order).
The contracts already existed. The registry file, the velite schema, the SEO doc, the taxonomy doc. I just hadn't wired them to a thing that runs them.
The agent doesn't add new rules. It runs existing rules. That's a different shape than building a multi-agent debate system, but for pre-publication checks it's the right shape. The debate system was good for getting a draft to publication-quality. The publisher agent is good for catching the things that slip past on the way.