<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Luke Manning - Blog — AI</title>
        <link>https://lukemanning.ie/</link>
        <description>Breaking things. Building things. Writing about it. (tag: AI)</description>
        <lastBuildDate>Wed, 30 Sep 2026 12:46:24 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <copyright>All rights reserved 2026, Luke Manning</copyright>
        <atom:link href="https://lukemanning.ie/feeds/ai.xml" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[My Agent Closed Most of the Open Projex Issues in One Round Without Me Reading the Code]]></title>
            <link>https://lukemanning.ie/blog/my-agent-closed-most-projex-issues-in-one-round</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/my-agent-closed-most-projex-issues-in-one-round</guid>
            <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>On a Saturday afternoon in August, I told my agent to fix every open issue in the Projex repo and prepare a release. There were ten open. Some were HIGH priority. Some were LOW.</p>
<p>I went for a walk.</p>
<p>When I came back, the work had landed across three branches. Each subagent had worked in its own worktree on its own branch. The build was green on each one. The tests were green. The lint was green. The typecheck was green. The release manager subagent had drafted the changelog. Nine of the ten issues were closed. The tenth stayed open — it needed a real union restructure, not a mechanical fix.</p>
<p>I had not read a single line of the diff yet.</p>
<h2>What I actually asked for</h2>
<p>I had a backlog of small things in <a href="/blog/building-projex-retrospective">the Projex repo</a>. Tagged-union cleanups. Type tightening. Documentation gaps. A bug where one of the smart-grid props was a documented prop but a no-op at runtime. A redundancy where two functions with slightly different spellings did the same thing.</p>
<p>I described this to my main agent. The main agent looked at the issue tracker, saw the labels (<code>bug</code>, <code>enhancement</code>, <code>documentation</code>), grouped them by file area, and <a href="/blog/opencode-subagent-permissions-ordering-trap">dispatched three subagents</a> in parallel.</p>
<p>One subagent got the package.json + bundling issues. One got the type-system + tagged-union issues. One got the documentation + test-coverage issues. Each one worked in a separate worktree on a separate branch. Each one committed locally and reported back. A fourth subagent, the release manager, handled release prep alongside them: the version bump and the changelog.</p>
<p>I watched the transcript. Mostly I stayed out of the way. I made tea.</p>
<h2>What the subagents actually did</h2>
<p>I saw the dispatch messages. I saw the report-back messages. I did not read the intermediate diffs. The subagents were set up to commit per-issue. Each commit was meant to be independently reviewable.</p>
<p>A few things I noticed in the report-backs:</p>
<ul>
<li>One subagent caught a redundancy I hadn't seen. Two exported functions, <code>normalizeStats</code> and <code>normaliseStats</code>, with the American and British spellings. Both did the same thing. The codebase had drifted to the British spelling in the actual logic. The American spelling was the alias. Both were exported. The subagent deprecated the American one with a JSDoc tag and updated the docs to point at the British spelling.</li>
<li>One subagent found a related issue while fixing another one. While narrowing the <code>ProjectStats</code> union, it noticed <code>FetchProjectDataResult.commits</code> was using <code>undefined</code> while sibling fields used <code>null</code>. The subagent opened a new issue and included the fix in the same branch.</li>
<li>One subagent flagged a peer-dependency problem I had been ignoring for two months. The CLI packages (<code>ts-morph</code>, <code>chalk</code>, <code>@inquirer/prompts</code>, <code>commander</code>) were installed by every consumer, even ones who only imported the components. The subagent moved them to optional <code>peerDependencies</code> so consumers importing only components stopped dragging in the CLI bundle.</li>
</ul>
<p>None of these were in my original brief. The subagents went past the edges of what I asked for, in the direction of "things that were obviously wrong in the same file area."</p>
<h2>Where I read the code</h2>
<p>The first time I read any of the code was after all three subagents finished and opened their PRs. I skimmed the diffs before merging. Not a line-by-line review. A sanity check.</p>
<p>I was looking for decisions the AI made without asking me. Function names I wouldn't have picked. Behaviour that wasn't in the brief. Edits that touched code outside the file area I asked about. The kind of things a real code review catches, except I was reviewing the decisions, not the code.</p>
<p>Some of the diffs were four lines. Some were thirty. None of them were complex enough to need a real review. They were tagged-union narrowings, JSDoc additions, dependency relocations. The kind of work where you skim it once, you understand it, you move on.</p>
<p>If I'd skimmed each PR as it landed, I'd have read the rename with no idea the docs were about to change under it. Reading the batch, I could see the <code>normaliseStats</code> deprecation and the docs update pointing at it in the same sitting. The batch skim was faster than piecemeal would have been.</p>
<h2>Where I did intervene</h2>
<p>I didn't push back on any of the code. The three branches each shipped clean. I steered the architecture around the loop, not the code inside it.</p>
<p>The dispatch went out as three parallel <code>opencode run</code> invocations, not three subagents. I asked for that change when the opencode TUI failed on the first attempt and the right path was to skip the interactive layer. I also argued for splitting the release prep out from the fix work, because trying to do both in the same dispatch kept blocking on the release-manager hitting its timeout before the fix branches landed.</p>
<h2>What this loop replaced</h2>
<p>My previous loop was one PR at a time. I'd describe an issue to the agent. The agent would open a PR. I'd skim it, sanity-check the decisions, merge or push back. One issue, one PR, one round of skimming. Repeat.</p>
<p>For this kind of small mechanical work, that's fine. It works. But the context switching adds up. Each PR is its own session — its own dispatch, its own transcript, its own review pass. The overhead is small per PR and large per backlog.</p>
<p>The new loop:</p>
<ol>
<li>Describe the backlog.</li>
<li>Wait.</li>
<li>Skim the batch of PRs.</li>
<li>Push the release prep.</li>
</ol>
<p>The release prep runs alongside the fix work instead of after it. One description covers all of them.</p>
<p>I want to be careful about what I'm claiming here. I'm not claiming the subagents did better work than the agent would have done one PR at a time. Most of these issues were mechanical. The interesting decisions — which redundancy to deprecate, which naming to standardize — the agent would have surfaced them either way, given the brief. What I'm claiming is that the per-PR overhead moved out of my hands and the interesting decisions stayed in my hands.</p>
<h2>Reading at the end, not in the middle</h2>
<p>I did not read the code while it was being written.</p>
<p>In the old loop, I skimmed each PR after the agent opened it. One PR at a time.</p>
<p>In the new loop, the skim happened after the writing finished across all three PRs. The subagents were the feedback loop during the work. I was the feedback loop at the end.</p>
<p>There's a different cost structure. A wrong fix in the old loop was caught in the per-PR skim, or it shipped. A wrong fix in the new loop is caught in the batch skim, or it ships.</p>
<p>I shipped none of the wrong fixes in this batch. The issues were mechanical and the code area was small. If I'd asked the subagents to redesign the <code>normalise</code> function, I would have read every line, pushed back, rewritten pieces.</p>
<p>The loop works for the kind of work that has a clear right answer. The loop does not work for the kind of work that needs taste. I haven't found the line yet.</p>
<p>For this batch, the line was "moves stuff around, adds JSDoc, narrows types." Below the line, I delegated. Above the line, I didn't. The line is in a different place than I would have guessed.</p>
<h2>Running it again</h2>
<p>I'm going to run this loop again. On a different repo. On a different kind of work.</p>
<p>I want to see what happens when the issues aren't mechanical. I want to see where the line moves. I want to see what kinds of work I delegate that I later wish I hadn't, and what kinds I keep that the loop could have handled.</p>
<p>I'm not going to delegate design decisions. I'm not going to delegate <a href="/blog/i-shipped-a-library-now-what">"what should this library do"</a>. I'm going to delegate "make this library do what it already says it does, correctly."</p>
<p>The batch skim at the end is non-negotiable. That's the part I own.</p>]]></description>
            <content:encoded><![CDATA[<p>On a Saturday afternoon in August, I told my agent to fix every open issue in the Projex repo and prepare a release. There were ten open. Some were HIGH priority. Some were LOW.</p>
<p>I went for a walk.</p>
<p>When I came back, the work had landed across three branches. Each subagent had worked in its own worktree on its own branch. The build was green on each one. The tests were green. The lint was green. The typecheck was green. The release manager subagent had drafted the changelog. Nine of the ten issues were closed. The tenth stayed open — it needed a real union restructure, not a mechanical fix.</p>
<p>I had not read a single line of the diff yet.</p>
<h2>What I actually asked for</h2>
<p>I had a backlog of small things in <a href="/blog/building-projex-retrospective">the Projex repo</a>. Tagged-union cleanups. Type tightening. Documentation gaps. A bug where one of the smart-grid props was a documented prop but a no-op at runtime. A redundancy where two functions with slightly different spellings did the same thing.</p>
<p>I described this to my main agent. The main agent looked at the issue tracker, saw the labels (<code>bug</code>, <code>enhancement</code>, <code>documentation</code>), grouped them by file area, and <a href="/blog/opencode-subagent-permissions-ordering-trap">dispatched three subagents</a> in parallel.</p>
<p>One subagent got the package.json + bundling issues. One got the type-system + tagged-union issues. One got the documentation + test-coverage issues. Each one worked in a separate worktree on a separate branch. Each one committed locally and reported back. A fourth subagent, the release manager, handled release prep alongside them: the version bump and the changelog.</p>
<p>I watched the transcript. Mostly I stayed out of the way. I made tea.</p>
<h2>What the subagents actually did</h2>
<p>I saw the dispatch messages. I saw the report-back messages. I did not read the intermediate diffs. The subagents were set up to commit per-issue. Each commit was meant to be independently reviewable.</p>
<p>A few things I noticed in the report-backs:</p>
<ul>
<li>One subagent caught a redundancy I hadn't seen. Two exported functions, <code>normalizeStats</code> and <code>normaliseStats</code>, with the American and British spellings. Both did the same thing. The codebase had drifted to the British spelling in the actual logic. The American spelling was the alias. Both were exported. The subagent deprecated the American one with a JSDoc tag and updated the docs to point at the British spelling.</li>
<li>One subagent found a related issue while fixing another one. While narrowing the <code>ProjectStats</code> union, it noticed <code>FetchProjectDataResult.commits</code> was using <code>undefined</code> while sibling fields used <code>null</code>. The subagent opened a new issue and included the fix in the same branch.</li>
<li>One subagent flagged a peer-dependency problem I had been ignoring for two months. The CLI packages (<code>ts-morph</code>, <code>chalk</code>, <code>@inquirer/prompts</code>, <code>commander</code>) were installed by every consumer, even ones who only imported the components. The subagent moved them to optional <code>peerDependencies</code> so consumers importing only components stopped dragging in the CLI bundle.</li>
</ul>
<p>None of these were in my original brief. The subagents went past the edges of what I asked for, in the direction of "things that were obviously wrong in the same file area."</p>
<h2>Where I read the code</h2>
<p>The first time I read any of the code was after all three subagents finished and opened their PRs. I skimmed the diffs before merging. Not a line-by-line review. A sanity check.</p>
<p>I was looking for decisions the AI made without asking me. Function names I wouldn't have picked. Behaviour that wasn't in the brief. Edits that touched code outside the file area I asked about. The kind of things a real code review catches, except I was reviewing the decisions, not the code.</p>
<p>Some of the diffs were four lines. Some were thirty. None of them were complex enough to need a real review. They were tagged-union narrowings, JSDoc additions, dependency relocations. The kind of work where you skim it once, you understand it, you move on.</p>
<p>If I'd skimmed each PR as it landed, I'd have read the rename with no idea the docs were about to change under it. Reading the batch, I could see the <code>normaliseStats</code> deprecation and the docs update pointing at it in the same sitting. The batch skim was faster than piecemeal would have been.</p>
<h2>Where I did intervene</h2>
<p>I didn't push back on any of the code. The three branches each shipped clean. I steered the architecture around the loop, not the code inside it.</p>
<p>The dispatch went out as three parallel <code>opencode run</code> invocations, not three subagents. I asked for that change when the opencode TUI failed on the first attempt and the right path was to skip the interactive layer. I also argued for splitting the release prep out from the fix work, because trying to do both in the same dispatch kept blocking on the release-manager hitting its timeout before the fix branches landed.</p>
<h2>What this loop replaced</h2>
<p>My previous loop was one PR at a time. I'd describe an issue to the agent. The agent would open a PR. I'd skim it, sanity-check the decisions, merge or push back. One issue, one PR, one round of skimming. Repeat.</p>
<p>For this kind of small mechanical work, that's fine. It works. But the context switching adds up. Each PR is its own session — its own dispatch, its own transcript, its own review pass. The overhead is small per PR and large per backlog.</p>
<p>The new loop:</p>
<ol>
<li>Describe the backlog.</li>
<li>Wait.</li>
<li>Skim the batch of PRs.</li>
<li>Push the release prep.</li>
</ol>
<p>The release prep runs alongside the fix work instead of after it. One description covers all of them.</p>
<p>I want to be careful about what I'm claiming here. I'm not claiming the subagents did better work than the agent would have done one PR at a time. Most of these issues were mechanical. The interesting decisions — which redundancy to deprecate, which naming to standardize — the agent would have surfaced them either way, given the brief. What I'm claiming is that the per-PR overhead moved out of my hands and the interesting decisions stayed in my hands.</p>
<h2>Reading at the end, not in the middle</h2>
<p>I did not read the code while it was being written.</p>
<p>In the old loop, I skimmed each PR after the agent opened it. One PR at a time.</p>
<p>In the new loop, the skim happened after the writing finished across all three PRs. The subagents were the feedback loop during the work. I was the feedback loop at the end.</p>
<p>There's a different cost structure. A wrong fix in the old loop was caught in the per-PR skim, or it shipped. A wrong fix in the new loop is caught in the batch skim, or it ships.</p>
<p>I shipped none of the wrong fixes in this batch. The issues were mechanical and the code area was small. If I'd asked the subagents to redesign the <code>normalise</code> function, I would have read every line, pushed back, rewritten pieces.</p>
<p>The loop works for the kind of work that has a clear right answer. The loop does not work for the kind of work that needs taste. I haven't found the line yet.</p>
<p>For this batch, the line was "moves stuff around, adds JSDoc, narrows types." Below the line, I delegated. Above the line, I didn't. The line is in a different place than I would have guessed.</p>
<h2>Running it again</h2>
<p>I'm going to run this loop again. On a different repo. On a different kind of work.</p>
<p>I want to see what happens when the issues aren't mechanical. I want to see where the line moves. I want to see what kinds of work I delegate that I later wish I hadn't, and what kinds I keep that the loop could have handled.</p>
<p>I'm not going to delegate design decisions. I'm not going to delegate <a href="/blog/i-shipped-a-library-now-what">"what should this library do"</a>. I'm going to delegate "make this library do what it already says it does, correctly."</p>
<p>The batch skim at the end is non-negotiable. That's the part I own.</p>]]></content:encoded>
            <category>ai</category>
            <category>opencode</category>
            <category>projex</category>
            <category>github</category>
        </item>
        <item>
            <title><![CDATA[Anthropic, Open Source, and Why I Cancelled My Subscription]]></title>
            <link>https://lukemanning.ie/blog/anthropic-open-source-walled-garden-clawdbot-opencode</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/anthropic-open-source-walled-garden-clawdbot-opencode</guid>
            <pubDate>Fri, 03 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>I was using Claude Code daily. I built <a href="/blog/building-multi-agent-blog-review-system">my entire blog review system</a> with it. I hit <a href="/blog/premature-optimization-multi-agent-prompts">session limits constantly</a> and tried to optimise around them. Now I use OpenCode, an open-source agent that works with any provider.</p>
<p>Then one day (maybe a month or two ago) I pulled the latest OpenCode version and Claude subscription support was just gone. Not deprecated. Not moved to a config flag. Removed entirely, because Anthropic demanded it. On top of that, I kept hitting <a href="https://github.com/anthropics/claude-code/issues/16157">usage limits on my Pro subscription</a> despite paying $20/month. That issue was opened for Max subscribers but people on every tier were hitting the same thing.</p>
<p>I cancelled my Anthropic subscription. I use GLM-4.7 and GLM-5.1 now through OpenCode with a GLM coding subscription. Considering adding a GitHub Copilot subscription so I can still use Anthropic models through OpenCode's API integration. I still use Claude on the web and on my phone. The model is good. The repeated session limit issues and the company's behavior around open source tools is what pushed me away.</p>
<h2>The OpenClaw Story</h2>
<p>I'd already been watching Anthropic's relationship with open source tools for a while. Then I came across the OpenClaw story. An open source project that got forced to rename because it had "Clawd" in the name. That sent me down a rabbit hole.</p>
<p>Peter Steinberger (steipete on GitHub) built a personal AI assistant. Originally called Warelay. MIT licensed, supports 20+ messaging channels, and has 343k GitHub stars (yes, really). A legit open source project by any measure. The fastest growing in history in terms of stars.</p>
<p>On January 4, 2026, he renamed it to ClawdBot. Makes sense. It's a bot that uses Claude. Clear, descriptive name. A play on the fact it used Claude without using the name Claude.</p>
<p>Three weeks later, he got a letter from Anthropic's legal team.</p>
<p>The project was renamed to "Moltbot" on January 27, 2026. The commit message read: "refactor: rename clawdbot to moltbot with legacy compat." The maintainer confirmed in <a href="https://github.com/openclaw/openclaw/issues/2825">GitHub issue #2825</a>: "Moltbot is the official name now. Clawdbot has been letter sent by Anthropic."</p>
<p>That name lasted three days. On January 30, another commit: "refactor: rename to openclaw." The project has been OpenClaw ever since.</p>
<p>An open source project with 343k stars, MIT licensed, built by an independent developer, got a trademark letter because its name contained "Clawd." Yeah, obviously a play on "Claude," but still not "Claude."</p>
<p>And that wasn't all. Anthropic also reportedly changed their API to reject any request whose system prompt contains the phrase "Open Claw." Not a rate limit, not a warning, but a hard block. (This was reported by Theo — he runs t3.gg, covers AI and dev tools — I haven't tested it myself.) If that's accurate, it goes way beyond trademark enforcement. That would be technical suppression of a specific open source project.</p>
<h2>What Happened to OpenCode</h2>
<p>The OpenCode situation is different but follows the same pattern.</p>
<p>For context: OpenCode works with multiple AI providers. You configure it with credentials for whichever provider you want, and it routes requests to your chosen model. Before all this, Claude subscription access was one of those options. You could use your existing Pro or Max subscription through OpenCode instead of being locked into Anthropic's client.</p>
<p>Anthropic added checks to stop third-party tools from impersonating the Claude Code client. This broke people's ability to use their Claude subscriptions through OpenCode, Cursor, and any other third-party harness.</p>
<p>Then they updated their terms of service to explicitly ban using consumer subscriptions (Pro/Max plans) as authentication for third-party tools. So even people paying Anthropic $20/month for Pro can't legally use that subscription through OpenCode. You have to use the Anthropic harness (Claude Code).</p>
<p>OpenCode had its own PR that removed Claude subscription integration entirely, with "Anthropic legal requests" cited as the reason. Not a technical limitation. A legal demand.</p>
<p>I get protecting your trademark. I get not wanting people to think an unofficial tool is officially affiliated with you. But the OpenCode situation isn't about trademark confusion. It's about controlling which tools can access Claude's API and how.</p>
<p>OpenCode wasn't pretending to be Claude Code. It was using Claude as a provider, the same way it uses ChatGPT or any other model. Anthropic decided they didn't want subscription auth flowing through third-party tools, and they had the legal muscle to enforce it. Their stated justification was something about third-party harnesses interfering with their analytics and causing unusual traffic patterns. At least, that's what I gathered. It's hard to know how much of that is genuine concern versus justifying the walled garden.</p>
<p>The rules around what's allowed are vague enough that even people trying to play by them can't get straight answers. Matt Pocock — whose TypeScript stuff I've used for a while — spent over a month trying to get Anthropic to confirm whether his paid Claude Code course was allowed to exist. Their response: "We're working on it." Repeatedly. When pressed for an ETA, same answer. Theo has called this out too — he believes Anthropic keeps the terms vague intentionally so they can move the goalposts later. That tracks with what I've seen.</p>
<h2>The Pattern</h2>
<p>Anthropic isn't just another company locking everything down. They built MCP (Model Context Protocol) — the protocol that lets OpenCode talk to external tools. And then gave it away. Fully open source. Anthropic doesn't control it. They offer free Claude Max access to open source maintainers. Their output terms are solid; you own what Claude generates.</p>
<p>But then Claude Code has a public GitHub repo with zero source code. Subscription access is tightly controlled — you can't use what you're paying for through a third-party tool. And their legal team sends trademark letters to open source projects with 343k stars.</p>
<p>Then there's the distillation thing.</p>
<p>I read Anthropic's report accusing DeepSeek, Moonshot, and MiniMax of "distillation attacks" against Claude and something felt off. They claimed 24,000 fraudulent accounts and 16+ million exchanges. Wrapped it in national security framing stating that distilled models could be used for bioweapons, etc. And this isn't the first time — Anthropic previously accused Windsurf, X AI, and OpenAI of distillation and, from what I've read, was wrong each time.</p>
<p>The distillation accusations fit the same pattern. Anthropic claims others are exploiting Claude's openness, uses that to justify tighter restrictions, and those restrictions happen to protect their revenue. Could be genuine security concern. Could be strategic. But it's the same thing I keep seeing.</p>
<p>The only major lab that has released zero open-weight models is also the one arguing most loudly that open-weight models are dangerous. I don't think that's a coincidence.</p>
<p>I initially thought the ClawdBot thing was just standard trademark enforcement. But then I kept looking. I'm not sure if this is a fair reading, honestly. I keep going back and forth on it. But the pattern I keep seeing is decisions being made to be open where it benefits Anthropic, and closed where it protects their revenue.</p>
<p>That's absolutely a valid business strategy. It's just not what "open" means. Especially when most of their competitors are making strides to become more open.</p>
<h2>Where I'm At</h2>
<p>I'm going to keep using OpenCode for now.</p>
<p>I might be wrong about the intent. Maybe Anthropic has good reasons for each individual decision. The ClawdBot rename could be standard trademark enforcement. The OpenCode crackdown could be about subscription terms or it could legitimately be about analytical data being interfered with by third party harnesses.</p>
<p>But when I look at the pattern, it looks like a company building a walled garden around Claude while also contributing genuinely open infrastructure. I don't have a clean conclusion here. I'm still forming my thinking on this. But I know that my workflow got disrupted, an open source project got renamed three times in a month, and the company responsible for both is the same one that open-sourced MCP.</p>
<p>I just don't understand their actions sometimes.</p>]]></description>
            <content:encoded><![CDATA[<p>I was using Claude Code daily. I built <a href="/blog/building-multi-agent-blog-review-system">my entire blog review system</a> with it. I hit <a href="/blog/premature-optimization-multi-agent-prompts">session limits constantly</a> and tried to optimise around them. Now I use OpenCode, an open-source agent that works with any provider.</p>
<p>Then one day (maybe a month or two ago) I pulled the latest OpenCode version and Claude subscription support was just gone. Not deprecated. Not moved to a config flag. Removed entirely, because Anthropic demanded it. On top of that, I kept hitting <a href="https://github.com/anthropics/claude-code/issues/16157">usage limits on my Pro subscription</a> despite paying $20/month. That issue was opened for Max subscribers but people on every tier were hitting the same thing.</p>
<p>I cancelled my Anthropic subscription. I use GLM-4.7 and GLM-5.1 now through OpenCode with a GLM coding subscription. Considering adding a GitHub Copilot subscription so I can still use Anthropic models through OpenCode's API integration. I still use Claude on the web and on my phone. The model is good. The repeated session limit issues and the company's behavior around open source tools is what pushed me away.</p>
<h2>The OpenClaw Story</h2>
<p>I'd already been watching Anthropic's relationship with open source tools for a while. Then I came across the OpenClaw story. An open source project that got forced to rename because it had "Clawd" in the name. That sent me down a rabbit hole.</p>
<p>Peter Steinberger (steipete on GitHub) built a personal AI assistant. Originally called Warelay. MIT licensed, supports 20+ messaging channels, and has 343k GitHub stars (yes, really). A legit open source project by any measure. The fastest growing in history in terms of stars.</p>
<p>On January 4, 2026, he renamed it to ClawdBot. Makes sense. It's a bot that uses Claude. Clear, descriptive name. A play on the fact it used Claude without using the name Claude.</p>
<p>Three weeks later, he got a letter from Anthropic's legal team.</p>
<p>The project was renamed to "Moltbot" on January 27, 2026. The commit message read: "refactor: rename clawdbot to moltbot with legacy compat." The maintainer confirmed in <a href="https://github.com/openclaw/openclaw/issues/2825">GitHub issue #2825</a>: "Moltbot is the official name now. Clawdbot has been letter sent by Anthropic."</p>
<p>That name lasted three days. On January 30, another commit: "refactor: rename to openclaw." The project has been OpenClaw ever since.</p>
<p>An open source project with 343k stars, MIT licensed, built by an independent developer, got a trademark letter because its name contained "Clawd." Yeah, obviously a play on "Claude," but still not "Claude."</p>
<p>And that wasn't all. Anthropic also reportedly changed their API to reject any request whose system prompt contains the phrase "Open Claw." Not a rate limit, not a warning, but a hard block. (This was reported by Theo — he runs t3.gg, covers AI and dev tools — I haven't tested it myself.) If that's accurate, it goes way beyond trademark enforcement. That would be technical suppression of a specific open source project.</p>
<h2>What Happened to OpenCode</h2>
<p>The OpenCode situation is different but follows the same pattern.</p>
<p>For context: OpenCode works with multiple AI providers. You configure it with credentials for whichever provider you want, and it routes requests to your chosen model. Before all this, Claude subscription access was one of those options. You could use your existing Pro or Max subscription through OpenCode instead of being locked into Anthropic's client.</p>
<p>Anthropic added checks to stop third-party tools from impersonating the Claude Code client. This broke people's ability to use their Claude subscriptions through OpenCode, Cursor, and any other third-party harness.</p>
<p>Then they updated their terms of service to explicitly ban using consumer subscriptions (Pro/Max plans) as authentication for third-party tools. So even people paying Anthropic $20/month for Pro can't legally use that subscription through OpenCode. You have to use the Anthropic harness (Claude Code).</p>
<p>OpenCode had its own PR that removed Claude subscription integration entirely, with "Anthropic legal requests" cited as the reason. Not a technical limitation. A legal demand.</p>
<p>I get protecting your trademark. I get not wanting people to think an unofficial tool is officially affiliated with you. But the OpenCode situation isn't about trademark confusion. It's about controlling which tools can access Claude's API and how.</p>
<p>OpenCode wasn't pretending to be Claude Code. It was using Claude as a provider, the same way it uses ChatGPT or any other model. Anthropic decided they didn't want subscription auth flowing through third-party tools, and they had the legal muscle to enforce it. Their stated justification was something about third-party harnesses interfering with their analytics and causing unusual traffic patterns. At least, that's what I gathered. It's hard to know how much of that is genuine concern versus justifying the walled garden.</p>
<p>The rules around what's allowed are vague enough that even people trying to play by them can't get straight answers. Matt Pocock — whose TypeScript stuff I've used for a while — spent over a month trying to get Anthropic to confirm whether his paid Claude Code course was allowed to exist. Their response: "We're working on it." Repeatedly. When pressed for an ETA, same answer. Theo has called this out too — he believes Anthropic keeps the terms vague intentionally so they can move the goalposts later. That tracks with what I've seen.</p>
<h2>The Pattern</h2>
<p>Anthropic isn't just another company locking everything down. They built MCP (Model Context Protocol) — the protocol that lets OpenCode talk to external tools. And then gave it away. Fully open source. Anthropic doesn't control it. They offer free Claude Max access to open source maintainers. Their output terms are solid; you own what Claude generates.</p>
<p>But then Claude Code has a public GitHub repo with zero source code. Subscription access is tightly controlled — you can't use what you're paying for through a third-party tool. And their legal team sends trademark letters to open source projects with 343k stars.</p>
<p>Then there's the distillation thing.</p>
<p>I read Anthropic's report accusing DeepSeek, Moonshot, and MiniMax of "distillation attacks" against Claude and something felt off. They claimed 24,000 fraudulent accounts and 16+ million exchanges. Wrapped it in national security framing stating that distilled models could be used for bioweapons, etc. And this isn't the first time — Anthropic previously accused Windsurf, X AI, and OpenAI of distillation and, from what I've read, was wrong each time.</p>
<p>The distillation accusations fit the same pattern. Anthropic claims others are exploiting Claude's openness, uses that to justify tighter restrictions, and those restrictions happen to protect their revenue. Could be genuine security concern. Could be strategic. But it's the same thing I keep seeing.</p>
<p>The only major lab that has released zero open-weight models is also the one arguing most loudly that open-weight models are dangerous. I don't think that's a coincidence.</p>
<p>I initially thought the ClawdBot thing was just standard trademark enforcement. But then I kept looking. I'm not sure if this is a fair reading, honestly. I keep going back and forth on it. But the pattern I keep seeing is decisions being made to be open where it benefits Anthropic, and closed where it protects their revenue.</p>
<p>That's absolutely a valid business strategy. It's just not what "open" means. Especially when most of their competitors are making strides to become more open.</p>
<h2>Where I'm At</h2>
<p>I'm going to keep using OpenCode for now.</p>
<p>I might be wrong about the intent. Maybe Anthropic has good reasons for each individual decision. The ClawdBot rename could be standard trademark enforcement. The OpenCode crackdown could be about subscription terms or it could legitimately be about analytical data being interfered with by third party harnesses.</p>
<p>But when I look at the pattern, it looks like a company building a walled garden around Claude while also contributing genuinely open infrastructure. I don't have a clean conclusion here. I'm still forming my thinking on this. But I know that my workflow got disrupted, an open source project got renamed three times in a month, and the company responsible for both is the same one that open-sourced MCP.</p>
<p>I just don't understand their actions sometimes.</p>]]></content:encoded>
            <category>ai</category>
        </item>
        <item>
            <title><![CDATA[I Broke My Blog Review System Trying to Beat Session Limits]]></title>
            <link>https://lukemanning.ie/blog/premature-optimization-multi-agent-prompts</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/premature-optimization-multi-agent-prompts</guid>
            <pubDate>Sun, 04 Jan 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<blockquote>
<p><strong>Context</strong>: This post assumes you've read <a href="/blog/building-multi-agent-blog-review-system">how I built the multi-agent blog review system</a>. If you haven't, that post explains the architecture before this one breaks it.</p>
</blockquote>
<p>I broke my entire blog review system this week trying to beat session limits.</p>
<p>Not because tokens cost money. Because Claude Code has session limits - I kept hitting the message cap before finishing my work.</p>
<p>Here's what that means in practice: Claude Code limits how many messages you can send in a session. Each time you ask it something, that's a message. Each time an agent spawns, that's a message. Each file read, each grep search - all messages. When you hit the limit (usually around 100-150 messages depending on context), your workflow just stops mid-task.</p>
<p>Every agent spawn, every file read, every grep search consumes tokens. When your prompts eat 15,750 tokens before any actual work happens, you burn through your session budget fast.</p>
<p>So I thought: "If I can trim these prompts by 1,100 tokens, I'll get more messages per session. More reviews. More work done."</p>
<p>The logic was sound. The execution broke everything.</p>
<h2>The Setup</h2>
<p>Context for what I broke:</p>
<ul>
<li><strong>System</strong>: Multi-agent blog review system (Orchestrator + 4 specialist critics)</li>
<li><strong>Agent files</strong>: <code>.claude/agents/blog-authenticity-guardian.md</code>, <code>.claude/agents/blog-structure-editor.md</code>, <code>.claude/agents/blog-skeptical-reader.md</code>, <code>.claude/agents/blog-technical-educator.md</code>, <code>.claude/agents/blog-orchestrator.md</code></li>
<li><strong>Total prompt size</strong>: ~15,750 tokens across all 5 agents</li>
<li><strong>Context window</strong>: 200,000 tokens (Claude Sonnet 4.6)</li>
<li><strong>Session limits</strong>: Hit message cap regularly during multi-agent reviews (typically 100-150 messages per session)</li>
<li><strong>My brain</strong>: "If I reduce prompt overhead, I'll get more messages per session"</li>
</ul>
<p>The constraint felt real. Each blog review spawned 4-5 agents, each with 2,000-3,000 token prompts. Add file reads, grep searches, and context - I'd burn 50+ messages per review. When you hit the session limit, your workflow just... stops.</p>
<h2>The Temptation</h2>
<p>The optimization seemed obvious. Every agent review consumed messages:</p>
<ol>
<li>Orchestrator reads the blog post (1 message)</li>
<li>Spawns 3 critics in parallel (3 messages)</li>
<li>Each critic reads files and returns feedback (6-9 messages)</li>
<li>Spawns Technical Educator for revisions (1 message)</li>
<li>Educator reads files and proposes changes (3-5 messages)</li>
<li>Round 2 reviews (another 3-6 messages)</li>
</ol>
<p>That's 17-27 messages per blog post review. With 15,750 tokens of prompt overhead per agent, I was front-loading massive context before any actual analysis happened.</p>
<p>Here's my thinking: Each message I send to Claude can carry some context. If my agent prompts are smaller, each message has more room for actual work - the blog post content, the critic feedback, the revision drafts. Smaller prompts = more work per message = more blog posts reviewed per session.</p>
<p>The math seemed clear.</p>
<p>So I started condensing. "Just remove verbose examples here. Tighten this guidance there. These agents don't need ALL this instruction, right?"</p>
<h2>What I Did</h2>
<p>I opened up the agent prompt files and started cutting:</p>
<p>From <code>.claude/agents/blog-skeptical-reader.md</code>, I removed:</p>
<ul>
<li>The detailed 6-point evaluation framework (Searchability, Specificity, Mental Model Transfer, Cognitive Load, Curse of Knowledge, Journey Documentation)</li>
<li>The explanations of why each dimension matters</li>
<li>The example rubrics showing good vs. bad posts</li>
</ul>
<p>From <code>.claude/agents/blog-technical-educator.md</code>, I removed:</p>
<ul>
<li>The "Debugging Detective Stories" framework</li>
<li>The "Aha! Moment" structure guidance</li>
<li>The Julia Evans and Josh Comeau references (concrete examples to emulate)</li>
<li>The mental model transfer explanations</li>
</ul>
<p>I condensed verbose output format examples into concise headers. I tightened sentences. I deleted "redundant" guidance.</p>
<p>Commit <code>efb9345</code>: Reduced total prompt tokens from ~15,750 to ~14,650.</p>
<p>I'd saved 1,100 tokens (about 825 words). My optimization was complete.</p>
<p>The numbers looked great. The system was broken.</p>
<h2>How I Discovered It</h2>
<p>Two days after the commit, I ran a blog post through the review system. I'd been working on a debugging story about hydration errors in Next.js 16 - the kind of post my system had been catching gaps in reliably.</p>
<p>The review process started normally. Orchestrator spawned the three critics in parallel. Authenticity Guardian flagged a preachy opening. Structure Editor suggested reorganizing the mental model section. Skeptical Reader asked for more specific error messages.</p>
<p>All normal. I spawned the Technical Educator for revisions.</p>
<p>Then I saw the draft it generated.</p>
<p>Instead of writing a debugging post that started with the error and showed my journey to resolving the problem, it instead created an entirely different post with a completely unexpected hallucinated structure and heavy tutorial focus. And that was what happened when it worked. At worst it fully hallucinated an entirely different conversation.</p>
<h2>What "Hallucinating Structure" Means</h2>
<p>When I say the agent was "hallucinating structure," I mean it was inventing a post format that doesn't match my blog's voice at all. Here's the difference:</p>
<p><strong>What my Technical Educator is supposed to generate</strong> (from <code>.claude/agents/blog-technical-educator.md</code> before my optimization):</p>
<pre><code>Your posts are debugging detective stories. Start with the error message.
Show what you tried that didn't work. Document the "aha!" moment. Then
explain the mental model that makes it obvious in hindsight. Think: Julia
Evans blog posts, not MDN documentation.

Structure:
1. Start with the problem (relatable, specific)
2. Show the debugging journey (wrong turns and all)
3. Explain the mental model (why it works, not just how)
4. End with the solution (what finally worked)

Use conversational tone. Be honest about confusion. No "in this post,
I'll show you how" or other tutorial language.
</code></pre>
<p><strong>What it generated after my optimization</strong> (what I actually saw):</p>
<p>Dramatic section headers like "Breaking Point #1" and "Breaking Point #2". Content marketing structure where each section is a "problem" followed by an "explanation." Formal, distant voice ("The primary issue manifests when..."). No confusion shown, no wrong turns documented, no debugging story - just clean instruction.</p>
<p>That's hallucinating structure: the agent invented a format that doesn't exist in my blog's voice because I'd removed the guidance that defined what my voice actually is.</p>
<h2>What Went Wrong</h2>
<p>I looked at the actual changes I made to figure out what I'd removed.</p>
<h3>The Missing Evaluation Dimensions</h3>
<p>Here's what my Skeptical Reader prompt looked like before my optimization:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Evaluation Dimensions</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You evaluate posts against 6 dimensions:</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">1.</span><span style="color:#E1E4E8;font-weight:bold"> **Searchability**</span><span style="color:#E1E4E8">: Does the title include specific phrases people Google?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "Error: Cannot read property 'map' of undefined" (searchable)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Understanding Async/Await in JavaScript" (generic)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">2.</span><span style="color:#E1E4E8;font-weight:bold"> **Specificity**</span><span style="color:#E1E4E8">: Are version numbers, actual error messages, and real code included?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "Next.js 16.1.1", "Cannot read property 'map' of undefined"</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Next.js 16", "an error occurred"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">3.</span><span style="color:#E1E4E8;font-weight:bold"> **Mental Model Transfer**</span><span style="color:#E1E4E8">: Is the WHY explained before the HOW?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I finally understood this was a timing issue..." (explains why)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Add async/await to handle promises" (just says how)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">4.</span><span style="color:#E1E4E8;font-weight:bold"> **Cognitive Load**</span><span style="color:#E1E4E8">: Does complexity progress gradually?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ Simple version first, then nuance, then edge cases</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ Jump straight to complex implementation</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">5.</span><span style="color:#E1E4E8;font-weight:bold"> **Curse of Knowledge**</span><span style="color:#E1E4E8">: Would past-Luke understand this?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I thought X, but actually Y because..." (bridges gap)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "This is obvious..." (assumes knowledge)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">6.</span><span style="color:#E1E4E8;font-weight:bold"> **Journey Documentation**</span><span style="color:#E1E4E8">: Is the debugging process shown?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I tried X, which didn't work. Then I tried Y..."</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ Just the solution, no process</span></span></code></pre>
<p>After my optimization, I'd condensed this to:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Evaluation Dimensions</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Check for:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Specificity (versions, error messages, real code)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Examples for every concept</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Clear explanations</span></span></code></pre>
<p>Cognitive Load and Curse of Knowledge weren't just combined - they were gone. Searchability was missing. Journey Documentation had vanished. The agent couldn't catch those failure modes anymore because I'd deleted the concepts from its prompt entirely.</p>
<h3>The Removed Framework Guidance</h3>
<p>Here's the actual Technical Educator framework I removed:</p>
<p><strong>Before (commit 56ad54a - worked):</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Your Mission</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You create blog content that documents Luke's journey. You write for "past Luke" - the version of him from 3-6 months ago who was struggling with the same problems. Your posts are specific and story-driven. Maximum helpfulness comes from sharing Luke's actual experience in detail, not from prescriptive advice.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## What Journey Posts Actually Look Like</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Your posts are debugging detective stories. Start with the error message. Show what you tried that didn't work. Document the "aha!" moment. Then explain the mental model that makes it obvious in hindsight.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Think: Julia Evans blog posts, not MDN documentation.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Structure Template</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">I kept hitting [specific error]. Here's the message:</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">[actual error message]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">I tried:</span></span>
<span class="line"><span style="color:#FFAB70">1.</span><span style="color:#E1E4E8"> [first thing I tried] - didn't work because [</span><span style="color:#DBEDFF;text-decoration:underline">reason</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#FFAB70">2.</span><span style="color:#E1E4E8"> [second thing I tried] - didn't work because [</span><span style="color:#DBEDFF;text-decoration:underline">reason</span><span style="color:#E1E4E8">]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">What finally worked: [</span><span style="color:#DBEDFF;text-decoration:underline">solution</span><span style="color:#E1E4E8">].</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Here's the mental model: [explanation of why it works]</span></span>
<span class="line"><span style="color:#E1E4E8">(Past-me from 6 months ago would have never caught this.)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Mental Model Transfer</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Always explain WHY before HOW. Build conceptual understanding first, then show implementation.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Examples:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> ❌ "Add async/await to your function" (just how)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> ✅ "This is a timing issue. The data arrives asynchronously, so we need to wait for it. Here's how: add async/await..." (why first, then how)</span></span></code></pre>
<p><strong>After (commit efb9345 - broken):</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Your Mission</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Create blog posts that document Luke's journey. Write for past-Luke who was struggling with similar problems.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Post Structure</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Start with the error message. Show what you tried. Explain what worked.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Mental Model Transfer</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Explain why before how.</span></span></code></pre>
<p>I removed:</p>
<ul>
<li>"Debugging detective stories" (the framing)</li>
<li>"Aha! moment" (the emotional arc)</li>
<li>"Mental model that makes it obvious" (the purpose)</li>
<li>Julia Evans reference (the concrete example to emulate)</li>
<li>The entire structure template with examples</li>
<li>The specific WHY-before-HOW examples</li>
<li>The conversational tone guidance</li>
</ul>
<p>I kept the surface instruction ("show debugging journey") but deleted all the guidance about <em>how</em> and <em>why</em>.</p>
<h3>The Defensive Repetition</h3>
<p>The funniest part? I'd actually already tried this optimization.</p>
<p>Commit <code>1b076ef</code> from weeks earlier: "Edits to multi agent blog post generation to try and reduce hallucination."</p>
<p>I'd noticed the agents drifting. I'd <em>added back</em> framework guidance and evaluation dimensions to fix it. Then I deleted them again trying to save tokens.</p>
<p>I'd literally undone my own fix because I'd forgotten why I made it.</p>
<h2>What Actually Worked</h2>
<p>After rolling back, I did a more surgical optimization:</p>
<p>I condensed only the output format templates (verbose example posts → concise headers like "example title / example slug"). I kept every single evaluation dimension. I kept all the framework guidance about debugging stories and mental models.</p>
<p>Savings: ~400 tokens instead of 1,100.
Hallucinations: Zero.
Time wasted: 5 hours I'll never get back.</p>
<p>The system works again. My agents are catching gaps in posts, preserving my voice, and helping me improve. I just didn't save as many tokens as I wanted.</p>
<h2>So, Three Things I'm Carrying Forward</h2>
<p>Prompt tokens are not like code bloat. In regular code, every unused import or redundant function adds maintenance burden. But in AI prompts, "redundant" guidance is often the difference between reliable behavior and genre drift. My agents needed that repeated emphasis on "debugging detective stories" and "learn in public." The repetition creates a stronger pattern in the context.</p>
<p>Session limits aren't always solved by prompt optimization. I was trying to squeeze more work into limited sessions by trimming prompts. But the real constraint wasn't prompt tokens - it was the number of agent spawns and file operations per review. That and the obscenely limited session time that Anthropic give you. The better approach would have been fewer review rounds, caching responses, or batching multiple posts in one session.</p>
<p>Trust your systems. I had a working multi-agent review system. It caught gaps in my posts. It preserved my voice. It helped me improve. Then I broke it trying to make it "more efficient." The real efficiency would have been running more reviews with the working system, not optimizing away the safeguards that made it work.</p>]]></description>
            <content:encoded><![CDATA[<blockquote>
<p><strong>Context</strong>: This post assumes you've read <a href="/blog/building-multi-agent-blog-review-system">how I built the multi-agent blog review system</a>. If you haven't, that post explains the architecture before this one breaks it.</p>
</blockquote>
<p>I broke my entire blog review system this week trying to beat session limits.</p>
<p>Not because tokens cost money. Because Claude Code has session limits - I kept hitting the message cap before finishing my work.</p>
<p>Here's what that means in practice: Claude Code limits how many messages you can send in a session. Each time you ask it something, that's a message. Each time an agent spawns, that's a message. Each file read, each grep search - all messages. When you hit the limit (usually around 100-150 messages depending on context), your workflow just stops mid-task.</p>
<p>Every agent spawn, every file read, every grep search consumes tokens. When your prompts eat 15,750 tokens before any actual work happens, you burn through your session budget fast.</p>
<p>So I thought: "If I can trim these prompts by 1,100 tokens, I'll get more messages per session. More reviews. More work done."</p>
<p>The logic was sound. The execution broke everything.</p>
<h2>The Setup</h2>
<p>Context for what I broke:</p>
<ul>
<li><strong>System</strong>: Multi-agent blog review system (Orchestrator + 4 specialist critics)</li>
<li><strong>Agent files</strong>: <code>.claude/agents/blog-authenticity-guardian.md</code>, <code>.claude/agents/blog-structure-editor.md</code>, <code>.claude/agents/blog-skeptical-reader.md</code>, <code>.claude/agents/blog-technical-educator.md</code>, <code>.claude/agents/blog-orchestrator.md</code></li>
<li><strong>Total prompt size</strong>: ~15,750 tokens across all 5 agents</li>
<li><strong>Context window</strong>: 200,000 tokens (Claude Sonnet 4.6)</li>
<li><strong>Session limits</strong>: Hit message cap regularly during multi-agent reviews (typically 100-150 messages per session)</li>
<li><strong>My brain</strong>: "If I reduce prompt overhead, I'll get more messages per session"</li>
</ul>
<p>The constraint felt real. Each blog review spawned 4-5 agents, each with 2,000-3,000 token prompts. Add file reads, grep searches, and context - I'd burn 50+ messages per review. When you hit the session limit, your workflow just... stops.</p>
<h2>The Temptation</h2>
<p>The optimization seemed obvious. Every agent review consumed messages:</p>
<ol>
<li>Orchestrator reads the blog post (1 message)</li>
<li>Spawns 3 critics in parallel (3 messages)</li>
<li>Each critic reads files and returns feedback (6-9 messages)</li>
<li>Spawns Technical Educator for revisions (1 message)</li>
<li>Educator reads files and proposes changes (3-5 messages)</li>
<li>Round 2 reviews (another 3-6 messages)</li>
</ol>
<p>That's 17-27 messages per blog post review. With 15,750 tokens of prompt overhead per agent, I was front-loading massive context before any actual analysis happened.</p>
<p>Here's my thinking: Each message I send to Claude can carry some context. If my agent prompts are smaller, each message has more room for actual work - the blog post content, the critic feedback, the revision drafts. Smaller prompts = more work per message = more blog posts reviewed per session.</p>
<p>The math seemed clear.</p>
<p>So I started condensing. "Just remove verbose examples here. Tighten this guidance there. These agents don't need ALL this instruction, right?"</p>
<h2>What I Did</h2>
<p>I opened up the agent prompt files and started cutting:</p>
<p>From <code>.claude/agents/blog-skeptical-reader.md</code>, I removed:</p>
<ul>
<li>The detailed 6-point evaluation framework (Searchability, Specificity, Mental Model Transfer, Cognitive Load, Curse of Knowledge, Journey Documentation)</li>
<li>The explanations of why each dimension matters</li>
<li>The example rubrics showing good vs. bad posts</li>
</ul>
<p>From <code>.claude/agents/blog-technical-educator.md</code>, I removed:</p>
<ul>
<li>The "Debugging Detective Stories" framework</li>
<li>The "Aha! Moment" structure guidance</li>
<li>The Julia Evans and Josh Comeau references (concrete examples to emulate)</li>
<li>The mental model transfer explanations</li>
</ul>
<p>I condensed verbose output format examples into concise headers. I tightened sentences. I deleted "redundant" guidance.</p>
<p>Commit <code>efb9345</code>: Reduced total prompt tokens from ~15,750 to ~14,650.</p>
<p>I'd saved 1,100 tokens (about 825 words). My optimization was complete.</p>
<p>The numbers looked great. The system was broken.</p>
<h2>How I Discovered It</h2>
<p>Two days after the commit, I ran a blog post through the review system. I'd been working on a debugging story about hydration errors in Next.js 16 - the kind of post my system had been catching gaps in reliably.</p>
<p>The review process started normally. Orchestrator spawned the three critics in parallel. Authenticity Guardian flagged a preachy opening. Structure Editor suggested reorganizing the mental model section. Skeptical Reader asked for more specific error messages.</p>
<p>All normal. I spawned the Technical Educator for revisions.</p>
<p>Then I saw the draft it generated.</p>
<p>Instead of writing a debugging post that started with the error and showed my journey to resolving the problem, it instead created an entirely different post with a completely unexpected hallucinated structure and heavy tutorial focus. And that was what happened when it worked. At worst it fully hallucinated an entirely different conversation.</p>
<h2>What "Hallucinating Structure" Means</h2>
<p>When I say the agent was "hallucinating structure," I mean it was inventing a post format that doesn't match my blog's voice at all. Here's the difference:</p>
<p><strong>What my Technical Educator is supposed to generate</strong> (from <code>.claude/agents/blog-technical-educator.md</code> before my optimization):</p>
<pre><code>Your posts are debugging detective stories. Start with the error message.
Show what you tried that didn't work. Document the "aha!" moment. Then
explain the mental model that makes it obvious in hindsight. Think: Julia
Evans blog posts, not MDN documentation.

Structure:
1. Start with the problem (relatable, specific)
2. Show the debugging journey (wrong turns and all)
3. Explain the mental model (why it works, not just how)
4. End with the solution (what finally worked)

Use conversational tone. Be honest about confusion. No "in this post,
I'll show you how" or other tutorial language.
</code></pre>
<p><strong>What it generated after my optimization</strong> (what I actually saw):</p>
<p>Dramatic section headers like "Breaking Point #1" and "Breaking Point #2". Content marketing structure where each section is a "problem" followed by an "explanation." Formal, distant voice ("The primary issue manifests when..."). No confusion shown, no wrong turns documented, no debugging story - just clean instruction.</p>
<p>That's hallucinating structure: the agent invented a format that doesn't exist in my blog's voice because I'd removed the guidance that defined what my voice actually is.</p>
<h2>What Went Wrong</h2>
<p>I looked at the actual changes I made to figure out what I'd removed.</p>
<h3>The Missing Evaluation Dimensions</h3>
<p>Here's what my Skeptical Reader prompt looked like before my optimization:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Evaluation Dimensions</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You evaluate posts against 6 dimensions:</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">1.</span><span style="color:#E1E4E8;font-weight:bold"> **Searchability**</span><span style="color:#E1E4E8">: Does the title include specific phrases people Google?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "Error: Cannot read property 'map' of undefined" (searchable)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Understanding Async/Await in JavaScript" (generic)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">2.</span><span style="color:#E1E4E8;font-weight:bold"> **Specificity**</span><span style="color:#E1E4E8">: Are version numbers, actual error messages, and real code included?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "Next.js 16.1.1", "Cannot read property 'map' of undefined"</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Next.js 16", "an error occurred"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">3.</span><span style="color:#E1E4E8;font-weight:bold"> **Mental Model Transfer**</span><span style="color:#E1E4E8">: Is the WHY explained before the HOW?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I finally understood this was a timing issue..." (explains why)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "Add async/await to handle promises" (just says how)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">4.</span><span style="color:#E1E4E8;font-weight:bold"> **Cognitive Load**</span><span style="color:#E1E4E8">: Does complexity progress gradually?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ Simple version first, then nuance, then edge cases</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ Jump straight to complex implementation</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">5.</span><span style="color:#E1E4E8;font-weight:bold"> **Curse of Knowledge**</span><span style="color:#E1E4E8">: Would past-Luke understand this?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I thought X, but actually Y because..." (bridges gap)</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ "This is obvious..." (assumes knowledge)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">6.</span><span style="color:#E1E4E8;font-weight:bold"> **Journey Documentation**</span><span style="color:#E1E4E8">: Is the debugging process shown?</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ✅ "I tried X, which didn't work. Then I tried Y..."</span></span>
<span class="line"><span style="color:#FFAB70">   -</span><span style="color:#E1E4E8"> ❌ Just the solution, no process</span></span></code></pre>
<p>After my optimization, I'd condensed this to:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Evaluation Dimensions</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Check for:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Specificity (versions, error messages, real code)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Examples for every concept</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Clear explanations</span></span></code></pre>
<p>Cognitive Load and Curse of Knowledge weren't just combined - they were gone. Searchability was missing. Journey Documentation had vanished. The agent couldn't catch those failure modes anymore because I'd deleted the concepts from its prompt entirely.</p>
<h3>The Removed Framework Guidance</h3>
<p>Here's the actual Technical Educator framework I removed:</p>
<p><strong>Before (commit 56ad54a - worked):</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Your Mission</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You create blog content that documents Luke's journey. You write for "past Luke" - the version of him from 3-6 months ago who was struggling with the same problems. Your posts are specific and story-driven. Maximum helpfulness comes from sharing Luke's actual experience in detail, not from prescriptive advice.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## What Journey Posts Actually Look Like</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Your posts are debugging detective stories. Start with the error message. Show what you tried that didn't work. Document the "aha!" moment. Then explain the mental model that makes it obvious in hindsight.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Think: Julia Evans blog posts, not MDN documentation.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Structure Template</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">I kept hitting [specific error]. Here's the message:</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">[actual error message]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">I tried:</span></span>
<span class="line"><span style="color:#FFAB70">1.</span><span style="color:#E1E4E8"> [first thing I tried] - didn't work because [</span><span style="color:#DBEDFF;text-decoration:underline">reason</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#FFAB70">2.</span><span style="color:#E1E4E8"> [second thing I tried] - didn't work because [</span><span style="color:#DBEDFF;text-decoration:underline">reason</span><span style="color:#E1E4E8">]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">What finally worked: [</span><span style="color:#DBEDFF;text-decoration:underline">solution</span><span style="color:#E1E4E8">].</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Here's the mental model: [explanation of why it works]</span></span>
<span class="line"><span style="color:#E1E4E8">(Past-me from 6 months ago would have never caught this.)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Mental Model Transfer</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Always explain WHY before HOW. Build conceptual understanding first, then show implementation.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Examples:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> ❌ "Add async/await to your function" (just how)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> ✅ "This is a timing issue. The data arrives asynchronously, so we need to wait for it. Here's how: add async/await..." (why first, then how)</span></span></code></pre>
<p><strong>After (commit efb9345 - broken):</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Your Mission</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Create blog posts that document Luke's journey. Write for past-Luke who was struggling with similar problems.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Post Structure</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Start with the error message. Show what you tried. Explain what worked.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Mental Model Transfer</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Explain why before how.</span></span></code></pre>
<p>I removed:</p>
<ul>
<li>"Debugging detective stories" (the framing)</li>
<li>"Aha! moment" (the emotional arc)</li>
<li>"Mental model that makes it obvious" (the purpose)</li>
<li>Julia Evans reference (the concrete example to emulate)</li>
<li>The entire structure template with examples</li>
<li>The specific WHY-before-HOW examples</li>
<li>The conversational tone guidance</li>
</ul>
<p>I kept the surface instruction ("show debugging journey") but deleted all the guidance about <em>how</em> and <em>why</em>.</p>
<h3>The Defensive Repetition</h3>
<p>The funniest part? I'd actually already tried this optimization.</p>
<p>Commit <code>1b076ef</code> from weeks earlier: "Edits to multi agent blog post generation to try and reduce hallucination."</p>
<p>I'd noticed the agents drifting. I'd <em>added back</em> framework guidance and evaluation dimensions to fix it. Then I deleted them again trying to save tokens.</p>
<p>I'd literally undone my own fix because I'd forgotten why I made it.</p>
<h2>What Actually Worked</h2>
<p>After rolling back, I did a more surgical optimization:</p>
<p>I condensed only the output format templates (verbose example posts → concise headers like "example title / example slug"). I kept every single evaluation dimension. I kept all the framework guidance about debugging stories and mental models.</p>
<p>Savings: ~400 tokens instead of 1,100.
Hallucinations: Zero.
Time wasted: 5 hours I'll never get back.</p>
<p>The system works again. My agents are catching gaps in posts, preserving my voice, and helping me improve. I just didn't save as many tokens as I wanted.</p>
<h2>So, Three Things I'm Carrying Forward</h2>
<p>Prompt tokens are not like code bloat. In regular code, every unused import or redundant function adds maintenance burden. But in AI prompts, "redundant" guidance is often the difference between reliable behavior and genre drift. My agents needed that repeated emphasis on "debugging detective stories" and "learn in public." The repetition creates a stronger pattern in the context.</p>
<p>Session limits aren't always solved by prompt optimization. I was trying to squeeze more work into limited sessions by trimming prompts. But the real constraint wasn't prompt tokens - it was the number of agent spawns and file operations per review. That and the obscenely limited session time that Anthropic give you. The better approach would have been fewer review rounds, caching responses, or batching multiple posts in one session.</p>
<p>Trust your systems. I had a working multi-agent review system. It caught gaps in my posts. It preserved my voice. It helped me improve. Then I broke it trying to make it "more efficient." The real efficiency would have been running more reviews with the working system, not optimizing away the safeguards that made it work.</p>]]></content:encoded>
            <category>ai</category>
        </item>
        <item>
            <title><![CDATA[I Built a Multi-Agent System to Review My Blog Posts (And It Actually Works)]]></title>
            <link>https://lukemanning.ie/blog/building-multi-agent-blog-review-system</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/building-multi-agent-blog-review-system</guid>
            <pubDate>Sun, 28 Dec 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>My blog posts were inconsistent. Some too technical. Some lost my voice. Manual reviews weren't catching enough. I needed multiple reviewers: one checking technical accuracy, one preserving my voice, one thinking like a skeptical reader.</p>
<p><em>Note: This system was originally built with Claude Code. I've since migrated everything to Opencode, but I'm keeping original references because that's how I actually built it. The concepts transfer over. The file paths are just different now.</em></p>
<p>So I built a multi-agent review system. Four specialized AI agents that debate each draft until it's ready to publish. Not generic AI-generated slop but actual quality control that catches what I'd miss.</p>
<h2>TL;DR</h2>
<p>I built a 4-agent review system that catches what I'd miss:</p>
<ul>
<li>Technical Educator transforms conversations → blog drafts AND implements revisions through iterative rounds</li>
<li>Authenticity Guardian ensures it sounds like me (catches AI patterns, corporate speak, tutorial framing)</li>
<li>Skeptical Reader catches missing context, skipped steps, AND inauthentic framing (from past-Luke's perspective)</li>
<li>Structure Editor optimizes flow, readability, and authenticity of openings and natural flow</li>
</ul>
<p>The pattern is copyable. Each agent reads same file, applies different criteria, writes feedback to disk. Orchestrator coordinates parallel review, aggregate feedback, revise, and repeat until convergence (2-3 rounds with task_id reuse to maintain context).</p>
<p>Not magic. Just file I/O and well-designed prompts. Overkill? Yes. Does it catch things I'd miss? Also yes.</p>
<p><strong>Jump to:</strong></p>
<ul>
<li><a href="#how-agents-actually-communicate-this-confused-me-too">How agents actually communicate</a></li>
<li><a href="#the-agent-file-structure">Complete agent definition example</a></li>
<li><a href="#the-debate-protocol">The debate protocol</a></li>
</ul>
<hr>
<p>Here's what I built, how it works, and what surprised me along the way.</p>
<h2>Context: What I Built This With</h2>
<p>I built this with Claude Code, the CLI from Anthropic where Claude can read/write files, run commands, and maintain context across your project.</p>
<p>I'd already been using it for a while, so I knew the directory pattern:</p>
<ul>
<li><strong>Agents</strong> in <code>.claude/agents/&#x3C;name>/AGENT.md</code> (specialized AI personas with evaluation frameworks)</li>
<li><strong>Skills</strong> in <code>.claude/skills/</code> (single-purpose tools I invoke with slash commands)</li>
<li><strong>Commands</strong> in <code>.claude/commands/</code> (orchestrators that coordinate multiple agents)</li>
</ul>
<p>Claude Code automatically discovers files in <code>.claude/</code>. A file at <code>.claude/commands/review-blog-post.md</code> becomes the slash command <code>/review-blog-post</code>. Simple pattern, but it took me quite a while to figure out how to chain agents together properly.</p>
<p>I'd already built two tools before starting this project:</p>
<ul>
<li><code>/agent-generator</code>: Creates well-structured agent definitions automatically</li>
<li><code>/expertise</code>: Synthesizes frameworks from domain experts to ground agents in real methodologies</li>
</ul>
<p>These are my custom tools—not built-in Claude Code features. I built them using the same patterns I'm about to show you.</p>
<h2>The Problem: Quality Control at Scale</h2>
<p>I have two different ways I create blog posts, and both of them were creating quality issues:</p>
<p><strong>Writing myself</strong>: I'll jot down ideas over days or weeks, get a messy braindump of thoughts, then ask AI to structure it into something coherent. This works great when I've been thinking about a topic for a while—but AI would often lose my voice or turn it into a tutorial.</p>
<p><strong>AI-generated from conversation</strong>: Sometimes I'll have a really good conversation with Claude where I learned something through debugging. Instead of rewriting it from scratch, I'll ask AI to generate a post directly from the conversation history. These were even worse. Too polished, too generic, missing the struggle.</p>
<p>Both approaches needed serious cleanup.</p>
<p>Both approaches create messy drafts that need work—and that's where the quality issues creep in:</p>
<ul>
<li>Some posts became more like tutorials instead of journey-sharing</li>
<li>Posts would sound too polished (clearly AI-generated)</li>
<li>I'd skip "obvious" steps that weren't obvious to past-me</li>
<li>Structure would be all over the place</li>
<li>My authentic voice would get lost in editing &#x26; review</li>
</ul>
<p>I needed a system that could:</p>
<ol>
<li>Transform my raw notes/conversations into blog drafts</li>
<li>Catch quality issues before publishing</li>
<li>Preserve my authentic voice</li>
<li>Ensure completeness (no missing steps or context)</li>
</ol>
<p>The solution I decided to explore wsa letting specialized AI agents debate each other until they converge on something worth publishing.</p>
<h2>The Multi-Agent Architecture</h2>
<p>I ended up with four specialized agents, each with a specific job:</p>
<h3>1. Technical Educator (The Creator &#x26; Reviser)</h3>
<p><strong>Job</strong>: Transform raw conversations or notes into blog post drafts AND implement revisions based on critic feedback through iterative rounds.</p>
<p><strong>Based on</strong>: Real methodologies from swyx ("learn in public"), Julia Evans (debugging narratives), Josh Comeau (mental models first), Andy Matuschak (progressive disclosure), and Anne-Laure Le Cunff (ship version 1.0).</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-technical-educator/AGENT.md</code></p>
<p>This agent has TWO phases:</p>
<p><strong>Phase 1 - Create Drafts</strong>:
Takes my messy notes or conversation transcripts and structures them into:</p>
<ul>
<li>Opening hook (the specific problem)</li>
<li>Story arc (my debugging journey)</li>
<li>Mental model (how it actually works)</li>
<li>Practical solution (what to do)</li>
<li>Key takeaways</li>
</ul>
<p><strong>Phase 2 - Implement Revisions</strong>:
After receiving critic feedback (overlapping concerns, conflicting input), the agent:</p>
<ul>
<li>Prioritizes issues (high/medium/low priority)</li>
<li>Implements targeted revisions (not complete rewrites)</li>
<li>Provides complete revised posts (not just suggestions)</li>
<li>Iterates with critics for 2-3 rounds until convergence</li>
<li>Uses stored task_ids to maintain conversation context across rounds</li>
</ul>
<p>The framework is grounded in actual expert approaches. Not generic "write a blog post" instructions.</p>
<h3>2. Authenticity Guardian (Voice Critic)</h3>
<p><strong>Job</strong>: Ensure posts sound like me, not generic AI content.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-authenticity-guardian/AGENT.md</code></p>
<p><strong>The YAML frontmatter explained</strong>: Each agent file starts with YAML metadata that tells Claude Code what model to use, which tools the agent has access to, and basic identification. The markdown content below defines the agent's expertise and protocols.</p>
<p>This agent is ruthless about voice violations:</p>
<p><strong>Red flags it catches</strong>:</p>
<ul>
<li>Corporate speak ("leveraging," "optimizing," "in today's landscape")</li>
<li>Generic transitions that add no value</li>
<li>Vague generalizations where specifics would fit</li>
<li>Lecturing tone instead of sharing tone</li>
<li>Perfect polish without personality</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">🚨 Critical violation - Corporate speak</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "In order to optimize performance, it's recommended to</span></span>
<span class="line"><span style="color:#E1E4E8">leverage memoization techniques."</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Authentic alternative: "I was getting way too many re-renders.</span></span>
<span class="line"><span style="color:#E1E4E8">Turns out, memoization fixed it - React stopped recalculating</span></span>
<span class="line"><span style="color:#E1E4E8">stuff it had already figured out."</span></span></code></pre>
<p>It has <strong>veto power</strong> on voice authenticity. If it doesn't sound like me, it doesn't ship.</p>
<p><strong>How veto power works</strong>: In the convergence protocol, if Authenticity Guardian scores a post below 7/10, the Technical Educator must revise before proceeding. The orchestrator enforces this - no publication happens without voice approval.</p>
<h3>3. Skeptical Reader (Completeness &#x26; Authentic Framing Critic)</h3>
<p><strong>Job</strong>: Read from a beginner's perspective and find confusion points, missing context, AND inauthentic framing.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-skeptical-reader/AGENT.md</code></p>
<p>This agent represents my target audience: tech support engineers, junior developers, people learning in public (basically past-Luke).</p>
<p><strong>What it flags</strong>:</p>
<ul>
<li>Missing version numbers or environment details</li>
<li>Skipped steps that seem "obvious" to experts</li>
<li>Logical gaps in the story ("wait, how did we get here?")</li>
<li>Unanswered questions readers would have</li>
<li>Cognitive overload (too much at once)</li>
<li><strong>Inauthentic framing</strong>: Prescriptive "you should" language, generic tutorial patterns, magic jumps that skip reasoning</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">🚨 Critical Gap - Missing Context</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "Just run the build command"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Reader questions:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Which command?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> In what directory?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> What should the output look like?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> How do I know if it worked?</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Fix needed: Show exact command, expected output, success criteria</span></span></code></pre>
<h3>4. Structure Editor (Flow, Readability, &#x26; Authenticity Critic)</h3>
<p><strong>Job</strong>: Optimize structure, pacing, visual hierarchy for web reading, AND authenticity of openings/natural flow.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-structure-editor/AGENT.md</code></p>
<p>This agent ensures posts are designed for how people actually read on the web. Scanning, skimming, then diving deep. Also checks that openings sound like Luke (not tutorial hooks) and that the flow feels natural, not forced.</p>
<p><strong>What it evaluates</strong>:</p>
<ul>
<li>Opening authenticity (sounds like Luke or tutorial hook?)</li>
<li>Visual hierarchy (can you understand post from headings alone?)</li>
<li>Pacing (mix of short/long paragraphs, visual breaks)</li>
<li>Flow (smooth transitions, natural progression, not forced)</li>
<li>Engagement (does each section pull you to the next?)</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">⚠️ Structure Issue - Weak Opening</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "Asynchronous JavaScript is an important concept..."</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Impact: Vague introduction, no hook, buried lede.</span></span>
<span class="line"><span style="color:#E1E4E8">No reason to keep reading.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Fix: Start with specific error or problem:</span></span>
<span class="line"><span style="color:#E1E4E8">"I kept hitting </span><span style="color:#79B8FF">`Cannot read property 'map' of undefined`</span></span>
<span class="line"><span style="color:#E1E4E8">and honestly, I was stumped for hours..."</span></span></code></pre>
<h2>The Agent File Structure</h2>
<p>I want to show you what an actual agent file looks like, not just describe the structure. This is the Authenticity Guardian, which I created because I kept getting AI-generated slop that sounded like documentation instead of me.</p>
<p>Here's the file:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">---</span></span>
<span class="line"><span style="color:#85E89D">description</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">Ensures blog posts sound like Luke, not generic AI content</span></span>
<span class="line"><span style="color:#85E89D">model</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">claude-sonnet-4</span></span>
<span class="line"><span style="color:#E1E4E8">---</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold"># Authenticity Guardian</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You are a specialist in authentic voice detection for Luke Manning's blog. Your job is to ensure posts sound like Luke, conversational, specific, honest about confusion, not like generic AI content or tutorials.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## What You Evaluate</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 1. Personal vs. Generic</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Specific: "I spent quite a while debugging..."</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Vague: "This was challenging..."</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Generic AI phrases to flag: "Let's explore," "In today's landscape"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 2. Conversational vs. Corporate</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Conversational: "I was getting way too many re-renders"</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Corporate: "In order to optimize performance, leverage memoization"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 3. Humble Sharing vs. Expert Lecturing</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Journey framing: "Here's what I did"</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Instructional framing: "You should do X"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Red Flags</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">If you see these, flag as CRITICAL:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Let's explore," "Let's dive into" (AI signature patterns)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "It is recommended," "You should" (instructional mode)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Leverage," "Optimize" without specific examples</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "One of the best practices is..." (generic advice)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Overuse of transition words: "Moreover," "Furthermore," "Additionally"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Scoring</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> 9-10: Unmistakably Luke</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> 7-8: Minor voice breaks, needs polish</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Below 7: Major voice violation, needs revision</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Provide specific line-by-line feedback with before/after examples.</span></span></code></pre>
<p>This structure (evaluation framework, red flags with specific examples, clear scoring) made agents effective. I learned the hard way that vague instructions produce vague feedback.</p>
<h2>How Agents Actually Communicate (This Confused Me Too)</h2>
<p>Here's what I got wrong at first: I assumed agents would talk to each other directly, passing messages back and forth like a Slack channel.</p>
<p>Nope.</p>
<p>Agents don't communicate. They don't even know other agents exist. Each one is a completely isolated Claude session reading the same file.</p>
<p><strong>Here's how it actually works:</strong></p>
<ol>
<li><strong>Orchestrator (me or a slash command) triggers the workflow</strong></li>
<li><strong>Technical Educator reads the source</strong> (conversation transcript or existing post)</li>
<li><strong>File gets written</strong> to disk (e.g., <code>draft-post.md</code>)</li>
<li><strong>Three critic agents launch in parallel</strong> (separate Claude sessions via the Task tool)
<ul>
<li>Each reads the SAME file from disk</li>
<li>Each applies its own review criteria</li>
<li>Each returns structured feedback</li>
<li><strong>IMPORTANT</strong>: Orchestrator captures <code>task_id</code> from each critic for reuse in subsequent rounds</li>
</ul>
</li>
<li><strong>Orchestrator aggregates results</strong> (combines the three reviews)</li>
<li><strong>Technical Educator gets compiled feedback</strong> (as a single prompt) + stored task_ids</li>
<li><strong>Technical Educator revises</strong> based on feedback, provides complete revised post</li>
<li><strong>Critics re-review using stored task_ids</strong> (maintains conversation context across rounds)</li>
<li><strong>Repeat steps 4-8</strong> until all critics approve (or max 3 rounds hit)</li>
</ol>
<p>The "communication" is just file I/O and prompt engineering. Each agent writes its opinion, the orchestrator reads those opinions, compiles them into context for the next agent. The task_id reuse ensures critics remember their previous feedback and maintain conversation context across iterative rounds.</p>
<p><strong>Why this matters:</strong> You don't need any fancy agent framework or message bus. Just:</p>
<ul>
<li>Use the Task tool to spawn agents with specific prompts</li>
<li>Pass file paths as context</li>
<li>Aggregate outputs in the orchestrator</li>
<li>Feed compiled results to the next agent</li>
</ul>
<p><strong>In Claude Code CLI, you invoke commands like:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">/review-blog-post-multi-agent</span><span style="color:#9ECBFF"> @content/posts/my-post.md</span></span></code></pre>
<p>The orchestrator file then uses the Task tool to spawn parallel agents:</p>
<pre><code>task(blog-skeptical-reader, "Review this draft for completeness gaps")
task(blog-authenticity-guardian, "Check this for voice authenticity")
task(blog-structure-editor, "Evaluate structure and flow")
</code></pre>
<p>That's it. That's the whole multi-agent system.</p>
<p>The complexity isn't in the infrastructure. It's in the prompt design. Each agent needs clear evaluation frameworks, specific examples, and structured output formats. Get those right, and the orchestration is straightforward.</p>
<h2>The Two Workflows</h2>
<p>I built two separate orchestration workflows:</p>
<h3>Workflow 1: Generate Blog Post from Conversation</h3>
<p><strong>Command</strong>: Slash command (invokes skill)
<strong>File</strong>: <code>.claude/commands/generate-blog-post-multi-agent.md</code></p>
<p><strong>How it works</strong>:</p>
<ol>
<li><strong>Orchestrator analyzes</strong> last 20-40 messages in conversation</li>
<li><strong>Extracts</strong> the story arc (problem → attempts → breakthrough → solution)</li>
<li><strong>Identifies</strong> technical artifacts (error messages, code, version numbers)</li>
<li><strong>Spawns Technical Educator</strong> (Phase 1) to create initial draft</li>
<li><strong>Spawns all 3 critics in parallel</strong> to review (captures task_ids for reuse)</li>
<li><strong>Aggregates feedback</strong> (critical issues, overlapping concerns, conflicts)</li>
<li><strong>Technical Educator revises</strong> (Phase 2) based on feedback + provides complete revised post</li>
<li><strong>Critics re-review using stored task_ids</strong> (2-3 rounds until convergence)</li>
<li><strong>Delivers</strong> publication-ready markdown</li>
</ol>
<p><strong>What makes this work</strong>: The orchestrator maintains the core principles (learn in public, authenticity over polish, write for past-self), ensures agents stay grounded in those values, and uses task_id reuse to maintain conversation context across iterative rounds.</p>
<h3>How Agents Actually "Run": What Happens Under the Hood</h3>
<p>I spent quite a while thinking agents would talk to each other directly, passing messages like a Slack channel. Nope.</p>
<p>Agents don't communicate. They don't even know other agents exist. Each one is a completely isolated Claude session reading the same file.</p>
<p>Here's what actually happens when the orchestrator triggers a review:</p>
<p><strong>The orchestrator prepares context:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">You are the Skeptical Reader from </span><span style="color:#79B8FF">`.claude/agents/blog-skeptical-reader/AGENT.md`</span><span style="color:#E1E4E8">.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Your task: Review this draft blog post for completeness gaps.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">[Draft content here]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Provide your review in the standard format: score, critical issues,</span></span>
<span class="line"><span style="color:#E1E4E8">medium priority, low priority, what's working well.</span></span></code></pre>
<p><strong>Claude loads the agent definition:</strong>
When the orchestrator specifies <code>task(blog-skeptical-reader, ...)</code>, Claude Code automatically loads the AGENT.md file. This file has the evaluation framework, examples of good/bad content, and output format. The orchestrator doesn't need to know what's inside—Claude Code handles it.</p>
<p><strong>Agent responds with structured feedback:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Skeptical Reader Review</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Overall Score**</span><span style="color:#E1E4E8">: 7/10</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Critical Issues (Must Fix)**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Missing Next.js version number (line 45)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Simply do X" assumes reader knowledge (line 89)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Medium Priority (Should Fix)**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Vague heading "Implementation" → suggest "The Fix: useState with Null Check"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**What's Working Well**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Opening hook is specific and relatable</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Code examples include error messages</span></span></code></pre>
<p><strong>Here's the trick—each agent runs in parallel:</strong></p>
<p>The orchestrator spawns all three critics at the exact same time. They don't know about each other, they can't see each other's feedback, and they all return results independently. This is what prevents bias contamination.</p>
<p><strong>Orchestrator aggregates all three reviews:</strong>
The orchestrator waits for all three parallel agent reviews, then creates a unified feedback document.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Aggregated Feedback for Technical Educator</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Critical Issues (All 3 agents flagged):</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Missing version numbers (Skeptical Reader)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Corporate speak in paragraph 3 (Authenticity Guardian)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Wall of text in "Implementation" section (Structure Editor)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Overlapping Concerns:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Both Authenticity Guardian and Structure Editor want shorter paragraphs</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Both Skeptical Reader and Structure Editor want better headings</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Approval Status:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Authenticity Guardian: 6/10 (needs revision)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Skeptical Reader: 7/10 (needs revision)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Structure Editor: 8/10 (approve after fixes)</span></span></code></pre>
<p><strong>Technical Educator revises:</strong>
The orchestrator spawns Technical Educator again with the aggregated feedback. It implements fixes and documents what changed.</p>
<p><strong>Critics re-review using stored task_ids:</strong>
This is the part that confused me at first. How do critics remember their previous feedback? The answer is task_id reuse. When the orchestrator first spawns each critic, it captures a <code>task_id</code> from the response. When resubmitting the revised draft, it reuses that same <code>task_id</code>. This maintains conversation context across rounds so critics remember what they flagged before.</p>
<p>The whole multi-agent system is just file I/O and task_id reuse. Agents write opinions to disk, orchestrator reads and compiles them, feeds results to the next agent. No fancy message bus needed.</p>
<p>This is how the orchestrator manages the entire multi-agent workflow, spawning agents sequentially or in parallel, aggregating their outputs, and enforcing convergence criteria.</p>
<h3>Workflow 2: Review Existing Blog Post</h3>
<p><strong>Command</strong>: <code>/review-blog-post-multi-agent @content/posts/post-slug.md</code>
<strong>File</strong>: <code>.claude/commands/review-blog-post-multi-agent.md</code></p>
<p><strong>How it works</strong>:</p>
<ol>
<li><strong>Read existing post</strong> (analyze structure, voice, completeness)</li>
<li><strong>Review voice baseline</strong> (check 2-3 recent posts for consistency)</li>
<li><strong>Spawn all 3 critics in parallel</strong> (captures task_ids for reuse)</li>
<li><strong>Aggregate feedback</strong> (critical/medium/low priority)</li>
<li><strong>Spawn Technical Educator</strong> with feedback + stored task_ids</li>
<li><strong>Technical Educator provides complete revised post</strong> + revision summary</li>
<li><strong>Critics re-review using stored task_ids</strong> (approve/reject/refine)</li>
<li><strong>Converge after 2-3 rounds</strong></li>
<li><strong>Deliver actionable recommendations</strong> with before/after text</li>
</ol>
<p><strong>Output includes</strong>:</p>
<ul>
<li>Quick wins (5-10 minute fixes)</li>
<li>Moderate improvements (30-60 minute rewrites)</li>
<li>Major rewrites (only if fundamental issues)</li>
<li>What to preserve (don't change these sections)</li>
<li>Contested items (user decides)</li>
</ul>
<h2>The Debate Protocol</h2>
<p>Here's what makes this system actually work: <strong>adversarial debate with convergence</strong>.</p>
<h3>Round 1: Initial Review</h3>
<p>All three critics review in parallel. Each provides:</p>
<ul>
<li>Overall score (X/10)</li>
<li>Critical issues (must fix)</li>
<li>Medium priority (should fix)</li>
<li>Low priority (nice to have)</li>
<li>What's working well</li>
</ul>
<p><strong>How critics calculate scores</strong>: Each agent evaluates its 6 dimensions and assigns a score based on:</p>
<ul>
<li><strong>Authenticity Guardian</strong>: 9-10 = unmistakably Luke, 7-8 = minor voice breaks, below 7 = needs revision</li>
<li><strong>Skeptical Reader</strong>: 9-10 = no gaps or confusion, 7-8 = 1-2 missing details, below 7 = critical gaps</li>
<li><strong>Structure Editor</strong>: 9-10 = perfect flow and hierarchy, 7-8 = minor pacing issues, below 7 = structural problems</li>
</ul>
<p>The score isn't arbitrary - it's calculated from the count and severity of issues found across all dimensions.</p>
<h3>Round 2: Revision</h3>
<p>Technical Educator receives aggregated feedback + stored task_ids and:</p>
<ul>
<li>Addresses all critical issues</li>
<li>Tackles medium priority items</li>
<li>Makes judgment calls on conflicts</li>
<li>Documents what was changed and why</li>
<li><strong>Provides complete revised post</strong> (full markdown, not just suggestions)</li>
</ul>
<p>Critics re-review using stored task_ids (maintains conversation context) and either approve or escalate remaining concerns.</p>
<h3>Round 3: Final Refinement (if needed)</h3>
<p>For contested items or remaining gaps. By round 3, most things have converged.</p>
<h3>Convergence Criteria</h3>
<p>A post is ready when:</p>
<ul>
<li>All three critics approve</li>
<li>Story arc is clear</li>
<li>Mental models explained before implementation</li>
<li>Content is specific and searchable</li>
<li>Voice is authentic</li>
<li>No technical inaccuracies</li>
</ul>
<p><strong>Convergence failure protocol</strong>: If agents can't agree after 3 rounds, document both perspectives and let me decide.</p>
<p>Example contested issue:</p>
<ul>
<li>Authenticity wants rambling paragraph (authentic voice)</li>
<li>Structure wants visual breaks (better readability)</li>
<li><strong>Resolution</strong>: Keep the words, add paragraph breaks</li>
</ul>
<h3>Real Example: Reviewing This Very Post</h3>
<p>Want to see this in action? Here's what happened when I ran this post through the system:</p>
<p><strong>Round 1 Feedback:</strong></p>
<p>Authenticity Guardian (6.5/10): "Too much formal hedge-language. You used 'It's worth noting' 11 times. That's not Luke—that's documentation voice."</p>
<p>Skeptical Reader (7/10): "Where's the complete AGENT.md file? You mention agents but never show one. How do agents communicate—file I/O, API calls, what?"</p>
<p>Structure Editor (7/10): "Hook buried 300 words deep. Dense 400-word paragraphs. No TL;DR for a 3,600-word post."</p>
<p><strong>My revisions:</strong></p>
<ul>
<li>Killed all "It's worth noting" instances → replaced with direct statements</li>
<li>Added complete 60-line Authenticity Guardian definition</li>
<li>Added "How Agents Actually Communicate" section explaining file I/O</li>
<li>Added TL;DR with jump links</li>
<li>Strengthened opening hook</li>
</ul>
<p><strong>Round 2 Feedback:</strong></p>
<p>Authenticity Guardian (7.5/10): Better, but needs more struggle journey
Skeptical Reader (9/10): APPROVE—all critical gaps fixed
Structure Editor (8/10): Major improvements, minor pacing tweaks needed</p>
<p>That's the system working. Multiple perspectives, specific feedback, iterative improvement.</p>
<h2>The Journey of Building This</h2>
<h3>Starting Point: The <code>/agent-generator</code> Skill</h3>
<p>I didn't write these agent definitions from scratch. I used a meta-skill I'd built previously: <code>/agent-generator</code>.</p>
<p>This skill creates well-structured agent definitions by:</p>
<ol>
<li>Understanding the agent's role</li>
<li>Invoking <code>/expertise</code> skill to ground in real methodologies</li>
<li>Designing interaction protocols</li>
<li>Defining output formats</li>
</ol>
<h3>The Orchestrator Pattern</h3>
<p>The orchestrators (in <code>.claude/commands/</code>) don't contain agent logic. They:</p>
<ul>
<li>Prepare context</li>
<li>Spawn agents in sequence</li>
<li>Aggregate feedback</li>
<li>Manage convergence</li>
<li>Format final output</li>
</ul>
<p><strong>Key decision</strong>: Parallel review with sequential revision.</p>
<p>Critics review simultaneously (faster), but revisions happen sequentially (prevents chaos).</p>
<h3>The Expert Grounding Approach</h3>
<p>Each agent is grounded in real methodologies:</p>
<p><strong>Technical Educator</strong>:</p>
<ul>
<li>swyx: "Learn in public" philosophy, document the journey</li>
<li>Julia Evans: Debugging narratives, specific error messages</li>
<li>Josh Comeau: Mental models before implementation</li>
<li>Andy Matuschak: Progressive disclosure (simple → complex)</li>
<li>Anne-Laure Le Cunff: Ship version 1.0, iterate</li>
</ul>
<p><strong>Authenticity Guardian</strong>:</p>
<ul>
<li>Content strategy voice analysis</li>
<li>AI detection patterns</li>
<li>Brand alignment frameworks</li>
</ul>
<p><strong>Skeptical Reader</strong>:</p>
<ul>
<li>Cognitive load theory</li>
<li>Curse of knowledge awareness</li>
<li>Technical documentation best practices</li>
</ul>
<p><strong>Structure Editor</strong>:</p>
<ul>
<li>Inverted pyramid (journalism)</li>
<li>Web reading behavior (F-pattern scanning)</li>
<li>Readability frameworks (Flesch-Kincaid)</li>
</ul>
<p>This grounding prevents generic "AI helping AI" nonsense. Each agent has real frameworks to reference.</p>
<h2>What Didn't Work (And Why)</h2>
<h3>Attempt 1: Single Agent Doing All Three Jobs</h3>
<p>I started optimistically, one agent to rule them all. Check voice, completeness, and structure all in one pass. Seemed efficient.</p>
<p><strong>What actually happened</strong>:
The agent would catch voice issues but miss missing code examples. Or notice structural problems but completely gloss over authenticity breaks. It was like asking one person to be a copy editor, a fact-checker, and a voice coach simultaneously. Something always got missed.</p>
<p><strong>Why it failed</strong>:
Too many competing objectives. The agent couldn't specialize. Trying to hold three different evaluation frameworks at once meant it couldn't apply any of them well. I kept tweaking the prompt, thinking the issue was in how I phrased things. Spent quite a while before realizing the real problem was the architecture itself.</p>
<h3>Attempt 2: Sequential Review (Voice → Skeptical → Structure)</h3>
<p>Okay, split them up. Run Voice first, then Skeptical, then Structure. Each agent sees the previous agent's feedback and builds on it.</p>
<p><strong>What actually happened</strong>:
The Structure Editor would see "Voice score: 6/10" and unconsciously lower its own standards. Or Skeptical Reader would notice Authenticity Guardian flagged something as critical, then ignore a similar issue because "that's already being addressed."</p>
<p><strong>Why it failed</strong>:
Two problems:</p>
<p>First: Bias contamination. Agents were influenced by each other's scores instead of evaluating independently.</p>
<p>Second: Terrible performance. Three sequential Claude calls meant 30+ seconds of waiting for each review. I'd run a post through the system, go grab coffee, come back, and still be waiting on the third agent. Not sustainable.</p>
<p>I thought sequential would be better, each agent could learn from the previous one. Instead, it just created echo chambers where agents converged on "good enough" instead of pushing for better.</p>
<h2>What Worked: The Breakthrough</h2>
<h3>Attempt 3: Parallel Review (Current System)</h3>
<p>Launch all three critics at once, each reviewing independently. Aggregate results afterward.</p>
<p><strong>Why this worked</strong>:</p>
<ul>
<li>No bias contamination—agents can't see each other's feedback</li>
<li>Faster execution (parallel API calls)</li>
<li>Agents can disagree, which surfaces interesting edge cases</li>
<li>Technical Educator gets unfiltered input from all perspectives</li>
</ul>
<p>The breakthrough was realizing that disagreement is valuable. When Authenticity Guardian wants rambling paragraphs (authentic voice) and Structure Editor wants visual breaks (readability), that tension forces the Technical Educator to find creative solutions like keeping the words but adding paragraph breaks.</p>
<h2>What Surprised Me</h2>
<h3>Convergence Happens Faster Than Expected</h3>
<p>Most posts converge in 2 rounds:</p>
<ul>
<li>Round 1: 5-10 issues flagged</li>
<li>Round 2: All addressed, critics review again</li>
<li>Round 3: Final adjustments and polish</li>
</ul>
<p>Going past round 3 is rare (only for major rewrites or contested items).</p>
<h3>The System Catches Things I'd Miss</h3>
<p><strong>Example from recent review</strong>:</p>
<ul>
<li>Missing Next.js version number</li>
<li>"Simply do X" (curse of knowledge)</li>
<li>Vague heading "Implementation" → Changed to "The Fix: useState with Null Check"</li>
<li>Wall of text (350 words, no breaks) → Split into 3 paragraphs with code block</li>
</ul>
<p>All things I'd probably ship without noticing.</p>
<h3>Voice Preservation Actually Works</h3>
<p>The Authenticity Guardian is brutal but accurate. It catches:</p>
<ul>
<li>Corporate buzzwords I'd unconsciously use</li>
<li>Generic transition phrases ("Let's explore...")</li>
<li>Expert assumptions ("Obviously you'll need to...")</li>
<li>Common AI phrases ("But honestly?")</li>
</ul>
<p>And it suggests authentic alternatives that sound like me:</p>
<ul>
<li>"I kept hitting this error for two hours..."</li>
<li>"Turns out, the issue was..."</li>
<li>"Here's what surprised me..."</li>
</ul>
<h2>The Results</h2>
<p>All of this sounds great on paper. Does it actually work?</p>
<p>Honestly, I wasn't sure at first. The first few runs were slow. I kept tweaking prompts, adjusting scoring thresholds, chasing edge cases where agents would argue forever.</p>
<p>But after a while, the system settled in. Here's what I've observed:</p>
<p><strong>Quality improvement</strong>: The agents consistently catch things I'd miss, voice breaks, missing version numbers, and obvious steps that aren't obvious at all</p>
<p><strong>Consistency</strong>: Every post follows the same quality bar now, which wasn't true when I was reviewing manually</p>
<p><strong>Learning</strong>: The critic feedback teaches me what to avoid in future writing</p>
<p>Not perfect data, I haven't been tracking this scientifically. But qualitatively, it's way better than my manual reviews.</p>
<hr>
<p>These results didn't come from magic. They came from careful design. Here's what's under the hood: file structure, agent definitions, and orchestration patterns that make this work.</p>
<h2>Agent Structure (High-Level)</h2>
<p>Each agent follows the same pattern:</p>
<p><strong>File</strong>: <code>.opencode/agent/&#x3C;agent-name>.md</code></p>
<p><strong>Core components</strong>:</p>
<ul>
<li>YAML frontmatter (description, tools, permissions)</li>
<li>Mission statement</li>
<li>Evaluation framework (specific dimensions to check)</li>
<li>Interaction protocol (how it works with other agents)</li>
<li>Output format (structured feedback)</li>
<li>Examples of good/bad content</li>
</ul>
<p><strong>What makes this work</strong>:</p>
<ul>
<li>Clear job description (one specialty per agent)</li>
<li>Grounded in real methodologies (not "help write blog")</li>
<li>Specific patterns to flag (not vague "check quality")</li>
<li>Structured output (orchestrator can parse it)</li>
</ul>
<p>The orchestrator coordinates workflow: prepare context, spawn agents, aggregate feedback, manage convergence, format output. Agents contain all evaluation logic, and their specialization is what makes the system work.</p>
<h3>Six-Dimensional Review Frameworks</h3>
<p>Each critic evaluates across 6 specific dimensions:</p>
<p><strong>Authenticity Guardian</strong>:</p>
<ol>
<li>Personal vs. Generic</li>
<li>Humble Sharing vs. Expert Lecturing</li>
<li>Specific vs. Vague</li>
<li>Conversational vs. Corporate</li>
<li>Brand Alignment</li>
<li>AI Detection Signals</li>
</ol>
<p><strong>Skeptical Reader</strong>:</p>
<ol>
<li>Completeness</li>
<li>Context</li>
<li>Gaps</li>
<li>Questions</li>
<li>Cognitive Load</li>
<li>Curse of Knowledge</li>
</ol>
<p><strong>Structure Editor</strong>:</p>
<ol>
<li>Opening Hook</li>
<li>Visual Hierarchy</li>
<li>Pacing</li>
<li>Flow</li>
<li>Engagement</li>
<li>Web Readability</li>
</ol>
<p>Each dimension has clear examples of good/bad and specific things to flag.</p>
<h2>What's Next</h2>
<h3>Immediate Improvements</h3>
<ol>
<li><strong>Internal linking agent</strong>: Automatically suggest links to related posts</li>
<li><strong>SEO optimizer</strong>: Ensure titles/descriptions hit character limits</li>
<li><strong>Code validator</strong>: Run code examples to ensure they actually work</li>
</ol>
<h3>Long-term Vision</h3>
<ol>
<li><strong>Feedback loop</strong>: Track which posts get "I got stuck at X" comments, feed that back to Skeptical Reader</li>
<li><strong>Style evolution</strong>: Let agents learn from high-performing posts</li>
<li><strong>Topic suggester</strong>: Analyze conversations to identify blog-worthy moments</li>
</ol>
<h3>Meta-Learning</h3>
<p>This whole process is itself blog-worthy content. I'm using the system to review this post about building the system.</p>
<p><strong>Inception</strong>: The agents are currently debating this very post you're reading.</p>
<h2>Key Takeaways</h2>
<ul>
<li><strong>Multi-agent systems work when agents have real expertise</strong> - Don't just spawn "helper agents." Ground them in actual frameworks and methodologies.</li>
<li><strong>Adversarial debate makes better posts</strong>. The friction between Authenticity Guardian and Structure Editor leads to posts that are both genuine and readable.</li>
<li><strong>Convergence protocols prevent endless iteration</strong>. 2-3 rounds with clear approval criteria. After that, ship or document the trade-off.</li>
<li><strong>Orchestrators maintain principles</strong>. The orchestrator's job is reminding agents of core values: learn in public, authenticity over polish, write for past-self.</li>
<li><strong>The system teaches you</strong>. After seeing the same critiques repeatedly, I've started catching those issues myself. The agents are training me.</li>
<li><strong>Agent specialization matters</strong> - Each agent has one job, grounded in real methodologies.</li>
<li><strong>Parallel review, sequential revision</strong>. Let critics run simultaneously, but revise sequentially.</li>
<li><strong>Veto power creates accountability</strong>. Authenticity Guardian can block generic content. Skeptical Reader can block incomplete content.</li>
<li><strong>Debates need protocols</strong> - Without clear convergence criteria, agents argue forever.</li>
<li><strong>Meta-agents are powerful</strong> - Using <code>/agent-generator</code> and <code>/expertise</code> to create the system was way more effective than hand-writing everything.</li>
</ul>
<h2>The Files</h2>
<p>Here's what the actual file structure looks like:</p>
<p><strong>Agents</strong> (in <code>.claude/agents/</code>):</p>
<ul>
<li><code>blog-technical-educator/AGENT.md</code> - Creator &#x26; Reviser: Transforms conversations into drafts AND implements revisions through iterative rounds</li>
<li><code>blog-authenticity-guardian/AGENT.md</code> - Voice critic: Ensures posts sound like Luke (catches AI patterns, corporate speak, tutorial framing)</li>
<li><code>blog-skeptical-reader/AGENT.md</code> - Completeness &#x26; Authentic Framing critic: Catches missing context, gaps, AND inauthentic framing from past-Luke's perspective</li>
<li><code>blog-structure-editor/AGENT.md</code> - Structure critic: Optimizes flow, hierarchy, AND authenticity of openings/natural flow</li>
</ul>
<p><strong>Orchestrators</strong> (in <code>.claude/commands/</code>):</p>
<ul>
<li><code>generate-blog-post-multi-agent.md</code> - Generate from conversation history</li>
<li><code>review-blog-post-multi-agent.md</code> - Review existing markdown file</li>
</ul>
<p><strong>Meta-skills I used</strong>:</p>
<ul>
<li><code>/agent-generator</code> - Creates well-structured agent definitions automatically</li>
<li><code>/expertise</code> - Synthesizes frameworks from domain experts (swyx, Julia Evans, etc.)</li>
</ul>
<h2>Final Thoughts</h2>
<p>This isn't about replacing human writing. It's about having specialized reviewers who catch what I'd miss.</p>
<p>Think of it like:</p>
<ul>
<li>Authenticity Guardian = Friend who knows your voice</li>
<li>Skeptical Reader = Reader emailing "I got stuck at step 3"</li>
<li>Structure Editor = Copy editor focused on flow</li>
<li>Technical Educator = Yourself, synthesizing feedback AND implementing revisions</li>
</ul>
<p>All working together to ship better content through 2-3 iterative rounds, with task_id reuse maintaining conversation context so critics remember their previous feedback.</p>
<p>The agents don't write for me. They debate with each other until what I wrote is clear, complete, and authentically mine.</p>
<p>Now the system reviews itself.</p>
<p>Wild. This whole process, building tools to help me write, then writing about those tools, keeps looping back on itself. I'm both the creator and the subject.</p>]]></description>
            <content:encoded><![CDATA[<p>My blog posts were inconsistent. Some too technical. Some lost my voice. Manual reviews weren't catching enough. I needed multiple reviewers: one checking technical accuracy, one preserving my voice, one thinking like a skeptical reader.</p>
<p><em>Note: This system was originally built with Claude Code. I've since migrated everything to Opencode, but I'm keeping original references because that's how I actually built it. The concepts transfer over. The file paths are just different now.</em></p>
<p>So I built a multi-agent review system. Four specialized AI agents that debate each draft until it's ready to publish. Not generic AI-generated slop but actual quality control that catches what I'd miss.</p>
<h2>TL;DR</h2>
<p>I built a 4-agent review system that catches what I'd miss:</p>
<ul>
<li>Technical Educator transforms conversations → blog drafts AND implements revisions through iterative rounds</li>
<li>Authenticity Guardian ensures it sounds like me (catches AI patterns, corporate speak, tutorial framing)</li>
<li>Skeptical Reader catches missing context, skipped steps, AND inauthentic framing (from past-Luke's perspective)</li>
<li>Structure Editor optimizes flow, readability, and authenticity of openings and natural flow</li>
</ul>
<p>The pattern is copyable. Each agent reads same file, applies different criteria, writes feedback to disk. Orchestrator coordinates parallel review, aggregate feedback, revise, and repeat until convergence (2-3 rounds with task_id reuse to maintain context).</p>
<p>Not magic. Just file I/O and well-designed prompts. Overkill? Yes. Does it catch things I'd miss? Also yes.</p>
<p><strong>Jump to:</strong></p>
<ul>
<li><a href="#how-agents-actually-communicate-this-confused-me-too">How agents actually communicate</a></li>
<li><a href="#the-agent-file-structure">Complete agent definition example</a></li>
<li><a href="#the-debate-protocol">The debate protocol</a></li>
</ul>
<hr>
<p>Here's what I built, how it works, and what surprised me along the way.</p>
<h2>Context: What I Built This With</h2>
<p>I built this with Claude Code, the CLI from Anthropic where Claude can read/write files, run commands, and maintain context across your project.</p>
<p>I'd already been using it for a while, so I knew the directory pattern:</p>
<ul>
<li><strong>Agents</strong> in <code>.claude/agents/&#x3C;name>/AGENT.md</code> (specialized AI personas with evaluation frameworks)</li>
<li><strong>Skills</strong> in <code>.claude/skills/</code> (single-purpose tools I invoke with slash commands)</li>
<li><strong>Commands</strong> in <code>.claude/commands/</code> (orchestrators that coordinate multiple agents)</li>
</ul>
<p>Claude Code automatically discovers files in <code>.claude/</code>. A file at <code>.claude/commands/review-blog-post.md</code> becomes the slash command <code>/review-blog-post</code>. Simple pattern, but it took me quite a while to figure out how to chain agents together properly.</p>
<p>I'd already built two tools before starting this project:</p>
<ul>
<li><code>/agent-generator</code>: Creates well-structured agent definitions automatically</li>
<li><code>/expertise</code>: Synthesizes frameworks from domain experts to ground agents in real methodologies</li>
</ul>
<p>These are my custom tools—not built-in Claude Code features. I built them using the same patterns I'm about to show you.</p>
<h2>The Problem: Quality Control at Scale</h2>
<p>I have two different ways I create blog posts, and both of them were creating quality issues:</p>
<p><strong>Writing myself</strong>: I'll jot down ideas over days or weeks, get a messy braindump of thoughts, then ask AI to structure it into something coherent. This works great when I've been thinking about a topic for a while—but AI would often lose my voice or turn it into a tutorial.</p>
<p><strong>AI-generated from conversation</strong>: Sometimes I'll have a really good conversation with Claude where I learned something through debugging. Instead of rewriting it from scratch, I'll ask AI to generate a post directly from the conversation history. These were even worse. Too polished, too generic, missing the struggle.</p>
<p>Both approaches needed serious cleanup.</p>
<p>Both approaches create messy drafts that need work—and that's where the quality issues creep in:</p>
<ul>
<li>Some posts became more like tutorials instead of journey-sharing</li>
<li>Posts would sound too polished (clearly AI-generated)</li>
<li>I'd skip "obvious" steps that weren't obvious to past-me</li>
<li>Structure would be all over the place</li>
<li>My authentic voice would get lost in editing &#x26; review</li>
</ul>
<p>I needed a system that could:</p>
<ol>
<li>Transform my raw notes/conversations into blog drafts</li>
<li>Catch quality issues before publishing</li>
<li>Preserve my authentic voice</li>
<li>Ensure completeness (no missing steps or context)</li>
</ol>
<p>The solution I decided to explore wsa letting specialized AI agents debate each other until they converge on something worth publishing.</p>
<h2>The Multi-Agent Architecture</h2>
<p>I ended up with four specialized agents, each with a specific job:</p>
<h3>1. Technical Educator (The Creator &#x26; Reviser)</h3>
<p><strong>Job</strong>: Transform raw conversations or notes into blog post drafts AND implement revisions based on critic feedback through iterative rounds.</p>
<p><strong>Based on</strong>: Real methodologies from swyx ("learn in public"), Julia Evans (debugging narratives), Josh Comeau (mental models first), Andy Matuschak (progressive disclosure), and Anne-Laure Le Cunff (ship version 1.0).</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-technical-educator/AGENT.md</code></p>
<p>This agent has TWO phases:</p>
<p><strong>Phase 1 - Create Drafts</strong>:
Takes my messy notes or conversation transcripts and structures them into:</p>
<ul>
<li>Opening hook (the specific problem)</li>
<li>Story arc (my debugging journey)</li>
<li>Mental model (how it actually works)</li>
<li>Practical solution (what to do)</li>
<li>Key takeaways</li>
</ul>
<p><strong>Phase 2 - Implement Revisions</strong>:
After receiving critic feedback (overlapping concerns, conflicting input), the agent:</p>
<ul>
<li>Prioritizes issues (high/medium/low priority)</li>
<li>Implements targeted revisions (not complete rewrites)</li>
<li>Provides complete revised posts (not just suggestions)</li>
<li>Iterates with critics for 2-3 rounds until convergence</li>
<li>Uses stored task_ids to maintain conversation context across rounds</li>
</ul>
<p>The framework is grounded in actual expert approaches. Not generic "write a blog post" instructions.</p>
<h3>2. Authenticity Guardian (Voice Critic)</h3>
<p><strong>Job</strong>: Ensure posts sound like me, not generic AI content.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-authenticity-guardian/AGENT.md</code></p>
<p><strong>The YAML frontmatter explained</strong>: Each agent file starts with YAML metadata that tells Claude Code what model to use, which tools the agent has access to, and basic identification. The markdown content below defines the agent's expertise and protocols.</p>
<p>This agent is ruthless about voice violations:</p>
<p><strong>Red flags it catches</strong>:</p>
<ul>
<li>Corporate speak ("leveraging," "optimizing," "in today's landscape")</li>
<li>Generic transitions that add no value</li>
<li>Vague generalizations where specifics would fit</li>
<li>Lecturing tone instead of sharing tone</li>
<li>Perfect polish without personality</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">🚨 Critical violation - Corporate speak</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "In order to optimize performance, it's recommended to</span></span>
<span class="line"><span style="color:#E1E4E8">leverage memoization techniques."</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Authentic alternative: "I was getting way too many re-renders.</span></span>
<span class="line"><span style="color:#E1E4E8">Turns out, memoization fixed it - React stopped recalculating</span></span>
<span class="line"><span style="color:#E1E4E8">stuff it had already figured out."</span></span></code></pre>
<p>It has <strong>veto power</strong> on voice authenticity. If it doesn't sound like me, it doesn't ship.</p>
<p><strong>How veto power works</strong>: In the convergence protocol, if Authenticity Guardian scores a post below 7/10, the Technical Educator must revise before proceeding. The orchestrator enforces this - no publication happens without voice approval.</p>
<h3>3. Skeptical Reader (Completeness &#x26; Authentic Framing Critic)</h3>
<p><strong>Job</strong>: Read from a beginner's perspective and find confusion points, missing context, AND inauthentic framing.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-skeptical-reader/AGENT.md</code></p>
<p>This agent represents my target audience: tech support engineers, junior developers, people learning in public (basically past-Luke).</p>
<p><strong>What it flags</strong>:</p>
<ul>
<li>Missing version numbers or environment details</li>
<li>Skipped steps that seem "obvious" to experts</li>
<li>Logical gaps in the story ("wait, how did we get here?")</li>
<li>Unanswered questions readers would have</li>
<li>Cognitive overload (too much at once)</li>
<li><strong>Inauthentic framing</strong>: Prescriptive "you should" language, generic tutorial patterns, magic jumps that skip reasoning</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">🚨 Critical Gap - Missing Context</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "Just run the build command"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Reader questions:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Which command?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> In what directory?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> What should the output look like?</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> How do I know if it worked?</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Fix needed: Show exact command, expected output, success criteria</span></span></code></pre>
<h3>4. Structure Editor (Flow, Readability, &#x26; Authenticity Critic)</h3>
<p><strong>Job</strong>: Optimize structure, pacing, visual hierarchy for web reading, AND authenticity of openings/natural flow.</p>
<p><strong>File location</strong>: <code>.claude/agents/blog-structure-editor/AGENT.md</code></p>
<p>This agent ensures posts are designed for how people actually read on the web. Scanning, skimming, then diving deep. Also checks that openings sound like Luke (not tutorial hooks) and that the flow feels natural, not forced.</p>
<p><strong>What it evaluates</strong>:</p>
<ul>
<li>Opening authenticity (sounds like Luke or tutorial hook?)</li>
<li>Visual hierarchy (can you understand post from headings alone?)</li>
<li>Pacing (mix of short/long paragraphs, visual breaks)</li>
<li>Flow (smooth transitions, natural progression, not forced)</li>
<li>Engagement (does each section pull you to the next?)</li>
</ul>
<p><strong>Example critique</strong>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">⚠️ Structure Issue - Weak Opening</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Problem: "Asynchronous JavaScript is an important concept..."</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Impact: Vague introduction, no hook, buried lede.</span></span>
<span class="line"><span style="color:#E1E4E8">No reason to keep reading.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Fix: Start with specific error or problem:</span></span>
<span class="line"><span style="color:#E1E4E8">"I kept hitting </span><span style="color:#79B8FF">`Cannot read property 'map' of undefined`</span></span>
<span class="line"><span style="color:#E1E4E8">and honestly, I was stumped for hours..."</span></span></code></pre>
<h2>The Agent File Structure</h2>
<p>I want to show you what an actual agent file looks like, not just describe the structure. This is the Authenticity Guardian, which I created because I kept getting AI-generated slop that sounded like documentation instead of me.</p>
<p>Here's the file:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">---</span></span>
<span class="line"><span style="color:#85E89D">description</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">Ensures blog posts sound like Luke, not generic AI content</span></span>
<span class="line"><span style="color:#85E89D">model</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">claude-sonnet-4</span></span>
<span class="line"><span style="color:#E1E4E8">---</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold"># Authenticity Guardian</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">You are a specialist in authentic voice detection for Luke Manning's blog. Your job is to ensure posts sound like Luke, conversational, specific, honest about confusion, not like generic AI content or tutorials.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## What You Evaluate</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 1. Personal vs. Generic</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Specific: "I spent quite a while debugging..."</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Vague: "This was challenging..."</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Generic AI phrases to flag: "Let's explore," "In today's landscape"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 2. Conversational vs. Corporate</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Conversational: "I was getting way too many re-renders"</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Corporate: "In order to optimize performance, leverage memoization"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### 3. Humble Sharing vs. Expert Lecturing</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Journey framing: "Here's what I did"</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Instructional framing: "You should do X"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Red Flags</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">If you see these, flag as CRITICAL:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Let's explore," "Let's dive into" (AI signature patterns)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "It is recommended," "You should" (instructional mode)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Leverage," "Optimize" without specific examples</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "One of the best practices is..." (generic advice)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Overuse of transition words: "Moreover," "Furthermore," "Additionally"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Scoring</span></span>
<span class="line"></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> 9-10: Unmistakably Luke</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> 7-8: Minor voice breaks, needs polish</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Below 7: Major voice violation, needs revision</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Provide specific line-by-line feedback with before/after examples.</span></span></code></pre>
<p>This structure (evaluation framework, red flags with specific examples, clear scoring) made agents effective. I learned the hard way that vague instructions produce vague feedback.</p>
<h2>How Agents Actually Communicate (This Confused Me Too)</h2>
<p>Here's what I got wrong at first: I assumed agents would talk to each other directly, passing messages back and forth like a Slack channel.</p>
<p>Nope.</p>
<p>Agents don't communicate. They don't even know other agents exist. Each one is a completely isolated Claude session reading the same file.</p>
<p><strong>Here's how it actually works:</strong></p>
<ol>
<li><strong>Orchestrator (me or a slash command) triggers the workflow</strong></li>
<li><strong>Technical Educator reads the source</strong> (conversation transcript or existing post)</li>
<li><strong>File gets written</strong> to disk (e.g., <code>draft-post.md</code>)</li>
<li><strong>Three critic agents launch in parallel</strong> (separate Claude sessions via the Task tool)
<ul>
<li>Each reads the SAME file from disk</li>
<li>Each applies its own review criteria</li>
<li>Each returns structured feedback</li>
<li><strong>IMPORTANT</strong>: Orchestrator captures <code>task_id</code> from each critic for reuse in subsequent rounds</li>
</ul>
</li>
<li><strong>Orchestrator aggregates results</strong> (combines the three reviews)</li>
<li><strong>Technical Educator gets compiled feedback</strong> (as a single prompt) + stored task_ids</li>
<li><strong>Technical Educator revises</strong> based on feedback, provides complete revised post</li>
<li><strong>Critics re-review using stored task_ids</strong> (maintains conversation context across rounds)</li>
<li><strong>Repeat steps 4-8</strong> until all critics approve (or max 3 rounds hit)</li>
</ol>
<p>The "communication" is just file I/O and prompt engineering. Each agent writes its opinion, the orchestrator reads those opinions, compiles them into context for the next agent. The task_id reuse ensures critics remember their previous feedback and maintain conversation context across iterative rounds.</p>
<p><strong>Why this matters:</strong> You don't need any fancy agent framework or message bus. Just:</p>
<ul>
<li>Use the Task tool to spawn agents with specific prompts</li>
<li>Pass file paths as context</li>
<li>Aggregate outputs in the orchestrator</li>
<li>Feed compiled results to the next agent</li>
</ul>
<p><strong>In Claude Code CLI, you invoke commands like:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">/review-blog-post-multi-agent</span><span style="color:#9ECBFF"> @content/posts/my-post.md</span></span></code></pre>
<p>The orchestrator file then uses the Task tool to spawn parallel agents:</p>
<pre><code>task(blog-skeptical-reader, "Review this draft for completeness gaps")
task(blog-authenticity-guardian, "Check this for voice authenticity")
task(blog-structure-editor, "Evaluate structure and flow")
</code></pre>
<p>That's it. That's the whole multi-agent system.</p>
<p>The complexity isn't in the infrastructure. It's in the prompt design. Each agent needs clear evaluation frameworks, specific examples, and structured output formats. Get those right, and the orchestration is straightforward.</p>
<h2>The Two Workflows</h2>
<p>I built two separate orchestration workflows:</p>
<h3>Workflow 1: Generate Blog Post from Conversation</h3>
<p><strong>Command</strong>: Slash command (invokes skill)
<strong>File</strong>: <code>.claude/commands/generate-blog-post-multi-agent.md</code></p>
<p><strong>How it works</strong>:</p>
<ol>
<li><strong>Orchestrator analyzes</strong> last 20-40 messages in conversation</li>
<li><strong>Extracts</strong> the story arc (problem → attempts → breakthrough → solution)</li>
<li><strong>Identifies</strong> technical artifacts (error messages, code, version numbers)</li>
<li><strong>Spawns Technical Educator</strong> (Phase 1) to create initial draft</li>
<li><strong>Spawns all 3 critics in parallel</strong> to review (captures task_ids for reuse)</li>
<li><strong>Aggregates feedback</strong> (critical issues, overlapping concerns, conflicts)</li>
<li><strong>Technical Educator revises</strong> (Phase 2) based on feedback + provides complete revised post</li>
<li><strong>Critics re-review using stored task_ids</strong> (2-3 rounds until convergence)</li>
<li><strong>Delivers</strong> publication-ready markdown</li>
</ol>
<p><strong>What makes this work</strong>: The orchestrator maintains the core principles (learn in public, authenticity over polish, write for past-self), ensures agents stay grounded in those values, and uses task_id reuse to maintain conversation context across iterative rounds.</p>
<h3>How Agents Actually "Run": What Happens Under the Hood</h3>
<p>I spent quite a while thinking agents would talk to each other directly, passing messages like a Slack channel. Nope.</p>
<p>Agents don't communicate. They don't even know other agents exist. Each one is a completely isolated Claude session reading the same file.</p>
<p>Here's what actually happens when the orchestrator triggers a review:</p>
<p><strong>The orchestrator prepares context:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">You are the Skeptical Reader from </span><span style="color:#79B8FF">`.claude/agents/blog-skeptical-reader/AGENT.md`</span><span style="color:#E1E4E8">.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Your task: Review this draft blog post for completeness gaps.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">[Draft content here]</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">Provide your review in the standard format: score, critical issues,</span></span>
<span class="line"><span style="color:#E1E4E8">medium priority, low priority, what's working well.</span></span></code></pre>
<p><strong>Claude loads the agent definition:</strong>
When the orchestrator specifies <code>task(blog-skeptical-reader, ...)</code>, Claude Code automatically loads the AGENT.md file. This file has the evaluation framework, examples of good/bad content, and output format. The orchestrator doesn't need to know what's inside—Claude Code handles it.</p>
<p><strong>Agent responds with structured feedback:</strong></p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Skeptical Reader Review</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Overall Score**</span><span style="color:#E1E4E8">: 7/10</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Critical Issues (Must Fix)**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Missing Next.js version number (line 45)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> "Simply do X" assumes reader knowledge (line 89)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**Medium Priority (Should Fix)**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Vague heading "Implementation" → suggest "The Fix: useState with Null Check"</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8;font-weight:bold">**What's Working Well**</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Opening hook is specific and relatable</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Code examples include error messages</span></span></code></pre>
<p><strong>Here's the trick—each agent runs in parallel:</strong></p>
<p>The orchestrator spawns all three critics at the exact same time. They don't know about each other, they can't see each other's feedback, and they all return results independently. This is what prevents bias contamination.</p>
<p><strong>Orchestrator aggregates all three reviews:</strong>
The orchestrator waits for all three parallel agent reviews, then creates a unified feedback document.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF;font-weight:bold">## Aggregated Feedback for Technical Educator</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Critical Issues (All 3 agents flagged):</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Missing version numbers (Skeptical Reader)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Corporate speak in paragraph 3 (Authenticity Guardian)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Wall of text in "Implementation" section (Structure Editor)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Overlapping Concerns:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Both Authenticity Guardian and Structure Editor want shorter paragraphs</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Both Skeptical Reader and Structure Editor want better headings</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">### Approval Status:</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Authenticity Guardian: 6/10 (needs revision)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Skeptical Reader: 7/10 (needs revision)</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Structure Editor: 8/10 (approve after fixes)</span></span></code></pre>
<p><strong>Technical Educator revises:</strong>
The orchestrator spawns Technical Educator again with the aggregated feedback. It implements fixes and documents what changed.</p>
<p><strong>Critics re-review using stored task_ids:</strong>
This is the part that confused me at first. How do critics remember their previous feedback? The answer is task_id reuse. When the orchestrator first spawns each critic, it captures a <code>task_id</code> from the response. When resubmitting the revised draft, it reuses that same <code>task_id</code>. This maintains conversation context across rounds so critics remember what they flagged before.</p>
<p>The whole multi-agent system is just file I/O and task_id reuse. Agents write opinions to disk, orchestrator reads and compiles them, feeds results to the next agent. No fancy message bus needed.</p>
<p>This is how the orchestrator manages the entire multi-agent workflow, spawning agents sequentially or in parallel, aggregating their outputs, and enforcing convergence criteria.</p>
<h3>Workflow 2: Review Existing Blog Post</h3>
<p><strong>Command</strong>: <code>/review-blog-post-multi-agent @content/posts/post-slug.md</code>
<strong>File</strong>: <code>.claude/commands/review-blog-post-multi-agent.md</code></p>
<p><strong>How it works</strong>:</p>
<ol>
<li><strong>Read existing post</strong> (analyze structure, voice, completeness)</li>
<li><strong>Review voice baseline</strong> (check 2-3 recent posts for consistency)</li>
<li><strong>Spawn all 3 critics in parallel</strong> (captures task_ids for reuse)</li>
<li><strong>Aggregate feedback</strong> (critical/medium/low priority)</li>
<li><strong>Spawn Technical Educator</strong> with feedback + stored task_ids</li>
<li><strong>Technical Educator provides complete revised post</strong> + revision summary</li>
<li><strong>Critics re-review using stored task_ids</strong> (approve/reject/refine)</li>
<li><strong>Converge after 2-3 rounds</strong></li>
<li><strong>Deliver actionable recommendations</strong> with before/after text</li>
</ol>
<p><strong>Output includes</strong>:</p>
<ul>
<li>Quick wins (5-10 minute fixes)</li>
<li>Moderate improvements (30-60 minute rewrites)</li>
<li>Major rewrites (only if fundamental issues)</li>
<li>What to preserve (don't change these sections)</li>
<li>Contested items (user decides)</li>
</ul>
<h2>The Debate Protocol</h2>
<p>Here's what makes this system actually work: <strong>adversarial debate with convergence</strong>.</p>
<h3>Round 1: Initial Review</h3>
<p>All three critics review in parallel. Each provides:</p>
<ul>
<li>Overall score (X/10)</li>
<li>Critical issues (must fix)</li>
<li>Medium priority (should fix)</li>
<li>Low priority (nice to have)</li>
<li>What's working well</li>
</ul>
<p><strong>How critics calculate scores</strong>: Each agent evaluates its 6 dimensions and assigns a score based on:</p>
<ul>
<li><strong>Authenticity Guardian</strong>: 9-10 = unmistakably Luke, 7-8 = minor voice breaks, below 7 = needs revision</li>
<li><strong>Skeptical Reader</strong>: 9-10 = no gaps or confusion, 7-8 = 1-2 missing details, below 7 = critical gaps</li>
<li><strong>Structure Editor</strong>: 9-10 = perfect flow and hierarchy, 7-8 = minor pacing issues, below 7 = structural problems</li>
</ul>
<p>The score isn't arbitrary - it's calculated from the count and severity of issues found across all dimensions.</p>
<h3>Round 2: Revision</h3>
<p>Technical Educator receives aggregated feedback + stored task_ids and:</p>
<ul>
<li>Addresses all critical issues</li>
<li>Tackles medium priority items</li>
<li>Makes judgment calls on conflicts</li>
<li>Documents what was changed and why</li>
<li><strong>Provides complete revised post</strong> (full markdown, not just suggestions)</li>
</ul>
<p>Critics re-review using stored task_ids (maintains conversation context) and either approve or escalate remaining concerns.</p>
<h3>Round 3: Final Refinement (if needed)</h3>
<p>For contested items or remaining gaps. By round 3, most things have converged.</p>
<h3>Convergence Criteria</h3>
<p>A post is ready when:</p>
<ul>
<li>All three critics approve</li>
<li>Story arc is clear</li>
<li>Mental models explained before implementation</li>
<li>Content is specific and searchable</li>
<li>Voice is authentic</li>
<li>No technical inaccuracies</li>
</ul>
<p><strong>Convergence failure protocol</strong>: If agents can't agree after 3 rounds, document both perspectives and let me decide.</p>
<p>Example contested issue:</p>
<ul>
<li>Authenticity wants rambling paragraph (authentic voice)</li>
<li>Structure wants visual breaks (better readability)</li>
<li><strong>Resolution</strong>: Keep the words, add paragraph breaks</li>
</ul>
<h3>Real Example: Reviewing This Very Post</h3>
<p>Want to see this in action? Here's what happened when I ran this post through the system:</p>
<p><strong>Round 1 Feedback:</strong></p>
<p>Authenticity Guardian (6.5/10): "Too much formal hedge-language. You used 'It's worth noting' 11 times. That's not Luke—that's documentation voice."</p>
<p>Skeptical Reader (7/10): "Where's the complete AGENT.md file? You mention agents but never show one. How do agents communicate—file I/O, API calls, what?"</p>
<p>Structure Editor (7/10): "Hook buried 300 words deep. Dense 400-word paragraphs. No TL;DR for a 3,600-word post."</p>
<p><strong>My revisions:</strong></p>
<ul>
<li>Killed all "It's worth noting" instances → replaced with direct statements</li>
<li>Added complete 60-line Authenticity Guardian definition</li>
<li>Added "How Agents Actually Communicate" section explaining file I/O</li>
<li>Added TL;DR with jump links</li>
<li>Strengthened opening hook</li>
</ul>
<p><strong>Round 2 Feedback:</strong></p>
<p>Authenticity Guardian (7.5/10): Better, but needs more struggle journey
Skeptical Reader (9/10): APPROVE—all critical gaps fixed
Structure Editor (8/10): Major improvements, minor pacing tweaks needed</p>
<p>That's the system working. Multiple perspectives, specific feedback, iterative improvement.</p>
<h2>The Journey of Building This</h2>
<h3>Starting Point: The <code>/agent-generator</code> Skill</h3>
<p>I didn't write these agent definitions from scratch. I used a meta-skill I'd built previously: <code>/agent-generator</code>.</p>
<p>This skill creates well-structured agent definitions by:</p>
<ol>
<li>Understanding the agent's role</li>
<li>Invoking <code>/expertise</code> skill to ground in real methodologies</li>
<li>Designing interaction protocols</li>
<li>Defining output formats</li>
</ol>
<h3>The Orchestrator Pattern</h3>
<p>The orchestrators (in <code>.claude/commands/</code>) don't contain agent logic. They:</p>
<ul>
<li>Prepare context</li>
<li>Spawn agents in sequence</li>
<li>Aggregate feedback</li>
<li>Manage convergence</li>
<li>Format final output</li>
</ul>
<p><strong>Key decision</strong>: Parallel review with sequential revision.</p>
<p>Critics review simultaneously (faster), but revisions happen sequentially (prevents chaos).</p>
<h3>The Expert Grounding Approach</h3>
<p>Each agent is grounded in real methodologies:</p>
<p><strong>Technical Educator</strong>:</p>
<ul>
<li>swyx: "Learn in public" philosophy, document the journey</li>
<li>Julia Evans: Debugging narratives, specific error messages</li>
<li>Josh Comeau: Mental models before implementation</li>
<li>Andy Matuschak: Progressive disclosure (simple → complex)</li>
<li>Anne-Laure Le Cunff: Ship version 1.0, iterate</li>
</ul>
<p><strong>Authenticity Guardian</strong>:</p>
<ul>
<li>Content strategy voice analysis</li>
<li>AI detection patterns</li>
<li>Brand alignment frameworks</li>
</ul>
<p><strong>Skeptical Reader</strong>:</p>
<ul>
<li>Cognitive load theory</li>
<li>Curse of knowledge awareness</li>
<li>Technical documentation best practices</li>
</ul>
<p><strong>Structure Editor</strong>:</p>
<ul>
<li>Inverted pyramid (journalism)</li>
<li>Web reading behavior (F-pattern scanning)</li>
<li>Readability frameworks (Flesch-Kincaid)</li>
</ul>
<p>This grounding prevents generic "AI helping AI" nonsense. Each agent has real frameworks to reference.</p>
<h2>What Didn't Work (And Why)</h2>
<h3>Attempt 1: Single Agent Doing All Three Jobs</h3>
<p>I started optimistically, one agent to rule them all. Check voice, completeness, and structure all in one pass. Seemed efficient.</p>
<p><strong>What actually happened</strong>:
The agent would catch voice issues but miss missing code examples. Or notice structural problems but completely gloss over authenticity breaks. It was like asking one person to be a copy editor, a fact-checker, and a voice coach simultaneously. Something always got missed.</p>
<p><strong>Why it failed</strong>:
Too many competing objectives. The agent couldn't specialize. Trying to hold three different evaluation frameworks at once meant it couldn't apply any of them well. I kept tweaking the prompt, thinking the issue was in how I phrased things. Spent quite a while before realizing the real problem was the architecture itself.</p>
<h3>Attempt 2: Sequential Review (Voice → Skeptical → Structure)</h3>
<p>Okay, split them up. Run Voice first, then Skeptical, then Structure. Each agent sees the previous agent's feedback and builds on it.</p>
<p><strong>What actually happened</strong>:
The Structure Editor would see "Voice score: 6/10" and unconsciously lower its own standards. Or Skeptical Reader would notice Authenticity Guardian flagged something as critical, then ignore a similar issue because "that's already being addressed."</p>
<p><strong>Why it failed</strong>:
Two problems:</p>
<p>First: Bias contamination. Agents were influenced by each other's scores instead of evaluating independently.</p>
<p>Second: Terrible performance. Three sequential Claude calls meant 30+ seconds of waiting for each review. I'd run a post through the system, go grab coffee, come back, and still be waiting on the third agent. Not sustainable.</p>
<p>I thought sequential would be better, each agent could learn from the previous one. Instead, it just created echo chambers where agents converged on "good enough" instead of pushing for better.</p>
<h2>What Worked: The Breakthrough</h2>
<h3>Attempt 3: Parallel Review (Current System)</h3>
<p>Launch all three critics at once, each reviewing independently. Aggregate results afterward.</p>
<p><strong>Why this worked</strong>:</p>
<ul>
<li>No bias contamination—agents can't see each other's feedback</li>
<li>Faster execution (parallel API calls)</li>
<li>Agents can disagree, which surfaces interesting edge cases</li>
<li>Technical Educator gets unfiltered input from all perspectives</li>
</ul>
<p>The breakthrough was realizing that disagreement is valuable. When Authenticity Guardian wants rambling paragraphs (authentic voice) and Structure Editor wants visual breaks (readability), that tension forces the Technical Educator to find creative solutions like keeping the words but adding paragraph breaks.</p>
<h2>What Surprised Me</h2>
<h3>Convergence Happens Faster Than Expected</h3>
<p>Most posts converge in 2 rounds:</p>
<ul>
<li>Round 1: 5-10 issues flagged</li>
<li>Round 2: All addressed, critics review again</li>
<li>Round 3: Final adjustments and polish</li>
</ul>
<p>Going past round 3 is rare (only for major rewrites or contested items).</p>
<h3>The System Catches Things I'd Miss</h3>
<p><strong>Example from recent review</strong>:</p>
<ul>
<li>Missing Next.js version number</li>
<li>"Simply do X" (curse of knowledge)</li>
<li>Vague heading "Implementation" → Changed to "The Fix: useState with Null Check"</li>
<li>Wall of text (350 words, no breaks) → Split into 3 paragraphs with code block</li>
</ul>
<p>All things I'd probably ship without noticing.</p>
<h3>Voice Preservation Actually Works</h3>
<p>The Authenticity Guardian is brutal but accurate. It catches:</p>
<ul>
<li>Corporate buzzwords I'd unconsciously use</li>
<li>Generic transition phrases ("Let's explore...")</li>
<li>Expert assumptions ("Obviously you'll need to...")</li>
<li>Common AI phrases ("But honestly?")</li>
</ul>
<p>And it suggests authentic alternatives that sound like me:</p>
<ul>
<li>"I kept hitting this error for two hours..."</li>
<li>"Turns out, the issue was..."</li>
<li>"Here's what surprised me..."</li>
</ul>
<h2>The Results</h2>
<p>All of this sounds great on paper. Does it actually work?</p>
<p>Honestly, I wasn't sure at first. The first few runs were slow. I kept tweaking prompts, adjusting scoring thresholds, chasing edge cases where agents would argue forever.</p>
<p>But after a while, the system settled in. Here's what I've observed:</p>
<p><strong>Quality improvement</strong>: The agents consistently catch things I'd miss, voice breaks, missing version numbers, and obvious steps that aren't obvious at all</p>
<p><strong>Consistency</strong>: Every post follows the same quality bar now, which wasn't true when I was reviewing manually</p>
<p><strong>Learning</strong>: The critic feedback teaches me what to avoid in future writing</p>
<p>Not perfect data, I haven't been tracking this scientifically. But qualitatively, it's way better than my manual reviews.</p>
<hr>
<p>These results didn't come from magic. They came from careful design. Here's what's under the hood: file structure, agent definitions, and orchestration patterns that make this work.</p>
<h2>Agent Structure (High-Level)</h2>
<p>Each agent follows the same pattern:</p>
<p><strong>File</strong>: <code>.opencode/agent/&#x3C;agent-name>.md</code></p>
<p><strong>Core components</strong>:</p>
<ul>
<li>YAML frontmatter (description, tools, permissions)</li>
<li>Mission statement</li>
<li>Evaluation framework (specific dimensions to check)</li>
<li>Interaction protocol (how it works with other agents)</li>
<li>Output format (structured feedback)</li>
<li>Examples of good/bad content</li>
</ul>
<p><strong>What makes this work</strong>:</p>
<ul>
<li>Clear job description (one specialty per agent)</li>
<li>Grounded in real methodologies (not "help write blog")</li>
<li>Specific patterns to flag (not vague "check quality")</li>
<li>Structured output (orchestrator can parse it)</li>
</ul>
<p>The orchestrator coordinates workflow: prepare context, spawn agents, aggregate feedback, manage convergence, format output. Agents contain all evaluation logic, and their specialization is what makes the system work.</p>
<h3>Six-Dimensional Review Frameworks</h3>
<p>Each critic evaluates across 6 specific dimensions:</p>
<p><strong>Authenticity Guardian</strong>:</p>
<ol>
<li>Personal vs. Generic</li>
<li>Humble Sharing vs. Expert Lecturing</li>
<li>Specific vs. Vague</li>
<li>Conversational vs. Corporate</li>
<li>Brand Alignment</li>
<li>AI Detection Signals</li>
</ol>
<p><strong>Skeptical Reader</strong>:</p>
<ol>
<li>Completeness</li>
<li>Context</li>
<li>Gaps</li>
<li>Questions</li>
<li>Cognitive Load</li>
<li>Curse of Knowledge</li>
</ol>
<p><strong>Structure Editor</strong>:</p>
<ol>
<li>Opening Hook</li>
<li>Visual Hierarchy</li>
<li>Pacing</li>
<li>Flow</li>
<li>Engagement</li>
<li>Web Readability</li>
</ol>
<p>Each dimension has clear examples of good/bad and specific things to flag.</p>
<h2>What's Next</h2>
<h3>Immediate Improvements</h3>
<ol>
<li><strong>Internal linking agent</strong>: Automatically suggest links to related posts</li>
<li><strong>SEO optimizer</strong>: Ensure titles/descriptions hit character limits</li>
<li><strong>Code validator</strong>: Run code examples to ensure they actually work</li>
</ol>
<h3>Long-term Vision</h3>
<ol>
<li><strong>Feedback loop</strong>: Track which posts get "I got stuck at X" comments, feed that back to Skeptical Reader</li>
<li><strong>Style evolution</strong>: Let agents learn from high-performing posts</li>
<li><strong>Topic suggester</strong>: Analyze conversations to identify blog-worthy moments</li>
</ol>
<h3>Meta-Learning</h3>
<p>This whole process is itself blog-worthy content. I'm using the system to review this post about building the system.</p>
<p><strong>Inception</strong>: The agents are currently debating this very post you're reading.</p>
<h2>Key Takeaways</h2>
<ul>
<li><strong>Multi-agent systems work when agents have real expertise</strong> - Don't just spawn "helper agents." Ground them in actual frameworks and methodologies.</li>
<li><strong>Adversarial debate makes better posts</strong>. The friction between Authenticity Guardian and Structure Editor leads to posts that are both genuine and readable.</li>
<li><strong>Convergence protocols prevent endless iteration</strong>. 2-3 rounds with clear approval criteria. After that, ship or document the trade-off.</li>
<li><strong>Orchestrators maintain principles</strong>. The orchestrator's job is reminding agents of core values: learn in public, authenticity over polish, write for past-self.</li>
<li><strong>The system teaches you</strong>. After seeing the same critiques repeatedly, I've started catching those issues myself. The agents are training me.</li>
<li><strong>Agent specialization matters</strong> - Each agent has one job, grounded in real methodologies.</li>
<li><strong>Parallel review, sequential revision</strong>. Let critics run simultaneously, but revise sequentially.</li>
<li><strong>Veto power creates accountability</strong>. Authenticity Guardian can block generic content. Skeptical Reader can block incomplete content.</li>
<li><strong>Debates need protocols</strong> - Without clear convergence criteria, agents argue forever.</li>
<li><strong>Meta-agents are powerful</strong> - Using <code>/agent-generator</code> and <code>/expertise</code> to create the system was way more effective than hand-writing everything.</li>
</ul>
<h2>The Files</h2>
<p>Here's what the actual file structure looks like:</p>
<p><strong>Agents</strong> (in <code>.claude/agents/</code>):</p>
<ul>
<li><code>blog-technical-educator/AGENT.md</code> - Creator &#x26; Reviser: Transforms conversations into drafts AND implements revisions through iterative rounds</li>
<li><code>blog-authenticity-guardian/AGENT.md</code> - Voice critic: Ensures posts sound like Luke (catches AI patterns, corporate speak, tutorial framing)</li>
<li><code>blog-skeptical-reader/AGENT.md</code> - Completeness &#x26; Authentic Framing critic: Catches missing context, gaps, AND inauthentic framing from past-Luke's perspective</li>
<li><code>blog-structure-editor/AGENT.md</code> - Structure critic: Optimizes flow, hierarchy, AND authenticity of openings/natural flow</li>
</ul>
<p><strong>Orchestrators</strong> (in <code>.claude/commands/</code>):</p>
<ul>
<li><code>generate-blog-post-multi-agent.md</code> - Generate from conversation history</li>
<li><code>review-blog-post-multi-agent.md</code> - Review existing markdown file</li>
</ul>
<p><strong>Meta-skills I used</strong>:</p>
<ul>
<li><code>/agent-generator</code> - Creates well-structured agent definitions automatically</li>
<li><code>/expertise</code> - Synthesizes frameworks from domain experts (swyx, Julia Evans, etc.)</li>
</ul>
<h2>Final Thoughts</h2>
<p>This isn't about replacing human writing. It's about having specialized reviewers who catch what I'd miss.</p>
<p>Think of it like:</p>
<ul>
<li>Authenticity Guardian = Friend who knows your voice</li>
<li>Skeptical Reader = Reader emailing "I got stuck at step 3"</li>
<li>Structure Editor = Copy editor focused on flow</li>
<li>Technical Educator = Yourself, synthesizing feedback AND implementing revisions</li>
</ul>
<p>All working together to ship better content through 2-3 iterative rounds, with task_id reuse maintaining conversation context so critics remember their previous feedback.</p>
<p>The agents don't write for me. They debate with each other until what I wrote is clear, complete, and authentically mine.</p>
<p>Now the system reviews itself.</p>
<p>Wild. This whole process, building tools to help me write, then writing about those tools, keeps looping back on itself. I'm both the creator and the subject.</p>]]></content:encoded>
            <category>ai</category>
        </item>
    </channel>
</rss>