<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Luke Manning - Blog — Docker</title>
        <link>https://lukemanning.ie/</link>
        <description>Breaking things. Building things. Writing about it. (tag: Docker)</description>
        <lastBuildDate>Wed, 30 Sep 2026 12:46:24 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <copyright>All rights reserved 2026, Luke Manning</copyright>
        <atom:link href="https://lukemanning.ie/feeds/docker.xml" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[My Docker Containers Were Working. The Bug Was in a Python Loop.]]></title>
            <link>https://lukemanning.ie/blog/hermes-profiles-to-docker-part-two</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/hermes-profiles-to-docker-part-two</guid>
            <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>This is a follow-up to <a href="/blog/hermes-profiles-to-docker">my post on running Hermes profiles in Docker containers</a>. In that post I described mounting each profile to its own container, setting up systemd to start them on boot, and declaring victory.</p>
<p>I was wrong.</p>
<p>Not wrong about the architecture. Wrong about whether it was actually working.</p>
<p><strong>TL;DR:</strong> I declared the Docker setup working. Then cron jobs vanished into the wrong database, sage was running two Hermes instances, and the actual bug turned out to be a Python loop that had nothing to do with Docker.</p>
<h2>The Problem I Didn't Know I Had (Again)</h2>
<p>I ran <code>docker exec hermes-sage ps aux</code> to verify sage was healthy. It showed one process. Good. I fired a test cron on sage. It said it created successfully. I fired a test cron on rex. It said it created successfully. I waited.</p>
<p>Nothing arrived.</p>
<p>I assumed the Telegram bots weren't set up correctly. I went down a whole path of checking bot tokens, allowed user lists, <code>TELEGRAM_HOME_CHANNEL</code> settings. Then I checked the logs and found something stranger.</p>
<p>Sage's cron engine wasn't running the jobs. Neither was rex's. They were being created (the CLI returned success), but they were going into <em>the host's</em> state database, not the container's.</p>
<p>And underneath that, I discovered something worse: sage wasn't running one Hermes instance. It was running two.</p>
<h2>Why 'ps aux' Lied to Me</h2>
<p>The Docker entrypoint for the Hermes image is <code>/init</code>, which is <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a>. s6-overlay scans <code>/run/service/</code> and starts everything it finds there. My compose file passed <code>--profile sage</code> to the gateway command, which <em>should</em> have meant "run sage profile only."</p>
<p>What actually happened: the compose <code>command:</code> was ignored. The ENTRYPOINT (<code>/init</code>) runs as PID 1, starts s6, and s6 starts <em>all</em> the services in <code>/run/service/</code>, <code>gateway-default</code> and <code>gateway-sage</code> both. Two Hermes instances, same process tree.</p>
<p>The <code>ps aux</code> I'd run to "verify" sage was healthy showed one gateway at the top of the output, so I stopped reading. The output ran longer than one screen. When I finally ran <code>docker exec hermes-sage ps aux | grep gateway</code>, two gateway processes showed up: the one I expected, and a second one nested under s6 that I'd scrolled straight past.</p>
<p>The compose <code>command:</code> instruction doesn't override an ENTRYPOINT. It gets handed to the entrypoint as arguments instead. That's in Docker's docs under how ENTRYPOINT and CMD interact. I'd never internalised it.</p>
<h2>The Cron Database Shell Game</h2>
<p>Once I grasped that sage was running two instances, the cron mystery made more sense. The CLI command <code>hermes cron create</code> was running inside the container, but it reads its config from environment variables, <code>HOME</code> especially, since that's where Hermes looks for its config directory.</p>
<p>I checked what HOME actually was inside the container:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">$ docker exec hermes-sage env </span><span style="color:#F97583">|</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> HOME</span></span>
<span class="line"><span style="color:#79B8FF">HOME=/home/luke</span></span></code></pre>
<p>That's the hermes user's passwd entry, not <code>/opt/data</code>. So when <code>hermes cron create</code> ran without <code>--profile sage</code>, it looked for config at <code>/home/luke/.hermes</code>, which doesn't exist in the container, and fell back to writing jobs to <code>/opt/data/cron/jobs.json</code> instead of <code>/opt/data/profiles/sage/cron/jobs.json</code>.</p>
<p>The sage gateway was reading from the right place. The cron CLI was writing to the wrong place. Jobs created successfully. Jobs never fired.</p>
<p>The fix was passing <code>--profile sage</code> to every cron command inside the container. But I also had to fix HOME, which led to the next problem.</p>
<h2>The su -m Trap</h2>
<p>I wanted the hermes process to run with <code>HOME=/opt/data</code> so it read from the right config. The entrypoint script ran as root, then used <code>su hermes</code> to drop privileges. Easy enough.</p>
<p><code>su -m</code> is meant to preserve the parent environment, so I set <code>HOME=/opt/data</code> before calling <code>su</code> and expected it to carry through. It didn't. The shell <code>su</code> spawned was resetting HOME back to the hermes user's passwd entry (<code>/home/luke</code>) on the way up, undoing what <code>su -m</code> had preserved. By the time hermes actually ran, HOME had been clobbered again.</p>
<p>The reliable fix was <code>env -i</code>, which wipes the environment entirely before setting only what I want. Nothing left for the shell to reset:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF">exec</span><span style="color:#9ECBFF"> env</span><span style="color:#79B8FF"> -i</span><span style="color:#9ECBFF"> HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_PROFILE=sage</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">  su</span><span style="color:#79B8FF"> -m</span><span style="color:#79B8FF"> -s</span><span style="color:#9ECBFF"> /bin/sh</span><span style="color:#9ECBFF"> hermes</span><span style="color:#79B8FF"> -c</span><span style="color:#9ECBFF"> "exec hermes gateway run --profile sage"</span></span></code></pre>
<p><code>env -i</code> clears everything, the explicit vars set what I need, and <code>su -m</code> preserves that clean state. hermes starts with exactly the HOME I specify.</p>
<h2>Custom Entrypoint: Replacing PID 1</h2>
<p>To stop <code>gateway-default</code> from starting, I needed to keep s6 from scanning it. The cleanest approach was replacing PID 1 entirely with a custom script that:</p>
<ol>
<li>Runs <code>stage2-hook.sh</code> for Hermes bootstrap (UID remapping, chown)</li>
<li>Waits for s6 to register its services</li>
<li>Kills <code>gateway-default</code> via <code>s6-svc -d</code></li>
<li>Starts only the sage gateway, using that same <code>env -i</code> exec line from above</li>
</ol>
<p>Then the compose file binds the script as the entrypoint:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#85E89D">entrypoint</span><span style="color:#E1E4E8">: [</span><span style="color:#9ECBFF">"/entrypoint.sh"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#85E89D">volumes</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/docker/entrypoints/sage-only.sh:/entrypoint.sh:ro</span></span></code></pre>
<p>Now sage starts exactly one Hermes instance. No gateway-default. No confusion.</p>
<h2>The Telegram 'Chat Not Found' Problem</h2>
<p>Once I had real isolation working, the crons fired. The sage cron engine ran its job. The rex cron engine ran its job. Both logged <code>completed successfully</code>.</p>
<p>No Telegram messages arrived.</p>
<p>The error: <code>Telegram send failed: Chat not found</code>.</p>
<p>Sage has its own Telegram bot. Rex has its own Telegram bot. Separate bots, separate tokens. A Telegram bot can only send messages to people who have messaged that specific bot first. I'd only ever messaged the default Hermes bot, which knew about my chat. Sage's bot had never seen me.</p>
<p>The second issue was <code>TELEGRAM_HOME_CHANNEL=RealLukeManning</code> in the <code>.env</code>. That's a username, not a chat ID. The Telegram Bot API accepts numeric <code>chat_id</code> for any chat, but only public channels and supergroups can be addressed via <code>@username</code> — for private chats (the home chat here), the API silently fails on a username. Sage was using the username and failing silently.</p>
<p>The fix: message each bot directly first, or use numeric chat IDs in the delivery target (<code>telegram:491962736</code>, not <code>telegram:RealLukeManning</code>).</p>
<h2>The env_file Trap</h2>
<p>I thought the Telegram issue was just setup. Then I noticed sage's logs showed its Telegram module attempting to connect, but the bot token wasn't in the container's config.</p>
<p>The <code>.env</code> with the bot token was on the host at <code>~/.hermes/profiles/sage/.env</code>. The container mounts the profile directory to <code>/opt/data/profiles/sage</code>. The hermes process runs from <code>/opt/data</code>. No <code>.env</code> at <code>/opt/data</code>.</p>
<p>The compose file wasn't passing it through:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Missing</span></span>
<span class="line"><span style="color:#85E89D">env_file</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/sage/.env</span></span></code></pre>
<p>Without the token, the Telegram connection failed gracefully, logged a warning, and didn't die. Messages never arrived.</p>
<h2>The Final Test (Or So I Thought)</h2>
<p>With all fixes in place, I fired simultaneous test crons on both containers. Both ran. Both Telegram bots delivered. Five seconds apart, completely independent.</p>
<p>Sage: one instance, sage profile, sage bot, sage cron. Rex: one instance, rex profile, rex bot, rex cron. Each container starting independently via systemd, each with its <code>.env</code> passed through, cron jobs created with <code>--profile</code> and numeric chat IDs.</p>
<p>I thought that was it. The containers were finally doing what I'd built them to do.</p>
<p>Then last Tuesday I got a message from Sage at 8am.</p>
<p>The daily system health check I'd set up months ago, before Docker and before any of this, was delivering through Sage's Telegram bot instead of the default gateway.</p>
<p>I swore up and down I'd fixed this. I had not fixed this.</p>
<h2>What the Health Check Was Actually Doing</h2>
<p>The health check runs every morning at 8am. Disk, swap, fail2ban, SSH failures. Then it delivers the result to Telegram. Crucially, it runs on the bare-metal default gateway, the one I never put in a container. Not sage, not rex.</p>
<p>The job itself was fine. The delivery was the problem. I opened the health-check cron job to see what it actually called, and found it was running a skill called <code>hermes-telegram-send</code>, one I didn't write, generated by an LLM using a template. That skill had one job: send a Telegram message from a cron context.</p>
<p>The skill worked by loading environment variables from profile <code>.env</code> files, in this order:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"default"</span><span style="color:#E1E4E8">]:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"/home/luke/.hermes/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#6A737D">            # load the token</span></span></code></pre>
<p><code>rex/.env</code> exists, loads token. <code>sage/.env</code> exists, overwrites the token. <code>default/.env</code> doesn't exist, skipped.</p>
<p>Sage's Telegram bot token was the last one loaded. So when the health check cron ran, it used Sage's bot to send the message.</p>
<p>The message came from Sage. Because the skill hardcoded <code>["rex", "sage", "default"]</code> and <code>default</code> didn't exist.</p>
<h2>The Docker Containers Were Working Fine</h2>
<p>This is the part that felt stupid to realise.</p>
<p>Sage was never misconfigured. She wasn't "replying when she shouldn't." She was doing exactly what her container was set up to do. The Docker isolation I'd spent a whole post documenting and debugging was completely correct. Separate containers, separate bots, separate processes. All working.</p>
<p>The bug wasn't in Docker. It wasn't in the architecture. It was in a Python loop in a skill that was generated without thinking about what happens when one of the hardcoded profile names doesn't exist.</p>
<p>The loop was supposed to try multiple profiles in case one didn't have a token. It would load <code>rex</code>, then <code>sage</code>, then <code>default</code>. The last one loaded would be the one used.</p>
<p>But there's no <code>default</code> profile directory. There's a bare-metal Hermes gateway running directly on the host, with its own <code>.env</code> at <code>~/.hermes/.env</code>. That file has the actual Telegram token for the main bot.</p>
<p>The skill didn't know about <code>~/.hermes/.env</code>. It only knew about the profile subdirectories.</p>
<h2>The Real Fix: Baseline First, Never Overwrite</h2>
<p>The skill now loads <code>~/.hermes/.env</code> <strong>first</strong> as a baseline (the bare-metal gateway's token), then profile <code>.env</code> files only fill in empty slots. They never overwrite what's already set.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># ~/.hermes/.env loads FIRST as baseline</span></span>
<span class="line"><span style="color:#E1E4E8">main_env </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#E1E4E8"> main_env.exists():</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> main_env.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">        if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">            continue</span></span>
<span class="line"><span style="color:#E1E4E8">        k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">        os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v  </span><span style="color:#6A737D"># no guard, this is the baseline</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Profile files only fill empty slots, never overwrite</span></span>
<span class="line"><span style="color:#E1E4E8">profile_order </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> ([current_profile] </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> current_profile </span><span style="color:#F97583">else</span><span style="color:#E1E4E8"> []) </span><span style="color:#F97583">+</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> profile_order:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">                continue</span></span>
<span class="line"><span style="color:#E1E4E8">            k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#E1E4E8"> k </span><span style="color:#F97583">not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> os.environ:  </span><span style="color:#6A737D"># only fill empty slots</span></span>
<span class="line"><span style="color:#E1E4E8">                os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v</span></span></code></pre>
<p>The old code had two problems. First, when <code>HERMES_PROFILE</code> is empty (bare-metal default), the current profile list is empty and <code>["rex", "sage"]</code> loads in that order, so sage overwrites rex. Second, the profile loop had <strong>no guard</strong>. <code>os.environ[k] = v</code> always overwrites, so even if the main <code>.env</code> loaded last, it was already too late. The fix: main <code>.env</code> first (no guard), profile files second (with <code>if k not in os.environ</code> guard).</p>
<p>The health check at 8am tomorrow should come from the right bot. For real this time.</p>
<h2>What I Actually Learned</h2>
<p>The implementation had three layers of subtle breakage. The compose <code>command:</code> doesn't override an ENTRYPOINT; if the image has one, the command gets passed to it as arguments. HOME matters more than I expected when <code>su</code>-ing to another user; <code>env -i</code> is the reliable way to lock it. And separate containers don't mean isolated if the entrypoint starts everything inside them regardless.</p>
<p>I fixed all of that. It was real work. And then the bug that actually mattered turned out to be none of it.</p>
<p>Three things I keep having to re-learn:</p>
<p><strong>Infrastructure work doesn't protect you from code bugs.</strong> I spent weeks on Docker containers, systemd services, separate Telegram bots. All correct, all irrelevant to the actual symptom. A hardcoded list of profile names in a Python loop. The containers couldn't catch that, because containers don't catch Python logic errors.</p>
<p><strong>LLM-generated skills inherit LLM failure modes.</strong> The skill came from a template with <code>["rex", "sage", "default"]</code> baked in. Nobody stopped to ask what happens when <code>default</code> doesn't exist. That kind of assumption sits in code for months before anyone notices.</p>
<p><strong>A fix I didn't verify was never really a fix.</strong> I thought I'd solved the Sage problem once isolation was working. The real fix had two subtle bugs hiding in it: an empty <code>HERMES_PROFILE</code> that silently disabled the "current profile first" logic, and a profile loop that overwrote tokens without checking if one was already set. I should have verified instead of trusting it.</p>
<p>The Docker setup is better now than before I started. The isolation is real. But the real fix required iterating on a Python loop that had nothing to do with Docker.</p>]]></description>
            <content:encoded><![CDATA[<p>This is a follow-up to <a href="/blog/hermes-profiles-to-docker">my post on running Hermes profiles in Docker containers</a>. In that post I described mounting each profile to its own container, setting up systemd to start them on boot, and declaring victory.</p>
<p>I was wrong.</p>
<p>Not wrong about the architecture. Wrong about whether it was actually working.</p>
<p><strong>TL;DR:</strong> I declared the Docker setup working. Then cron jobs vanished into the wrong database, sage was running two Hermes instances, and the actual bug turned out to be a Python loop that had nothing to do with Docker.</p>
<h2>The Problem I Didn't Know I Had (Again)</h2>
<p>I ran <code>docker exec hermes-sage ps aux</code> to verify sage was healthy. It showed one process. Good. I fired a test cron on sage. It said it created successfully. I fired a test cron on rex. It said it created successfully. I waited.</p>
<p>Nothing arrived.</p>
<p>I assumed the Telegram bots weren't set up correctly. I went down a whole path of checking bot tokens, allowed user lists, <code>TELEGRAM_HOME_CHANNEL</code> settings. Then I checked the logs and found something stranger.</p>
<p>Sage's cron engine wasn't running the jobs. Neither was rex's. They were being created (the CLI returned success), but they were going into <em>the host's</em> state database, not the container's.</p>
<p>And underneath that, I discovered something worse: sage wasn't running one Hermes instance. It was running two.</p>
<h2>Why 'ps aux' Lied to Me</h2>
<p>The Docker entrypoint for the Hermes image is <code>/init</code>, which is <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a>. s6-overlay scans <code>/run/service/</code> and starts everything it finds there. My compose file passed <code>--profile sage</code> to the gateway command, which <em>should</em> have meant "run sage profile only."</p>
<p>What actually happened: the compose <code>command:</code> was ignored. The ENTRYPOINT (<code>/init</code>) runs as PID 1, starts s6, and s6 starts <em>all</em> the services in <code>/run/service/</code>, <code>gateway-default</code> and <code>gateway-sage</code> both. Two Hermes instances, same process tree.</p>
<p>The <code>ps aux</code> I'd run to "verify" sage was healthy showed one gateway at the top of the output, so I stopped reading. The output ran longer than one screen. When I finally ran <code>docker exec hermes-sage ps aux | grep gateway</code>, two gateway processes showed up: the one I expected, and a second one nested under s6 that I'd scrolled straight past.</p>
<p>The compose <code>command:</code> instruction doesn't override an ENTRYPOINT. It gets handed to the entrypoint as arguments instead. That's in Docker's docs under how ENTRYPOINT and CMD interact. I'd never internalised it.</p>
<h2>The Cron Database Shell Game</h2>
<p>Once I grasped that sage was running two instances, the cron mystery made more sense. The CLI command <code>hermes cron create</code> was running inside the container, but it reads its config from environment variables, <code>HOME</code> especially, since that's where Hermes looks for its config directory.</p>
<p>I checked what HOME actually was inside the container:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">$ docker exec hermes-sage env </span><span style="color:#F97583">|</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> HOME</span></span>
<span class="line"><span style="color:#79B8FF">HOME=/home/luke</span></span></code></pre>
<p>That's the hermes user's passwd entry, not <code>/opt/data</code>. So when <code>hermes cron create</code> ran without <code>--profile sage</code>, it looked for config at <code>/home/luke/.hermes</code>, which doesn't exist in the container, and fell back to writing jobs to <code>/opt/data/cron/jobs.json</code> instead of <code>/opt/data/profiles/sage/cron/jobs.json</code>.</p>
<p>The sage gateway was reading from the right place. The cron CLI was writing to the wrong place. Jobs created successfully. Jobs never fired.</p>
<p>The fix was passing <code>--profile sage</code> to every cron command inside the container. But I also had to fix HOME, which led to the next problem.</p>
<h2>The su -m Trap</h2>
<p>I wanted the hermes process to run with <code>HOME=/opt/data</code> so it read from the right config. The entrypoint script ran as root, then used <code>su hermes</code> to drop privileges. Easy enough.</p>
<p><code>su -m</code> is meant to preserve the parent environment, so I set <code>HOME=/opt/data</code> before calling <code>su</code> and expected it to carry through. It didn't. The shell <code>su</code> spawned was resetting HOME back to the hermes user's passwd entry (<code>/home/luke</code>) on the way up, undoing what <code>su -m</code> had preserved. By the time hermes actually ran, HOME had been clobbered again.</p>
<p>The reliable fix was <code>env -i</code>, which wipes the environment entirely before setting only what I want. Nothing left for the shell to reset:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF">exec</span><span style="color:#9ECBFF"> env</span><span style="color:#79B8FF"> -i</span><span style="color:#9ECBFF"> HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_PROFILE=sage</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">  su</span><span style="color:#79B8FF"> -m</span><span style="color:#79B8FF"> -s</span><span style="color:#9ECBFF"> /bin/sh</span><span style="color:#9ECBFF"> hermes</span><span style="color:#79B8FF"> -c</span><span style="color:#9ECBFF"> "exec hermes gateway run --profile sage"</span></span></code></pre>
<p><code>env -i</code> clears everything, the explicit vars set what I need, and <code>su -m</code> preserves that clean state. hermes starts with exactly the HOME I specify.</p>
<h2>Custom Entrypoint: Replacing PID 1</h2>
<p>To stop <code>gateway-default</code> from starting, I needed to keep s6 from scanning it. The cleanest approach was replacing PID 1 entirely with a custom script that:</p>
<ol>
<li>Runs <code>stage2-hook.sh</code> for Hermes bootstrap (UID remapping, chown)</li>
<li>Waits for s6 to register its services</li>
<li>Kills <code>gateway-default</code> via <code>s6-svc -d</code></li>
<li>Starts only the sage gateway, using that same <code>env -i</code> exec line from above</li>
</ol>
<p>Then the compose file binds the script as the entrypoint:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#85E89D">entrypoint</span><span style="color:#E1E4E8">: [</span><span style="color:#9ECBFF">"/entrypoint.sh"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#85E89D">volumes</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/docker/entrypoints/sage-only.sh:/entrypoint.sh:ro</span></span></code></pre>
<p>Now sage starts exactly one Hermes instance. No gateway-default. No confusion.</p>
<h2>The Telegram 'Chat Not Found' Problem</h2>
<p>Once I had real isolation working, the crons fired. The sage cron engine ran its job. The rex cron engine ran its job. Both logged <code>completed successfully</code>.</p>
<p>No Telegram messages arrived.</p>
<p>The error: <code>Telegram send failed: Chat not found</code>.</p>
<p>Sage has its own Telegram bot. Rex has its own Telegram bot. Separate bots, separate tokens. A Telegram bot can only send messages to people who have messaged that specific bot first. I'd only ever messaged the default Hermes bot, which knew about my chat. Sage's bot had never seen me.</p>
<p>The second issue was <code>TELEGRAM_HOME_CHANNEL=RealLukeManning</code> in the <code>.env</code>. That's a username, not a chat ID. The Telegram Bot API accepts numeric <code>chat_id</code> for any chat, but only public channels and supergroups can be addressed via <code>@username</code> — for private chats (the home chat here), the API silently fails on a username. Sage was using the username and failing silently.</p>
<p>The fix: message each bot directly first, or use numeric chat IDs in the delivery target (<code>telegram:491962736</code>, not <code>telegram:RealLukeManning</code>).</p>
<h2>The env_file Trap</h2>
<p>I thought the Telegram issue was just setup. Then I noticed sage's logs showed its Telegram module attempting to connect, but the bot token wasn't in the container's config.</p>
<p>The <code>.env</code> with the bot token was on the host at <code>~/.hermes/profiles/sage/.env</code>. The container mounts the profile directory to <code>/opt/data/profiles/sage</code>. The hermes process runs from <code>/opt/data</code>. No <code>.env</code> at <code>/opt/data</code>.</p>
<p>The compose file wasn't passing it through:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Missing</span></span>
<span class="line"><span style="color:#85E89D">env_file</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/sage/.env</span></span></code></pre>
<p>Without the token, the Telegram connection failed gracefully, logged a warning, and didn't die. Messages never arrived.</p>
<h2>The Final Test (Or So I Thought)</h2>
<p>With all fixes in place, I fired simultaneous test crons on both containers. Both ran. Both Telegram bots delivered. Five seconds apart, completely independent.</p>
<p>Sage: one instance, sage profile, sage bot, sage cron. Rex: one instance, rex profile, rex bot, rex cron. Each container starting independently via systemd, each with its <code>.env</code> passed through, cron jobs created with <code>--profile</code> and numeric chat IDs.</p>
<p>I thought that was it. The containers were finally doing what I'd built them to do.</p>
<p>Then last Tuesday I got a message from Sage at 8am.</p>
<p>The daily system health check I'd set up months ago, before Docker and before any of this, was delivering through Sage's Telegram bot instead of the default gateway.</p>
<p>I swore up and down I'd fixed this. I had not fixed this.</p>
<h2>What the Health Check Was Actually Doing</h2>
<p>The health check runs every morning at 8am. Disk, swap, fail2ban, SSH failures. Then it delivers the result to Telegram. Crucially, it runs on the bare-metal default gateway, the one I never put in a container. Not sage, not rex.</p>
<p>The job itself was fine. The delivery was the problem. I opened the health-check cron job to see what it actually called, and found it was running a skill called <code>hermes-telegram-send</code>, one I didn't write, generated by an LLM using a template. That skill had one job: send a Telegram message from a cron context.</p>
<p>The skill worked by loading environment variables from profile <code>.env</code> files, in this order:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"default"</span><span style="color:#E1E4E8">]:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"/home/luke/.hermes/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#6A737D">            # load the token</span></span></code></pre>
<p><code>rex/.env</code> exists, loads token. <code>sage/.env</code> exists, overwrites the token. <code>default/.env</code> doesn't exist, skipped.</p>
<p>Sage's Telegram bot token was the last one loaded. So when the health check cron ran, it used Sage's bot to send the message.</p>
<p>The message came from Sage. Because the skill hardcoded <code>["rex", "sage", "default"]</code> and <code>default</code> didn't exist.</p>
<h2>The Docker Containers Were Working Fine</h2>
<p>This is the part that felt stupid to realise.</p>
<p>Sage was never misconfigured. She wasn't "replying when she shouldn't." She was doing exactly what her container was set up to do. The Docker isolation I'd spent a whole post documenting and debugging was completely correct. Separate containers, separate bots, separate processes. All working.</p>
<p>The bug wasn't in Docker. It wasn't in the architecture. It was in a Python loop in a skill that was generated without thinking about what happens when one of the hardcoded profile names doesn't exist.</p>
<p>The loop was supposed to try multiple profiles in case one didn't have a token. It would load <code>rex</code>, then <code>sage</code>, then <code>default</code>. The last one loaded would be the one used.</p>
<p>But there's no <code>default</code> profile directory. There's a bare-metal Hermes gateway running directly on the host, with its own <code>.env</code> at <code>~/.hermes/.env</code>. That file has the actual Telegram token for the main bot.</p>
<p>The skill didn't know about <code>~/.hermes/.env</code>. It only knew about the profile subdirectories.</p>
<h2>The Real Fix: Baseline First, Never Overwrite</h2>
<p>The skill now loads <code>~/.hermes/.env</code> <strong>first</strong> as a baseline (the bare-metal gateway's token), then profile <code>.env</code> files only fill in empty slots. They never overwrite what's already set.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># ~/.hermes/.env loads FIRST as baseline</span></span>
<span class="line"><span style="color:#E1E4E8">main_env </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#E1E4E8"> main_env.exists():</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> main_env.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">        if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">            continue</span></span>
<span class="line"><span style="color:#E1E4E8">        k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">        os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v  </span><span style="color:#6A737D"># no guard, this is the baseline</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Profile files only fill empty slots, never overwrite</span></span>
<span class="line"><span style="color:#E1E4E8">profile_order </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> ([current_profile] </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> current_profile </span><span style="color:#F97583">else</span><span style="color:#E1E4E8"> []) </span><span style="color:#F97583">+</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> profile_order:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">                continue</span></span>
<span class="line"><span style="color:#E1E4E8">            k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#E1E4E8"> k </span><span style="color:#F97583">not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> os.environ:  </span><span style="color:#6A737D"># only fill empty slots</span></span>
<span class="line"><span style="color:#E1E4E8">                os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v</span></span></code></pre>
<p>The old code had two problems. First, when <code>HERMES_PROFILE</code> is empty (bare-metal default), the current profile list is empty and <code>["rex", "sage"]</code> loads in that order, so sage overwrites rex. Second, the profile loop had <strong>no guard</strong>. <code>os.environ[k] = v</code> always overwrites, so even if the main <code>.env</code> loaded last, it was already too late. The fix: main <code>.env</code> first (no guard), profile files second (with <code>if k not in os.environ</code> guard).</p>
<p>The health check at 8am tomorrow should come from the right bot. For real this time.</p>
<h2>What I Actually Learned</h2>
<p>The implementation had three layers of subtle breakage. The compose <code>command:</code> doesn't override an ENTRYPOINT; if the image has one, the command gets passed to it as arguments. HOME matters more than I expected when <code>su</code>-ing to another user; <code>env -i</code> is the reliable way to lock it. And separate containers don't mean isolated if the entrypoint starts everything inside them regardless.</p>
<p>I fixed all of that. It was real work. And then the bug that actually mattered turned out to be none of it.</p>
<p>Three things I keep having to re-learn:</p>
<p><strong>Infrastructure work doesn't protect you from code bugs.</strong> I spent weeks on Docker containers, systemd services, separate Telegram bots. All correct, all irrelevant to the actual symptom. A hardcoded list of profile names in a Python loop. The containers couldn't catch that, because containers don't catch Python logic errors.</p>
<p><strong>LLM-generated skills inherit LLM failure modes.</strong> The skill came from a template with <code>["rex", "sage", "default"]</code> baked in. Nobody stopped to ask what happens when <code>default</code> doesn't exist. That kind of assumption sits in code for months before anyone notices.</p>
<p><strong>A fix I didn't verify was never really a fix.</strong> I thought I'd solved the Sage problem once isolation was working. The real fix had two subtle bugs hiding in it: an empty <code>HERMES_PROFILE</code> that silently disabled the "current profile first" logic, and a profile loop that overwrote tokens without checking if one was already set. I should have verified instead of trusting it.</p>
<p>The Docker setup is better now than before I started. The isolation is real. But the real fix required iterating on a Python loop that had nothing to do with Docker.</p>]]></content:encoded>
            <category>hermes</category>
            <category>homelab</category>
            <category>docker</category>
        </item>
        <item>
            <title><![CDATA[Why I Ditched Hermes Profiles for Docker Containers]]></title>
            <link>https://lukemanning.ie/blog/hermes-profiles-to-docker</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/hermes-profiles-to-docker</guid>
            <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>I ran Hermes (the self-hosted agent I use for <a href="/blog/setting-up-camoufox-with-hermes">reading and research</a> and a few other jobs) with three profiles: default, sage, and rex. Each was a separate persona with its own skills and its own Telegram bot.</p>
<p>They were gateway profiles. In Hermes, the gateway is the long-running process that polls Telegram, runs scheduled jobs, and executes skills. "Profiles" means one gateway binary, one process manager, just different configs passed in via <code>--profile</code>.</p>
<p>One thing worth saying up front. This post is the "why I tried containers" half of the story. When I went to verify the setup, it wasn't actually working yet, and that's covered in <a href="/blog/hermes-profiles-to-docker-part-two">part two of the series</a>. I'm keeping this one focused on the reasoning and the container plumbing, because that part still stands even though my diagnosis turned out to be off.</p>
<h2>The Problem I Didn't Know I Had</h2>
<p>The symptom was cron replies coming from the wrong profile, and it was inconsistent, which is what made it hard to pin down. Sometimes a scheduled job would fire and default and sage would both reply. Sometimes only rex would send the update, when it wasn't his job to. It never happened with messages I sent directly — only cron.</p>
<p>My theory at the time was that the profiles shared too much. One scheduler, one process tree, no hard boundary between them, so jobs bled across. Three profiles all live in the same process, so when a cron job fired, I figured whichever gateway instance was free picked it up.</p>
<p>That theory turned out to be mostly wrong, which is a whole <a href="/blog/hermes-profiles-to-docker-part-two">separate story</a>. I'm laying it out anyway, because it's what drove me to containers, and the architecture reasoning is sound even if the diagnosis wasn't.</p>
<p>What I didn't realise at the time: gateway profiles aren't separate services. They're separate config directories, separate skill directories, separate <code>.env</code> files, and a <code>--profile</code> flag handed to the same gateway binary. Same process tree. Same supervisor. No hard boundary when one of them misbehaves.</p>
<p>I'd been treating them like independent services. They're not.</p>
<h2>What the Docs Recommended (and Why I Went Further)</h2>
<p>The Hermes docs recommend one container hosting all profiles, with <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a> (a process supervisor designed for containers) managing each profile as a first-class service. Mount the whole <code>~/.hermes</code> directory to <code>/opt/data</code>, create profiles with <code>hermes profile create</code>, let s6 start and stop them.</p>
<p>I didn't think that would solve what I was seeing. With one container and s6 running every profile, sage and rex are still co-located processes sharing a network namespace and a data directory. If the problem was jobs bleeding across profiles, bundling them into one container wouldn't stop any of it. They'd just be supervised versions of the same shared setup.</p>
<p>So I ran separate containers for sage and rex, each mounting only its own profile directory. Real process isolation: separate network namespaces, separate PID 1, separate Telegram polling loops. At the process level they genuinely cannot interfere with each other.</p>
<p>The architecture was right, it just wasn't the cause of what I was seeing. That part's in <a href="/blog/hermes-profiles-to-docker-part-two">the verification post</a>.</p>
<h2>The Mount Path Mistake</h2>
<p>The mount paths. I spent far too long on this.</p>
<p>The profile needs to live at <code>/opt/data/profiles/&#x3C;name></code> inside the container, not at <code>/opt/data</code>. I kept mounting the rex profile to <code>/opt/data</code>, and Hermes would log <code>Error: Profile 'rex' does not exist. Create it with: hermes profile create rex</code>. Same files on the host, same mount command, but Hermes couldn't find the profile because it was looking in the wrong place.</p>
<p>The fix was obvious once I saw it:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Wrong — mounts to /opt/data, Hermes expects /opt/data/profiles/rex</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Right</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data/profiles/rex</span></span></code></pre>
<p>The containers weren't the wrong architecture. The paths were wrong. A simple mistake that cost more time than it should have.</p>
<h2>The Boot Problem</h2>
<p>Containers don't start themselves. Bare-metal services auto-start through systemd. I needed the Docker containers to come up on boot too.</p>
<p><code>docker compose</code> has no native systemd integration. The cleanest approach was a systemd user service that runs <code>docker compose up -d</code> for each profile's compose file. It feels slightly off, systemd managing containers instead of services, but it's straightforward and it works.</p>
<h2>What I'd Tell Myself</h2>
<p>When I was running three gateway profiles on bare metal, I was already past what profiles are designed for. They're great for development. Spinning up a new persona with different skills takes seconds. For anything that needs to actually stay separated, the isolation story falls apart.</p>
<p>Containers were the right architecture for that. The operational complexity is real (two compose files, a custom systemd service, separate log streams), but it's the right kind of complexity. Explicit, manageable, debuggable. When rex goes wrong, I look at the rex container, not a shared process tree.</p>
<p>I still run default bare-metal. Default is the daily driver; it has no scheduled cron that fires critical replies. Sage and rex are the workers. They get containers.</p>
<p>What I got wrong was assuming that building the right architecture meant the problem was solved. It didn't. When I went to verify the containers were actually doing their job, nothing worked. Sage was running two Hermes instances, cron jobs were vanishing into the wrong database, and the symptom that started all of this was still happening.</p>
<p>That's <a href="/blog/hermes-profiles-to-docker-part-two">the debugging follow-up</a>.</p>]]></description>
            <content:encoded><![CDATA[<p>I ran Hermes (the self-hosted agent I use for <a href="/blog/setting-up-camoufox-with-hermes">reading and research</a> and a few other jobs) with three profiles: default, sage, and rex. Each was a separate persona with its own skills and its own Telegram bot.</p>
<p>They were gateway profiles. In Hermes, the gateway is the long-running process that polls Telegram, runs scheduled jobs, and executes skills. "Profiles" means one gateway binary, one process manager, just different configs passed in via <code>--profile</code>.</p>
<p>One thing worth saying up front. This post is the "why I tried containers" half of the story. When I went to verify the setup, it wasn't actually working yet, and that's covered in <a href="/blog/hermes-profiles-to-docker-part-two">part two of the series</a>. I'm keeping this one focused on the reasoning and the container plumbing, because that part still stands even though my diagnosis turned out to be off.</p>
<h2>The Problem I Didn't Know I Had</h2>
<p>The symptom was cron replies coming from the wrong profile, and it was inconsistent, which is what made it hard to pin down. Sometimes a scheduled job would fire and default and sage would both reply. Sometimes only rex would send the update, when it wasn't his job to. It never happened with messages I sent directly — only cron.</p>
<p>My theory at the time was that the profiles shared too much. One scheduler, one process tree, no hard boundary between them, so jobs bled across. Three profiles all live in the same process, so when a cron job fired, I figured whichever gateway instance was free picked it up.</p>
<p>That theory turned out to be mostly wrong, which is a whole <a href="/blog/hermes-profiles-to-docker-part-two">separate story</a>. I'm laying it out anyway, because it's what drove me to containers, and the architecture reasoning is sound even if the diagnosis wasn't.</p>
<p>What I didn't realise at the time: gateway profiles aren't separate services. They're separate config directories, separate skill directories, separate <code>.env</code> files, and a <code>--profile</code> flag handed to the same gateway binary. Same process tree. Same supervisor. No hard boundary when one of them misbehaves.</p>
<p>I'd been treating them like independent services. They're not.</p>
<h2>What the Docs Recommended (and Why I Went Further)</h2>
<p>The Hermes docs recommend one container hosting all profiles, with <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a> (a process supervisor designed for containers) managing each profile as a first-class service. Mount the whole <code>~/.hermes</code> directory to <code>/opt/data</code>, create profiles with <code>hermes profile create</code>, let s6 start and stop them.</p>
<p>I didn't think that would solve what I was seeing. With one container and s6 running every profile, sage and rex are still co-located processes sharing a network namespace and a data directory. If the problem was jobs bleeding across profiles, bundling them into one container wouldn't stop any of it. They'd just be supervised versions of the same shared setup.</p>
<p>So I ran separate containers for sage and rex, each mounting only its own profile directory. Real process isolation: separate network namespaces, separate PID 1, separate Telegram polling loops. At the process level they genuinely cannot interfere with each other.</p>
<p>The architecture was right, it just wasn't the cause of what I was seeing. That part's in <a href="/blog/hermes-profiles-to-docker-part-two">the verification post</a>.</p>
<h2>The Mount Path Mistake</h2>
<p>The mount paths. I spent far too long on this.</p>
<p>The profile needs to live at <code>/opt/data/profiles/&#x3C;name></code> inside the container, not at <code>/opt/data</code>. I kept mounting the rex profile to <code>/opt/data</code>, and Hermes would log <code>Error: Profile 'rex' does not exist. Create it with: hermes profile create rex</code>. Same files on the host, same mount command, but Hermes couldn't find the profile because it was looking in the wrong place.</p>
<p>The fix was obvious once I saw it:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Wrong — mounts to /opt/data, Hermes expects /opt/data/profiles/rex</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Right</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data/profiles/rex</span></span></code></pre>
<p>The containers weren't the wrong architecture. The paths were wrong. A simple mistake that cost more time than it should have.</p>
<h2>The Boot Problem</h2>
<p>Containers don't start themselves. Bare-metal services auto-start through systemd. I needed the Docker containers to come up on boot too.</p>
<p><code>docker compose</code> has no native systemd integration. The cleanest approach was a systemd user service that runs <code>docker compose up -d</code> for each profile's compose file. It feels slightly off, systemd managing containers instead of services, but it's straightforward and it works.</p>
<h2>What I'd Tell Myself</h2>
<p>When I was running three gateway profiles on bare metal, I was already past what profiles are designed for. They're great for development. Spinning up a new persona with different skills takes seconds. For anything that needs to actually stay separated, the isolation story falls apart.</p>
<p>Containers were the right architecture for that. The operational complexity is real (two compose files, a custom systemd service, separate log streams), but it's the right kind of complexity. Explicit, manageable, debuggable. When rex goes wrong, I look at the rex container, not a shared process tree.</p>
<p>I still run default bare-metal. Default is the daily driver; it has no scheduled cron that fires critical replies. Sage and rex are the workers. They get containers.</p>
<p>What I got wrong was assuming that building the right architecture meant the problem was solved. It didn't. When I went to verify the containers were actually doing their job, nothing worked. Sage was running two Hermes instances, cron jobs were vanishing into the wrong database, and the symptom that started all of this was still happening.</p>
<p>That's <a href="/blog/hermes-profiles-to-docker-part-two">the debugging follow-up</a>.</p>]]></content:encoded>
            <category>hermes</category>
            <category>homelab</category>
            <category>docker</category>
        </item>
        <item>
            <title><![CDATA[Fixing Unraid Docker Containers After Upgrading - The Missing Label Problem]]></title>
            <link>https://lukemanning.ie/blog/unraid-docker-label-fix</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/unraid-docker-label-fix</guid>
            <pubDate>Wed, 26 Nov 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>So I just upgraded my Unraid server from a very old 6.9 installation to 7.0.1, and immediately ran into a fun little problem: all my Docker containers were suddenly marked as "3rd party." Couldn't edit them, couldn't check for updates, couldn't do anything except stare at them in frustration.</p>
<h2>The Problem</h2>
<p>There's a similar thread documented here in the <a href="https://forums.unraid.net/topic/178736-docker-container-now-shows-3rd-party/">Unraid Forums</a> which I found helpful.</p>
<p>Turns out, Dockerman in Unraid 7.0+ (Unraid's Docker management plugin) uses a container label <code>net.unraid.docker.managed=dockerman</code> to determine which containers it actually manages. My containers were created way back on an older version of Unraid, so they didn't have this label. Without it, Dockerman basically said "not my problem" and refused to touch them.</p>
<p>The nuclear option would be to recreate every single container from scratch, but that's tedious and error-prone when you have dozens of containers with specific configurations. I know I could reuse an existing template as well.. but it still felt like a tedious task. So I did what I normally do and implemented an overly engineered solution to a problem that I could fixed pretty quickly doing it manually.</p>
<h2>The Solution</h2>
<p>The good news is you can add the missing label using Docker's CLI without manually reconfiguring everything.</p>
<p>I started by testing this on one container first - Jackett, which is one of my torrent index containers. I wanted to make sure the whole process worked before batch-processing everything.</p>
<p>First, I generated the docker run command using <code>runlike</code> (it inspects a running container and outputs the equivalent <code>docker run</code> command):</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --rm</span><span style="color:#79B8FF"> -v</span><span style="color:#9ECBFF"> /var/run/docker.sock:/var/run/docker.sock</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">    assaflavie/runlike</span><span style="color:#9ECBFF"> Jackett</span><span style="color:#F97583"> ></span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>(Note: you'll need root privileges to run docker commands in the Unraid terminal - otherwise you'll get "permission denied" errors.)</p>
<p>Then I reviewed what runlike generated:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">cat</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>This showed me the full docker run command with all the volumes, ports, environment variables - exact configuration for my Jackett container.</p>
<p>Next, I needed to add the missing label. I opened the file with nano:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">nano</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>And added <code>--label net.unraid.docker.managed=dockerman</code> and <code>--detach=true</code> right after <code>docker run</code>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --label</span><span style="color:#9ECBFF"> net.unraid.docker.managed=dockerman</span><span style="color:#79B8FF"> --detach=true</span><span style="color:#79B8FF"> --name=Jackett</span><span style="color:#9ECBFF"> ...</span></span></code></pre>
<p>Then I stopped the old container, removed it, and recreated it with the new configuration:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> stop</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> rm</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">bash</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>Finally, I needed to make Unraid fully recognize the container. In the Docker tab, I clicked on Jackett to open its dropdown menu and selected "Force Update." This tells Unraid to add its other management labels.</p>
<p>After doing this, Jackett showed up properly in Unraid instead of as "3rd party." Success!</p>
<h2>The Automated Script</h2>
<p>Once I verified the manual process worked, I created a script to batch-process all my containers. The script loops through all running containers, generates the run command, injects the label, recreates the container, and queues a Force Update:</p>
<p><a href="https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e">https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e</a></p>
<p>To use this, I added it to Unraid's User Scripts plugin (you can install it from Community Applications). This lets you run one-off scripts safely instead of running them directly in the Bash Shell.</p>
<p>If you're comfortable with scripts, this will process all your containers at once. Otherwise, the manual steps above work fine for one-off fixes.</p>
<h2>Important Notes</h2>
<p>By the way, if you run into different data recovery issues — like <a href="/blog/git-corrupt-object-recovery">Git object corruption</a> — the same instinct applies: move don't delete, have a rollback plan.</p>
<ul>
<li><strong>Your data is safe</strong> - This process only removes and recreates the container definitions, not your actual data in appdata or volumes</li>
<li><strong>Test on one container first</strong> - Make sure the manual steps work for your setup before running the bash script</li>
<li><strong>The Force Update step matters</strong> - Don't skip it, as it adds the additional Unraid management labels</li>
<li><strong>Back up your flash drive</strong> - Before any major changes, it's always good practice to back up your Unraid configuration</li>
</ul>
<h2>Why This Happens</h2>
<p>Unraid made this change to better track which containers it's managing versus containers you might have created manually or through other tools. It's actually a good change for container management, but it does mean old containers need this label added retroactively.</p>
<p>Hopefully this saves someone else the frustration of staring at dozens of "3rd party" containers after an upgrade! If this saved you an hour, the Gist is there to share.</p>]]></description>
            <content:encoded><![CDATA[<p>So I just upgraded my Unraid server from a very old 6.9 installation to 7.0.1, and immediately ran into a fun little problem: all my Docker containers were suddenly marked as "3rd party." Couldn't edit them, couldn't check for updates, couldn't do anything except stare at them in frustration.</p>
<h2>The Problem</h2>
<p>There's a similar thread documented here in the <a href="https://forums.unraid.net/topic/178736-docker-container-now-shows-3rd-party/">Unraid Forums</a> which I found helpful.</p>
<p>Turns out, Dockerman in Unraid 7.0+ (Unraid's Docker management plugin) uses a container label <code>net.unraid.docker.managed=dockerman</code> to determine which containers it actually manages. My containers were created way back on an older version of Unraid, so they didn't have this label. Without it, Dockerman basically said "not my problem" and refused to touch them.</p>
<p>The nuclear option would be to recreate every single container from scratch, but that's tedious and error-prone when you have dozens of containers with specific configurations. I know I could reuse an existing template as well.. but it still felt like a tedious task. So I did what I normally do and implemented an overly engineered solution to a problem that I could fixed pretty quickly doing it manually.</p>
<h2>The Solution</h2>
<p>The good news is you can add the missing label using Docker's CLI without manually reconfiguring everything.</p>
<p>I started by testing this on one container first - Jackett, which is one of my torrent index containers. I wanted to make sure the whole process worked before batch-processing everything.</p>
<p>First, I generated the docker run command using <code>runlike</code> (it inspects a running container and outputs the equivalent <code>docker run</code> command):</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --rm</span><span style="color:#79B8FF"> -v</span><span style="color:#9ECBFF"> /var/run/docker.sock:/var/run/docker.sock</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">    assaflavie/runlike</span><span style="color:#9ECBFF"> Jackett</span><span style="color:#F97583"> ></span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>(Note: you'll need root privileges to run docker commands in the Unraid terminal - otherwise you'll get "permission denied" errors.)</p>
<p>Then I reviewed what runlike generated:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">cat</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>This showed me the full docker run command with all the volumes, ports, environment variables - exact configuration for my Jackett container.</p>
<p>Next, I needed to add the missing label. I opened the file with nano:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">nano</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>And added <code>--label net.unraid.docker.managed=dockerman</code> and <code>--detach=true</code> right after <code>docker run</code>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --label</span><span style="color:#9ECBFF"> net.unraid.docker.managed=dockerman</span><span style="color:#79B8FF"> --detach=true</span><span style="color:#79B8FF"> --name=Jackett</span><span style="color:#9ECBFF"> ...</span></span></code></pre>
<p>Then I stopped the old container, removed it, and recreated it with the new configuration:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> stop</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> rm</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">bash</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>Finally, I needed to make Unraid fully recognize the container. In the Docker tab, I clicked on Jackett to open its dropdown menu and selected "Force Update." This tells Unraid to add its other management labels.</p>
<p>After doing this, Jackett showed up properly in Unraid instead of as "3rd party." Success!</p>
<h2>The Automated Script</h2>
<p>Once I verified the manual process worked, I created a script to batch-process all my containers. The script loops through all running containers, generates the run command, injects the label, recreates the container, and queues a Force Update:</p>
<p><a href="https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e">https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e</a></p>
<p>To use this, I added it to Unraid's User Scripts plugin (you can install it from Community Applications). This lets you run one-off scripts safely instead of running them directly in the Bash Shell.</p>
<p>If you're comfortable with scripts, this will process all your containers at once. Otherwise, the manual steps above work fine for one-off fixes.</p>
<h2>Important Notes</h2>
<p>By the way, if you run into different data recovery issues — like <a href="/blog/git-corrupt-object-recovery">Git object corruption</a> — the same instinct applies: move don't delete, have a rollback plan.</p>
<ul>
<li><strong>Your data is safe</strong> - This process only removes and recreates the container definitions, not your actual data in appdata or volumes</li>
<li><strong>Test on one container first</strong> - Make sure the manual steps work for your setup before running the bash script</li>
<li><strong>The Force Update step matters</strong> - Don't skip it, as it adds the additional Unraid management labels</li>
<li><strong>Back up your flash drive</strong> - Before any major changes, it's always good practice to back up your Unraid configuration</li>
</ul>
<h2>Why This Happens</h2>
<p>Unraid made this change to better track which containers it's managing versus containers you might have created manually or through other tools. It's actually a good change for container management, but it does mean old containers need this label added retroactively.</p>
<p>Hopefully this saves someone else the frustration of staring at dozens of "3rd party" containers after an upgrade! If this saved you an hour, the Gist is there to share.</p>]]></content:encoded>
            <category>homelab</category>
            <category>docker</category>
        </item>
    </channel>
</rss>