<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Luke Manning - Blog — Homelab</title>
        <link>https://lukemanning.ie/</link>
        <description>Breaking things. Building things. Writing about it. (tag: Homelab)</description>
        <lastBuildDate>Wed, 30 Sep 2026 12:46:24 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <copyright>All rights reserved 2026, Luke Manning</copyright>
        <atom:link href="https://lukemanning.ie/feeds/homelab.xml" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[My Docker Containers Were Working. The Bug Was in a Python Loop.]]></title>
            <link>https://lukemanning.ie/blog/hermes-profiles-to-docker-part-two</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/hermes-profiles-to-docker-part-two</guid>
            <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>This is a follow-up to <a href="/blog/hermes-profiles-to-docker">my post on running Hermes profiles in Docker containers</a>. In that post I described mounting each profile to its own container, setting up systemd to start them on boot, and declaring victory.</p>
<p>I was wrong.</p>
<p>Not wrong about the architecture. Wrong about whether it was actually working.</p>
<p><strong>TL;DR:</strong> I declared the Docker setup working. Then cron jobs vanished into the wrong database, sage was running two Hermes instances, and the actual bug turned out to be a Python loop that had nothing to do with Docker.</p>
<h2>The Problem I Didn't Know I Had (Again)</h2>
<p>I ran <code>docker exec hermes-sage ps aux</code> to verify sage was healthy. It showed one process. Good. I fired a test cron on sage. It said it created successfully. I fired a test cron on rex. It said it created successfully. I waited.</p>
<p>Nothing arrived.</p>
<p>I assumed the Telegram bots weren't set up correctly. I went down a whole path of checking bot tokens, allowed user lists, <code>TELEGRAM_HOME_CHANNEL</code> settings. Then I checked the logs and found something stranger.</p>
<p>Sage's cron engine wasn't running the jobs. Neither was rex's. They were being created (the CLI returned success), but they were going into <em>the host's</em> state database, not the container's.</p>
<p>And underneath that, I discovered something worse: sage wasn't running one Hermes instance. It was running two.</p>
<h2>Why 'ps aux' Lied to Me</h2>
<p>The Docker entrypoint for the Hermes image is <code>/init</code>, which is <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a>. s6-overlay scans <code>/run/service/</code> and starts everything it finds there. My compose file passed <code>--profile sage</code> to the gateway command, which <em>should</em> have meant "run sage profile only."</p>
<p>What actually happened: the compose <code>command:</code> was ignored. The ENTRYPOINT (<code>/init</code>) runs as PID 1, starts s6, and s6 starts <em>all</em> the services in <code>/run/service/</code>, <code>gateway-default</code> and <code>gateway-sage</code> both. Two Hermes instances, same process tree.</p>
<p>The <code>ps aux</code> I'd run to "verify" sage was healthy showed one gateway at the top of the output, so I stopped reading. The output ran longer than one screen. When I finally ran <code>docker exec hermes-sage ps aux | grep gateway</code>, two gateway processes showed up: the one I expected, and a second one nested under s6 that I'd scrolled straight past.</p>
<p>The compose <code>command:</code> instruction doesn't override an ENTRYPOINT. It gets handed to the entrypoint as arguments instead. That's in Docker's docs under how ENTRYPOINT and CMD interact. I'd never internalised it.</p>
<h2>The Cron Database Shell Game</h2>
<p>Once I grasped that sage was running two instances, the cron mystery made more sense. The CLI command <code>hermes cron create</code> was running inside the container, but it reads its config from environment variables, <code>HOME</code> especially, since that's where Hermes looks for its config directory.</p>
<p>I checked what HOME actually was inside the container:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">$ docker exec hermes-sage env </span><span style="color:#F97583">|</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> HOME</span></span>
<span class="line"><span style="color:#79B8FF">HOME=/home/luke</span></span></code></pre>
<p>That's the hermes user's passwd entry, not <code>/opt/data</code>. So when <code>hermes cron create</code> ran without <code>--profile sage</code>, it looked for config at <code>/home/luke/.hermes</code>, which doesn't exist in the container, and fell back to writing jobs to <code>/opt/data/cron/jobs.json</code> instead of <code>/opt/data/profiles/sage/cron/jobs.json</code>.</p>
<p>The sage gateway was reading from the right place. The cron CLI was writing to the wrong place. Jobs created successfully. Jobs never fired.</p>
<p>The fix was passing <code>--profile sage</code> to every cron command inside the container. But I also had to fix HOME, which led to the next problem.</p>
<h2>The su -m Trap</h2>
<p>I wanted the hermes process to run with <code>HOME=/opt/data</code> so it read from the right config. The entrypoint script ran as root, then used <code>su hermes</code> to drop privileges. Easy enough.</p>
<p><code>su -m</code> is meant to preserve the parent environment, so I set <code>HOME=/opt/data</code> before calling <code>su</code> and expected it to carry through. It didn't. The shell <code>su</code> spawned was resetting HOME back to the hermes user's passwd entry (<code>/home/luke</code>) on the way up, undoing what <code>su -m</code> had preserved. By the time hermes actually ran, HOME had been clobbered again.</p>
<p>The reliable fix was <code>env -i</code>, which wipes the environment entirely before setting only what I want. Nothing left for the shell to reset:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF">exec</span><span style="color:#9ECBFF"> env</span><span style="color:#79B8FF"> -i</span><span style="color:#9ECBFF"> HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_PROFILE=sage</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">  su</span><span style="color:#79B8FF"> -m</span><span style="color:#79B8FF"> -s</span><span style="color:#9ECBFF"> /bin/sh</span><span style="color:#9ECBFF"> hermes</span><span style="color:#79B8FF"> -c</span><span style="color:#9ECBFF"> "exec hermes gateway run --profile sage"</span></span></code></pre>
<p><code>env -i</code> clears everything, the explicit vars set what I need, and <code>su -m</code> preserves that clean state. hermes starts with exactly the HOME I specify.</p>
<h2>Custom Entrypoint: Replacing PID 1</h2>
<p>To stop <code>gateway-default</code> from starting, I needed to keep s6 from scanning it. The cleanest approach was replacing PID 1 entirely with a custom script that:</p>
<ol>
<li>Runs <code>stage2-hook.sh</code> for Hermes bootstrap (UID remapping, chown)</li>
<li>Waits for s6 to register its services</li>
<li>Kills <code>gateway-default</code> via <code>s6-svc -d</code></li>
<li>Starts only the sage gateway, using that same <code>env -i</code> exec line from above</li>
</ol>
<p>Then the compose file binds the script as the entrypoint:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#85E89D">entrypoint</span><span style="color:#E1E4E8">: [</span><span style="color:#9ECBFF">"/entrypoint.sh"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#85E89D">volumes</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/docker/entrypoints/sage-only.sh:/entrypoint.sh:ro</span></span></code></pre>
<p>Now sage starts exactly one Hermes instance. No gateway-default. No confusion.</p>
<h2>The Telegram 'Chat Not Found' Problem</h2>
<p>Once I had real isolation working, the crons fired. The sage cron engine ran its job. The rex cron engine ran its job. Both logged <code>completed successfully</code>.</p>
<p>No Telegram messages arrived.</p>
<p>The error: <code>Telegram send failed: Chat not found</code>.</p>
<p>Sage has its own Telegram bot. Rex has its own Telegram bot. Separate bots, separate tokens. A Telegram bot can only send messages to people who have messaged that specific bot first. I'd only ever messaged the default Hermes bot, which knew about my chat. Sage's bot had never seen me.</p>
<p>The second issue was <code>TELEGRAM_HOME_CHANNEL=RealLukeManning</code> in the <code>.env</code>. That's a username, not a chat ID. The Telegram Bot API accepts numeric <code>chat_id</code> for any chat, but only public channels and supergroups can be addressed via <code>@username</code> — for private chats (the home chat here), the API silently fails on a username. Sage was using the username and failing silently.</p>
<p>The fix: message each bot directly first, or use numeric chat IDs in the delivery target (<code>telegram:491962736</code>, not <code>telegram:RealLukeManning</code>).</p>
<h2>The env_file Trap</h2>
<p>I thought the Telegram issue was just setup. Then I noticed sage's logs showed its Telegram module attempting to connect, but the bot token wasn't in the container's config.</p>
<p>The <code>.env</code> with the bot token was on the host at <code>~/.hermes/profiles/sage/.env</code>. The container mounts the profile directory to <code>/opt/data/profiles/sage</code>. The hermes process runs from <code>/opt/data</code>. No <code>.env</code> at <code>/opt/data</code>.</p>
<p>The compose file wasn't passing it through:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Missing</span></span>
<span class="line"><span style="color:#85E89D">env_file</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/sage/.env</span></span></code></pre>
<p>Without the token, the Telegram connection failed gracefully, logged a warning, and didn't die. Messages never arrived.</p>
<h2>The Final Test (Or So I Thought)</h2>
<p>With all fixes in place, I fired simultaneous test crons on both containers. Both ran. Both Telegram bots delivered. Five seconds apart, completely independent.</p>
<p>Sage: one instance, sage profile, sage bot, sage cron. Rex: one instance, rex profile, rex bot, rex cron. Each container starting independently via systemd, each with its <code>.env</code> passed through, cron jobs created with <code>--profile</code> and numeric chat IDs.</p>
<p>I thought that was it. The containers were finally doing what I'd built them to do.</p>
<p>Then last Tuesday I got a message from Sage at 8am.</p>
<p>The daily system health check I'd set up months ago, before Docker and before any of this, was delivering through Sage's Telegram bot instead of the default gateway.</p>
<p>I swore up and down I'd fixed this. I had not fixed this.</p>
<h2>What the Health Check Was Actually Doing</h2>
<p>The health check runs every morning at 8am. Disk, swap, fail2ban, SSH failures. Then it delivers the result to Telegram. Crucially, it runs on the bare-metal default gateway, the one I never put in a container. Not sage, not rex.</p>
<p>The job itself was fine. The delivery was the problem. I opened the health-check cron job to see what it actually called, and found it was running a skill called <code>hermes-telegram-send</code>, one I didn't write, generated by an LLM using a template. That skill had one job: send a Telegram message from a cron context.</p>
<p>The skill worked by loading environment variables from profile <code>.env</code> files, in this order:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"default"</span><span style="color:#E1E4E8">]:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"/home/luke/.hermes/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#6A737D">            # load the token</span></span></code></pre>
<p><code>rex/.env</code> exists, loads token. <code>sage/.env</code> exists, overwrites the token. <code>default/.env</code> doesn't exist, skipped.</p>
<p>Sage's Telegram bot token was the last one loaded. So when the health check cron ran, it used Sage's bot to send the message.</p>
<p>The message came from Sage. Because the skill hardcoded <code>["rex", "sage", "default"]</code> and <code>default</code> didn't exist.</p>
<h2>The Docker Containers Were Working Fine</h2>
<p>This is the part that felt stupid to realise.</p>
<p>Sage was never misconfigured. She wasn't "replying when she shouldn't." She was doing exactly what her container was set up to do. The Docker isolation I'd spent a whole post documenting and debugging was completely correct. Separate containers, separate bots, separate processes. All working.</p>
<p>The bug wasn't in Docker. It wasn't in the architecture. It was in a Python loop in a skill that was generated without thinking about what happens when one of the hardcoded profile names doesn't exist.</p>
<p>The loop was supposed to try multiple profiles in case one didn't have a token. It would load <code>rex</code>, then <code>sage</code>, then <code>default</code>. The last one loaded would be the one used.</p>
<p>But there's no <code>default</code> profile directory. There's a bare-metal Hermes gateway running directly on the host, with its own <code>.env</code> at <code>~/.hermes/.env</code>. That file has the actual Telegram token for the main bot.</p>
<p>The skill didn't know about <code>~/.hermes/.env</code>. It only knew about the profile subdirectories.</p>
<h2>The Real Fix: Baseline First, Never Overwrite</h2>
<p>The skill now loads <code>~/.hermes/.env</code> <strong>first</strong> as a baseline (the bare-metal gateway's token), then profile <code>.env</code> files only fill in empty slots. They never overwrite what's already set.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># ~/.hermes/.env loads FIRST as baseline</span></span>
<span class="line"><span style="color:#E1E4E8">main_env </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#E1E4E8"> main_env.exists():</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> main_env.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">        if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">            continue</span></span>
<span class="line"><span style="color:#E1E4E8">        k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">        os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v  </span><span style="color:#6A737D"># no guard, this is the baseline</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Profile files only fill empty slots, never overwrite</span></span>
<span class="line"><span style="color:#E1E4E8">profile_order </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> ([current_profile] </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> current_profile </span><span style="color:#F97583">else</span><span style="color:#E1E4E8"> []) </span><span style="color:#F97583">+</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> profile_order:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">                continue</span></span>
<span class="line"><span style="color:#E1E4E8">            k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#E1E4E8"> k </span><span style="color:#F97583">not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> os.environ:  </span><span style="color:#6A737D"># only fill empty slots</span></span>
<span class="line"><span style="color:#E1E4E8">                os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v</span></span></code></pre>
<p>The old code had two problems. First, when <code>HERMES_PROFILE</code> is empty (bare-metal default), the current profile list is empty and <code>["rex", "sage"]</code> loads in that order, so sage overwrites rex. Second, the profile loop had <strong>no guard</strong>. <code>os.environ[k] = v</code> always overwrites, so even if the main <code>.env</code> loaded last, it was already too late. The fix: main <code>.env</code> first (no guard), profile files second (with <code>if k not in os.environ</code> guard).</p>
<p>The health check at 8am tomorrow should come from the right bot. For real this time.</p>
<h2>What I Actually Learned</h2>
<p>The implementation had three layers of subtle breakage. The compose <code>command:</code> doesn't override an ENTRYPOINT; if the image has one, the command gets passed to it as arguments. HOME matters more than I expected when <code>su</code>-ing to another user; <code>env -i</code> is the reliable way to lock it. And separate containers don't mean isolated if the entrypoint starts everything inside them regardless.</p>
<p>I fixed all of that. It was real work. And then the bug that actually mattered turned out to be none of it.</p>
<p>Three things I keep having to re-learn:</p>
<p><strong>Infrastructure work doesn't protect you from code bugs.</strong> I spent weeks on Docker containers, systemd services, separate Telegram bots. All correct, all irrelevant to the actual symptom. A hardcoded list of profile names in a Python loop. The containers couldn't catch that, because containers don't catch Python logic errors.</p>
<p><strong>LLM-generated skills inherit LLM failure modes.</strong> The skill came from a template with <code>["rex", "sage", "default"]</code> baked in. Nobody stopped to ask what happens when <code>default</code> doesn't exist. That kind of assumption sits in code for months before anyone notices.</p>
<p><strong>A fix I didn't verify was never really a fix.</strong> I thought I'd solved the Sage problem once isolation was working. The real fix had two subtle bugs hiding in it: an empty <code>HERMES_PROFILE</code> that silently disabled the "current profile first" logic, and a profile loop that overwrote tokens without checking if one was already set. I should have verified instead of trusting it.</p>
<p>The Docker setup is better now than before I started. The isolation is real. But the real fix required iterating on a Python loop that had nothing to do with Docker.</p>]]></description>
            <content:encoded><![CDATA[<p>This is a follow-up to <a href="/blog/hermes-profiles-to-docker">my post on running Hermes profiles in Docker containers</a>. In that post I described mounting each profile to its own container, setting up systemd to start them on boot, and declaring victory.</p>
<p>I was wrong.</p>
<p>Not wrong about the architecture. Wrong about whether it was actually working.</p>
<p><strong>TL;DR:</strong> I declared the Docker setup working. Then cron jobs vanished into the wrong database, sage was running two Hermes instances, and the actual bug turned out to be a Python loop that had nothing to do with Docker.</p>
<h2>The Problem I Didn't Know I Had (Again)</h2>
<p>I ran <code>docker exec hermes-sage ps aux</code> to verify sage was healthy. It showed one process. Good. I fired a test cron on sage. It said it created successfully. I fired a test cron on rex. It said it created successfully. I waited.</p>
<p>Nothing arrived.</p>
<p>I assumed the Telegram bots weren't set up correctly. I went down a whole path of checking bot tokens, allowed user lists, <code>TELEGRAM_HOME_CHANNEL</code> settings. Then I checked the logs and found something stranger.</p>
<p>Sage's cron engine wasn't running the jobs. Neither was rex's. They were being created (the CLI returned success), but they were going into <em>the host's</em> state database, not the container's.</p>
<p>And underneath that, I discovered something worse: sage wasn't running one Hermes instance. It was running two.</p>
<h2>Why 'ps aux' Lied to Me</h2>
<p>The Docker entrypoint for the Hermes image is <code>/init</code>, which is <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a>. s6-overlay scans <code>/run/service/</code> and starts everything it finds there. My compose file passed <code>--profile sage</code> to the gateway command, which <em>should</em> have meant "run sage profile only."</p>
<p>What actually happened: the compose <code>command:</code> was ignored. The ENTRYPOINT (<code>/init</code>) runs as PID 1, starts s6, and s6 starts <em>all</em> the services in <code>/run/service/</code>, <code>gateway-default</code> and <code>gateway-sage</code> both. Two Hermes instances, same process tree.</p>
<p>The <code>ps aux</code> I'd run to "verify" sage was healthy showed one gateway at the top of the output, so I stopped reading. The output ran longer than one screen. When I finally ran <code>docker exec hermes-sage ps aux | grep gateway</code>, two gateway processes showed up: the one I expected, and a second one nested under s6 that I'd scrolled straight past.</p>
<p>The compose <code>command:</code> instruction doesn't override an ENTRYPOINT. It gets handed to the entrypoint as arguments instead. That's in Docker's docs under how ENTRYPOINT and CMD interact. I'd never internalised it.</p>
<h2>The Cron Database Shell Game</h2>
<p>Once I grasped that sage was running two instances, the cron mystery made more sense. The CLI command <code>hermes cron create</code> was running inside the container, but it reads its config from environment variables, <code>HOME</code> especially, since that's where Hermes looks for its config directory.</p>
<p>I checked what HOME actually was inside the container:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#E1E4E8">$ docker exec hermes-sage env </span><span style="color:#F97583">|</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> HOME</span></span>
<span class="line"><span style="color:#79B8FF">HOME=/home/luke</span></span></code></pre>
<p>That's the hermes user's passwd entry, not <code>/opt/data</code>. So when <code>hermes cron create</code> ran without <code>--profile sage</code>, it looked for config at <code>/home/luke/.hermes</code>, which doesn't exist in the container, and fell back to writing jobs to <code>/opt/data/cron/jobs.json</code> instead of <code>/opt/data/profiles/sage/cron/jobs.json</code>.</p>
<p>The sage gateway was reading from the right place. The cron CLI was writing to the wrong place. Jobs created successfully. Jobs never fired.</p>
<p>The fix was passing <code>--profile sage</code> to every cron command inside the container. But I also had to fix HOME, which led to the next problem.</p>
<h2>The su -m Trap</h2>
<p>I wanted the hermes process to run with <code>HOME=/opt/data</code> so it read from the right config. The entrypoint script ran as root, then used <code>su hermes</code> to drop privileges. Easy enough.</p>
<p><code>su -m</code> is meant to preserve the parent environment, so I set <code>HOME=/opt/data</code> before calling <code>su</code> and expected it to carry through. It didn't. The shell <code>su</code> spawned was resetting HOME back to the hermes user's passwd entry (<code>/home/luke</code>) on the way up, undoing what <code>su -m</code> had preserved. By the time hermes actually ran, HOME had been clobbered again.</p>
<p>The reliable fix was <code>env -i</code>, which wipes the environment entirely before setting only what I want. Nothing left for the shell to reset:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#79B8FF">exec</span><span style="color:#9ECBFF"> env</span><span style="color:#79B8FF"> -i</span><span style="color:#9ECBFF"> HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_HOME=/opt/data</span><span style="color:#9ECBFF"> HERMES_PROFILE=sage</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">  su</span><span style="color:#79B8FF"> -m</span><span style="color:#79B8FF"> -s</span><span style="color:#9ECBFF"> /bin/sh</span><span style="color:#9ECBFF"> hermes</span><span style="color:#79B8FF"> -c</span><span style="color:#9ECBFF"> "exec hermes gateway run --profile sage"</span></span></code></pre>
<p><code>env -i</code> clears everything, the explicit vars set what I need, and <code>su -m</code> preserves that clean state. hermes starts with exactly the HOME I specify.</p>
<h2>Custom Entrypoint: Replacing PID 1</h2>
<p>To stop <code>gateway-default</code> from starting, I needed to keep s6 from scanning it. The cleanest approach was replacing PID 1 entirely with a custom script that:</p>
<ol>
<li>Runs <code>stage2-hook.sh</code> for Hermes bootstrap (UID remapping, chown)</li>
<li>Waits for s6 to register its services</li>
<li>Kills <code>gateway-default</code> via <code>s6-svc -d</code></li>
<li>Starts only the sage gateway, using that same <code>env -i</code> exec line from above</li>
</ol>
<p>Then the compose file binds the script as the entrypoint:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#85E89D">entrypoint</span><span style="color:#E1E4E8">: [</span><span style="color:#9ECBFF">"/entrypoint.sh"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#85E89D">volumes</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/docker/entrypoints/sage-only.sh:/entrypoint.sh:ro</span></span></code></pre>
<p>Now sage starts exactly one Hermes instance. No gateway-default. No confusion.</p>
<h2>The Telegram 'Chat Not Found' Problem</h2>
<p>Once I had real isolation working, the crons fired. The sage cron engine ran its job. The rex cron engine ran its job. Both logged <code>completed successfully</code>.</p>
<p>No Telegram messages arrived.</p>
<p>The error: <code>Telegram send failed: Chat not found</code>.</p>
<p>Sage has its own Telegram bot. Rex has its own Telegram bot. Separate bots, separate tokens. A Telegram bot can only send messages to people who have messaged that specific bot first. I'd only ever messaged the default Hermes bot, which knew about my chat. Sage's bot had never seen me.</p>
<p>The second issue was <code>TELEGRAM_HOME_CHANNEL=RealLukeManning</code> in the <code>.env</code>. That's a username, not a chat ID. The Telegram Bot API accepts numeric <code>chat_id</code> for any chat, but only public channels and supergroups can be addressed via <code>@username</code> — for private chats (the home chat here), the API silently fails on a username. Sage was using the username and failing silently.</p>
<p>The fix: message each bot directly first, or use numeric chat IDs in the delivery target (<code>telegram:491962736</code>, not <code>telegram:RealLukeManning</code>).</p>
<h2>The env_file Trap</h2>
<p>I thought the Telegram issue was just setup. Then I noticed sage's logs showed its Telegram module attempting to connect, but the bot token wasn't in the container's config.</p>
<p>The <code>.env</code> with the bot token was on the host at <code>~/.hermes/profiles/sage/.env</code>. The container mounts the profile directory to <code>/opt/data/profiles/sage</code>. The hermes process runs from <code>/opt/data</code>. No <code>.env</code> at <code>/opt/data</code>.</p>
<p>The compose file wasn't passing it through:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Missing</span></span>
<span class="line"><span style="color:#85E89D">env_file</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">  - </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/sage/.env</span></span></code></pre>
<p>Without the token, the Telegram connection failed gracefully, logged a warning, and didn't die. Messages never arrived.</p>
<h2>The Final Test (Or So I Thought)</h2>
<p>With all fixes in place, I fired simultaneous test crons on both containers. Both ran. Both Telegram bots delivered. Five seconds apart, completely independent.</p>
<p>Sage: one instance, sage profile, sage bot, sage cron. Rex: one instance, rex profile, rex bot, rex cron. Each container starting independently via systemd, each with its <code>.env</code> passed through, cron jobs created with <code>--profile</code> and numeric chat IDs.</p>
<p>I thought that was it. The containers were finally doing what I'd built them to do.</p>
<p>Then last Tuesday I got a message from Sage at 8am.</p>
<p>The daily system health check I'd set up months ago, before Docker and before any of this, was delivering through Sage's Telegram bot instead of the default gateway.</p>
<p>I swore up and down I'd fixed this. I had not fixed this.</p>
<h2>What the Health Check Was Actually Doing</h2>
<p>The health check runs every morning at 8am. Disk, swap, fail2ban, SSH failures. Then it delivers the result to Telegram. Crucially, it runs on the bare-metal default gateway, the one I never put in a container. Not sage, not rex.</p>
<p>The job itself was fine. The delivery was the problem. I opened the health-check cron job to see what it actually called, and found it was running a skill called <code>hermes-telegram-send</code>, one I didn't write, generated by an LLM using a template. That skill had one job: send a Telegram message from a cron context.</p>
<p>The skill worked by loading environment variables from profile <code>.env</code> files, in this order:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"default"</span><span style="color:#E1E4E8">]:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"/home/luke/.hermes/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#6A737D">            # load the token</span></span></code></pre>
<p><code>rex/.env</code> exists, loads token. <code>sage/.env</code> exists, overwrites the token. <code>default/.env</code> doesn't exist, skipped.</p>
<p>Sage's Telegram bot token was the last one loaded. So when the health check cron ran, it used Sage's bot to send the message.</p>
<p>The message came from Sage. Because the skill hardcoded <code>["rex", "sage", "default"]</code> and <code>default</code> didn't exist.</p>
<h2>The Docker Containers Were Working Fine</h2>
<p>This is the part that felt stupid to realise.</p>
<p>Sage was never misconfigured. She wasn't "replying when she shouldn't." She was doing exactly what her container was set up to do. The Docker isolation I'd spent a whole post documenting and debugging was completely correct. Separate containers, separate bots, separate processes. All working.</p>
<p>The bug wasn't in Docker. It wasn't in the architecture. It was in a Python loop in a skill that was generated without thinking about what happens when one of the hardcoded profile names doesn't exist.</p>
<p>The loop was supposed to try multiple profiles in case one didn't have a token. It would load <code>rex</code>, then <code>sage</code>, then <code>default</code>. The last one loaded would be the one used.</p>
<p>But there's no <code>default</code> profile directory. There's a bare-metal Hermes gateway running directly on the host, with its own <code>.env</code> at <code>~/.hermes/.env</code>. That file has the actual Telegram token for the main bot.</p>
<p>The skill didn't know about <code>~/.hermes/.env</code>. It only knew about the profile subdirectories.</p>
<h2>The Real Fix: Baseline First, Never Overwrite</h2>
<p>The skill now loads <code>~/.hermes/.env</code> <strong>first</strong> as a baseline (the bare-metal gateway's token), then profile <code>.env</code> files only fill in empty slots. They never overwrite what's already set.</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># ~/.hermes/.env loads FIRST as baseline</span></span>
<span class="line"><span style="color:#E1E4E8">main_env </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#E1E4E8"> main_env.exists():</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> main_env.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">        if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">            continue</span></span>
<span class="line"><span style="color:#E1E4E8">        k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">        os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v  </span><span style="color:#6A737D"># no guard, this is the baseline</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Profile files only fill empty slots, never overwrite</span></span>
<span class="line"><span style="color:#E1E4E8">profile_order </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> ([current_profile] </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> current_profile </span><span style="color:#F97583">else</span><span style="color:#E1E4E8"> []) </span><span style="color:#F97583">+</span><span style="color:#E1E4E8"> [</span><span style="color:#9ECBFF">"rex"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"sage"</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">for</span><span style="color:#E1E4E8"> profile </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> profile_order:</span></span>
<span class="line"><span style="color:#E1E4E8">    env_file </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> Path(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">hermes_home</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/profiles/</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">profile</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">/.env"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> env_file.exists():</span></span>
<span class="line"><span style="color:#F97583">        for</span><span style="color:#E1E4E8"> line </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> env_file.read_text().splitlines():</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> line.strip() </span><span style="color:#F97583">or</span><span style="color:#E1E4E8"> line.lstrip().startswith(</span><span style="color:#9ECBFF">"#"</span><span style="color:#E1E4E8">) </span><span style="color:#F97583">or</span><span style="color:#9ECBFF"> "="</span><span style="color:#F97583"> not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> line:</span></span>
<span class="line"><span style="color:#F97583">                continue</span></span>
<span class="line"><span style="color:#E1E4E8">            k, v </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> line.split(</span><span style="color:#9ECBFF">"="</span><span style="color:#E1E4E8">, </span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">            if</span><span style="color:#E1E4E8"> k </span><span style="color:#F97583">not</span><span style="color:#F97583"> in</span><span style="color:#E1E4E8"> os.environ:  </span><span style="color:#6A737D"># only fill empty slots</span></span>
<span class="line"><span style="color:#E1E4E8">                os.environ[k] </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> v</span></span></code></pre>
<p>The old code had two problems. First, when <code>HERMES_PROFILE</code> is empty (bare-metal default), the current profile list is empty and <code>["rex", "sage"]</code> loads in that order, so sage overwrites rex. Second, the profile loop had <strong>no guard</strong>. <code>os.environ[k] = v</code> always overwrites, so even if the main <code>.env</code> loaded last, it was already too late. The fix: main <code>.env</code> first (no guard), profile files second (with <code>if k not in os.environ</code> guard).</p>
<p>The health check at 8am tomorrow should come from the right bot. For real this time.</p>
<h2>What I Actually Learned</h2>
<p>The implementation had three layers of subtle breakage. The compose <code>command:</code> doesn't override an ENTRYPOINT; if the image has one, the command gets passed to it as arguments. HOME matters more than I expected when <code>su</code>-ing to another user; <code>env -i</code> is the reliable way to lock it. And separate containers don't mean isolated if the entrypoint starts everything inside them regardless.</p>
<p>I fixed all of that. It was real work. And then the bug that actually mattered turned out to be none of it.</p>
<p>Three things I keep having to re-learn:</p>
<p><strong>Infrastructure work doesn't protect you from code bugs.</strong> I spent weeks on Docker containers, systemd services, separate Telegram bots. All correct, all irrelevant to the actual symptom. A hardcoded list of profile names in a Python loop. The containers couldn't catch that, because containers don't catch Python logic errors.</p>
<p><strong>LLM-generated skills inherit LLM failure modes.</strong> The skill came from a template with <code>["rex", "sage", "default"]</code> baked in. Nobody stopped to ask what happens when <code>default</code> doesn't exist. That kind of assumption sits in code for months before anyone notices.</p>
<p><strong>A fix I didn't verify was never really a fix.</strong> I thought I'd solved the Sage problem once isolation was working. The real fix had two subtle bugs hiding in it: an empty <code>HERMES_PROFILE</code> that silently disabled the "current profile first" logic, and a profile loop that overwrote tokens without checking if one was already set. I should have verified instead of trusting it.</p>
<p>The Docker setup is better now than before I started. The isolation is real. But the real fix required iterating on a Python loop that had nothing to do with Docker.</p>]]></content:encoded>
            <category>hermes</category>
            <category>homelab</category>
            <category>docker</category>
        </item>
        <item>
            <title><![CDATA[Why I Ditched Hermes Profiles for Docker Containers]]></title>
            <link>https://lukemanning.ie/blog/hermes-profiles-to-docker</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/hermes-profiles-to-docker</guid>
            <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>I ran Hermes (the self-hosted agent I use for <a href="/blog/setting-up-camoufox-with-hermes">reading and research</a> and a few other jobs) with three profiles: default, sage, and rex. Each was a separate persona with its own skills and its own Telegram bot.</p>
<p>They were gateway profiles. In Hermes, the gateway is the long-running process that polls Telegram, runs scheduled jobs, and executes skills. "Profiles" means one gateway binary, one process manager, just different configs passed in via <code>--profile</code>.</p>
<p>One thing worth saying up front. This post is the "why I tried containers" half of the story. When I went to verify the setup, it wasn't actually working yet, and that's covered in <a href="/blog/hermes-profiles-to-docker-part-two">part two of the series</a>. I'm keeping this one focused on the reasoning and the container plumbing, because that part still stands even though my diagnosis turned out to be off.</p>
<h2>The Problem I Didn't Know I Had</h2>
<p>The symptom was cron replies coming from the wrong profile, and it was inconsistent, which is what made it hard to pin down. Sometimes a scheduled job would fire and default and sage would both reply. Sometimes only rex would send the update, when it wasn't his job to. It never happened with messages I sent directly — only cron.</p>
<p>My theory at the time was that the profiles shared too much. One scheduler, one process tree, no hard boundary between them, so jobs bled across. Three profiles all live in the same process, so when a cron job fired, I figured whichever gateway instance was free picked it up.</p>
<p>That theory turned out to be mostly wrong, which is a whole <a href="/blog/hermes-profiles-to-docker-part-two">separate story</a>. I'm laying it out anyway, because it's what drove me to containers, and the architecture reasoning is sound even if the diagnosis wasn't.</p>
<p>What I didn't realise at the time: gateway profiles aren't separate services. They're separate config directories, separate skill directories, separate <code>.env</code> files, and a <code>--profile</code> flag handed to the same gateway binary. Same process tree. Same supervisor. No hard boundary when one of them misbehaves.</p>
<p>I'd been treating them like independent services. They're not.</p>
<h2>What the Docs Recommended (and Why I Went Further)</h2>
<p>The Hermes docs recommend one container hosting all profiles, with <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a> (a process supervisor designed for containers) managing each profile as a first-class service. Mount the whole <code>~/.hermes</code> directory to <code>/opt/data</code>, create profiles with <code>hermes profile create</code>, let s6 start and stop them.</p>
<p>I didn't think that would solve what I was seeing. With one container and s6 running every profile, sage and rex are still co-located processes sharing a network namespace and a data directory. If the problem was jobs bleeding across profiles, bundling them into one container wouldn't stop any of it. They'd just be supervised versions of the same shared setup.</p>
<p>So I ran separate containers for sage and rex, each mounting only its own profile directory. Real process isolation: separate network namespaces, separate PID 1, separate Telegram polling loops. At the process level they genuinely cannot interfere with each other.</p>
<p>The architecture was right, it just wasn't the cause of what I was seeing. That part's in <a href="/blog/hermes-profiles-to-docker-part-two">the verification post</a>.</p>
<h2>The Mount Path Mistake</h2>
<p>The mount paths. I spent far too long on this.</p>
<p>The profile needs to live at <code>/opt/data/profiles/&#x3C;name></code> inside the container, not at <code>/opt/data</code>. I kept mounting the rex profile to <code>/opt/data</code>, and Hermes would log <code>Error: Profile 'rex' does not exist. Create it with: hermes profile create rex</code>. Same files on the host, same mount command, but Hermes couldn't find the profile because it was looking in the wrong place.</p>
<p>The fix was obvious once I saw it:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Wrong — mounts to /opt/data, Hermes expects /opt/data/profiles/rex</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Right</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data/profiles/rex</span></span></code></pre>
<p>The containers weren't the wrong architecture. The paths were wrong. A simple mistake that cost more time than it should have.</p>
<h2>The Boot Problem</h2>
<p>Containers don't start themselves. Bare-metal services auto-start through systemd. I needed the Docker containers to come up on boot too.</p>
<p><code>docker compose</code> has no native systemd integration. The cleanest approach was a systemd user service that runs <code>docker compose up -d</code> for each profile's compose file. It feels slightly off, systemd managing containers instead of services, but it's straightforward and it works.</p>
<h2>What I'd Tell Myself</h2>
<p>When I was running three gateway profiles on bare metal, I was already past what profiles are designed for. They're great for development. Spinning up a new persona with different skills takes seconds. For anything that needs to actually stay separated, the isolation story falls apart.</p>
<p>Containers were the right architecture for that. The operational complexity is real (two compose files, a custom systemd service, separate log streams), but it's the right kind of complexity. Explicit, manageable, debuggable. When rex goes wrong, I look at the rex container, not a shared process tree.</p>
<p>I still run default bare-metal. Default is the daily driver; it has no scheduled cron that fires critical replies. Sage and rex are the workers. They get containers.</p>
<p>What I got wrong was assuming that building the right architecture meant the problem was solved. It didn't. When I went to verify the containers were actually doing their job, nothing worked. Sage was running two Hermes instances, cron jobs were vanishing into the wrong database, and the symptom that started all of this was still happening.</p>
<p>That's <a href="/blog/hermes-profiles-to-docker-part-two">the debugging follow-up</a>.</p>]]></description>
            <content:encoded><![CDATA[<p>I ran Hermes (the self-hosted agent I use for <a href="/blog/setting-up-camoufox-with-hermes">reading and research</a> and a few other jobs) with three profiles: default, sage, and rex. Each was a separate persona with its own skills and its own Telegram bot.</p>
<p>They were gateway profiles. In Hermes, the gateway is the long-running process that polls Telegram, runs scheduled jobs, and executes skills. "Profiles" means one gateway binary, one process manager, just different configs passed in via <code>--profile</code>.</p>
<p>One thing worth saying up front. This post is the "why I tried containers" half of the story. When I went to verify the setup, it wasn't actually working yet, and that's covered in <a href="/blog/hermes-profiles-to-docker-part-two">part two of the series</a>. I'm keeping this one focused on the reasoning and the container plumbing, because that part still stands even though my diagnosis turned out to be off.</p>
<h2>The Problem I Didn't Know I Had</h2>
<p>The symptom was cron replies coming from the wrong profile, and it was inconsistent, which is what made it hard to pin down. Sometimes a scheduled job would fire and default and sage would both reply. Sometimes only rex would send the update, when it wasn't his job to. It never happened with messages I sent directly — only cron.</p>
<p>My theory at the time was that the profiles shared too much. One scheduler, one process tree, no hard boundary between them, so jobs bled across. Three profiles all live in the same process, so when a cron job fired, I figured whichever gateway instance was free picked it up.</p>
<p>That theory turned out to be mostly wrong, which is a whole <a href="/blog/hermes-profiles-to-docker-part-two">separate story</a>. I'm laying it out anyway, because it's what drove me to containers, and the architecture reasoning is sound even if the diagnosis wasn't.</p>
<p>What I didn't realise at the time: gateway profiles aren't separate services. They're separate config directories, separate skill directories, separate <code>.env</code> files, and a <code>--profile</code> flag handed to the same gateway binary. Same process tree. Same supervisor. No hard boundary when one of them misbehaves.</p>
<p>I'd been treating them like independent services. They're not.</p>
<h2>What the Docs Recommended (and Why I Went Further)</h2>
<p>The Hermes docs recommend one container hosting all profiles, with <a href="https://github.com/just-containers/s6-overlay">s6-overlay</a> (a process supervisor designed for containers) managing each profile as a first-class service. Mount the whole <code>~/.hermes</code> directory to <code>/opt/data</code>, create profiles with <code>hermes profile create</code>, let s6 start and stop them.</p>
<p>I didn't think that would solve what I was seeing. With one container and s6 running every profile, sage and rex are still co-located processes sharing a network namespace and a data directory. If the problem was jobs bleeding across profiles, bundling them into one container wouldn't stop any of it. They'd just be supervised versions of the same shared setup.</p>
<p>So I ran separate containers for sage and rex, each mounting only its own profile directory. Real process isolation: separate network namespaces, separate PID 1, separate Telegram polling loops. At the process level they genuinely cannot interfere with each other.</p>
<p>The architecture was right, it just wasn't the cause of what I was seeing. That part's in <a href="/blog/hermes-profiles-to-docker-part-two">the verification post</a>.</p>
<h2>The Mount Path Mistake</h2>
<p>The mount paths. I spent far too long on this.</p>
<p>The profile needs to live at <code>/opt/data/profiles/&#x3C;name></code> inside the container, not at <code>/opt/data</code>. I kept mounting the rex profile to <code>/opt/data</code>, and Hermes would log <code>Error: Profile 'rex' does not exist. Create it with: hermes profile create rex</code>. Same files on the host, same mount command, but Hermes couldn't find the profile because it was looking in the wrong place.</p>
<p>The fix was obvious once I saw it:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#6A737D"># Wrong — mounts to /opt/data, Hermes expects /opt/data/profiles/rex</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Right</span></span>
<span class="line"><span style="color:#E1E4E8">- </span><span style="color:#9ECBFF">/home/luke/.hermes/profiles/rex:/opt/data/profiles/rex</span></span></code></pre>
<p>The containers weren't the wrong architecture. The paths were wrong. A simple mistake that cost more time than it should have.</p>
<h2>The Boot Problem</h2>
<p>Containers don't start themselves. Bare-metal services auto-start through systemd. I needed the Docker containers to come up on boot too.</p>
<p><code>docker compose</code> has no native systemd integration. The cleanest approach was a systemd user service that runs <code>docker compose up -d</code> for each profile's compose file. It feels slightly off, systemd managing containers instead of services, but it's straightforward and it works.</p>
<h2>What I'd Tell Myself</h2>
<p>When I was running three gateway profiles on bare metal, I was already past what profiles are designed for. They're great for development. Spinning up a new persona with different skills takes seconds. For anything that needs to actually stay separated, the isolation story falls apart.</p>
<p>Containers were the right architecture for that. The operational complexity is real (two compose files, a custom systemd service, separate log streams), but it's the right kind of complexity. Explicit, manageable, debuggable. When rex goes wrong, I look at the rex container, not a shared process tree.</p>
<p>I still run default bare-metal. Default is the daily driver; it has no scheduled cron that fires critical replies. Sage and rex are the workers. They get containers.</p>
<p>What I got wrong was assuming that building the right architecture meant the problem was solved. It didn't. When I went to verify the containers were actually doing their job, nothing worked. Sage was running two Hermes instances, cron jobs were vanishing into the wrong database, and the symptom that started all of this was still happening.</p>
<p>That's <a href="/blog/hermes-profiles-to-docker-part-two">the debugging follow-up</a>.</p>]]></content:encoded>
            <category>hermes</category>
            <category>homelab</category>
            <category>docker</category>
        </item>
        <item>
            <title><![CDATA[Dropping Obsidian Sync for Syncthing: What I Learned Setting Up a Headless Linux Sync Hub]]></title>
            <link>https://lukemanning.ie/blog/syncthing-obsidian-vault-sync</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/syncthing-obsidian-vault-sync</guid>
            <pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>Obsidian Sync costs up to $120 a year if you're on Sync Plus, and paying monthly. That's the number that made me actually do something about it.</p>
<p>I was curious about using Obsidian Sync because people generally were happy that it worked. At least people on reddit were. However I did not want yet another subscription. As much as I like the work that Obsidian do, I just couldn't justify it.</p>
<p>I went with Syncthing because it's peer-to-peer (no third-party server holding my files), it's free, and I've read several threads and watched a few YouTube videos where people swear by it. My always-on NucBox (Ubuntu 24.04, headless) acts as the hub. Windows desktop and Android phone sync to it. I genuinely forget it's there most of the time.</p>
<p>This post isn't a tutorial. The Syncthing docs are fine. What I kept running into was the specific problem of configuring a headless Linux machine without a browser, and nobody really spelled that part out clearly. So here's my setup, including the bits I had to figure out myself.</p>
<h2>Getting to the Web UI</h2>
<p>Most Syncthing guides assume you can just open <code>localhost:8384</code> in a browser on the same machine running Syncthing. My NucBox doesn't have a browser. Or a screen. It's just a small fanless PC sitting under my desk running Ubuntu without a Desktop environment.</p>
<p>My first thought was to install a VNC server or something. I wasn't a fan of that approach and after some more time Googling around I eventually found the SSH tunnel approach, which is way simpler than I expected.</p>
<p>Syncthing's web UI runs on port 8384 on the NucBox. I mapped that to a local port on my Windows machine over SSH:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">ssh</span><span style="color:#79B8FF"> -L</span><span style="color:#9ECBFF"> 9384:localhost:8384</span><span style="color:#9ECBFF"> luke@</span><span style="color:#F97583">&#x3C;</span><span style="color:#9ECBFF">nucbox-i</span><span style="color:#E1E4E8">p</span><span style="color:#F97583">></span><span style="color:#79B8FF"> -N</span></span></code></pre>
<p>Then opened <code>localhost:9384</code> in my browser on my desktop. Syncthing defaults to no auth on the web UI, and I didn't love the idea of an unauthenticated interface accessible over the network. However given this was all running on localhost, I figured that was a problem for another day.</p>
<p>I was already running Syncthing on Windows, which was using port 8384. That's why I mapped to 9384 instead. Anything unused works.</p>
<h2>What I Ran on the NucBox</h2>
<p>Three commands:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">sudo</span><span style="color:#9ECBFF"> apt</span><span style="color:#9ECBFF"> install</span><span style="color:#9ECBFF"> syncthing</span></span>
<span class="line"><span style="color:#B392F0">systemctl</span><span style="color:#79B8FF"> --user</span><span style="color:#9ECBFF"> enable</span><span style="color:#9ECBFF"> syncthing</span></span>
<span class="line"><span style="color:#B392F0">systemctl</span><span style="color:#79B8FF"> --user</span><span style="color:#9ECBFF"> start</span><span style="color:#9ECBFF"> syncthing</span></span></code></pre>
<p>After those, I checked the service was actually running with <code>systemctl --user status syncthing</code>. Active and running. That was my "okay, it's alive" moment.</p>
<p>Actually, one thing I assumed I'd have to deal with. <code>systemctl --user</code> services on a headless machine don't survive logout unless you enable "lingering." Without it, Syncthing stops the second you close your SSH session. I checked mine:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">loginctl</span><span style="color:#9ECBFF"> show-user</span><span style="color:#9ECBFF"> luke</span><span style="color:#F97583"> |</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> Linger</span></span></code></pre>
<p>Turns out it was already enabled (<code>Linger=yes</code>). I'm guessing something during the Ubuntu install set it up — I definitely never ran <code>loginctl enable-linger</code> myself. But if yours shows <code>Linger=no</code>, you'll need:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">loginctl</span><span style="color:#9ECBFF"> enable-linger</span><span style="color:#9ECBFF"> luke</span></span></code></pre>
<p>This is probably the most headless-specific gotcha in the whole setup. And I suspect a lot of guides skip it because they assume you have a desktop login session.</p>
<p>The Ubuntu package is fine. I didn't need any third-party repos.</p>
<p>Once the tunnel was up it was the standard Syncthing workflow. I added a folder, pointed it at my vault, and shared it with my other devices. I installed Syncthing-Fork on Android from the F-Droid store. Paired devices using the Device ID from <strong>Actions → Show ID</strong>.</p>
<p>The Android app can scan a QR code from that same screen, which is much easier than typing out a 60-character device ID by hand. It can also discover other Syncthing devices on the same network, which was super useful for me.</p>
<h2>The .stignore Thing</h2>
<p>I'd been running Syncthing for maybe a day when I started seeing orange warning triangles in the UI. Low-level conflicts on <code>.obsidian/workspace.json</code> and <code>.obsidian/workspace-mobile.json</code>. Nothing catastrophic, but it was noise I didn't need.</p>
<p>Turns out Obsidian rewrites those files every time you open it — they store open tabs, cursor positions, that kind of thing. Different on every device, changing constantly. No wonder Syncthing was confused.</p>
<p>A bit of searching told me about <code>.stignore</code>. I added this at the root of the vault:</p>
<pre><code>.obsidian/workspace.json
.obsidian/workspace-mobile.json
.trash/
</code></pre>
<p>The rest of <code>.obsidian/</code> (plugins, themes, settings) syncs fine. That's the stuff I actually want consistent across machines.</p>
<h2>A Few Weeks In</h2>
<p>On the same LAN it syncs in seconds. Over the internet it handles NAT traversal automatically. I didn't have to configure any port forwarding or open firewall ports. The NucBox isn't running ufw, so there was nothing to touch there. It just worked.</p>
<p>Syncthing apparently used to use relay servers for this, but they were removed a while back (I think around v1.27). Whatever it's doing now, I didn't have to think about it.</p>
<p>I've had this running for a few days now across the NucBox, my Windows desktop, and my Android phone. I haven't thought about it once since setting it up, which is exactly what I wanted.</p>
<p>The $120 a year I was considering paying Obsidian Sync? Going toward something else now.</p>]]></description>
            <content:encoded><![CDATA[<p>Obsidian Sync costs up to $120 a year if you're on Sync Plus, and paying monthly. That's the number that made me actually do something about it.</p>
<p>I was curious about using Obsidian Sync because people generally were happy that it worked. At least people on reddit were. However I did not want yet another subscription. As much as I like the work that Obsidian do, I just couldn't justify it.</p>
<p>I went with Syncthing because it's peer-to-peer (no third-party server holding my files), it's free, and I've read several threads and watched a few YouTube videos where people swear by it. My always-on NucBox (Ubuntu 24.04, headless) acts as the hub. Windows desktop and Android phone sync to it. I genuinely forget it's there most of the time.</p>
<p>This post isn't a tutorial. The Syncthing docs are fine. What I kept running into was the specific problem of configuring a headless Linux machine without a browser, and nobody really spelled that part out clearly. So here's my setup, including the bits I had to figure out myself.</p>
<h2>Getting to the Web UI</h2>
<p>Most Syncthing guides assume you can just open <code>localhost:8384</code> in a browser on the same machine running Syncthing. My NucBox doesn't have a browser. Or a screen. It's just a small fanless PC sitting under my desk running Ubuntu without a Desktop environment.</p>
<p>My first thought was to install a VNC server or something. I wasn't a fan of that approach and after some more time Googling around I eventually found the SSH tunnel approach, which is way simpler than I expected.</p>
<p>Syncthing's web UI runs on port 8384 on the NucBox. I mapped that to a local port on my Windows machine over SSH:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">ssh</span><span style="color:#79B8FF"> -L</span><span style="color:#9ECBFF"> 9384:localhost:8384</span><span style="color:#9ECBFF"> luke@</span><span style="color:#F97583">&#x3C;</span><span style="color:#9ECBFF">nucbox-i</span><span style="color:#E1E4E8">p</span><span style="color:#F97583">></span><span style="color:#79B8FF"> -N</span></span></code></pre>
<p>Then opened <code>localhost:9384</code> in my browser on my desktop. Syncthing defaults to no auth on the web UI, and I didn't love the idea of an unauthenticated interface accessible over the network. However given this was all running on localhost, I figured that was a problem for another day.</p>
<p>I was already running Syncthing on Windows, which was using port 8384. That's why I mapped to 9384 instead. Anything unused works.</p>
<h2>What I Ran on the NucBox</h2>
<p>Three commands:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">sudo</span><span style="color:#9ECBFF"> apt</span><span style="color:#9ECBFF"> install</span><span style="color:#9ECBFF"> syncthing</span></span>
<span class="line"><span style="color:#B392F0">systemctl</span><span style="color:#79B8FF"> --user</span><span style="color:#9ECBFF"> enable</span><span style="color:#9ECBFF"> syncthing</span></span>
<span class="line"><span style="color:#B392F0">systemctl</span><span style="color:#79B8FF"> --user</span><span style="color:#9ECBFF"> start</span><span style="color:#9ECBFF"> syncthing</span></span></code></pre>
<p>After those, I checked the service was actually running with <code>systemctl --user status syncthing</code>. Active and running. That was my "okay, it's alive" moment.</p>
<p>Actually, one thing I assumed I'd have to deal with. <code>systemctl --user</code> services on a headless machine don't survive logout unless you enable "lingering." Without it, Syncthing stops the second you close your SSH session. I checked mine:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">loginctl</span><span style="color:#9ECBFF"> show-user</span><span style="color:#9ECBFF"> luke</span><span style="color:#F97583"> |</span><span style="color:#B392F0"> grep</span><span style="color:#9ECBFF"> Linger</span></span></code></pre>
<p>Turns out it was already enabled (<code>Linger=yes</code>). I'm guessing something during the Ubuntu install set it up — I definitely never ran <code>loginctl enable-linger</code> myself. But if yours shows <code>Linger=no</code>, you'll need:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">loginctl</span><span style="color:#9ECBFF"> enable-linger</span><span style="color:#9ECBFF"> luke</span></span></code></pre>
<p>This is probably the most headless-specific gotcha in the whole setup. And I suspect a lot of guides skip it because they assume you have a desktop login session.</p>
<p>The Ubuntu package is fine. I didn't need any third-party repos.</p>
<p>Once the tunnel was up it was the standard Syncthing workflow. I added a folder, pointed it at my vault, and shared it with my other devices. I installed Syncthing-Fork on Android from the F-Droid store. Paired devices using the Device ID from <strong>Actions → Show ID</strong>.</p>
<p>The Android app can scan a QR code from that same screen, which is much easier than typing out a 60-character device ID by hand. It can also discover other Syncthing devices on the same network, which was super useful for me.</p>
<h2>The .stignore Thing</h2>
<p>I'd been running Syncthing for maybe a day when I started seeing orange warning triangles in the UI. Low-level conflicts on <code>.obsidian/workspace.json</code> and <code>.obsidian/workspace-mobile.json</code>. Nothing catastrophic, but it was noise I didn't need.</p>
<p>Turns out Obsidian rewrites those files every time you open it — they store open tabs, cursor positions, that kind of thing. Different on every device, changing constantly. No wonder Syncthing was confused.</p>
<p>A bit of searching told me about <code>.stignore</code>. I added this at the root of the vault:</p>
<pre><code>.obsidian/workspace.json
.obsidian/workspace-mobile.json
.trash/
</code></pre>
<p>The rest of <code>.obsidian/</code> (plugins, themes, settings) syncs fine. That's the stuff I actually want consistent across machines.</p>
<h2>A Few Weeks In</h2>
<p>On the same LAN it syncs in seconds. Over the internet it handles NAT traversal automatically. I didn't have to configure any port forwarding or open firewall ports. The NucBox isn't running ufw, so there was nothing to touch there. It just worked.</p>
<p>Syncthing apparently used to use relay servers for this, but they were removed a while back (I think around v1.27). Whatever it's doing now, I didn't have to think about it.</p>
<p>I've had this running for a few days now across the NucBox, my Windows desktop, and my Android phone. I haven't thought about it once since setting it up, which is exactly what I wanted.</p>
<p>The $120 a year I was considering paying Obsidian Sync? Going toward something else now.</p>]]></content:encoded>
            <category>homelab</category>
            <category>obsidian</category>
            <category>ubuntu</category>
        </item>
        <item>
            <title><![CDATA[I Nuked My X History with a ZeroWork TaskBot]]></title>
            <link>https://lukemanning.ie/blog/nuking-x-history-zerowork-taskbot</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/nuking-x-history-zerowork-taskbot</guid>
            <pubDate>Sun, 22 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>I wanted to start using X more actively. Not for "branding" or "growth hacking" - just personal engagement, building relationships, connecting with likeminded people, sharing what I'm building. You know, building in public.</p>
<h2>I Just Wanted to Delete Old Tweets</h2>
<p>My X history was flooded with hundreds of tweets and retweets after years of neglect. There were so many corporate social media sharing campaign posts, random competition retweets, and so on. None of that history represented the new image I wanted to portray on X.</p>
<p>But I've had my X profile for years, and I didn't want to delete it entirely.</p>
<p>So I needed to nuke my history.</p>
<p>This is where it got annoying. I looked around for tools to do this and found... third-party options. Sketchy-looking websites asking for my credentials. Paid subscriptions just to delete tweets. Nothing that felt safe or worth using.</p>
<p>(I don't know if I'm overly paranoid, but giving my X login to a random website doesn't sit right with me.)</p>
<p>I'd used ZeroWork before. I was pretty active in their community for a few months. Hadn't touched it in a while, but this problem sparked my interest in using it again.</p>
<p>So I built my own solution.</p>
<hr>
<h2>Automating with ZeroWork</h2>
<p>This automation completely cleans up your Twitter/X profile by automatically deleting all your original tweets and undoing all your retweets. Instead of manually clicking delete on hundreds of posts, this tool does the entire job for you in one go.</p>
<p>The process works like this:</p>
<ol>
<li>Opens your profile - It starts by visiting your Twitter profile page</li>
<li>Checks your Posts tab - It looks at your main Posts feed first to see if there are any tweets to clean up</li>
<li>Deletes/undoes every tweet:
<ul>
<li>For retweets: It finds the "Undo Retweet" button and clicks it, removing your retweet</li>
<li>For your own tweets: It clicks the more menu (⋮), selects "Delete," and confirms the deletion</li>
<li>It tracks how many tweets it deleted and how many retweets it undid</li>
</ul>
</li>
<li>Automatically switches to Replies - If you don't have posts, it seamlessly moves to your Replies tab and cleans those up too</li>
<li>Reports results - When it finishes, it tells you exactly what happened: "Posts Feed Done | X Tweets Deleted | Y Retweets Undone"</li>
</ol>
<p>Key features:</p>
<ul>
<li>Handles both posts and replies automatically</li>
<li>Distinguishes between your tweets and retweets</li>
<li>Safe, deliberate process with confirmation steps</li>
<li>Auto-scrolls to find more tweets as it goes</li>
</ul>
<div class="border border-terminal-dim/15 p-4 my-6">
  <div class="text-xs text-terminal-dim mb-2">
    screenshot: <span class="text-terminal-accent">twitter_taskbot.png</span>
  </div>
  <img src="/images/twitter_taskbot.png" alt="ZeroWork TaskBot showing the delete tweets automation workflow" class="rounded-sm w-full border border-terminal-dim/10">
</div>
<p>Getting the TaskBot to work reliably took way more trial and error than I expected.</p>
<p>The main thing that tripped me up was loop logic. If I set the loop iterations too high, it would randomly fail.</p>
<p>So I realized the best thing to do was limit the loop to a single tweet, then add an "After Repeat" block to check if there were still more tweets. A single loop per single tweet was way more reliable than trying to process multiple tweets per loop.</p>
<p>I also had to identify whether each object was a post or a retweet - the "delete" and "un-repost" actions are different, so I needed separate branches.</p>
<p>But that wasn't the weird part.</p>
<h2>The Weird Stuff</h2>
<p>Older reposts/retweets were confusing. For some reason, they showed up in my Replies feed as me having retweeted them, but there was no "Undo Retweet" button. This broke my workflow since I couldn't just undo them directly.</p>
<p>I figured it out relatively quickly: I had to repost the item again, then undo <em>that</em> repost. Only then would it finally get removed from the Replies page. This makes zero sense to me, but it worked.</p>
<p>Then there were replies. Handling tweets where either I was replying to someone, or someone was replying to me - these got mixed together in the feed. The automation would try to delete tweets that weren't mine, which obviously failed.</p>
<p>I solved this with a counter. My loop would always review the tweet at position zero. If that tweet belonged to someone else, I'd increment the counter, so it would review the tweet at position one instead. This way I could skip over other people's replies and only target mine.</p>
<p>It wasn't plug-and-play. I was pretty rusty on ZeroWork, and I spent many hours tweaking and re-running the TaskBot before it finally worked smoothly.</p>
<p>I let it run while I went to do other things. When I came back about 3 hours later, it had finished.</p>
<p>It had removed about 400 reposts and a few tweets.</p>
<p>Now I have a fresh X profile with 0 posts and 0 retweets.</p>
<p>That's it. Clean slate. Ready to start posting for real this time.</p>]]></description>
            <content:encoded><![CDATA[<p>I wanted to start using X more actively. Not for "branding" or "growth hacking" - just personal engagement, building relationships, connecting with likeminded people, sharing what I'm building. You know, building in public.</p>
<h2>I Just Wanted to Delete Old Tweets</h2>
<p>My X history was flooded with hundreds of tweets and retweets after years of neglect. There were so many corporate social media sharing campaign posts, random competition retweets, and so on. None of that history represented the new image I wanted to portray on X.</p>
<p>But I've had my X profile for years, and I didn't want to delete it entirely.</p>
<p>So I needed to nuke my history.</p>
<p>This is where it got annoying. I looked around for tools to do this and found... third-party options. Sketchy-looking websites asking for my credentials. Paid subscriptions just to delete tweets. Nothing that felt safe or worth using.</p>
<p>(I don't know if I'm overly paranoid, but giving my X login to a random website doesn't sit right with me.)</p>
<p>I'd used ZeroWork before. I was pretty active in their community for a few months. Hadn't touched it in a while, but this problem sparked my interest in using it again.</p>
<p>So I built my own solution.</p>
<hr>
<h2>Automating with ZeroWork</h2>
<p>This automation completely cleans up your Twitter/X profile by automatically deleting all your original tweets and undoing all your retweets. Instead of manually clicking delete on hundreds of posts, this tool does the entire job for you in one go.</p>
<p>The process works like this:</p>
<ol>
<li>Opens your profile - It starts by visiting your Twitter profile page</li>
<li>Checks your Posts tab - It looks at your main Posts feed first to see if there are any tweets to clean up</li>
<li>Deletes/undoes every tweet:
<ul>
<li>For retweets: It finds the "Undo Retweet" button and clicks it, removing your retweet</li>
<li>For your own tweets: It clicks the more menu (⋮), selects "Delete," and confirms the deletion</li>
<li>It tracks how many tweets it deleted and how many retweets it undid</li>
</ul>
</li>
<li>Automatically switches to Replies - If you don't have posts, it seamlessly moves to your Replies tab and cleans those up too</li>
<li>Reports results - When it finishes, it tells you exactly what happened: "Posts Feed Done | X Tweets Deleted | Y Retweets Undone"</li>
</ol>
<p>Key features:</p>
<ul>
<li>Handles both posts and replies automatically</li>
<li>Distinguishes between your tweets and retweets</li>
<li>Safe, deliberate process with confirmation steps</li>
<li>Auto-scrolls to find more tweets as it goes</li>
</ul>
<div class="border border-terminal-dim/15 p-4 my-6">
  <div class="text-xs text-terminal-dim mb-2">
    screenshot: <span class="text-terminal-accent">twitter_taskbot.png</span>
  </div>
  <img src="/images/twitter_taskbot.png" alt="ZeroWork TaskBot showing the delete tweets automation workflow" class="rounded-sm w-full border border-terminal-dim/10">
</div>
<p>Getting the TaskBot to work reliably took way more trial and error than I expected.</p>
<p>The main thing that tripped me up was loop logic. If I set the loop iterations too high, it would randomly fail.</p>
<p>So I realized the best thing to do was limit the loop to a single tweet, then add an "After Repeat" block to check if there were still more tweets. A single loop per single tweet was way more reliable than trying to process multiple tweets per loop.</p>
<p>I also had to identify whether each object was a post or a retweet - the "delete" and "un-repost" actions are different, so I needed separate branches.</p>
<p>But that wasn't the weird part.</p>
<h2>The Weird Stuff</h2>
<p>Older reposts/retweets were confusing. For some reason, they showed up in my Replies feed as me having retweeted them, but there was no "Undo Retweet" button. This broke my workflow since I couldn't just undo them directly.</p>
<p>I figured it out relatively quickly: I had to repost the item again, then undo <em>that</em> repost. Only then would it finally get removed from the Replies page. This makes zero sense to me, but it worked.</p>
<p>Then there were replies. Handling tweets where either I was replying to someone, or someone was replying to me - these got mixed together in the feed. The automation would try to delete tweets that weren't mine, which obviously failed.</p>
<p>I solved this with a counter. My loop would always review the tweet at position zero. If that tweet belonged to someone else, I'd increment the counter, so it would review the tweet at position one instead. This way I could skip over other people's replies and only target mine.</p>
<p>It wasn't plug-and-play. I was pretty rusty on ZeroWork, and I spent many hours tweaking and re-running the TaskBot before it finally worked smoothly.</p>
<p>I let it run while I went to do other things. When I came back about 3 hours later, it had finished.</p>
<p>It had removed about 400 reposts and a few tweets.</p>
<p>Now I have a fresh X profile with 0 posts and 0 retweets.</p>
<p>That's it. Clean slate. Ready to start posting for real this time.</p>]]></content:encoded>
            <category>homelab</category>
            <category>zerowork</category>
        </item>
        <item>
            <title><![CDATA[Setting Up an Asus Router Behind an ISP Modem: My Unexpected Gotchas]]></title>
            <link>https://lukemanning.ie/blog/asus-router-setup-journey</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/asus-router-setup-journey</guid>
            <pubDate>Thu, 27 Nov 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>I bought an Asus RT-BE88U router because my ISP-provided Fritzbox WiFi was, to put it mildly, absolute garbage. What I thought would be a simple "plug it in and go" situation turned into a two-hour troubleshooting adventure that taught me more about networking than I ever wanted to know.</p>
<p>This is part of a pattern I notice in my work: <a href="/blog/unraid-docker-label-fix">home server maintenance, recovery tasks, and infrastructure projects</a> don't always go smoothly, but they're worth documenting.</p>
<p>This is that story. No pretending I knew what I was doing. Just the actual messy process of figuring it out.</p>
<h2>The Starting Point</h2>
<p><strong>What I had:</strong></p>
<ul>
<li>Brand new Asus RT-BE88U still in the box</li>
<li>Fritzbox modem/router combo from Digiweb (my ISP)</li>
<li>Shockingly bad WiFi coverage</li>
<li>Confidence that this would be easy</li>
</ul>
<p><strong>What I knew about networking:</strong></p>
<ul>
<li>Enterprise troubleshooting, VLANs, routing protocols, complex network issues</li>
<li>Consumer router setups behind ISP modems? Not my area</li>
</ul>
<h2>Step 1: Understanding Double NAT (Or Not Understanding It)</h2>
<p>The first thing I learned is that my Fritzbox isn't just a modem - it's a router too. Which means when I add the Asus, I'm creating what's called "double NAT."</p>
<p>When someone first explained this to me, I got a whole technical breakdown about Network Address Translation and routing tables and... honestly, this is all enterprise networking stuff I deal with daily. But I needed it in consumer terms.</p>
<p>As I understand it now: both routers are translating network addresses. The Fritzbox translates from your ISP's public IP to its internal network (192.168.178.x), then the Asus translates again to its own network (192.168.50.x). Packets go through two address translations before reaching the internet.</p>
<p>At this point, I was worried this was going to cause problems. Everything I read said "double NAT is bad" but nobody could really explain why in a way I understood. Gaming? Servers? I don't do any of that.</p>
<p>I decided to just try it and see what happens.</p>
<h2>Step 2: Physical Setup</h2>
<p>This part was actually straightforward:</p>
<ol>
<li>Unbox the Asus (finally)</li>
<li>Connect ethernet cable from Fritzbox LAN port to Asus WAN port (the blue one)</li>
<li>Power everything up</li>
<li>Download the Asus Router app</li>
</ol>
<p>The app walked me through initial setup, asking me to choose between creating a new network or extending an existing one. I chose "new network" and selected DHCP as my WAN type, accepting the other defaults.</p>
<p>Then came the fun part: picking a WiFi name. After cycling through metal puns, Stephen King references, and various nerdy jokes, I landed on something that made me laugh. That part took longer than the actual setup.</p>
<h2>Step 3: Everything Seemed Fine (It Wasn't)</h2>
<p>The Asus app said setup was complete. I could see my new WiFi network broadcasting. The app showed everything as connected and happy.</p>
<p>But when I tried to actually use the internet... nothing. No connection.</p>
<h2>Step 4: Down the Troubleshooting Rabbit Hole</h2>
<p><strong>Problem #1: WAN IP showing 0.0.0.0</strong></p>
<p>The Asus router settings showed my WAN IP as <code>0.0.0.0</code>, which apparently means "I'm not getting an IP address from the Fritzbox."</p>
<p>First attempts to fix:</p>
<ul>
<li>Rebooted the Asus (didn't work)</li>
<li>Rebooted the Fritzbox (didn't work)</li>
<li>Checked the physical connections (all good)</li>
<li>Questioned my life choices (very effective, but didn't fix the internet)</li>
</ul>
<p><strong>Problem #2: The DHCP Mystery</strong></p>
<p>I logged into my Fritzbox (at <code>fritz.box</code> or <code>192.168.178.1</code>) and found something interesting. The Fritzbox could see a device connected to port 4 - listed as "PC-86-FB-ED-94-40-49" (definitely my Asus router based on the MAC address). But it wasn't assigning it an IP.</p>
<p>The DHCP server was enabled, with plenty of available addresses in the range. So why wasn't it working?</p>
<p>At this point I was stuck. The Asus should be getting an IP via DHCP automatically, but it just wasn't happening. I checked the Fritzbox settings - everything looked fine. No obvious DHCP conflicts, no MAC filtering blocking the Asus.</p>
<p>I thought: if DHCP isn't working, just assign a static IP manually. Standard troubleshooting step.</p>
<p>I picked 192.168.178.50 - well above the Fritzbox's default DHCP range (.20–.199) to avoid conflicts.</p>
<p><strong>The Fix: Static IP Configuration</strong></p>
<p>In the Asus router settings, I changed from DHCP to Static IP and manually entered:</p>
<ul>
<li>IP Address: <code>192.168.178.50</code></li>
<li>Subnet Mask: <code>255.255.255.0</code></li>
<li>Default Gateway: <code>192.168.178.1</code></li>
<li>DNS Server: <code>192.168.178.1</code></li>
</ul>
<p>This bypassed the DHCP issue entirely. The Asus app suddenly showed a valid WAN connection!</p>
<p>Progress! Except...</p>
<p><strong>Problem #3: Router Has Internet, Devices Don't</strong></p>
<p>My phone connected to the new WiFi just fine. It got an IP address in the <code>192.168.50.x</code> range. I could access the Asus router settings at its default address, <code>192.168.50.1</code>. Everything looked perfect.</p>
<p>But still no actual internet.</p>
<p><strong>The Diagnostic That Revealed Everything</strong></p>
<p>I tried accessing the Fritzbox (<code>192.168.178.1</code>) from my phone while connected to the Asus WiFi. It failed completely.</p>
<p>The Asus couldn't talk to the Fritzbox. Even though they were physically connected. Even though the Asus showed a valid WAN connection. The routing between them was completely broken.</p>
<p>I verified all the WAN settings were correct. I tried pinging from the router's diagnostic tools. Nothing worked.</p>
<p>Then I tried something stupid simple.</p>
<p><strong>The Actual Fix: Wrong Port</strong></p>
<p>I unplugged the ethernet cable from LAN port 4 on the Fritzbox and plugged it into port 3 instead.</p>
<p>It immediately worked.</p>
<p>Port 4 apparently had some restriction or configuration I never figured out. I checked Fritzbox logs, firewall rules, VLAN settings, port-specific configs - nothing obvious. Port 3 worked, so I moved on.</p>
<h2>What I Learned</h2>
<p><strong>1. Start Simple</strong>
When something doesn't work, start with the physical layer. Different port? Different cable? Sometimes the answer is that basic.</p>
<p><strong>2. Work Layer by Layer</strong></p>
<ul>
<li>Can the router see the modem? (Yes - device showed up in Fritzbox)</li>
<li>Can the router get an IP? (No - fixed with static IP)</li>
<li>Can the router reach the internet? (No - wrong port)</li>
<li>Can devices reach the router? (Yes)</li>
<li>Can devices reach the internet through the router? (Finally, yes)</li>
</ul>
<p><strong>3. Double NAT Isn't Scary</strong>
For 95% of home use, double NAT works perfectly fine. The internet loves to make it sound like a critical issue, but for browsing, streaming, and normal life, it's completely transparent.</p>
<p><strong>4. Learning in Public Means Sharing the Mess</strong>
I didn't know what I was doing. I made wrong assumptions. I tried things that didn't work. And that's the point - this is what real learning looks like. Not polished tutorials where everything works perfectly the first time.</p>
<h2>The Current Setup</h2>
<p>I now have:</p>
<ul>
<li>Fritzbox handling the ISP connection</li>
<li>Asus RT-BE88U connected via LAN port 3 (not 4, very important)</li>
<li>Static IP configuration on the WAN side</li>
<li>Separate 2.4GHz, 5GHz, and 6GHz WiFi networks</li>
<li>Coverage throughout my entire house</li>
<li>No more dead zones</li>
</ul>
<p>Is it optimal? Probably not. Could I further optimize by putting the Fritzbox in bridge mode? Sure. But it works, it's fast, and I'm not touching it again until something breaks.</p>
<h2>So... What Actually Worked?</h2>
<p>I went with static IP configuration in the end. DHCP just refused to cooperate on the Fritzbox (port 4 was the actual problem, but I didn't know that yet), and once I figured out how to set the Asus manually, everything connected immediately.</p>
<p>Double NAT? I ended up leaving it alone. Everything I need - browsing, streaming, regular stuff - works fine. If I ever need port forwarding for something, I'll deal with it then.</p>
<p>The whole thing took way longer than it should have, mostly because I was plugged into the wrong LAN port. But I got there.</p>
<hr>
<p><em>Having router troubles? Different experience? Let me know - I'm always learning.</em></p>]]></description>
            <content:encoded><![CDATA[<p>I bought an Asus RT-BE88U router because my ISP-provided Fritzbox WiFi was, to put it mildly, absolute garbage. What I thought would be a simple "plug it in and go" situation turned into a two-hour troubleshooting adventure that taught me more about networking than I ever wanted to know.</p>
<p>This is part of a pattern I notice in my work: <a href="/blog/unraid-docker-label-fix">home server maintenance, recovery tasks, and infrastructure projects</a> don't always go smoothly, but they're worth documenting.</p>
<p>This is that story. No pretending I knew what I was doing. Just the actual messy process of figuring it out.</p>
<h2>The Starting Point</h2>
<p><strong>What I had:</strong></p>
<ul>
<li>Brand new Asus RT-BE88U still in the box</li>
<li>Fritzbox modem/router combo from Digiweb (my ISP)</li>
<li>Shockingly bad WiFi coverage</li>
<li>Confidence that this would be easy</li>
</ul>
<p><strong>What I knew about networking:</strong></p>
<ul>
<li>Enterprise troubleshooting, VLANs, routing protocols, complex network issues</li>
<li>Consumer router setups behind ISP modems? Not my area</li>
</ul>
<h2>Step 1: Understanding Double NAT (Or Not Understanding It)</h2>
<p>The first thing I learned is that my Fritzbox isn't just a modem - it's a router too. Which means when I add the Asus, I'm creating what's called "double NAT."</p>
<p>When someone first explained this to me, I got a whole technical breakdown about Network Address Translation and routing tables and... honestly, this is all enterprise networking stuff I deal with daily. But I needed it in consumer terms.</p>
<p>As I understand it now: both routers are translating network addresses. The Fritzbox translates from your ISP's public IP to its internal network (192.168.178.x), then the Asus translates again to its own network (192.168.50.x). Packets go through two address translations before reaching the internet.</p>
<p>At this point, I was worried this was going to cause problems. Everything I read said "double NAT is bad" but nobody could really explain why in a way I understood. Gaming? Servers? I don't do any of that.</p>
<p>I decided to just try it and see what happens.</p>
<h2>Step 2: Physical Setup</h2>
<p>This part was actually straightforward:</p>
<ol>
<li>Unbox the Asus (finally)</li>
<li>Connect ethernet cable from Fritzbox LAN port to Asus WAN port (the blue one)</li>
<li>Power everything up</li>
<li>Download the Asus Router app</li>
</ol>
<p>The app walked me through initial setup, asking me to choose between creating a new network or extending an existing one. I chose "new network" and selected DHCP as my WAN type, accepting the other defaults.</p>
<p>Then came the fun part: picking a WiFi name. After cycling through metal puns, Stephen King references, and various nerdy jokes, I landed on something that made me laugh. That part took longer than the actual setup.</p>
<h2>Step 3: Everything Seemed Fine (It Wasn't)</h2>
<p>The Asus app said setup was complete. I could see my new WiFi network broadcasting. The app showed everything as connected and happy.</p>
<p>But when I tried to actually use the internet... nothing. No connection.</p>
<h2>Step 4: Down the Troubleshooting Rabbit Hole</h2>
<p><strong>Problem #1: WAN IP showing 0.0.0.0</strong></p>
<p>The Asus router settings showed my WAN IP as <code>0.0.0.0</code>, which apparently means "I'm not getting an IP address from the Fritzbox."</p>
<p>First attempts to fix:</p>
<ul>
<li>Rebooted the Asus (didn't work)</li>
<li>Rebooted the Fritzbox (didn't work)</li>
<li>Checked the physical connections (all good)</li>
<li>Questioned my life choices (very effective, but didn't fix the internet)</li>
</ul>
<p><strong>Problem #2: The DHCP Mystery</strong></p>
<p>I logged into my Fritzbox (at <code>fritz.box</code> or <code>192.168.178.1</code>) and found something interesting. The Fritzbox could see a device connected to port 4 - listed as "PC-86-FB-ED-94-40-49" (definitely my Asus router based on the MAC address). But it wasn't assigning it an IP.</p>
<p>The DHCP server was enabled, with plenty of available addresses in the range. So why wasn't it working?</p>
<p>At this point I was stuck. The Asus should be getting an IP via DHCP automatically, but it just wasn't happening. I checked the Fritzbox settings - everything looked fine. No obvious DHCP conflicts, no MAC filtering blocking the Asus.</p>
<p>I thought: if DHCP isn't working, just assign a static IP manually. Standard troubleshooting step.</p>
<p>I picked 192.168.178.50 - well above the Fritzbox's default DHCP range (.20–.199) to avoid conflicts.</p>
<p><strong>The Fix: Static IP Configuration</strong></p>
<p>In the Asus router settings, I changed from DHCP to Static IP and manually entered:</p>
<ul>
<li>IP Address: <code>192.168.178.50</code></li>
<li>Subnet Mask: <code>255.255.255.0</code></li>
<li>Default Gateway: <code>192.168.178.1</code></li>
<li>DNS Server: <code>192.168.178.1</code></li>
</ul>
<p>This bypassed the DHCP issue entirely. The Asus app suddenly showed a valid WAN connection!</p>
<p>Progress! Except...</p>
<p><strong>Problem #3: Router Has Internet, Devices Don't</strong></p>
<p>My phone connected to the new WiFi just fine. It got an IP address in the <code>192.168.50.x</code> range. I could access the Asus router settings at its default address, <code>192.168.50.1</code>. Everything looked perfect.</p>
<p>But still no actual internet.</p>
<p><strong>The Diagnostic That Revealed Everything</strong></p>
<p>I tried accessing the Fritzbox (<code>192.168.178.1</code>) from my phone while connected to the Asus WiFi. It failed completely.</p>
<p>The Asus couldn't talk to the Fritzbox. Even though they were physically connected. Even though the Asus showed a valid WAN connection. The routing between them was completely broken.</p>
<p>I verified all the WAN settings were correct. I tried pinging from the router's diagnostic tools. Nothing worked.</p>
<p>Then I tried something stupid simple.</p>
<p><strong>The Actual Fix: Wrong Port</strong></p>
<p>I unplugged the ethernet cable from LAN port 4 on the Fritzbox and plugged it into port 3 instead.</p>
<p>It immediately worked.</p>
<p>Port 4 apparently had some restriction or configuration I never figured out. I checked Fritzbox logs, firewall rules, VLAN settings, port-specific configs - nothing obvious. Port 3 worked, so I moved on.</p>
<h2>What I Learned</h2>
<p><strong>1. Start Simple</strong>
When something doesn't work, start with the physical layer. Different port? Different cable? Sometimes the answer is that basic.</p>
<p><strong>2. Work Layer by Layer</strong></p>
<ul>
<li>Can the router see the modem? (Yes - device showed up in Fritzbox)</li>
<li>Can the router get an IP? (No - fixed with static IP)</li>
<li>Can the router reach the internet? (No - wrong port)</li>
<li>Can devices reach the router? (Yes)</li>
<li>Can devices reach the internet through the router? (Finally, yes)</li>
</ul>
<p><strong>3. Double NAT Isn't Scary</strong>
For 95% of home use, double NAT works perfectly fine. The internet loves to make it sound like a critical issue, but for browsing, streaming, and normal life, it's completely transparent.</p>
<p><strong>4. Learning in Public Means Sharing the Mess</strong>
I didn't know what I was doing. I made wrong assumptions. I tried things that didn't work. And that's the point - this is what real learning looks like. Not polished tutorials where everything works perfectly the first time.</p>
<h2>The Current Setup</h2>
<p>I now have:</p>
<ul>
<li>Fritzbox handling the ISP connection</li>
<li>Asus RT-BE88U connected via LAN port 3 (not 4, very important)</li>
<li>Static IP configuration on the WAN side</li>
<li>Separate 2.4GHz, 5GHz, and 6GHz WiFi networks</li>
<li>Coverage throughout my entire house</li>
<li>No more dead zones</li>
</ul>
<p>Is it optimal? Probably not. Could I further optimize by putting the Fritzbox in bridge mode? Sure. But it works, it's fast, and I'm not touching it again until something breaks.</p>
<h2>So... What Actually Worked?</h2>
<p>I went with static IP configuration in the end. DHCP just refused to cooperate on the Fritzbox (port 4 was the actual problem, but I didn't know that yet), and once I figured out how to set the Asus manually, everything connected immediately.</p>
<p>Double NAT? I ended up leaving it alone. Everything I need - browsing, streaming, regular stuff - works fine. If I ever need port forwarding for something, I'll deal with it then.</p>
<p>The whole thing took way longer than it should have, mostly because I was plugged into the wrong LAN port. But I got there.</p>
<hr>
<p><em>Having router troubles? Different experience? Let me know - I'm always learning.</em></p>]]></content:encoded>
            <category>homelab</category>
        </item>
        <item>
            <title><![CDATA[Fixing Unraid Docker Containers After Upgrading - The Missing Label Problem]]></title>
            <link>https://lukemanning.ie/blog/unraid-docker-label-fix</link>
            <guid isPermaLink="true">https://lukemanning.ie/blog/unraid-docker-label-fix</guid>
            <pubDate>Wed, 26 Nov 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[<p>So I just upgraded my Unraid server from a very old 6.9 installation to 7.0.1, and immediately ran into a fun little problem: all my Docker containers were suddenly marked as "3rd party." Couldn't edit them, couldn't check for updates, couldn't do anything except stare at them in frustration.</p>
<h2>The Problem</h2>
<p>There's a similar thread documented here in the <a href="https://forums.unraid.net/topic/178736-docker-container-now-shows-3rd-party/">Unraid Forums</a> which I found helpful.</p>
<p>Turns out, Dockerman in Unraid 7.0+ (Unraid's Docker management plugin) uses a container label <code>net.unraid.docker.managed=dockerman</code> to determine which containers it actually manages. My containers were created way back on an older version of Unraid, so they didn't have this label. Without it, Dockerman basically said "not my problem" and refused to touch them.</p>
<p>The nuclear option would be to recreate every single container from scratch, but that's tedious and error-prone when you have dozens of containers with specific configurations. I know I could reuse an existing template as well.. but it still felt like a tedious task. So I did what I normally do and implemented an overly engineered solution to a problem that I could fixed pretty quickly doing it manually.</p>
<h2>The Solution</h2>
<p>The good news is you can add the missing label using Docker's CLI without manually reconfiguring everything.</p>
<p>I started by testing this on one container first - Jackett, which is one of my torrent index containers. I wanted to make sure the whole process worked before batch-processing everything.</p>
<p>First, I generated the docker run command using <code>runlike</code> (it inspects a running container and outputs the equivalent <code>docker run</code> command):</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --rm</span><span style="color:#79B8FF"> -v</span><span style="color:#9ECBFF"> /var/run/docker.sock:/var/run/docker.sock</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">    assaflavie/runlike</span><span style="color:#9ECBFF"> Jackett</span><span style="color:#F97583"> ></span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>(Note: you'll need root privileges to run docker commands in the Unraid terminal - otherwise you'll get "permission denied" errors.)</p>
<p>Then I reviewed what runlike generated:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">cat</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>This showed me the full docker run command with all the volumes, ports, environment variables - exact configuration for my Jackett container.</p>
<p>Next, I needed to add the missing label. I opened the file with nano:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">nano</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>And added <code>--label net.unraid.docker.managed=dockerman</code> and <code>--detach=true</code> right after <code>docker run</code>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --label</span><span style="color:#9ECBFF"> net.unraid.docker.managed=dockerman</span><span style="color:#79B8FF"> --detach=true</span><span style="color:#79B8FF"> --name=Jackett</span><span style="color:#9ECBFF"> ...</span></span></code></pre>
<p>Then I stopped the old container, removed it, and recreated it with the new configuration:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> stop</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> rm</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">bash</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>Finally, I needed to make Unraid fully recognize the container. In the Docker tab, I clicked on Jackett to open its dropdown menu and selected "Force Update." This tells Unraid to add its other management labels.</p>
<p>After doing this, Jackett showed up properly in Unraid instead of as "3rd party." Success!</p>
<h2>The Automated Script</h2>
<p>Once I verified the manual process worked, I created a script to batch-process all my containers. The script loops through all running containers, generates the run command, injects the label, recreates the container, and queues a Force Update:</p>
<p><a href="https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e">https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e</a></p>
<p>To use this, I added it to Unraid's User Scripts plugin (you can install it from Community Applications). This lets you run one-off scripts safely instead of running them directly in the Bash Shell.</p>
<p>If you're comfortable with scripts, this will process all your containers at once. Otherwise, the manual steps above work fine for one-off fixes.</p>
<h2>Important Notes</h2>
<p>By the way, if you run into different data recovery issues — like <a href="/blog/git-corrupt-object-recovery">Git object corruption</a> — the same instinct applies: move don't delete, have a rollback plan.</p>
<ul>
<li><strong>Your data is safe</strong> - This process only removes and recreates the container definitions, not your actual data in appdata or volumes</li>
<li><strong>Test on one container first</strong> - Make sure the manual steps work for your setup before running the bash script</li>
<li><strong>The Force Update step matters</strong> - Don't skip it, as it adds the additional Unraid management labels</li>
<li><strong>Back up your flash drive</strong> - Before any major changes, it's always good practice to back up your Unraid configuration</li>
</ul>
<h2>Why This Happens</h2>
<p>Unraid made this change to better track which containers it's managing versus containers you might have created manually or through other tools. It's actually a good change for container management, but it does mean old containers need this label added retroactively.</p>
<p>Hopefully this saves someone else the frustration of staring at dozens of "3rd party" containers after an upgrade! If this saved you an hour, the Gist is there to share.</p>]]></description>
            <content:encoded><![CDATA[<p>So I just upgraded my Unraid server from a very old 6.9 installation to 7.0.1, and immediately ran into a fun little problem: all my Docker containers were suddenly marked as "3rd party." Couldn't edit them, couldn't check for updates, couldn't do anything except stare at them in frustration.</p>
<h2>The Problem</h2>
<p>There's a similar thread documented here in the <a href="https://forums.unraid.net/topic/178736-docker-container-now-shows-3rd-party/">Unraid Forums</a> which I found helpful.</p>
<p>Turns out, Dockerman in Unraid 7.0+ (Unraid's Docker management plugin) uses a container label <code>net.unraid.docker.managed=dockerman</code> to determine which containers it actually manages. My containers were created way back on an older version of Unraid, so they didn't have this label. Without it, Dockerman basically said "not my problem" and refused to touch them.</p>
<p>The nuclear option would be to recreate every single container from scratch, but that's tedious and error-prone when you have dozens of containers with specific configurations. I know I could reuse an existing template as well.. but it still felt like a tedious task. So I did what I normally do and implemented an overly engineered solution to a problem that I could fixed pretty quickly doing it manually.</p>
<h2>The Solution</h2>
<p>The good news is you can add the missing label using Docker's CLI without manually reconfiguring everything.</p>
<p>I started by testing this on one container first - Jackett, which is one of my torrent index containers. I wanted to make sure the whole process worked before batch-processing everything.</p>
<p>First, I generated the docker run command using <code>runlike</code> (it inspects a running container and outputs the equivalent <code>docker run</code> command):</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --rm</span><span style="color:#79B8FF"> -v</span><span style="color:#9ECBFF"> /var/run/docker.sock:/var/run/docker.sock</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">    assaflavie/runlike</span><span style="color:#9ECBFF"> Jackett</span><span style="color:#F97583"> ></span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>(Note: you'll need root privileges to run docker commands in the Unraid terminal - otherwise you'll get "permission denied" errors.)</p>
<p>Then I reviewed what runlike generated:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">cat</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>This showed me the full docker run command with all the volumes, ports, environment variables - exact configuration for my Jackett container.</p>
<p>Next, I needed to add the missing label. I opened the file with nano:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">nano</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>And added <code>--label net.unraid.docker.managed=dockerman</code> and <code>--detach=true</code> right after <code>docker run</code>:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> run</span><span style="color:#79B8FF"> --label</span><span style="color:#9ECBFF"> net.unraid.docker.managed=dockerman</span><span style="color:#79B8FF"> --detach=true</span><span style="color:#79B8FF"> --name=Jackett</span><span style="color:#9ECBFF"> ...</span></span></code></pre>
<p>Then I stopped the old container, removed it, and recreated it with the new configuration:</p>
<pre class="shiki github-dark" style="background-color:#24292e;color:#e1e4e8" tabindex="0"><code><span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> stop</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">docker</span><span style="color:#9ECBFF"> rm</span><span style="color:#9ECBFF"> Jackett</span></span>
<span class="line"><span style="color:#B392F0">bash</span><span style="color:#9ECBFF"> /tmp/jackett_run.sh</span></span></code></pre>
<p>Finally, I needed to make Unraid fully recognize the container. In the Docker tab, I clicked on Jackett to open its dropdown menu and selected "Force Update." This tells Unraid to add its other management labels.</p>
<p>After doing this, Jackett showed up properly in Unraid instead of as "3rd party." Success!</p>
<h2>The Automated Script</h2>
<p>Once I verified the manual process worked, I created a script to batch-process all my containers. The script loops through all running containers, generates the run command, injects the label, recreates the container, and queues a Force Update:</p>
<p><a href="https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e">https://gist.github.com/ManningWorks/0f57db9c450d7b7dea741e585b31d23e</a></p>
<p>To use this, I added it to Unraid's User Scripts plugin (you can install it from Community Applications). This lets you run one-off scripts safely instead of running them directly in the Bash Shell.</p>
<p>If you're comfortable with scripts, this will process all your containers at once. Otherwise, the manual steps above work fine for one-off fixes.</p>
<h2>Important Notes</h2>
<p>By the way, if you run into different data recovery issues — like <a href="/blog/git-corrupt-object-recovery">Git object corruption</a> — the same instinct applies: move don't delete, have a rollback plan.</p>
<ul>
<li><strong>Your data is safe</strong> - This process only removes and recreates the container definitions, not your actual data in appdata or volumes</li>
<li><strong>Test on one container first</strong> - Make sure the manual steps work for your setup before running the bash script</li>
<li><strong>The Force Update step matters</strong> - Don't skip it, as it adds the additional Unraid management labels</li>
<li><strong>Back up your flash drive</strong> - Before any major changes, it's always good practice to back up your Unraid configuration</li>
</ul>
<h2>Why This Happens</h2>
<p>Unraid made this change to better track which containers it's managing versus containers you might have created manually or through other tools. It's actually a good change for container management, but it does mean old containers need this label added retroactively.</p>
<p>Hopefully this saves someone else the frustration of staring at dozens of "3rd party" containers after an upgrade! If this saved you an hour, the Gist is there to share.</p>]]></content:encoded>
            <category>homelab</category>
            <category>docker</category>
        </item>
    </channel>
</rss>