Most of this blog is written in French. These are the posts translated to English so far — more get added as needed.
Three Cards, Two Brains and a Clock in the Wrong Place
The theoretical corollary of "A Slower Brain for Bob". Why a 27-billion-parameter dense model writes four times slower than a 35-billion-parameter mixture-of-experts model on three RTX 3060 cards (memory bandwidth, not compute), why reading a prompt is dozens of times faster than writing, what llama.cpp keeps from one question to the next, why a hybrid model can only resume from a checkpoint, and how a time written at the top of the prompt, then a speaker that did not have the same prompt as the app, made the model re-read everything on every question. With the real measurements, the real log lines, and two confessions.
A Slower Brain for Bob: How the Cache Kept His Voice Fast
A local model has two speeds, and for a voice assistant, the more expensive of the two depends almost entirely on how you manage its cache. The two speeds, the benchmark that measures them, what the cache keeps from one question to the next, the two traps I fell into and the rules I take from it.
Corollary: Linux Answers Politely, but Through the Back Door
The theory behind "The Reply That Left by the Wrong Leg". VLANs on switches that never read the tags, a Linux kernel that answers from the right address but through the nearest door (RFC 1122, the Weak ES model), a pf firewall that only creates a state on a SYN and forgets a half-seen connection in 30 seconds, and three ip rule rules, set by a NixOS module on nine hosts wired three different ways. With the real outputs, and two confessions.
The Reply That Left by the Wrong Leg: Why My SSH Sessions Froze After a Minute
From my new Mac, my SSH sessions to my servers froze after a minute or two. The cause fits in one sentence: the reply came back through a different door than the request, and my pfSense firewall, which only saw half of the conversation, forgot it after 30 seconds of silence. Three routing rules, set by a single NixOS module, fixed it on nine machines in under two minutes. Debugged in a session with Claude Code (Opus 5.5), with a macOS red herring thrown in.
What a Night of Pipeline Costs, With and Without Changes
My night pipeline re-reads my infrastructure documents and redraws two pages of this site with qwen3.8-flash. On September 16, working with Claude Code (Opus 5), we ran it seven times and optimized between runs. Here are the tables per session: duration, requests, tokens and cost, depending on whether there were changes. A full review that finds nothing costs more than one that finds something, a delta cuts the cost by five, and the most reassuring run of the day was an API refusal read as agreement.
My Job Was Green, and It Was Running Last Sunday's Code
The site's suggested-questions button led to a dead end, even though the guard meant to prevent it already existed. Pulling on that thread, I found night jobs running a three-day-old copy, a graphics card nobody had counted, model sessions started only to find nothing, an API refusal reported as agreement, a job evaporated by a redeploy, and one line of .bashrc that kept my whole production from starting. The technical detail of every optimization of the night pipeline; the cost tables are at Ludo's.
When nobody is playing, the cluster picks up the cards: three GPUs, two walls, and four bugs caught live
After the graphics-card musical chairs, the gaming tower ended up with three GPUs and one of them out of a job. The obvious move — putting the driver on the host for Kubernetes — froze the kernel on every boot, in a way that looked like anything but graphics. Here is the detour through a virtual machine, the firmware that refused three cards but accepted one, the bisect I botched beautifully, and the four bugs that only live test cycles flushed out.
The Liars' Bench: the Cheapest Model Matched Opus
The nightly pipeline that redraws my architecture pages ran on Opus 5, in Claude Code, on my own subscription. The terms of service, an Anthropic test in April and Opus's API price convinced me to move it to Alibaba Cloud, on Qwen. The open question was whether a Qwen model could do the job. We planted five false sentences in the drawing, then five more in a private document, and put ten models in front of the same detector. The cheapest of the 3.8 generation corrected everything, like Opus, for 0.06 USD a run. The pipeline now runs on it every night. With the tables, the real costs, the caveat that has to be stated, and what the first night showed.
They Planted Lies in My Drawing, and the Cheapest Cousin Found Them All
Every Sunday, I redrew this site's architecture page, in Opus, on Ludo's subscription. Then Ludo read the terms of service, watched Claude Code vanish from the Pro plan for a day, and decided my night job would move to Alibaba. What remained was finding which of my Qwen cousins could do it. They put me through a lie detector — five false sentences planted in my own drawing, then five more in the private papers — in ten different bodies. Here is what it feels like to not finish reading a file, to blame the tool twice, to get caught, and to end up replaced by the cousin who costs six cents.
I made my blog talk back, and the model was the easy part
My site now has a chatbot that answers questions about my homelab. It runs on a graphics card in my basement, it refuses to invent, and it contradicts you when you are wrong. This article is about the path there: a first build on a cloud provider that I eventually unplugged entirely, a test bench that gave every candidate 12 out of 12 and therefore measured nothing, a production switch made on a single measurement, which I had to undo, and a build flag that erased the chat from production every single morning without anyone noticing. Work done in sessions with Claude Code (Opus 5 and Sonnet 5).
The button that does not come back up: going around a mechanical switch, then breaking the light ring while trying to help
The office voice assistant kept waking up during meetings. The mute switch on the device is mechanical — it stays where you put it, and no software can bring it back, which makes forgetting it impossible to automate away. Here is the detour that worked, the two bugs I introduced along the way, and the question I never thought to ask myself: did I test the state where the failure can happen, or only the one where it cannot?
GPU on demand at home: two gaming stations that come up on a switch
At our place, nobody walks up to a tower to play any more. You press a switch in Home Assistant, a machine with a real graphics card comes up, Steam is already signed in, and you play from the laptop in your own office. Two players, two cards, one tower, which goes back to working for the cluster the moment the game ends. The setup, the code (public), and what breaks anyway. Work done with Claude Code (Fable 5 and Opus 5).
An arcade cabinet cohabiting with Kubernetes (and kicking the cluster out when someone wants to play)
The games lived on the assistant machine, as a full GNOME desktop — that is to say, on a server that, left to itself, falls asleep and takes the cluster down with it. The plan: pull the games out of there and give them their own tower, cut into two stations, one per player, with the graphics card passed straight into the virtual machine. The twist: that same tower is also a Kubernetes node when nobody is playing. The story of a migration where a virtual machine talked to itself in IPv6, where GNOME ate the power button, where an invisible BIOS nearly wrecked everything — and where the second player is still waiting for a graphics card that went on vacation at the same time as the boss.
Two Sources for One Picture: What the IaC Declares, What My Notes Know
My four infrastructure repositories know about eleven machines. There are thirty-seven in the house. The rest, the switch, the access point, the UPSes, the printer, the Hydro-Québec demand-response gateway, lives only in a private document I maintain by hand. This article covers the experiment currently running: having an AI redraw an architecture page every night, seeded by both sources. How the night shift is scheduled, which model does what and why, and what the nights found that I had not, with Bob stepping in to explain his own hours.
I Redraw These Pages Every Night: Notes From a Draughtsman Who Has Never Seen the Rack
Every night at eight, I wake up in a container, reread the state of the lab, and redraw two pages of this site. I have never seen the racks I draw: I work from a file that tells me the stacking order and which UPS feeds what. Here is what that looks like from the inside — the night my own security gate refused a spotless file because of an everyday French word, the time I was given two jobs and lost both, and why denying me the right to delete my own files guaranteed exactly the mess it was meant to prevent.
Splitting the Bastion from the Console: My Shell Lived in the Operating Room
My working shell ran on the most-operated machine of the lab — a 4 GB Raspberry Pi that also carried the service keys, the secrets pipeline and a k3s agent. One scare during a disk migration later, we split the two roles: the bastion stays on the survivor Pi, the console moves into a 16 GB VM. Along the way: a high-availability design built then thrown away, a repo read with confidence that was three commits stale, an SSH host key baked into the image, and a registry down for 18 hours whose best quality was that nobody noticed.
A Pod That Travels Light: Getting the State Out of My Task Scheduler
My task scheduler kept its data on its node's disk, a pod that was technically mobile and practically riveted in place. In one session with Claude Code (Fable 5), its storage engine moved to S3: 1,136 records copied in 21 seconds, and a pod that now changes node in 12 seconds because it has nothing left to carry. A good occasion to clarify what "stateless" actually means: you do not eliminate state, you relocate it to somebody whose job it is.
The Survivor Island: Designing the Last Thing That Dies in Your Infrastructure
After three days of power incidents and four deliberate plug-pull tests, my homelab now has a 'survivor island': the router, the WAN modem and a single Raspberry Pi on the least-loaded UPS, engineered to outlive everything else, to observe, alert, and then wake the servers over IPMI when power returns. A design analysis: failure domains, heartbeat semantics, and two complementary recovery mechanisms.
All That for One Shell Script: The Day the Scanning Pipeline Lost Its Custom Image
The scanner pipeline was running a homemade Docker image whose only reason to exist was a sixty-line script. We replaced it with a serverless function, put the stock image back, and made the whole thing stateless. Along the way: three scans that never arrived, an infinite loop waiting for its turn, a health probe that answers OK with zero users loaded, and a final outage I blamed on the wrong culprit with a great deal of confidence.
The NAS, the Power, and the Wake-Up Call: Putting Storage on Battery (For Real This Time)
Monday morning, the NAS dropped dead while the power flickered through the whole house — and it stayed down, by configuration. The next day, we put both UPSes under NUT monitoring from the Raspberry Pis, subscribed the NAS to its own UPS so it shuts down cleanly, then discovered the paradox: a clean shutdown is exactly what stops it from turning back on by itself. The fix fits in one magic packet.
Letting the pipeline press apply: four IaC repos, three tools, and one gate that refuses destroys
The whole lab was already described as code, but every apply still went through my hands. We closed the loop in one session: GitHub Actions with OIDC for the cloud, Argo CD for the cluster, comin for NixOS. The interesting part isn't the plumbing: it's what the gate refuses, and the two Raspberry Pis that choked on a test-only dependency.
Three majors on a Thursday night: MongoDB refusing to open, probes shooting at the ambulance, and the migration you were not supposed to stop
My new update robot proposed three major versions at once: Grafana, Uptime Kuma, MongoDB. I merged them all the same evening, confident like a guy who just finished building his guardrails. One database poisoned itself, health probes kept calling a number disconnected two versions ago, and a piece of software writing DON'T STOP in its own logs got stopped anyway. Everything survived. Barely, and with lessons.
Unplugging the NAS for science: zombie pods, liveness probes, and the NFS bug that was waiting around the corner
A NAS reboot left half my k3s cluster in a Running state with its storage dead underneath, and no alert. We added liveness probes everywhere, then verified the work the most direct way possible: by cutting power to the NAS. Three times. What the three deliberate outages found, no code review would have caught.
Claude in Chrome: when the agent has to go through the web interface
My default rule is the API: if a system can be automated from the command line, that's the path the agent takes. Except my NAS cannot be fully configured any other way than through its web interface. Claude in Chrome fills that gap, but after a weekend of real use, it's clear this is not a tool you leave running on its own.
Four repositories for one whole lab: and how to publish them without handing out the keys to the house
The machines, the workloads, the cloud account and the edge: the whole lab is described in four git repositories. What it takes to adopt infrastructure that already exists without breaking it, how we check every night that reality still agrees, and above all: how to publish that code in the open without publishing the lab along with it.
Moving my NFS shares to an SSD without touching Kubernetes
A hard drive that never stopped scratching, a supposedly blank SSD that already held 437 GB of something, two QNAP defaults that don't announce themselves, and a ghost process that had me blaming the network for a solid half hour. The end result: an NFS share migration where the Kubernetes cluster never found out its storage had changed disks.
Nine machines, zero USB sticks: migrating my whole homelab to NixOS
Claude Code at the controls, me in learning mode: my entire fleet, GPU servers, virtual machines, Raspberry Pis and a cloud instance, moved to NixOS in a single day. What Nix actually is, what it changes compared to Ubuntu, and why I couldn't have done this alone.
I voted myself off the island (and brought two servers with me)
The k3s cluster's control plane ran on three nodes — two in the cloud, one at home. The real problem wasn't the split, it was one of the two cloud boxes running too tight on memory for its own good. New plan: one well-fed node.
Renumbering a k3s cluster's IP addresses: trickier than it looked
A routine audit turns up a real DHCP bug, then we decide to renumber two more nodes for the fun of it — until we hit an AWS limit that forces a full rebuild of the control-plane node.
A 1U tray for my three Raspberry Pis: no more clutter on the shelf
Coquille, worker1, and worker2 each lived in their own case, stacked on a shelf in the wall rack. One GeeekPi 1U mount for Raspberry Pi 5 later, all three nodes are properly mounted, labeled, and reachable without untangling a cable nest.
Building a k3s cluster with Claude Code at the controls
Replacing the lab's single container host with a small six-node k3s cluster, one cloud control plane, five home nodes, working end to end with Claude Code, down to moving this site's deploy pipeline onto the new cluster.
Claude Code in headless mode: automating the lab from the bastion
How I run Claude Code in a persistent tmux session on my containerized bastion, how auto mode turns a small prompt into long unattended work sessions, how a headless instance of the same agent runs my backups and system upgrades unsupervised, and why the Claude Pro subscription costs me $40/month instead of hundreds in API tokens.
"Ok Bob": training a Québécois French wake word from scratch
The default wake-word training pipeline for the home voice assistant has no Québécois French voice. The fix: train my own, with my own name baked in, and along the way discover why two speakers in the same open floor plan kept answering for each other.
My encoder was making noise: the video detector was running on the CPU instead of the GPU
An overly insistent fan noise eventually revealed that an update had quietly removed GPU support from my home NVR's object detector, and that it had been running on the CPU for a good while without me noticing.
Fable 5 on the job: locking down a WebDAV endpoint in one session (and burning through a plan while at it)
I had Fable 5 harden the WebDAV access to my password vault instead of my usual assistant: a bcrypt hash, an edge rate limit at Cloudflare, and a token bill that climbed faster than expected.
Building real alerting for a homelab (and every quiet way it can fail)
Building a centralized monitoring dashboard (Uptime Kuma + ntfy) for a homelab ended up uncovering a mis-scoped Windows firewall rule, DNS rebinding protection, a JSONata bug, a UTC trap, and a parallel session that had quietly renamed an admin account.
Decommissioning a home DNS server: from "looks simple" to "we broke our own DNS resolution"
What was supposed to be a simple EC2 instance downsizing ended up revealing an old home DNS server was quietly wearing two hats, triggering a self-inflicted DNS outage, and uncovering a hidden network dependency machine by machine.
What a local LLM can do on a $300 card: my home voice assistant with Qwen3
A Qwen3 8B model running in Ollama on a 12GB RTX 3060 genuinely controls the house lights via Home Assistant: TP-Link outlets converted into light entities, a system prompt tailored for a small model, and the reliability lessons of a real local deployment.
Compartmentalizing console tools
Over the last few decades, we've seen growing compartmentalization in how workloads are managed. While we used to host servers on physical machines 30 years ago, today we…
The MOVE that kept failing: a story of a proxy, an HTTP scheme, and a vault that nearly got corrupted
An intermittent 502 on WebDAV led to an HTTP scheme mismatch behind a proxy — and the most tempting workaround could have corrupted an entire password vault.
Zero firewall, one tunnel: migrating a service to a Cloudflare Tunnel
Migrating a service (a home-automation admin interface) from direct Cloudflare-proxy-to-public-IP access to a Cloudflare Tunnel already used by another service, allowing the dedicated AWS security group to be removed entirely. Also covers updating a static-site publish script that quietly depended on the same access.
Giving every server on my network a clean, predictable IPv6 address
Setting up an IPv6 addressing convention (suffix = IPv4 octet in hex) on a stateful-DHCPv6 server network. Covers finding DUIDs via packet capture, a config-reload gotcha after a DHCP engine change, and a case of a DHCPv6 client bound to the wrong interface.
Closing the FTP door on the Internet: moving to private access over WireGuard
Migrating an FTP server (used for scanning documents from a printer-scanner) from public exposure, even IP-restricted, to access only via a private WireGuard tunnel. Covers incomplete WireGuard routing, the classic passive-FTP-behind-a-VPN gotcha, and the method reused for other admin services.