<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Oreoft</title>
    <description>Oreoft&apos;s tech blog: hands-on notes on Java, databases and middleware, cloud, Linux and developer tools, plus write-ups of real production incidents.</description>
    <link>https://www.someget.cn/en/</link>
    <language>en</language>
    <atom:link href="https://www.someget.cn/en/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Sun, 04 Oct 2026 00:46:46 +0800</pubDate>
    <lastBuildDate>Sun, 04 Oct 2026 00:46:46 +0800</lastBuildDate>
    <generator>Jekyll v4.4.1</generator>
    
      <item>
        <title>Turning My Windows Mini PC Loaded with Random Chores into PVE</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;I have a Beelink SER mini PC at home equipped with an AMD Ryzen 7 5800H CPU (8 cores, 16 threads), two 16GB DDR4 3200 RAM sticks, and a 1TB NVMe SSD. Since my home broadband comes with a public IP, it has been running 24/7 as a home server.&lt;/p&gt;

&lt;p&gt;It had always been running Windows. Over time, I piled more and more tasks onto it, but the stability was never quite where I wanted it to be. This post documents why I eventually decided to reinstall everything with PVE (Proxmox VE), how I set it up, and how the virtual machines are partitioned now. Along the way, I also tore it down for some dust cleaning, so there are a few pictures included.&lt;/p&gt;

&lt;p&gt;p.s. This is more of an experience and mindset log rather than a step-by-step tutorial.&lt;/p&gt;

&lt;h2 id=&quot;1-more-and-more-work-piled-on-windows&quot;&gt;1. More and More Work Piled on Windows&lt;/h2&gt;

&lt;p&gt;Initially, it was just an ordinary Windows machine. Later, I assigned more and more jobs to it:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Various Services&lt;/strong&gt;: A few self-written services (like an AI service), ddns-go (which automatically updates my dynamic public IP to my domain), an Nginx gateway, and proxy utilities like Hysteria 2 (hy2).&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;NAS &amp;amp; Smart Home&lt;/strong&gt;: fnOS (Feiniu OS) and HAOS (Home Assistant OS).&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Windows-Exclusive Applications&lt;/strong&gt;: A few apps that only exist on Windows and must run in a Windows environment.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Among these, only the third category genuinely required Windows. The first two were just piggybacking on it.&lt;/p&gt;

&lt;h2 id=&quot;2-stability-issues-never-went-away&quot;&gt;2. Stability Issues Never Went Away&lt;/h2&gt;

&lt;p&gt;Using Windows as the host OS came with two major problems:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Windows itself is a resource hog.&lt;/strong&gt; The OS disk easily takes up dozens of gigabytes, a bunch of useless background services run constantly, and it eats up a large chunk of RAM.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It restarts on its own.&lt;/strong&gt; No matter how you tweak the Windows Update policies, it always finds an excuse to reboot every once in a while. When the host reboots, every single service goes down, and fnOS disconnects along with them.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;3-three-iterations-of-service-deployment&quot;&gt;3. Three Iterations of Service Deployment&lt;/h2&gt;

&lt;p&gt;To make those services run more stably, I cycled through three different approaches.&lt;/p&gt;

&lt;h3 id=&quot;1-the-hyper-v-era&quot;&gt;1. The Hyper-V Era&lt;/h3&gt;

&lt;p&gt;Initially, I spun up Linux VMs inside Hyper-V (Windows’ built-in hypervisor) to host services, and fnOS ran there too. That setup suffered from three issues:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Heavy resource overhead.&lt;/strong&gt; Each VM had to run a full Linux OS. As services grew, each VM added its own OS-level overhead.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Unreliable auto-start after power loss.&lt;/strong&gt; Sometimes after a power outage and recovery, the VMs would fail to start up properly.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Too bloated overall.&lt;/strong&gt; Later, I thought WSL would be lighter, so I switched to WSL.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3 id=&quot;2-the-wsl2-era&quot;&gt;2. The WSL2 Era&lt;/h3&gt;

&lt;p&gt;WSL2 is the subsystem for running Linux inside Windows. It is much lighter than traditional VMs, but brought a new set of headaches:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Networking pain.&lt;/strong&gt; Although WSL supports mirrored networking (sharing network adapters and IPs with the Windows host), configuring it reliably is a real pain.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;File sharing with the host was prone to errors.&lt;/strong&gt; Service logs and miscellaneous files resided on the Windows side. WSL’s default file sharing threw errors when files were modified concurrently from both sides. It only improved after switching over to SMB shares.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Randomly exiting.&lt;/strong&gt; Even with active background services, WSL would terminate itself out of nowhere. I had to write a keepalive script on Windows just to keep it alive.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3 id=&quot;3-all-in-docker-based-on-windows&quot;&gt;3. All In Docker based on Windows&lt;/h3&gt;

&lt;p&gt;Next, I installed Docker directly on Windows and shoved all services into containers. But Docker Desktop for Windows essentially runs on top of WSL anyway—it is just a dedicated WSL distro running the Linux version of Docker. So that extra layer remained, and all the quirks with it.&lt;/p&gt;

&lt;p&gt;The most memorable issue: UDP port forwarding in Docker on Windows goes through an extra layer, leading to cases where &lt;strong&gt;the container is clearly alive, but the port stops accepting traffic&lt;/strong&gt;. To work around this, I had to deploy a “sidecar” container specifically for hy2: every 30 seconds it performed a real handshake, and if it failed, it fired an alert and restarted the service. It worked, but needing such a hack proved the foundation was flawed. Along with various other messy glitches, it was simply exhausting.&lt;/p&gt;

&lt;p&gt;After cycling through three different solutions, the root cause remained the host layer: as long as the host was Windows, its reboots, resource hogging, and translation layer for the Linux ecosystem were unavoidable.&lt;/p&gt;

&lt;h2 id=&quot;4-why-switch-to-pve&quot;&gt;4. Why Switch to PVE&lt;/h2&gt;

&lt;p&gt;This machine actually packs plenty of power; there is no need to bottle it all up under a single Windows instance.&lt;/p&gt;

&lt;p&gt;Proxmox VE (PVE) is a Debian-based virtualization platform installed bare-metal, allowing you to manage VMs and containers via a web UI. After migrating to it, the strategy became clear:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;The host does nothing but virtualization.&lt;/strong&gt; It runs no user workloads and will not reboot out of the blue.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;When Windows is needed, spin up a lightweight Windows VM&lt;/strong&gt; solely to run Windows-exclusive apps.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Offload other services to a stable Linux VM&lt;/strong&gt;, running all Docker applications natively without translation layers.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Give systems like fnOS and HAOS their own dedicated VMs.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This way, CPU cores and RAM can be allocated on demand. If something breaks, only that specific VM needs a reboot without affecting others. Backups and snapshots are handled natively by PVE. It is a clean slate, and expanding in the future becomes seamless.&lt;/p&gt;

&lt;h2 id=&quot;5-cleaning-out-the-dust-first&quot;&gt;5. Cleaning Out the Dust First&lt;/h2&gt;

&lt;p&gt;Before getting to work, I took the machine apart for a quick dust cleanup.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003100309797.jpg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Two 16GB RAM sticks, with the 1TB SSD on the right&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Inside are the two 16GB RAM modules (bought back before RAM prices surged) and next to them is the 1TB SK Hynix NVMe SSD. Overall, the hardware specs are quite solid.&lt;/p&gt;

&lt;p&gt;When cleaning the bottom cover, I facepalmed: &lt;strong&gt;turns out I had never peeled off the protective film on the SSD’s thermal pad.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095647224.JPG?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;03-ssd-film&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Thermal pad on the bottom plate; that blue tab is the unpeeled protective film&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;While cleaning, I noticed a SATA connector on the bottom bracket, which reminded me of an old 2.5-inch mechanical HDD sitting around.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003100159596.jpg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;
&lt;center&gt;An old mechanical hard drive removed from an external enclosure&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;This was salvaged from an ancient Mac over a decade ago. Since then, it barely saw any use except as a spare, and I had bought an external enclosure for it back then. Now that I use NVMe external enclosures, this drive was totally obsolete. Might as well slap it into the mini PC as a backup/storage drive.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095643186.jpg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Installing the mechanical HDD into the bottom bracket slot&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Being an older drive, it is 9.5mm thick. Even though it is a 2.5-inch drive, it was a tight squeeze and made the bottom cover bulge slightly. But after tightening down the four corner screws, it did not stick out too much.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095629924.jpg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;
&lt;center&gt;The bottom cover bulges a bit after installation&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;It looks a bit funny with its potbelly, but it works without issues. I later formatted it specifically to store PVE VM backups.&lt;/p&gt;

&lt;h2 id=&quot;6-backing-up-the-data-first&quot;&gt;6. Backing Up the Data First&lt;/h2&gt;

&lt;p&gt;My Mac has an 8TB drive, which is more than enough for the migration, so I dumped everything onto the Mac first. The data I actually needed to move was pretty minimal:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;fnOS Virtual Disk&lt;/strong&gt;: Originally 250GB. I copied out the large files first, reclaimed unused space inside fnOS, shrunk the virtual disk, and got it down to just 33GB.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Configurations and data for each service&lt;/strong&gt;: Totaled around 2GB. Docker images and Git repos can always be pulled again, so no need to back those up.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;After copying and verifying checksums, I started working on the machine.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095548355.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;06-copy-fnos-files&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Copying large files from fnOS over to the Mac first, totaling over 200GB&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095544443.webp?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;07-fnos-fstrim&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Running `fstrim` inside fnOS to reclaim 257GB of free space before shrinking the virtual disk&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;7-bios-configuration-and-pve-installation&quot;&gt;7. BIOS Configuration and PVE Installation&lt;/h2&gt;

&lt;p&gt;First, I flashed the PVE ISO onto a USB drive on my Mac.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095452415.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;09-dd-iso&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Using `dd` to write the PVE 9.2 ISO to the USB drive&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Then, I went into the BIOS and adjusted a few settings:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Enable virtualization: SVM (AMD CPU virtualization) and IOMMU (required for device passthrough).&lt;/li&gt;
  &lt;li&gt;AC Power Loss recovery: Automatically powers on after an outage without needing someone to press the power button.&lt;/li&gt;
  &lt;li&gt;Reduce iGPU VRAM to 1GB (previously 3GB when Windows was the host; now I can lower it to free up more RAM).&lt;/li&gt;
  &lt;li&gt;Set System Mode to Performance Mode and enable CPPC (enabling finer-grained CPU frequency scaling, which the PVE governor relies on).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003094607877.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;10-bios-main&quot; /&gt;&lt;/p&gt;
&lt;center&gt;BIOS main screen: R7 5800H with 32GB RAM&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095355337.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;11-bios-svm&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Enable SVM&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095457700.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;12-bios-iommu&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Enable IOMMU&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095348920.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;13-bios-ac-loss&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Auto power-on upon AC recovery&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003101926989.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;15-bios-performance-mode&quot; /&gt;&lt;/p&gt;
&lt;center&gt;System Mode set to Performance Mode&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095421546.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;16-bios-cppc&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Enable CPPC&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;After applying settings, I booted from the USB drive to install PVE 9.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095312969.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;17-pve-installer&quot; /&gt;&lt;/p&gt;
&lt;center&gt;PVE installer screen&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095011684.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;18-pve-disk&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Installing the OS on the 1TB SSD with ext4 filesystem&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095033142.jpeg?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;19-pve-network&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Assigning a static LAN IP to the host&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;After installation, I made a small tweak: set the CPU scaling governor straight to performance mode. This machine runs 24/7 plugged in—&lt;strong&gt;I don’t care about power consumption, I only care about performance.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;8-allocating-virtual-machines&quot;&gt;8. Allocating Virtual Machines&lt;/h2&gt;

&lt;p&gt;There are currently four VMs running on PVE:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Virtual Machine&lt;/th&gt;
      &lt;th&gt;OS&lt;/th&gt;
      &lt;th&gt;Specs&lt;/th&gt;
      &lt;th&gt;Purpose&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;ser-docker&lt;/td&gt;
      &lt;td&gt;Debian 13&lt;/td&gt;
      &lt;td&gt;4 Cores / 4G&lt;/td&gt;
      &lt;td&gt;All Docker services&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ser-fnos&lt;/td&gt;
      &lt;td&gt;fnOS (Feiniu OS)&lt;/td&gt;
      &lt;td&gt;4 Cores / 4G&lt;/td&gt;
      &lt;td&gt;NAS, existing virtual disk imported directly&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ser-win&lt;/td&gt;
      &lt;td&gt;Windows Server 2025&lt;/td&gt;
      &lt;td&gt;4 Cores / 6G&lt;/td&gt;
      &lt;td&gt;Runs only Windows-exclusive apps&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ser-mac&lt;/td&gt;
      &lt;td&gt;macOS 15 (Hackintosh)&lt;/td&gt;
      &lt;td&gt;8 Cores / 8G&lt;/td&gt;
      &lt;td&gt;Dedicated to AI Agent use&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;After migrating to the Docker VM, I immediately deleted the “sidecar” watchdog container previously created for hy2. On Linux, UDP packets are routed directly by the kernel, so the old bug simply vanished.&lt;/p&gt;

&lt;p&gt;This time around, I installed Windows Server instead of Windows 11. The Server edition cuts out the bloatware and offers granular control over system updates, making it much better suited for long-running apps. There was one minor hiccup during installation: PVE assigns VirtIO virtual disks, which the Windows installer doesn’t recognize out of the box—no disks showed up until I manually loaded the VirtIO drivers.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095307993.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;20-winserver-no-disk&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Windows Server installer cannot find any drive; VirtIO drivers must be loaded first&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003095304834.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;22-winserver-autologon&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Using Sysinternals Autologon to configure automatic logon on startup&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Finally, I also set up a Hackintosh VM. Since the iGPU cannot be passed through or properly accelerated, GUI rendering falls back completely on CPU software rendering, making it a bit sluggish. However, this CPU is decent enough that without graphic-intensive workloads, pure CPU brute-forcing handles most compilation and web browsing tasks fine. &lt;strong&gt;It is more than enough for an AI Agent.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20261003090959062.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1791036590133&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Hackintosh running on PVE&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;My daily driver Mac is plenty powerful, so the division of labor is straightforward: compilation, iOS development, and graphical tasks happen on my local Mac; miscellaneous chores like web browsing are delegated to the Agent running on the virtual Hackintosh. Neither interferes with the other, and I do not have to surrender my own workstation to the Agent.&lt;/p&gt;

&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;

&lt;p&gt;After years of tinkering across Hyper-V, WSL2, and Windows Docker, I finally realized the problem wasn’t the individual solutions, but the host OS itself. Once the host switched to PVE, Windows retreated to what it does best—just a dedicated VM for Windows apps—while everything else runs natively on Linux.&lt;/p&gt;
</description>
        <pubDate>Sat, 03 Oct 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/other/2026/10/03/windows-mini-pc-to-pve.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/other/2026/10/03/windows-mini-pc-to-pve.html</guid>
        
        <category>en</category>
        
        <category>other</category>
        
      </item>
    
      <item>
        <title>Installing Tailscale on Alibaba Cloud Blocked All of Alibaba Cloud’s Own Internal Services</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;I’ve had an Alibaba Cloud machine sitting around for a long time. I bought it during a promo, and later renewed it for 20 years. The specs and bandwidth aren’t great (2 vCPU, 2 GB RAM, 3 Mbps), and since I usually manage more machines on GCP and use GCP more often, this one has mostly been idle.&lt;/p&gt;

&lt;p&gt;Recently I wanted to give it something useful to do. I have a few servers at home running various services, and they go down from time to time—sometimes because the machine rebooted, sometimes because memory blew up. When a service dies, I usually don’t notice until something in production starts breaking. I’ve never had a unified monitoring setup. Using the home machines to monitor themselves doesn’t feel very reliable if they’re the ones going down; this Alibaba Cloud box may be underpowered, but it’s stable, which makes it a good fit for health checks.&lt;/p&gt;

&lt;p&gt;My plan was to connect everything with Tailscale. Tailscale is a WireGuard-based networking tool that pulls machines scattered across different places into the same virtual private network. Each machine gets a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.x.x.x&lt;/code&gt; address, and they can talk to each other like they’re on the same LAN. I’ve installed it plenty of times on AWS and GCP machines and basically never had issues. But this time on Alibaba Cloud, right after installing it, DNS resolution, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; package installs, and pulling images from registries all started failing one after another.&lt;/p&gt;

&lt;p&gt;The reason can be summed up in one sentence: &lt;strong&gt;Alibaba Cloud puts its internal services in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt; range, and Tailscale uses that same range.&lt;/strong&gt; On Linux, Tailscale installs an anti-spoofing firewall rule that drops any packet that “claims to come from that range but didn’t arrive through the Tailscale interface.” Replies from Alibaba Cloud’s internal services got caught by that rule.&lt;/p&gt;

&lt;p&gt;This post is a record of how I traced it down step by step, a few workarounds I tried along the way, and the official Tailscale solution I ended up using.&lt;/p&gt;

&lt;p&gt;p.s. The system is Ubuntu 22.04, and Tailscale is 1.102.&lt;/p&gt;

&lt;h2 id=&quot;1-what-this-machine-is-supposed-to-do&quot;&gt;1. What this machine is supposed to do&lt;/h2&gt;

&lt;p&gt;First, here’s how I planned to use it, because a lot of the decisions later are tied to this.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Network access&lt;/strong&gt;: connect all services that need monitoring to the Tailscale private network. This machine will hit each service’s health check endpoint over the private network, then expose a dashboard.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Access management&lt;/strong&gt;: normally I’ll manage it directly over the Tailscale network; on the public internet, I’ll only expose SSH on port 22, so I can still get in when I’m not on the private network.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Dashboard access&lt;/strong&gt;: access the dashboard directly from inside the private network; when I’m outside it, use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ssh -L&lt;/code&gt; to forward the dashboard port locally.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260920223654647.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;Overall architecture&quot; /&gt;&lt;/p&gt;

&lt;center&gt;Overall architecture: health checks and management go through the Tailscale private network, and only port 22 is exposed publicly&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Health checks, management, and dashboard access all go through Tailscale. So whether Tailscale can work properly on this machine is the prerequisite for the whole setup.&lt;/p&gt;

&lt;h2 id=&quot;2-after-installing-tailscale-dns-resolution-completely-broke&quot;&gt;2. After installing Tailscale, DNS resolution completely broke&lt;/h2&gt;

&lt;p&gt;After installing Tailscale and joining the private network, domain resolution on this machine stopped working. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;getent hosts&lt;/code&gt; returned nothing:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;getent hosts github.com
&lt;span class=&quot;err&quot;&gt;$&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;My first suspicion was upstream DNS or outbound connectivity. If that were the case, querying a public DNS server directly should also fail. So I checked:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;dig +short github.com @8.8.8.8
20.205.243.166
&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;dig +short github.com
&lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt; communications error to 127.0.0.53#53: timed out
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Querying &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;8.8.8.8&lt;/code&gt; worked, but using the default resolver timed out. Outbound networking was fine, so &lt;strong&gt;the problem was in the local resolution path&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Ubuntu 22.04 uses &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemd-resolved&lt;/code&gt; for local DNS by default: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/resolv.conf&lt;/code&gt; points to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;127.0.0.53&lt;/code&gt;, which is a local forwarding stub, while the real upstream resolvers are configured elsewhere. So I checked which upstreams it was using:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;resolvectl status
Link 2 &lt;span class=&quot;o&quot;&gt;(&lt;/span&gt;eth0&lt;span class=&quot;o&quot;&gt;)&lt;/span&gt;
    DNS Servers: 100.100.2.136 100.100.2.138
Link 3 &lt;span class=&quot;o&quot;&gt;(&lt;/span&gt;tailscale0&lt;span class=&quot;o&quot;&gt;)&lt;/span&gt;
    DNS Servers: 100.100.100.100
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The two addresses on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt; are Alibaba Cloud internal DNS servers provided by DHCP. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.100.100&lt;/code&gt; on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt; is Tailscale’s own DNS, used to resolve machine names inside the tailnet. Tailscale calls this MagicDNS.&lt;/p&gt;

&lt;p&gt;The address &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.2.136&lt;/code&gt; is the important one here. Tailscale assigns machine addresses from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt;, which is the RFC 6598 reserved range for carrier-grade NAT (CGNAT), covering &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0&lt;/code&gt; through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.127.255.255&lt;/code&gt;. Alibaba Cloud’s internal DNS also falls inside that range.&lt;/p&gt;

&lt;p&gt;If Tailscale applies restrictions to that range, Alibaba Cloud DNS would get caught too. So I looked at the iptables rules Tailscale had installed:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;iptables &lt;span class=&quot;nt&quot;&gt;-S&lt;/span&gt; ts-input
&lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; ts-input &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; tailscale0 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ACCEPT
&lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; ts-input &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; udp &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; udp &lt;span class=&quot;nt&quot;&gt;--dport&lt;/span&gt; 41641 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ACCEPT
&lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; ts-input &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 100.64.0.0/10 &lt;span class=&quot;o&quot;&gt;!&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; tailscale0 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; DROP
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The last line means: &lt;strong&gt;drop any packet whose source address is in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt; unless it came in through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The intent of this rule is anti-spoofing. Every machine in the Tailscale network uses an address from this range, and legitimate Tailscale traffic should only arrive through the virtual &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt; interface. If a packet comes in from a physical NIC but claims to be from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.x&lt;/code&gt;, that looks like someone pretending to be a Tailscale machine, so dropping it makes sense.&lt;/p&gt;

&lt;p&gt;But Alibaba Cloud’s internal DNS also lives in that range. The machine sends a query to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.2.136&lt;/code&gt;, which goes out through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt; just fine; the reply comes back with source address &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.2.136&lt;/code&gt;, enters through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt;, and gets hit by that rule.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260920223706392.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;Rule matching: before the fix&quot; /&gt;&lt;/p&gt;

&lt;center&gt;How the same rule treats three kinds of packets: spoofed packets and Alibaba Cloud replies look identical to it&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I verified it directly:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;dig +short +tries&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 github.com @100.100.2.136
&lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt; communications error to 100.100.2.136#53: timed out
&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;dig +short +tries&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 github.com @223.5.5.5
20.205.243.166
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Alibaba Cloud internal DNS timed out, while Alibaba Cloud’s public DNS &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;223.5.5.5&lt;/code&gt; (outside that range) worked. That matched the theory.&lt;/p&gt;

&lt;p&gt;One extra note: this rule has nothing to do with Tailscale’s exit node feature (using one machine as the internet gateway for other devices). As long as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; is running, it installs this rule.&lt;/p&gt;

&lt;p&gt;Actually, before reinstalling the OS, when I was doing some initial checks on this machine, I had already noticed that I couldn’t reach Alibaba Cloud’s instance metadata service at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.100.200&lt;/code&gt; (used for things like instance ID and security group info). It was blocked by the same rule. Back then it only affected metadata. This time it was DNS, which made the impact global—anything that needed domain resolution stopped working.&lt;/p&gt;

&lt;p&gt;That also explains why I’d never seen this on AWS or GCP: their metadata service and DNS live in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;169.254.x.x&lt;/code&gt; link-local range, not in the CGNAT range, so they never collide with this rule.&lt;/p&gt;

&lt;h2 id=&quot;3-first-workaround-switch-dns-to-public-dns&quot;&gt;3. First workaround: switch DNS to public DNS&lt;/h2&gt;

&lt;p&gt;Once I knew the cause, there were a few possible directions:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Approach&lt;/th&gt;
      &lt;th&gt;Problem&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Add an allow rule in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; chain&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; is managed by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; itself and gets rebuilt on every start, so manually added rules get wiped&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale up --accept-dns=false&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Prevents Tailscale from taking over DNS, but then I lose the ability to access other machines in the tailnet by hostname&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Switch DNS to public resolvers outside this range&lt;/td&gt;
      &lt;td&gt;Works, but means no longer using Alibaba Cloud internal DNS&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;At the time, I chose the third option. Alibaba Cloud’s DNS is provided via DHCP, so replacing it required changing two things together: first tell netplan not to accept DNS from DHCP, then explicitly set upstream resolvers for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemd-resolved&lt;/code&gt;.&lt;/p&gt;

&lt;div class=&quot;language-yaml highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;# /etc/netplan/99-dns-override.yaml&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;network&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;version&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;m&quot;&gt;2&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;ethernets&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;na&quot;&gt;eth0&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
      &lt;span class=&quot;na&quot;&gt;match&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;na&quot;&gt;macaddress&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;00:16:3e:xx:xx:xx&lt;/span&gt;
      &lt;span class=&quot;na&quot;&gt;set-name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;eth0&lt;/span&gt;
      &lt;span class=&quot;na&quot;&gt;dhcp4-overrides&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;na&quot;&gt;use-dns&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;kc&quot;&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-ini highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# /etc/systemd/resolved.conf.d/99-public-dns.conf
&lt;/span&gt;&lt;span class=&quot;nn&quot;&gt;[Resolve]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;DNS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;223.5.5.5 119.29.29.29&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Resolution came back, and MagicDNS was unaffected because it goes through the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt; interface.&lt;/p&gt;

&lt;h2 id=&quot;4-the-same-problem-showed-up-again-with-apt&quot;&gt;4. The same problem showed up again with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt;&lt;/h2&gt;

&lt;p&gt;After fixing DNS, I went to install Docker, and then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt update&lt;/code&gt; failed too:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;W: Failed to fetch http://mirrors.cloud.aliyuncs.com/ubuntu/dists/jammy-security/InRelease
   Unable to connect to mirrors.cloud.aliyuncs.com:http:
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;On Alibaba Cloud Ubuntu images, the default package source is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mirrors.cloud.aliyuncs.com&lt;/code&gt;, which is an internal Alibaba Cloud mirror. I checked what it resolved to:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;getent hosts mirrors.cloud.aliyuncs.com
100.100.2.148   mirrors.cloud.aliyuncs.com
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Still &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.x.x&lt;/code&gt;, so it was the same root cause again.&lt;/p&gt;

&lt;p&gt;Following the same logic as DNS, I switched the source to the public mirror &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mirrors.aliyun.com&lt;/code&gt;. This time it worked, but it was slow:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;Time&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt update&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;366 seconds&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Install Docker&lt;/td&gt;
      &lt;td&gt;still not finished after 25 minutes&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The reason was bandwidth. &lt;strong&gt;Traffic over Alibaba Cloud’s internal network is free and not rate-limited; public internet traffic has to squeeze through this machine’s 3 Mbps bandwidth.&lt;/strong&gt; A full Ubuntu &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt update&lt;/code&gt; downloads over 100 MB of index files, which takes several minutes at 3 Mbps. Previously, using the internal mirror didn’t consume public bandwidth at all. Once that internal path was blocked, all traffic got pushed onto the public network.&lt;/p&gt;

&lt;p&gt;At this point, I had hit the same root cause for the third time: metadata, DNS, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; mirrors. Alibaba Cloud’s internal services mostly live in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.x.x&lt;/code&gt;. Replacing them one by one with public endpoints was just whack-a-mole, and every replacement consumed more public bandwidth. I needed to make the whole range work again.&lt;/p&gt;

&lt;h2 id=&quot;5-put-an-allow-rule-at-the-very-top-of-the-input-chain&quot;&gt;5. Put an allow rule at the very top of the INPUT chain&lt;/h2&gt;

&lt;p&gt;As mentioned above, I couldn’t add a rule directly to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt;, because that chain gets rebuilt. But iptables matches in order, and the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INPUT&lt;/code&gt; chain only jumps to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; on its first line. If I allow Alibaba Cloud packets in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INPUT&lt;/code&gt; &lt;em&gt;before&lt;/em&gt; that jump, they’ll never reach the DROP rule:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;iptables &lt;span class=&quot;nt&quot;&gt;-I&lt;/span&gt; INPUT 1 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; eth0 &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 100.100.0.0/16 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ACCEPT
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-i eth0&lt;/code&gt; restriction is the key: Tailscale’s own traffic comes through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt;, so this rule won’t accidentally allow spoofed Tailscale traffic.&lt;/p&gt;

&lt;p&gt;After adding it, internal DNS, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; mirrors, and the metadata service all started working again.&lt;/p&gt;

&lt;p&gt;Before making it persistent, I needed to confirm one thing: when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; restarts, does it reinsert its own jump back at the top? If it does, my rule would get pushed down to line 2 and become useless. So I tested:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# Before restart
1    ACCEPT     all  --  100.100.0.0/16
2    ts-input   all  --  0.0.0.0/0

# After systemctl restart tailscaled
1    ts-input   all  --  0.0.0.0/0
2    ACCEPT     all  --  100.100.0.0/16
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;It does.&lt;/strong&gt; Every time &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; starts, it uses &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-I&lt;/code&gt; to put its jump back at line 1, and my manual rule immediately stops being effective.&lt;/p&gt;

&lt;p&gt;So it wasn’t enough to just save an iptables rule and restore it at boot. I needed to move the rule back to line 1 every time &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; started. I wrote a systemd unit that follows &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-ini highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nn&quot;&gt;[Unit]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;After&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;tailscaled.service&lt;/span&gt;
&lt;span class=&quot;py&quot;&gt;PartOf&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;tailscaled.service&lt;/span&gt;
&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;nn&quot;&gt;[Service]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;Type&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;oneshot&lt;/span&gt;
&lt;span class=&quot;py&quot;&gt;RemainAfterExit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;yes&lt;/span&gt;
&lt;span class=&quot;py&quot;&gt;ExecStart&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;/usr/local/sbin/aliyun-internal-allow.sh&lt;/span&gt;
&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;nn&quot;&gt;[Install]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;WantedBy&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;tailscaled.service&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PartOf&lt;/code&gt; makes it restart when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; restarts, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;WantedBy&lt;/code&gt; makes it get pulled in when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; starts.&lt;/p&gt;

&lt;p&gt;There was also a timing issue in the script: when the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; service becomes active, its iptables rules may not have been installed yet. So the script first waits for the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; jump to appear, then moves my rule back to line 1:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;i &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;seq &lt;/span&gt;1 30&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;iptables &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; INPUT &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ts-input 2&amp;gt;/dev/null &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;break
  sleep &lt;/span&gt;1
&lt;span class=&quot;k&quot;&gt;done
&lt;/span&gt;iptables &lt;span class=&quot;nt&quot;&gt;-D&lt;/span&gt; INPUT &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; eth0 &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 100.100.0.0/16 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ACCEPT 2&amp;gt;/dev/null
iptables &lt;span class=&quot;nt&quot;&gt;-I&lt;/span&gt; INPUT 1 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; eth0 &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 100.100.0.0/16 &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; ACCEPT
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;After restarting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; again, the rule stayed at line 1. Then I switched both &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; and DNS back to Alibaba Cloud internal endpoints:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;Public mirror&lt;/th&gt;
      &lt;th&gt;Internal mirror&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt update&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;366 seconds&lt;/td&gt;
      &lt;td&gt;18 seconds&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Install Docker&lt;/td&gt;
      &lt;td&gt;not finished after 25 minutes&lt;/td&gt;
      &lt;td&gt;19 seconds&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;6-the-internal-address-for-the-image-registry-wasnt-in-that-range&quot;&gt;6. The internal address for the image registry wasn’t in that range&lt;/h2&gt;

&lt;p&gt;After Docker was installed, I needed to pull images. Accessing Docker Hub from mainland China is basically unreliable, so I planned to use Alibaba Cloud Container Registry (ACR) as a relay. I checked the ACR internal endpoint:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;getent hosts registry-vpc.cn-hangzhou.aliyuncs.com
100.103.7.180   registry-vpc.cn-hangzhou.aliyuncs.com
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.103.x.x&lt;/code&gt;—not inside the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.0.0/16&lt;/code&gt; range I had just allowed, so it would still be blocked.&lt;/p&gt;

&lt;p&gt;Alibaba Cloud internal services are not all in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.100.x.x&lt;/code&gt;. To cover everything, I’d have to widen the allow rule to the full &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt; range. Before doing that, I went to see how other people were handling this problem.&lt;/p&gt;

&lt;h2 id=&quot;7-community-workarounds-and-tailscales-official-solution&quot;&gt;7. Community workarounds, and Tailscale’s official solution&lt;/h2&gt;

&lt;p&gt;There’s an issue on Tailscale’s GitHub describing this exact problem: after installing Tailscale on Alibaba Cloud ECS, internal DNS and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; both stop working. It’s still open. There are also quite a few Chinese blog posts about it. Summarizing the approaches I found:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Approach&lt;/th&gt;
      &lt;th&gt;Notes&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Move DNS and mirrors to public endpoints&lt;/td&gt;
      &lt;td&gt;Works, but consumes public bandwidth and requires changing services one by one&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--accept-dns=false&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Only fixes DNS; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; and image registries still fail&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Manually insert allow rules&lt;/td&gt;
      &lt;td&gt;This is what I did above; you also have to deal with reordering when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; restarts&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--netfilter-mode=nodivert&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Tailscale creates rule chains but does not attach the jump; you manage the jump yourself&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-linux-cgnat-drop-rule&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;A node attribute provided officially by Tailscale&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Tailscale’s docs mention that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nodivert&lt;/code&gt; was the recommended approach before &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-linux-cgnat-drop-rule&lt;/code&gt; existed, but now the latter is the recommended one.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-linux-cgnat-drop-rule&lt;/code&gt; is a node attribute. If you assign it to a specific machine in the Tailscale policy file, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; changes that DROP rule in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; to RETURN: instead of dropping the packet, it hands it off to the following rules. Since this rule is generated by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; itself, there’s no issue with it being overwritten on restart, which means I no longer needed that systemd unit babysitting my custom rule.&lt;/p&gt;

&lt;h2 id=&quot;8-if-anti-spoofing-is-disabled-what-replaces-it&quot;&gt;8. If anti-spoofing is disabled, what replaces it?&lt;/h2&gt;

&lt;p&gt;The official docs also mention the tradeoff: removing that rule removes the anti-spoofing protection, so machines on the local network could send packets pretending to come from Tailscale addresses. The recommended compensation is to enable reverse path filtering (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter&lt;/code&gt; is a kernel check: when a packet arrives, the kernel asks, “if I were to send a reply to this source address, which interface would I use?” If that doesn’t match the interface the packet actually came in on, the packet gets dropped. In strict mode (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter=1&lt;/code&gt;), it has to be the exact same interface.&lt;/p&gt;

&lt;p&gt;There was one thing I needed to confirm first: if Tailscale had installed a route for the entire &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt; range via &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt;, then replies from Alibaba Cloud DNS would also be considered “supposed to come in through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt;”, and strict mode would drop them too, bringing me right back to square one.&lt;/p&gt;

&lt;p&gt;So I checked Tailscale’s routing table. It uses table 52:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;ip route show table 52
100.101.102.103 dev tailscale0
100.101.102.104 dev tailscale0
100.100.100.100 dev tailscale0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;These are per-host routes, not a route for the whole range. Then I checked the return path for a few addresses:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;ip route get 100.100.2.136      &lt;span class=&quot;c&quot;&gt;# Alibaba Cloud DNS&lt;/span&gt;
100.100.2.136 via 172.16.x.x dev eth0
&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;ip route get 100.103.7.180      &lt;span class=&quot;c&quot;&gt;# ACR internal&lt;/span&gt;
100.103.7.180 via 172.16.x.x dev eth0
&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;ip route get 100.101.102.103    &lt;span class=&quot;c&quot;&gt;# A machine in Tailscale&lt;/span&gt;
100.101.102.103 dev tailscale0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Alibaba Cloud internal services route back through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt;, and they also arrive through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt;, so they pass. If someone on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eth0&lt;/code&gt; spoofs a packet pretending to be a Tailscale machine, the return path should go through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscale0&lt;/code&gt;, which doesn’t match, so it gets dropped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strict mode turns out to be exactly the anti-spoofing protection I needed, and it’s more precise than the old blanket DROP rule: it only protects real Tailscale addresses and doesn’t accidentally hit Alibaba Cloud.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260920223718431.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;Rule matching: after the fix&quot; /&gt;&lt;/p&gt;

&lt;center&gt;After changing DROP to RETURN and enabling rp_filter, how the same three kinds of packets are handled&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;9-what-i-actually-changed&quot;&gt;9. What I actually changed&lt;/h2&gt;

&lt;h3 id=&quot;1-add-a-tag-and-a-node-attribute-in-the-policy-file&quot;&gt;1. Add a tag and a node attribute in the policy file&lt;/h3&gt;

&lt;div class=&quot;language-json highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nl&quot;&gt;&quot;tagOwners&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;tag:aliyun&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;autogroup:admin&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;nodeAttrs&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;target&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;tag:aliyun&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;attr&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;   &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;disable-linux-cgnat-drop-rule&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;I used a tag instead of applying it directly to my own account because this attribute affects all Linux machines, and I also have some Linux machines outside Alibaba Cloud joining the tailnet. There’s no reason to disable anti-spoofing on those. Using a tag lets me target only the Alibaba Cloud machines.&lt;/p&gt;

&lt;p&gt;Before changing it, I backed up the original policy file. After editing it, I first called Tailscale’s validation API to make sure it was fine, and only then applied it for real.&lt;/p&gt;

&lt;h3 id=&quot;2-tag-this-machine-with-tagaliyun&quot;&gt;2. Tag this machine with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tag:aliyun&lt;/code&gt;&lt;/h3&gt;

&lt;p&gt;You can do this from the machine list in the admin console, or via the API. As soon as I applied the tag, the rule on the machine changed immediately:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;-A ts-input -s 100.64.0.0/10 ! -i tailscale0 -j DROP
-A ts-input -s 100.64.0.0/10 ! -i tailscale0 -j RETURN
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;3-enable-rp_filter&quot;&gt;3. Enable &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter&lt;/code&gt;&lt;/h3&gt;

&lt;div class=&quot;language-ini highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# /etc/sysctl.d/999-ali-hz-tuning.conf
&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;net.ipv4.conf.all.rp_filter&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;1&lt;/span&gt;
&lt;span class=&quot;py&quot;&gt;net.ipv4.conf.default.rp_filter&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The filename needs to start with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;999-&lt;/code&gt;—this is a weird Alibaba Cloud-specific pitfall&lt;/strong&gt;, which I’ll explain in the next section.&lt;/p&gt;

&lt;h3 id=&quot;4-remove-the-old-systemd-unit&quot;&gt;4. Remove the old systemd unit&lt;/h3&gt;

&lt;p&gt;The previous workaround was no longer needed. After removing it, the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INPUT&lt;/code&gt; chain only had the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ts-input&lt;/code&gt; jump left.&lt;/p&gt;

&lt;h3 id=&quot;verification&quot;&gt;Verification&lt;/h3&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Internal DNS   100.100.2.136     OK
apt mirror     100.100.2.148     HTTP 200
Metadata       100.100.100.200   OK
ACR internal   100.103.7.180     HTTP 401      ← previously it timed out
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The 401 from ACR is expected—an image registry returns 401 if you haven’t logged in. The important part is that it responded; before, it just timed out. Docker container networking, SSH over the Tailscale network, and public SSH all worked normally too. I restarted &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; one more time, and the RETURN rule was still there.&lt;/p&gt;

&lt;h2 id=&quot;10-alibaba-clouds-built-in-sysctl-config-can-override-yours&quot;&gt;10. Alibaba Cloud’s built-in sysctl config can override yours&lt;/h2&gt;

&lt;p&gt;Files under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/sysctl.d/&lt;/code&gt; are loaded in filename order, and &lt;strong&gt;later-loaded files override earlier ones&lt;/strong&gt;. Alibaba Cloud images ship with a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-apsara-sysctl.conf&lt;/code&gt; file that contains these lines:&lt;/p&gt;

&lt;div class=&quot;language-ini highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;py&quot;&gt;vm.swappiness&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;py&quot;&gt;net.ipv4.conf.all.rp_filter&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;net.ipv4.conf.*&lt;/span&gt;&lt;span class=&quot;py&quot;&gt;.rp_filter&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It disables &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter&lt;/code&gt; entirely. If your config file is named something like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-xxx.conf&lt;/code&gt;, it may sort before &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-apsara&lt;/code&gt;, and then get overridden back to 0 at boot.&lt;/p&gt;

&lt;p&gt;I actually discovered this pitfall through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;swappiness&lt;/code&gt; first. In an earlier tuning pass, I had written &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;vm.swappiness = 10&lt;/code&gt; in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-ali-hz-tuning.conf&lt;/code&gt;. If I loaded just that file with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sysctl -p&lt;/code&gt;, the runtime value was indeed 10. But when I used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sysctl --system&lt;/code&gt; to simulate the full boot-time load order, the result was 0—overridden by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-apsara&lt;/code&gt;. So every reboot silently undid it.&lt;/p&gt;

&lt;p&gt;Renaming the file to start with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;999-&lt;/code&gt; fixed it, because it sorts after all &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-*&lt;/code&gt; files (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;9&lt;/code&gt; has ASCII code &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0x39&lt;/code&gt;, which is greater than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-&lt;/code&gt; at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0x2d&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;There’s another detail here: if you just use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;999-&lt;/code&gt; may appear before &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;99-apsara&lt;/code&gt;, because UTF-8 collation can ignore punctuation. But &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemd-sysctl&lt;/code&gt;, which loads sysctl settings at boot, uses byte order. To see the real order, use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LC_ALL=C ls&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-shell highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ LC_ALL&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;C &lt;span class=&quot;nb&quot;&gt;ls&lt;/span&gt; /etc/sysctl.d/
...
99-apsara-sysctl.conf
99-sysctl.conf
99-tailscale.conf
999-ali-hz-tuning.conf
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;To verify sysctl config, use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sysctl --system&lt;/code&gt; and check the final value. Don’t rely only on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sysctl -p&lt;/code&gt; for a single file.&lt;/strong&gt; The latter only tells you your file is syntactically fine; it doesn’t tell you whether it wins in the end.&lt;/p&gt;

&lt;h2 id=&quot;11-retrospective&quot;&gt;11. Retrospective&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. If you install Tailscale on Alibaba Cloud, do these steps first&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;In the policy file, give Alibaba Cloud machines a tag and add &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-linux-cgnat-drop-rule&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;Enable &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter=1&lt;/code&gt;, and make the config filename start with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;999-&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;Keep Alibaba Cloud’s default internal DNS and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; mirror config; no need to change them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you do this before installation, you won’t run into the whole chain of problems above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. If the same root cause shows up a second time, it’s time to look for the real cause&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This time I hit the same issue four times in a row: metadata, DNS, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt;, and the image registry. DNS and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt; were each “fixed” by switching to public endpoints, and each fix worked in isolation—but together they were just whack-a-mole, and every change consumed more public bandwidth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. If you manually insert firewall rules, check whether some other program will rewrite them&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The rule I manually inserted into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INPUT&lt;/code&gt; got pushed down to line 2 after &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tailscaled&lt;/code&gt; restarted. If I hadn’t verified that, the rule would have silently stopped working after some future restart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Check the official docs first&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most community solutions involve writing your own script to insert rules, or using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nodivert&lt;/code&gt;, but Tailscale already has an official node attribute for exactly this case. If I had gone through the docs earlier, I could have skipped that whole detour in the middle.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;The root cause this time was simple: Alibaba Cloud puts internal services in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100.64.0.0/10&lt;/code&gt;, and Tailscale uses the same range. Tailscale’s anti-spoofing rule dropped all replies from Alibaba Cloud internal services. AWS and GCP put their internal services in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;169.254.x.x&lt;/code&gt;, which is why I’d never hit this before.&lt;/p&gt;

&lt;p&gt;In the end, I used Tailscale’s official &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-linux-cgnat-drop-rule&lt;/code&gt;, combined with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_filter=1&lt;/code&gt; to restore anti-spoofing protection. DNS, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;apt&lt;/code&gt;, and the image registry all went back to using Alibaba Cloud’s internal network, with no public bandwidth cost and no custom firewall rules to maintain. Along the way I tried public DNS, public mirrors, and manually inserted rules—each one worked at the time, but each only solved the immediate symptom for one service. The actual problem was a subnet conflict, and the final fix was also applied at the subnet level.&lt;/p&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://tailscale.com/docs/reference/cgnat-interoperability&quot;&gt;CGNAT interoperability · Tailscale Docs&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://tailscale.com/docs/reference/netfilter-modes&quot;&gt;Tailscale netfilter modes · Tailscale Docs&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://github.com/tailscale/tailscale/issues/21249&quot;&gt;Linux: ts-input DROP of source 100.64.0.0/10 breaks Alibaba Cloud ECS internal DNS and apt · tailscale/tailscale&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://blog.hellowood.dev/posts/resolve-alibaba-cloud-tailscale-100-network-conflict/&quot;&gt;解决阿里云和 Tailscale 的 100 网段冲突的问题&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
        <pubDate>Sun, 20 Sep 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/cloud/2026/09/20/aliyun-tailscale-cgnat-conflict.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/cloud/2026/09/20/aliyun-tailscale-cgnat-conflict.html</guid>
        
        <category>en</category>
        
        <category>cloud</category>
        
      </item>
    
      <item>
        <title>After Upgrading the Client, Redis Started Returning Tons of MOVED Errors Again: A Local Config That Had Been Lurking for Seven Months</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;At the end of the previous post (&lt;a href=&quot;/middleware/2026/09/03/redis-cluster-resharding-client-stuck.html&quot;&gt;We added a shard to Redis and broke production writes for over an hour&lt;/a&gt;), upgrading redis-py from 5.0.0 to 8.x stopped the client hang issue.&lt;/p&gt;

&lt;p&gt;What it stopped was only the hanging.&lt;/p&gt;

&lt;p&gt;The next day, when I looked at that cluster again, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INFO errorstats&lt;/code&gt; was sitting on tens of millions of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MOVED&lt;/code&gt; responses, and the rejection rate for read commands was hovering around 40%:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;zrevrange         42.5% MOVED
zrangebyscore     35.6%
hgetall           44.9%
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In CloudWatch, each shard was creating 440 new connections per second, but only 280 were connected at any given time. In other words, each connection lived for about half a second on average before being dropped and recreated.&lt;/p&gt;

&lt;p&gt;There was business impact too. In the recommendation service, one provider is responsible for reading users’ like, play, and favorite history. If the Redis read fails, the whole provider throws an exception. An outer &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;except Exception&lt;/code&gt; catches it, logs one critical line, and then continues. The API still returns 200 as usual. So the user profile data was silently dropped, and the request itself showed no obvious error. Worse, that critical log line never made it into Datadog, so this incident left almost no visible trace.&lt;/p&gt;

&lt;center&gt;(Image placeholder: Redis error curve in Datadog, or CloudWatch NewConnections curve staying high)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;This post records the next day and a half of investigation. The final cause wasn’t complicated, but I think the way those intermediate hypotheses were proposed and then ruled out is more useful than the conclusion itself.&lt;/p&gt;

&lt;h2 id=&quot;1-first-figure-out-which-machines-are-connecting&quot;&gt;1. First, figure out which machines are connecting&lt;/h2&gt;

&lt;p&gt;This time I didn’t start by guessing the cause. I started by figuring out &lt;strong&gt;who was connecting to this cluster&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLIENT LIST&lt;/code&gt; showed 21 client IPs. I mapped them back through ECS, and got this:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Service&lt;/th&gt;
      &lt;th&gt;Machine Count&lt;/th&gt;
      &lt;th&gt;redis-py&lt;/th&gt;
      &lt;th&gt;Connection Lifetime&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Recommendation service&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;8.1.0&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;1–3 seconds&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;API service&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;8.0.0&lt;/td&gt;
      &lt;td&gt;1.8–5 hours&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Event consumer service&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;7.1.0&lt;/td&gt;
      &lt;td&gt;68 minutes&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;All 8 recommendation service machines were constantly rebuilding connections. Right next to it, another backend API service (I’ll call it the API service below) was connecting to the same cluster, also using redis-py 8.x, and its connections lived for hours. Same cluster, same major library version, but one was on a seconds scale and the other on an hours scale. From this point on, the API service became the best control group.&lt;/p&gt;

&lt;p&gt;Then I looked at error attribution. The server-side &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;errorstats&lt;/code&gt; is a global counter and doesn’t distinguish clients, but the three services happened to use completely non-overlapping command sets: the recommendation service used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;zrevrange&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;zrangebyscore&lt;/code&gt;, the API service only used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;zrevrangebyscore&lt;/code&gt;, and the event consumer service only used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;zremrangebyscore&lt;/code&gt;. Matching that against the rejection rates:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;zrevrange          42.5% MOVED   ← exclusive to recommendation service
zrangebyscore      35.6%         ← exclusive to recommendation service
zrevrangebyscore    0.004%       ← exclusive to API service
zremrangebyscore    0.0%         ← exclusive to event consumer service
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Same cluster, same moment in time: commands used only by the recommendation service were getting 40% MOVED, while the other two services were at 0. &lt;strong&gt;All MOVEDs were coming from the recommendation service.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The problem was now compressed into one sentence: &lt;strong&gt;same library, same cluster — why does the recommendation service break while the API service doesn’t?&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;2-hypothesis-1-the-removed-socket_timeout&quot;&gt;2. Hypothesis 1: the removed &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;socket_timeout&lt;/code&gt;&lt;/h2&gt;

&lt;p&gt;The first thing that came to mind was socket timeout. There were two reasons.&lt;/p&gt;

&lt;p&gt;First, the only client config difference between the recommendation service and the API service was exactly this: the API service had &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;socket_timeout=3, socket_connect_timeout=2&lt;/code&gt;, while the recommendation service had removed those two parameters the day before. Second, the deleted comment was very explicit:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# Bounded connect/read/write: a hung socket times out and closes itself,
&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;#   instead of exposing a blocked await to external cancellation
&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;#   (which is exactly the window where pool slot leaks happen)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Without timeouts, a blocked await can only end through external cancellation, and cancellation is exactly the window where connection pool leaks can happen. It looked like a complete causal chain.&lt;/p&gt;

&lt;p&gt;But when I checked it against reality, two things didn’t line up. First, timeout is a fallback mechanism. The service itself had no visible errors, and a healthy service should not have 440 hung sockets per second. &lt;strong&gt;Missing a fallback doesn’t create a failure; it just fails to catch one once it happens.&lt;/strong&gt; Second, if this were really a leak, the number of live connections should keep climbing until it hit the limit of 300 and then start erroring. But it stayed stable at 280. &lt;strong&gt;Stable live connections plus hundreds of new connections per second means the connection objects themselves are not leaking — it’s the same set of objects whose sockets are repeatedly disconnecting and reconnecting.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So this hypothesis was ruled out.&lt;/p&gt;

&lt;h2 id=&quot;3-hypothesis-2-moved-and-connection-rebuilds-are-triggering-each-other&quot;&gt;3. Hypothesis 2: MOVED and connection rebuilds are triggering each other&lt;/h2&gt;

&lt;p&gt;The second idea came from reading the library source. redis-py handles MOVED like this:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;except&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;MovedError&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reinitialize_counter&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reinitialize_counter&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;%&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reinitialize_steps&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;   &lt;span class=&quot;c1&quot;&gt;# default 5
&lt;/span&gt;        &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;aclose&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;       &lt;span class=&quot;c1&quot;&gt;# close the whole client
&lt;/span&gt;    &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;nodes_manager&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;move_slot&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;   &lt;span class=&quot;c1&quot;&gt;# only fix this one slot
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Every 5 MOVEDs, the client decides its topology is stale, closes and rebuilds the entire client, and invalidates all connections. And during each rebuild, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;set_nodes()&lt;/code&gt; marks all connections on all nodes as needing reconnect — the library author’s comment says it very directly: “reconnect is lazy and cheap”, assuming topology changes are rare.&lt;/p&gt;

&lt;p&gt;So I got a chain that looked self-consistent: MOVED triggers rebuild, rebuild invalidates all connections, routing changes during rebuild, commands go to the wrong node, which creates more MOVEDs. The numbers even lined up: 208 MOVEDs per second, divided by 5, is about 41 rebuilds; measured &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLUSTER SLOTS&lt;/code&gt; was about 30 times per second, same order of magnitude.&lt;/p&gt;

&lt;p&gt;But one part of this chain didn’t make sense. If the connection is broken, the command should &lt;strong&gt;fail to send&lt;/strong&gt; — why would it &lt;strong&gt;go to the wrong place&lt;/strong&gt;? Connection health and routing correctness are two different things. So I went back to check the “routing table changes during rebuild” step: if redis-py can’t find a slot, it throws an exception; it doesn’t send randomly. And &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aclose()&lt;/code&gt; doesn’t clear the routing table either. That step was just an assumption.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This chain explains why connections were being rebuilt over and over, but it doesn’t explain where the MOVEDs came from.&lt;/strong&gt; MOVED is the cause, not the effect. I had to look elsewhere.&lt;/p&gt;

&lt;h2 id=&quot;4-couldnt-reproduce-it-on-the-bastion-host&quot;&gt;4. Couldn’t reproduce it on the bastion host&lt;/h2&gt;

&lt;p&gt;So I moved to reproduction. The bastion host was in the same VPC and could connect directly to the cluster. I only did read operations.&lt;/p&gt;

&lt;p&gt;I created a client with exactly the same constructor parameters as the recommendation service and sent a few hundred read commands: 0 MOVED. Increased concurrency to 256: still 0. Tried pipeline: 0. Added dependencies one by one, like hiredis and ddtrace: still 0.&lt;/p&gt;

&lt;p&gt;Then I pointed all 16,384 slots in the routing table to a single shard, simulating the old pre-expansion topology — after about 20 commands, the client repaired itself automatically. That was actually useful information: &lt;strong&gt;a stale routing table could not be the cause&lt;/strong&gt;, because the client can fix that on its own, while production had been stuck at 42% for two days.&lt;/p&gt;

&lt;h2 id=&quot;5-a-reproduction-that-didnt-hold-up&quot;&gt;5. A reproduction that didn’t hold up&lt;/h2&gt;

&lt;p&gt;After the bastion path stalled, I tried a different angle: force-refresh topology 33 times per second in the background while sending commands, and measure the increment using the server-side &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MOVED&lt;/code&gt; counter.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;A  no refresh                  MOVED per command = 0.00
B  topology refresh 33/s       MOVED per command = 0.22
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That looked like a signal. At the time I was ready to go change config in that direction.&lt;/p&gt;

&lt;p&gt;But before changing anything, I hooked the client method that handles MOVED, because I wanted to capture what the client thought the slot owner was versus what the server said it was. Result: across 3000 commands, that method was never called once. &lt;strong&gt;My client had not produced a single MOVED.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So where did the 0.22 in group B come from? What I had measured was the server’s global counter, and the production recommendation service itself already had a background rate of 390 MOVEDs per second. In my measurement window, the background value was over nine thousand, and my observed “increment” was only one tenth of that, while production’s own fluctuation was already on that scale. &lt;strong&gt;That 0.22 was completely within the noise floor.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I reran it using client-side counting instead, and both groups were 0.&lt;/p&gt;

&lt;p&gt;If you’re measuring your tiny signal in a place that already has huge background noise, the result is unreliable. And sometimes the noise gives you a number that looks very convincing.&lt;/p&gt;

&lt;h2 id=&quot;6-i-could-reproduce-the-symptom-but-not-the-production-cause&quot;&gt;6. I could reproduce the symptom, but not the production cause&lt;/h2&gt;

&lt;p&gt;Digging further into the source, I found that the routing strategy is decided before client initialization completes: if the client doesn’t yet have a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;default_node&lt;/code&gt;, the command is treated as a “no-key command”, doesn’t consult the routing table, and is sent to a randomly chosen node. And the first line of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aclose()&lt;/code&gt; is to set &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;default_node&lt;/code&gt; to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So I artificially forced &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;default_node&lt;/code&gt; to stay &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt;, ran at concurrency 300, and counted on the client side:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;normal                      0.00 / 0.00 / 0.00
default_node = None         0.91 / 0.89 / 0.89     ← persistent, no decay
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This time I really did reproduce it, and it was persistent. The mechanism itself was valid.&lt;/p&gt;

&lt;p&gt;The problem is that &lt;strong&gt;this only proves that “if you break X, you get this symptom” — it does not prove that X was actually broken in production.&lt;/strong&gt; So I checked the startup logs from the production containers:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;22:39:05 | app.core.database_manager:322
  Redis Cluster connection initialized - host=clustercfg.xxx...
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Initialization had succeeded, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;default_node&lt;/code&gt; is assigned at exactly that step. So production did have a valid &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;default_node&lt;/code&gt;; the state I had constructed did not exist online.&lt;/p&gt;

&lt;h2 id=&quot;7-one-info-line-in-the-startup-logs&quot;&gt;7. One INFO line in the startup logs&lt;/h2&gt;

&lt;p&gt;But reading the startup logs wasn’t wasted effort. Right before that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;initialized&lt;/code&gt; line was this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;22:39:05 | app.core.database_manager:383
  Redis Cluster address_remap enabled for local development
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;All 8 containers logged this line.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;address_remap&lt;/code&gt; is a redis-py parameter. Under normal circumstances, the cluster client first asks for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLUSTER SLOTS&lt;/code&gt;, the cluster tells it which slots belong to which nodes and what those node addresses are, and then the client sends commands according to that table. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;address_remap&lt;/code&gt; lets you rewrite the node address after receiving it. The implementation in the recommendation service was:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;# Address remapping: map cluster-returned private addresses to local port-forwarded addresses (for local development only)
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;address_remap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;address&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;nf&quot;&gt;return &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;      &lt;span class=&quot;c1&quot;&gt;# host = clustercfg
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;No matter which node address the cluster returned, it rewrote all of them to the configured &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt; address. &lt;strong&gt;In essence, it hardcoded all node addresses to the same one.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Its intended use was local development: a laptop connects to the cluster inside the VPC through SSM port forwarding, the cluster returns private IPs, the laptop can’t reach them, so everything has to be rewritten back to the tunnel address. The comment said “for local development only”, and the code default was &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;false&lt;/code&gt;. All of that was fine.&lt;/p&gt;

&lt;p&gt;So how did it get enabled in production? There was no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REDIS_USER_ADDRESS_REMAP&lt;/code&gt; in the task definition environment variables. Then I checked the repo:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;.env (committed in git):  REDIS_USER_ADDRESS_REMAP=true
no .dockerignore in repo
Dockerfile:               COPY . .           ← .env gets baked into the image
app/main.py:16            load_dotenv()      ← fills in values from .env for variables that are not set
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;load_dotenv()&lt;/code&gt; does not override existing environment variables, so the ones explicitly set in ECS were all correct. But ECS did not set &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REDIS_USER_ADDRESS_REMAP&lt;/code&gt;, so the local-development &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;true&lt;/code&gt; from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; took effect.&lt;/p&gt;

&lt;p&gt;Once enabled, when the cluster answered “slots 0-8191 belong to node A, 8192-16383 belong to node B”, remap rewrote both addresses to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt;. The routing table became:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;normal:  slot 0-8191 → node A,         slot 8192-16383 → node B
actual:  slot 0-8191 → clustercfg,     slot 8192-16383 → clustercfg
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The slot calculation was correct, but both results pointed to the same name. When a real connection was created, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt; DNS round-robined between the two nodes, so the effect was equivalent to randomly picking a shard — half the requests would go to the wrong one.&lt;/p&gt;

&lt;p&gt;This time I started with production evidence and then reproduced it:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;address_remap off    1500 commands   MOVED = 0        routing table = [node A, node B]
address_remap on     1500 commands   MOVED = 1459     routing table = [clustercfg]
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The two real nodes in the routing table had been collapsed into one.&lt;/p&gt;

&lt;h2 id=&quot;8-why-nothing-happened-for-seven-months-and-why-the-api-service-was-fine&quot;&gt;8. Why nothing happened for seven months, and why the API service was fine&lt;/h2&gt;

&lt;p&gt;I checked with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;git log -S&lt;/code&gt;, and both the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;address_remap&lt;/code&gt; code and the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;true&lt;/code&gt; in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; came in through the same commit on 2026-01-28. The incident happened seven months later.&lt;/p&gt;

&lt;p&gt;Nothing happened for seven months because &lt;strong&gt;before the scale-out there was only one shard&lt;/strong&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt; only resolved to that single node, so “rewrite all addresses to clustercfg” was effectively the same as not rewriting anything. On 09-02, when the second shard was added, this config changed from harmless into random routing.&lt;/p&gt;

&lt;p&gt;Then it got masked by a more obvious problem — in the previous post, the redis-py 5.0.0 client hang caused all writes to fail, which was much more visible than MOVED. Only after upgrading to 8.x on 09-03 and fixing the hang did MOVED become the remaining curve.&lt;/p&gt;

&lt;p&gt;Why was the API service fine? Its initialization code makes it obvious:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_is_cluster_host&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;              &lt;span class=&quot;c1&quot;&gt;# whether it starts with clustercfg
&lt;/span&gt;    &lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;RedisCluster&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;common&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;        &lt;span class=&quot;c1&quot;&gt;# prod/stage: cluster client
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;else&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;aioredis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;Redis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;common&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;      &lt;span class=&quot;c1&quot;&gt;# dev: regular single-node client
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In the API service, the dev environment connects to a single-node instance with cluster mode disabled, and the client treats it as plain Redis. It doesn’t do topology discovery, so the “can’t reach private node addresses” problem doesn’t exist there. &lt;strong&gt;It never needed the remap config in the first place.&lt;/strong&gt; The recommendation service didn’t have this branch, and used the cluster client locally too, so it needed remap to deal with private addresses.&lt;/p&gt;

&lt;p&gt;One easy point of confusion is worth clarifying here. In the AWS console, all ElastiCache instances are called clusters, but &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Cluster mode: Enabled&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Disabled&lt;/code&gt; are two different things: Disabled is a primary-replica replication group, while Enabled is the actual Redis Cluster protocol. A colleague mentioned that “other services have always connected to cluster through SSM tunnel just fine” — those instances were all Disabled.&lt;/p&gt;

&lt;h2 id=&quot;9-the-fix&quot;&gt;9. The fix&lt;/h2&gt;

&lt;p&gt;The fix was the most conservative possible: leave &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; untouched, and explicitly add &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REDIS_USER_ADDRESS_REMAP=false&lt;/code&gt; to the prod and stage task definitions. ECS environment variables have higher priority than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt;, so this was a four-line change.&lt;/p&gt;

&lt;p&gt;I didn’t remove &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; because I checked and found that 14 other variables from it were currently taking effect in production. Removing it could have side effects I couldn’t confidently predict; that needed to be handled separately.&lt;/p&gt;

&lt;p&gt;After deployment:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;Before Fix&lt;/th&gt;
      &lt;th&gt;After Fix&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Read command MOVED ratio&lt;/td&gt;
      &lt;td&gt;35–45%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;New connections per second&lt;/td&gt;
      &lt;td&gt;~880&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;0.1&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Provider crashes&lt;/td&gt;
      &lt;td&gt;11/minute&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;center&gt;(Image placeholder: error curves before and after the fix, dropping to 0 after deployment)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;One note on the last line. The direct cause of the provider crashes was a bug in redis-py on the reconnect path (&lt;a href=&quot;https://github.com/redis/redis-py/issues/4028&quot;&gt;#4028&lt;/a&gt;, fix already merged but not yet released): connections were being repeatedly disconnected and rebuilt, and during the reconnect handshake, a concurrent disconnect could race with it and throw &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AttributeError&lt;/code&gt;. It wasn’t an independent issue. Once connections stopped being rebuilt over and over, that race condition no longer had a chance to trigger, so it also dropped to zero after the fix.&lt;/p&gt;

&lt;h2 id=&quot;10-retrospective&quot;&gt;10. Retrospective&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Read the logs first, then run experiments&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most effective step in the whole investigation was reading the startup logs from the prod containers. Once the line &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;address_remap enabled&lt;/code&gt; showed up, the problem was basically settled. Before that, I had tried more than a dozen combinations on the bastion host, and every one of them was just guessing.&lt;/p&gt;

&lt;p&gt;The logs had actually been telling me the answer the whole time. I just hadn’t looked first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Being able to produce the symptom does not mean you’ve found the cause&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This time, two different hypotheses could both stably reproduce the same MOVED rate as production on the bastion host, and neither one was what was actually happening online. The difference is the direction of evidence: did you first see evidence in production and then reproduce it, or did you first construct a state and then try to match the symptom? There can be many different paths to the same symptom.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. In a noisy environment, first ask whether the signal is even separable&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That 0.22 “reproduction” was just noise in the server-side counter. Switching to client-side counting made it 0. In the future, when running experiments in a live-traffic environment, the first question should be: how big is the signal, how big is the noise, and can I separate them? If not, measure somewhere else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. How do you keep local config from leaking into production&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; committed into git, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;COPY . .&lt;/code&gt; baking it into the image, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;load_dotenv&lt;/code&gt; filling in missing values — each of those is common on its own. Together, they mean: any local-only switch will go to production unless ECS explicitly sets it. This time it was &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;address_remap&lt;/code&gt;; there are still 14 other variables in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; in the same state.&lt;/p&gt;

&lt;p&gt;The practical fixes are straightforward: exclude &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt; with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.dockerignore&lt;/code&gt;; explicitly list switches in production task definitions instead of relying on defaults; and add environment guards in code for “local only” switches.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;The root cause was actually very simple: a local-development config hardcoded all node addresses to the same one, and it got carried into production through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.env&lt;/code&gt;. With only one shard, it had no effect. After scaling out to two shards, it turned into random routing, and half the reads went to the wrong node. redis-py rebuilds the entire client every 5 MOVEDs, so connections couldn’t live for even 1 second, and the rebuild process then triggered another race condition, causing user profile data to be dropped too.&lt;/p&gt;

&lt;p&gt;Four lines of config fixed everything.&lt;/p&gt;

&lt;p&gt;But between seeing those tens of millions of MOVEDs and changing those four lines, there was a day and a half, seven or eight hypotheses, and a lot of dead ends. Every hypothesis looked internally consistent at the time, and each could produce some kind of symptom on the bastion host, but none of them matched production once checked against real evidence. The step that finally settled it was simply reading through the container startup logs.&lt;/p&gt;

&lt;p&gt;That’s probably the most practical lesson from this incident: when debugging a production issue, finish looking at what production is already telling you before you start running experiments.&lt;/p&gt;
</description>
        <pubDate>Sat, 05 Sep 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/middleware/2026/09/05/redis-cluster-address-remap-moved-storm.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/middleware/2026/09/05/redis-cluster-address-remap-moved-storm.html</guid>
        
        <category>en</category>
        
        <category>middleware</category>
        
      </item>
    
      <item>
        <title>Adding Redis Sharding Took Production Writes Down for Over an Hour</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;This started because the Redis instance used by our recommendation system ran out of memory again. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;used_memory&lt;/code&gt; hit &lt;strong&gt;79.34G / 79.36G (99.98%)&lt;/strong&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Evictions&lt;/code&gt; also started climbing, which meant it had already begun throwing data out.&lt;/p&gt;

&lt;p&gt;We first tried cleaning things up. The biggest chunk in this DB was exposure history (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;device_imp:*&lt;/code&gt; and the like, which records what content each device has seen for recommendation deduplication), taking up more than half the memory, so we started there and deleted keys that hadn’t been accessed for over 30 days based on LRU idle time.&lt;/p&gt;

&lt;p&gt;The result was pretty underwhelming. The reason was simple: keys untouched for over 30 days mostly belonged to churned users, and those users barely used our product anyway, so their exposure history was tiny to begin with. After all that work, we deleted 880k keys and only freed about 5 GB of memory. Basically a lot of effort for almost nothing.&lt;/p&gt;

&lt;p&gt;So my teammate proposed two action items: shorten the TTL, and also just add another shard.&lt;/p&gt;

&lt;p&gt;Adding a shard in ElastiCache is an online operation. AWS advertises it as not affecting traffic, and we were also connecting through the cluster-mode configuration endpoint (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg.xxx&lt;/code&gt;), so in theory the client should automatically detect topology changes.&lt;/p&gt;

&lt;p&gt;And that “in theory” was exactly what caused our online exposure writes to break for over an hour. So this post is about what happened at the application layer after adding the shard, and whether cluster move behaved normally. Even if the process is supposed to be transparent, production still deserves respect. Sure enough, things broke right after the scale-out.&lt;/p&gt;

&lt;h2 id=&quot;1-errors-started-12-seconds-after-the-shard-was-added&quot;&gt;1. Errors started 12 seconds after the shard was added&lt;/h2&gt;

&lt;p&gt;First, here’s the ElastiCache event timeline (all times are UTC):&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;05:11:29  Scaling out replication group from 1 node groups to 2 node groups
05:15:39  Modified replication group to add 1 new node groups - 0002
05:17:18  Migrating slots from node groups 0001 to 0002 to rebalance slots   ← slot migration starts
05:21:24  Moved a total of 8192 slots out of 8192 slots from shard 0001 to shard 0002
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The first application-side error appeared at &lt;strong&gt;05:17:30&lt;/strong&gt;, exactly 12 seconds after slot migration began.&lt;/p&gt;

&lt;p&gt;And during the 12 minutes before that—from the scale-out starting at 05:11 to the new shard being ready at 05:15—not a single error showed up. Adding nodes and creating the shard were both completely fine. It only blew up the moment slots actually started moving.&lt;/p&gt;

&lt;p&gt;The error looked like this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ERROR | app.service:215 - info_collect failed,
exception_type=RedisClusterException,
exception=Redis Cluster cannot be connected. Please provide at least one reachable node: None
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The Datadog graph made it especially obvious: it had been flat at 0, then suddenly shot straight up:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260904015652175.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1788505007127&quot; /&gt;&lt;/p&gt;

&lt;center&gt;The `step.exception` metric had been at 0 the whole time, then jumped straight to 2k–3k / 2min once slot migration started, and stayed there&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;At that rate, it meant 17–25 requests per second were failing to write exposure records.&lt;/p&gt;

&lt;p&gt;Quick side note: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;info_collect&lt;/code&gt; is one step in our recommendation pipeline. It does three things, and the order matters:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;items&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rsp_items&lt;/span&gt;        &lt;span class=&quot;c1&quot;&gt;# 1. First assemble the recommendation results to return to the user
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;                       &lt;span class=&quot;c1&quot;&gt;# 2. Emit metrics
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;PosterMgr&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;run_all&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;...)&lt;/span&gt;     &lt;span class=&quot;c1&quot;&gt;# 3. Finally write exposure records to Redis  ← this is where it blew up
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The response is already assembled in step 1, so step 3 failing doesn’t affect the user getting recommendations. So throughout this entire incident, &lt;strong&gt;the API kept returning 200s, with no timeouts, no 5xxs, and no user complaints at all&lt;/strong&gt;. It was completely silent. If I hadn’t happened to be looking at this Redis key distribution at the time, it probably would have stayed broken even longer.&lt;/p&gt;

&lt;h2 id=&quot;2-i-thought-a-restart-would-fix-it-it-didnt&quot;&gt;2. I thought a restart would fix it. It didn’t.&lt;/h2&gt;

&lt;p&gt;Seeing “cannot be connected”, my first reaction was very natural: when the service started, there was only one shard. Now there were two, and maybe it didn’t know that. The old process was holding stale topology info, so restarting it to fetch topology again should fix it, right?&lt;/p&gt;

&lt;p&gt;So I restarted it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After the restart, it was still failing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That made things more interesting. And even more confusing: the startup logs clearly showed initialization succeeding:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;06:01:41 | app.core.database_manager:299 - Redis Cluster connection initialized -
          host=clustercfg.xxx-redis.xxx.cache.amazonaws.com, port=6379,
          max_connections=300, use_tls=True
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;There wasn’t a single &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;initialization failed&lt;/code&gt; log. In other words, &lt;strong&gt;the process connected just fine at startup, then somehow broke while running&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Later I compared timestamps: initialization succeeded at 06:02, and errors started again at 06:09. So it relapsed after about 7 minutes. Restarting only bought us a few more minutes of life.&lt;/p&gt;

&lt;p&gt;At that point, the hypothesis that “the service didn’t detect the new node” no longer held up. The new process clearly did detect it. It just forgot again after running for a while.&lt;/p&gt;

&lt;h2 id=&quot;3-the-none-at-the-end-of-the-error-was-the-real-clue&quot;&gt;3. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt; at the end of the error was the real clue&lt;/h2&gt;

&lt;p&gt;Later I pulled the full traceback, and only then noticed an important detail:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;File &quot;/usr/local/lib/python3.11/site-packages/redis/asyncio/cluster.py&quot;, line 685, in execute_command
    await self.initialize()
File &quot;/usr/local/lib/python3.11/site-packages/redis/asyncio/cluster.py&quot;, line 392, in initialize
    await self.nodes_manager.initialize()
File &quot;/usr/local/lib/python3.11/site-packages/redis/asyncio/cluster.py&quot;, line 1300, in initialize
    raise RedisClusterException(
redis.exceptions.RedisClusterException: Redis Cluster cannot be connected.
Please provide at least one reachable node: None
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260904020221262.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1788505311066&quot; style=&quot;zoom:25%;&quot; /&gt;&lt;/p&gt;

&lt;p&gt;At first I thought this error message wasn’t very informative. “Please provide at least one reachable node” sounded like a plain network issue. But the network was clearly fine—I could connect to the cluster manually from the bastion host with no problem.&lt;/p&gt;

&lt;p&gt;The key was that final &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt;. I checked the redis-py source, and this is how that exception gets built:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;exception&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;startup_node&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;values&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;try&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;cluster_slots&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;startup_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;execute_command&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;CLUSTER SLOTS&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;startup_nodes_reachable&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;except&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;Exception&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;exception&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;          &lt;span class=&quot;c1&quot;&gt;# ← if connection fails, the error is stored here
&lt;/span&gt;        &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;startup_nodes_reachable&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;raise&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;RedisClusterException&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;...Please provide at least one reachable node: &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;str&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;exception&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If this were really a connectivity issue, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exception&lt;/code&gt; should have contained a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ConnectionError&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TimeoutError&lt;/code&gt;. But it was &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt;, which means the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;for&lt;/code&gt; loop never ran even once—in other words, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;self.startup_nodes&lt;/code&gt; was empty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It wasn’t that the client couldn’t connect to any node. It was that the client no longer had any nodes left to connect to.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;4-how-the-client-lost-all-its-nodes&quot;&gt;4. How the client lost all its nodes&lt;/h2&gt;

&lt;p&gt;Following that lead, I kept digging through redis-py and found this in the exception handling inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;execute_command&lt;/code&gt; (we were using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;redis==5.0.0&lt;/code&gt; in production):&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nf&quot;&gt;except &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;ConnectionError&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;TimeoutError&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;c1&quot;&gt;# Connection retries are being handled in the node&apos;s Retry object.
&lt;/span&gt;    &lt;span class=&quot;c1&quot;&gt;# Remove the failed node from the startup nodes before we try
&lt;/span&gt;    &lt;span class=&quot;c1&quot;&gt;# to reinitialize the cluster
&lt;/span&gt;    &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;nodes_manager&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;pop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;   &lt;span class=&quot;c1&quot;&gt;# ← this is the line
&lt;/span&gt;    &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;close&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;raise&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;As soon as a node times out, it gets removed from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;startup_nodes&lt;/code&gt;. The design assumption here is: “there are still other nodes available for rediscovering topology.”&lt;/p&gt;

&lt;p&gt;But there was a second trap too. I verified this on the bastion host:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;When constructing the client: startup_nodes = [clustercfg.xxx...]
After first initialization:   startup_nodes = [0001-001, 0002-001]   ← replaced by discovered real nodes
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt; configuration endpoint you provide gets discarded after the first successful initialization&lt;/strong&gt;, and replaced with the concrete nodes discovered at that time.&lt;/p&gt;

&lt;p&gt;Those two traps together were what made this fatal:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Before scale-out: single shard → after first initialization, the list only contains {0001-001}
                               (the clustercfg endpoint has already been overwritten and lost)
    ↓
05:17:18 slot migration starts, connection jitter happens, ConnectionError/TimeoutError is raised
    ↓
startup_nodes.pop(&quot;0001-001&quot;)   → the list becomes empty
    ↓
await self.close()              → marks the client as needing reinitialization
    ↓
Every command afterward → initialize() → iterates over an empty list → exception stays None
                       → raise &quot;...reachable node: None&quot;
    ↓
Permanently wedged; it never recovers unless the process restarts
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;So this is especially easy to hit when scaling from a single shard, because the list only has one node. One &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pop&lt;/code&gt;, and it’s empty. If the cluster had already had multiple shards, deleting one node would still leave others, and it probably could have recovered on its own. We just happened to be in the worst possible scenario.&lt;/p&gt;

&lt;p&gt;One thing that confused me at the time was that there were no timeout errors anywhere in the logs. Later it made sense: that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ConnectionError&lt;/code&gt; was swallowed inside redis-py’s internal retry loop. After the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pop&lt;/code&gt;, it does &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;raise&lt;/code&gt;, the outer retry logic catches it and calls &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;initialize()&lt;/code&gt; again, and this time it hits the empty node list, so what finally bubbles up is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RedisClusterException&lt;/code&gt;. That’s all the application sees. &lt;strong&gt;No timeout in the logs does not mean no timeout actually happened.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;5-this-isnt-newpeople-had-already-hit-it-on-github&quot;&gt;5. This isn’t new—people had already hit it on GitHub&lt;/h2&gt;

&lt;p&gt;I searched around and found an issue in the redis-py repo describing exactly the same problem:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://github.com/redis/redis-py/issues/3221&quot;&gt;RedisCluster becomes unrecoverable if all nodes timeout · Issue #3221&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The title was literally our symptom. It explicitly points to that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;startup_nodes.pop&lt;/code&gt; line as the root cause, and also mentions that it’s especially bad in single-node cluster setups, which matched our case perfectly.&lt;/p&gt;

&lt;p&gt;There was another related issue too:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://github.com/redis/redis-py/issues/2472&quot;&gt;async redis cluster should use initial startup nodes during reinitialization in case of failover · Issue #2472&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This one describes the second trap: the async version overwrites the startup nodes you originally configured during first initialization.&lt;/p&gt;

&lt;p&gt;Put those two issues together, and you get the full picture of our incident. The frustrating part is that &lt;strong&gt;#3221 is marked &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;closed as not planned&lt;/code&gt; + &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stale&lt;/code&gt;&lt;/strong&gt;. Upstream basically didn’t care and let it auto-close.&lt;/p&gt;

&lt;h2 id=&quot;6-why-another-service-was-unaffected-the-version-was-three-major-releases-newer&quot;&gt;6. Why another service was unaffected: the version was three major releases newer&lt;/h2&gt;

&lt;p&gt;There was one thing during the investigation that kept bothering me: another backend service of ours was connected to the same cluster, went through the exact same scale-out, and had zero issues.&lt;/p&gt;

&lt;p&gt;At first I thought it was due to different initialization styles:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;# The other service: pass host/port directly
&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;RedisCluster&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ssl&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;...,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;decode_responses&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;...)&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;# Recommendation service: pass startup_nodes
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;ClusterNode&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)]&lt;/span&gt;
&lt;span class=&quot;nc&quot;&gt;RedisCluster&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ssl&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;...,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;I ran a comparison test using the same redis-py version, and it turned out both styles behaved exactly the same: both got overwritten, both could get stuck. Inside redis-py, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;host=&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;port=&lt;/code&gt; just gets converted into a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ClusterNode&lt;/code&gt; and stuffed into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;startup_nodes&lt;/code&gt; anyway, so they end up on the same code path.&lt;/p&gt;

&lt;p&gt;That left only one variable: the version.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Other service         : redis[hiredis]~=8.0.0      → uv.lock pins 8.0.0
Recommendation service: redis==5.0.0               → confirmed by the traceback path
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That’s a full three major versions apart. I checked the same code path in 8.0.0, and that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pop&lt;/code&gt; line is gone:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;# 5.0.0 —— directly removes the node from the list
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;nodes_manager&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;startup_nodes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;pop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;# 8.0.0 —— no longer removes it; just moves the failed node to the end and retries later
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;update_active_connections_for_reconnect&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;disconnect_free_connections&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;nodes_manager&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;move_node_to_end_of_cached_nodes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;last_failed_node_name&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;target_node&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The list never becomes empty, so &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;reachable node: None&lt;/code&gt; can never happen.&lt;/p&gt;

&lt;p&gt;Same cluster, same slot migration, same scale-out, two services: one got wedged, one didn’t. The only variable was the client version. Honestly, the variable control was cleaner than most lab experiments.&lt;/p&gt;

&lt;p&gt;A quick side note on something that’s easy to mix up: &lt;strong&gt;the client version number does not map 1:1 to the server engine version&lt;/strong&gt;. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;8&lt;/code&gt; in redis-py 8.x has no direct relationship to the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;8&lt;/code&gt; in Valkey 8.2. They don’t need to match, and things still work fine. Our server was Valkey 8.2.0 (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INFO&lt;/code&gt; still reports &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;redis_version&lt;/code&gt; as 7.2.4, because Valkey intentionally pretends to be old Redis for compatibility), while our clients were 5.0.0 and 8.0.0, and both could connect and read/write normally.&lt;/p&gt;

&lt;p&gt;But “it works” doesn’t mean “it’s fine to leave it like that forever.” &lt;strong&gt;If the server engine version moves forward, the client should ideally move forward too, and not lag too far behind.&lt;/strong&gt; I think that’s the real takeaway here. There are two reasons:&lt;/p&gt;

&lt;p&gt;First, new engine features are simply unavailable to old clients. RESP3 support, newer commands, client-side caching—those all require client support too. Upgrading only the server is basically wasting the upgrade.&lt;/p&gt;

&lt;p&gt;Second, old clients accumulate bugs that were fixed long ago, but you still haven’t benefited from those fixes. &lt;strong&gt;This incident was exactly that second category.&lt;/strong&gt; Version 5.0.0 was from August 2023, and we missed every fix across three major releases, including the removal of that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pop&lt;/code&gt; line.&lt;/p&gt;

&lt;p&gt;So the reason to upgrade needs to be framed correctly: it’s not “the client version must match the server version,” it’s “the client shouldn’t stay frozen forever.” That distinction matters. Otherwise next time someone might draw a ridiculous conclusion like “then should we downgrade the server instead?” (No. The bug is in the client. It has nothing to do with the server version.)&lt;/p&gt;

&lt;h2 id=&quot;7-temporary-mitigation-add-a-watchdog&quot;&gt;7. Temporary mitigation: add a watchdog&lt;/h2&gt;

&lt;p&gt;At the time, we understood the root cause, but upgrading a client library across three major versions in the middle of the night felt too risky, so we first added a watchdog as a safety net.&lt;/p&gt;

&lt;p&gt;The idea was simple: if the client can get itself into a permanently unrecoverable state, then something outside it should monitor it and rebuild it when that happens:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;async&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_watchdog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;while&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;sleep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;WATCHDOG_INTERVAL_SECONDS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;   &lt;span class=&quot;c1&quot;&gt;# 10 seconds
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;ping_task&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;ensure_future&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;ping&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;done&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;wait&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;({&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ping_task&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;timeout&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;WATCHDOG_PING_TIMEOUT_SECONDS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ping_task&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;done&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;ping_task&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;cancel&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;      &lt;span class=&quot;c1&quot;&gt;# treat slow ping as a transient issue; don&apos;t rebuild
&lt;/span&gt;            &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;error&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ping_task&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;exception&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;error&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;or&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;isinstance&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;error&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;RedisClusterException&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
            &lt;span class=&quot;c1&quot;&gt;# Only RedisClusterException means the client entered a non-self-healing state;
&lt;/span&gt;            &lt;span class=&quot;c1&quot;&gt;# transient errors like timeouts can recover on their own, and rebuilding would just add noise
&lt;/span&gt;            &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;
        &lt;span class=&quot;c1&quot;&gt;# rebuild using the original configuration endpoint
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;new_client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;_build_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;old_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;new_client&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;old_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;close&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;There were two details I paid special attention to:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Rebuilding must use the original configuration endpoint (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;clustercfg&lt;/code&gt;), not whatever nodes the client currently has in memory, because its current node list is already corrupted.&lt;/li&gt;
  &lt;li&gt;Only rebuild on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RedisClusterException&lt;/code&gt;. My first instinct was “rebuild on any error,” but that would be bad—if the network hiccups, rebuilding can easily create a rebuild storm, and in those cases the client can recover by itself anyway.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We set the detection interval to 10 seconds, which means in the worst case we’d lose 10 seconds of exposure data. Compared to “permanently wedged until someone manually restarts it,” that was acceptable.&lt;/p&gt;

&lt;p&gt;During the day, we also cleaned things up on the Redis side and ran verification. Only then was the issue truly under control. The before/after comparison was very clear (20-second window):&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;During incident&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;After recovery&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;MOVED on shard 0001&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+5,567&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+1&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ZADD rejected on shard 0001&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+5,560&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+0&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ZADD succeeded on shard 0002&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+0&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+5,656&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Total commands on shard 0002&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+28&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;+16,902&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The new shard finally started receiving traffic, and writes were now being split normally across the two halves of the slot range.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260904015735296.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1788505052649&quot; style=&quot;zoom:33%;&quot; /&gt;&lt;/p&gt;

&lt;center&gt;After the watchdog was deployed, `step.exception` dropped back to a normal level&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;There was also a small side story here. The curve dropped, but not all the way to zero—it was still around 14%. After checking, we found one old-version instance hadn’t been replaced. Out of 7 instances, 1 was still old, and 1/7 = 14.3%, which matched the residual error ratio exactly. Once we stopped that instance, everything was clean.&lt;/p&gt;

&lt;h2 id=&quot;8-postmortem-a-few-things-more-worth-remembering-than-the-bug-itself&quot;&gt;8. Postmortem: a few things more worth remembering than the bug itself&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The most troublesome failures aren’t the ones that error—they’re the ones that don’t&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Throughout this incident, the API kept returning 200s. No timeouts, no 5xxs, no user complaints. Exposure writes were silently lost for over an hour, and luckily I happened to notice.&lt;/p&gt;

&lt;p&gt;The root cause was this code:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;try&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;step_ins&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;do_process&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;except&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;Exception&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;logger&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;error&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;{} failed, ...&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;step_name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;...)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;datadog_agent&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;increment&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;step.exception&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;tags&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;step:&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;step_name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;      &lt;span class=&quot;c1&quot;&gt;# ← swallow it and continue to the next step
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;except + continue&lt;/code&gt; pattern isn’t wrong by itself. If exposure writing fails, it really shouldn’t take down the whole recommendation API. &lt;strong&gt;The problem is that this degradation path had no alerting attached to it.&lt;/strong&gt; The metric had been emitted the whole time, but nobody had set a threshold on it. A failure that could swallow one-third of all writes just sat there in metrics for over an hour with nobody looking. That’s the real hole we need to patch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. “Try restarting it” can be misleading&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After the restart, the error count did go down for a while, which makes it very easy to think, “looks fixed, let’s keep watching.” In reality, the new process just started from a clean state and then fell into the exact same trap 7 minutes later. To tell whether something is really fixed, you can’t just look at whether errors dropped—you need to look at how far they dropped, whether they hit zero, and whether they bounce back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. When a cloud vendor says “online scale-out doesn’t affect traffic,” they mean the server side&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And to be fair, the server side really did its job. Throughout the whole process, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cluster_state: ok&lt;/code&gt;, all 16384 slots remained assigned, and no data was lost. &lt;strong&gt;But whether the client can keep up with topology changes is a separate matter, and that’s your responsibility.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Next time we do this kind of topology change, the runbook should include one extra step: after the change, check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INFO commandstats&lt;/code&gt; on every shard and confirm the new shard is actually receiving business traffic. This time, the new shard held half the slots and 4.33 million keys, yet its business command count was 0. That signal was already pretty obvious—we just didn’t think to look until after the fact.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;My biggest takeaway from this incident is that every link in the failure chain looked “reasonable” when viewed in isolation.&lt;/p&gt;

&lt;p&gt;The cloud vendor’s online scale-out worked fine; the server stayed healthy the whole time. We used a cluster client and connected through the configuration endpoint, which was also correct. Exceptions were caught and degraded gracefully, preventing the whole API from going down, which even sounds like a good practice. And yet all of those individually reasonable pieces, combined with one line in the client library that says “remove the node if it times out,” turned into a completely silent production incident that lasted over an hour.&lt;/p&gt;

&lt;p&gt;Another takeaway was about debugging method. Along the way I disproved four of my own hypotheses: “stale topology caused retries to exhaust,” “redis 5.0.0 is fundamentally broken for cluster,” “reinitialization clears the node list,” and “close clears the node list.” Every one of them sounded plausible at the time, and every one was disproven by experiments. In the end, the thing that really nailed the problem was that lonely little &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;None&lt;/code&gt; at the end of the error message.&lt;/p&gt;

&lt;p&gt;But after going in such a big circle, the final conclusion was actually pretty simple: &lt;strong&gt;this was a version problem&lt;/strong&gt;. That &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pop&lt;/code&gt; line in 5.0.0 is a real bug. It’s already gone in 8.0.0. Same cluster, same scale-out, and the service running 8.0.0 was unaffected.&lt;/p&gt;

&lt;p&gt;So the most practical lesson is still the same one from earlier: &lt;strong&gt;if the server engine version moves forward, don’t let the client lag too far behind&lt;/strong&gt;. Missing out on new features is one thing. What’s really dangerous is missing fixes you should have gotten for free just by upgrading. Most of the time you won’t notice the difference when everything is calm. But once you do topology changes like scale-out, scale-in, or failover, that technical debt gets collected all at once. We let ours sit untouched across three major versions, and this time we paid the interest too.&lt;/p&gt;
</description>
        <pubDate>Thu, 03 Sep 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/middleware/2026/09/03/redis-cluster-resharding-client-stuck.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/middleware/2026/09/03/redis-cluster-resharding-client-stuck.html</guid>
        
        <category>en</category>
        
        <category>middleware</category>
        
      </item>
    
      <item>
        <title>A Synchronous Blocking Call Froze the Entire Event Loop</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;There have been quite a few issues in the service lately, so I’ve been watching the dashboards much more closely than usual. One day, while looking through the memory metrics, I noticed a very strange segment: Min, Average, and even Max—which normally stayed steadily above 70%—all dropped together. All three lines fell all the way down to somewhere near the cold-start baseline, then spent more than half an hour slowly climbing back to normal.&lt;/p&gt;

&lt;p&gt;I checked the event history in the container orchestration platform, and during that period, almost all running containers for this service were marked as “unhealthy” and removed/rebuilt in bulk. This wasn’t one or two containers flapping a bit—it was a real batch-scale traffic removal event, which is why even Max got dragged down instead of only Min dropping.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822174144524.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787438497584&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: memory usage Min/Max/Average curves, with all three lines dropping together near the cold-start baseline, then slowly climbing back to normal)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I checked the deployment history too. During that time, &lt;strong&gt;nobody deployed anything, and there wasn’t a single release&lt;/strong&gt;. The containers were not restarted as part of a “normal new version rollout.” If you don’t rule that out completely, nothing you find afterward really stands—because a batch restart could very well have just been caused by an ordinary deployment, and there’d be no need to look elsewhere. Only after confirming this was not deployment-triggered did it make sense to keep digging.&lt;/p&gt;

&lt;p&gt;The most confusing part was: the database looked perfectly normal, and the application itself had no OOMs and no worker heartbeat timeout records. &lt;strong&gt;Nothing seemed broken, yet the service was still being judged unhealthy and removed in batches.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In the end, the culprit turned out to be a very easy-to-miss piece of code: a third-party SDK whose interface was written as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async def&lt;/code&gt;, but whose actual internal call was &lt;strong&gt;pure synchronous blocking&lt;/strong&gt;. That one call froze the entire service’s event loop for several seconds, and like knocking over the first domino, it cascaded into the database connection pool, health checks, and eventually large-scale traffic removal and rebuilds. I’m writing down the full investigation process and the underlying mechanics here.&lt;/p&gt;

&lt;h2 id=&quot;1-first-rule-out-the-two-most-common-suspects&quot;&gt;1. First, rule out the two most common suspects&lt;/h2&gt;

&lt;p&gt;The service uses gunicorn + uvicorn workers. When something like this happens, the usual first checks are these two:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;The system OOM Killer force-killed the process&lt;/strong&gt;: if the kernel kills a worker, the gunicorn master process logs a very specific message when it notices the worker disappeared (something like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Worker was sent SIGKILL! Perhaps out of memory?&lt;/code&gt;).&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The worker itself missed heartbeats and got declared dead&lt;/strong&gt;: uvicorn workers periodically “check in” with the gunicorn master process. If the master doesn’t receive heartbeats for a while, it assumes the worker is hung and force-kills/restarts it. The logs would contain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;WORKER TIMEOUT&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I checked how many times those two keywords appeared around the incident window—&lt;strong&gt;not once&lt;/strong&gt;. The two most common explanations for “why was the container replaced” both didn’t match.&lt;/p&gt;

&lt;h2 id=&quot;2-change-the-angle-who-removed-the-containers&quot;&gt;2. Change the angle: who removed the containers?&lt;/h2&gt;

&lt;p&gt;If the application logs weren’t getting me anywhere, then I needed a different angle—look at the &lt;strong&gt;service events maintained by the container orchestration platform itself&lt;/strong&gt; (this information is not in the application logs; it’s a separate operational event stream recorded by the orchestrator).&lt;/p&gt;

&lt;p&gt;Once I pulled that up, things started to make sense. In those few minutes, the orchestration platform marked almost all running containers for this service as “unhealthy,” and the reason was identical across the board:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;(task xxx) is unhealthy in (target-group ...)
due to (reason Request timed out).
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In other words, the load balancer’s active health checks timed out. The process hadn’t actually died—it was just &lt;strong&gt;responding too slowly to the health check probe, slow enough to exceed the load balancer’s threshold&lt;/strong&gt;, so it got removed and rebuilt. And it happened wave after wave: once a few containers were removed, the remaining ones had to absorb more traffic, got even slower, and then more containers were removed. It snowballed.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822174006543.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787438403325&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: orchestration platform event stream screenshot, showing many tasks marked unhealthy and rebuilt due to health check timeouts in a short period)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;At this point the nature of the issue was basically clear: &lt;strong&gt;the process itself most likely never died; it was just so slow to respond to health checks that an external system misjudged it as dead&lt;/strong&gt;. That lines up perfectly with “no OOM, no worker timeout,” because those are two completely separate detection mechanisms.&lt;/p&gt;

&lt;h2 id=&quot;3-trace-backward-what-exactly-was-happening-during-those-few-seconds&quot;&gt;3. Trace backward: what exactly was happening during those few seconds?&lt;/h2&gt;

&lt;p&gt;I looked through the business logs from the tens of seconds before this batch of containers got removed, and found a large number of database errors. The gist was: “all connections in the pool were occupied, waited in queue for 30 seconds and still didn’t get one, so gave up” (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;QueuePool limit ... connection timed out&lt;/code&gt;, the standard SQLAlchemy connection pool error).&lt;/p&gt;

&lt;p&gt;By itself, that error isn’t unusual. Under high concurrency, a saturated connection pool is nothing new. But the strange part was that &lt;strong&gt;the timestamps of these errors were almost all squeezed into the same sub-second window&lt;/strong&gt;, with dozens of them. Under normal circumstances, if the connection pool is exhausted and everyone waits and times out independently, the errors should be spread out—each request enters the queue at a different time, so their 30-second countdowns should expire at different moments. This kind of “all at once” pattern looked wrong. It also matched the earlier observation that “the database itself looked stable”: if the database really lacked capacity, the errors should have been more evenly distributed, not clustered.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822174345393.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787438622864&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: timestamp distribution of connection pool timeout errors in logs, with dozens of entries tightly clustered into a very short window)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;It looked much more like this: “these requests’ countdowns should actually have expired earlier, but the thing responsible for triggering those timeout callbacks stopped functioning normally for a while; then the moment it resumed, all the accumulated callbacks fired at once.” In other words: &lt;strong&gt;it wasn’t that the database was insufficient—it was that something froze the event loop that this whole timeout mechanism depended on.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;4-find-the-real-culprit-a-call-that-looked-async-but-was-actually-sync&quot;&gt;4. Find the real culprit: a call that “looked async but was actually sync”&lt;/h2&gt;

&lt;p&gt;With the hypothesis that “the event loop got frozen,” I went back to check what this service was doing at the time, focusing specifically on places where code was &lt;strong&gt;written inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async def&lt;/code&gt;, but the actual internal call was synchronous blocking&lt;/strong&gt; (this kind of code can fool the eye, but not the event loop—&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await&lt;/code&gt; does not automatically turn an ordinary synchronous function call into a non-blocking one).&lt;/p&gt;

&lt;p&gt;I found one. The project had integrated the official SDK of a third-party AB testing / feature flag platform a long time ago. Judging from the commit history, this wrapper code was written pretty early, back when traffic was probably much lower than it is now, so this kind of synchronous blocking call likely went unnoticed for a long time. The codebase came from multiple contributors, and the overall quality was uneven. A pitfall like this—”looks async, actually sync”—wasn’t something anyone intentionally left behind; it was more like historical baggage that nobody paid much attention to, until recent traffic growth amplified it into a real avalanche. The wrapper looked roughly like this:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;class&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;ExperimentManager&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;

    &lt;span class=&quot;nd&quot;&gt;@classmethod&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;async&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;fetch_variants&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;device_id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;user_id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&amp;gt;&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;dict&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;sh&quot;&gt;&quot;&quot;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;docstring 里写着&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;用线程避免阻塞事件循环&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;……&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&quot;&quot;&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{}&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;user&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;User&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;device_id&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;device_id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;user_properties&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{...})&lt;/span&gt;
        &lt;span class=&quot;c1&quot;&gt;# 这一行是纯同步调用，SDK 内部走的是同步 HTTP 客户端
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;variants&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;fetch_v2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;user&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;v&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;value&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;v&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;variants&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;or&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{}).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;items&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The docstring sounded great—”use a thread pool to avoid blocking the event loop”—but in the actual code there was no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.to_thread&lt;/code&gt;, nor any kind of thread scheduling at all. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cls._client.fetch_v2(user)&lt;/code&gt; was just a plain synchronous method call executed in place. The SDK itself was configured with “3-second timeout, retry once on failure, backoff 0.3–2 seconds before retry, then another 3 seconds for the retry,” so a single call could theoretically take several seconds.&lt;/p&gt;

&lt;p&gt;I checked the SDK’s own failure logs. Under normal conditions they were almost zero, but during the incident minute, failures shot up into the hundreds. And the timing of that spike happened to land &lt;strong&gt;before&lt;/strong&gt; the clustered database connection pool errors.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822174434716.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787438672246&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: bar chart of this third-party SDK’s failure logs aggregated by minute, spiking to hundreds in one minute while adjacent minutes stay in single digits)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The timing across the three evidence chains lined up: &lt;strong&gt;this synchronous call blocked the event loop first → during that period, everything relying on event-loop scheduling (waiting for DB connections, responding to health checks) got delayed → the moment the event loop resumed, a batch of timeouts exploded all at once → health checks couldn’t be answered → the load balancer judged them timed out → the orchestration platform started removing containers.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;5-one-remaining-question-why-did-the-downstream-storm-last-longer-than-the-trigger-itself&quot;&gt;5. One remaining question: why did the downstream storm last longer than the trigger itself?&lt;/h2&gt;

&lt;p&gt;When I put the two timelines side by side, one thing looked asymmetric: that SDK’s failure logs only exploded within a single minute, then disappeared completely after that; but the database connection pool timeout errors only really got started after that minute ended, and kept going for nearly ten minutes, eventually totaling more than 70,000 errors.&lt;/p&gt;

&lt;p&gt;One thing needs to be made clear first: &lt;strong&gt;the database itself did not have a problem during this period&lt;/strong&gt;—core metrics like connection count and CPU stayed stable the whole time. These errors were entirely from the application-side connection pool failing to get a slot and giving up on its own (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;QueuePool&lt;/code&gt; is maintained inside the SQLAlchemy application process itself; it is not the database’s physical connection count being exhausted). This had nothing to do with database capacity. &lt;strong&gt;It was an application-layer problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the question becomes: if the trigger only caused trouble for one minute, why did the downstream application-layer errors last ten times longer?&lt;/p&gt;

&lt;p&gt;Looking closely at the distribution of connection pool errors during that period, it wasn’t one single explosion and then done. There were 5 distinct peaks over roughly ten minutes, each separated by about 2 minutes. The overall shape was “small initial rise, then two or three waves reaching the highest point, then gradual decay.” It was rhythmic oscillation, not random noise. There were probably two mechanisms layered together behind it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First layer: avalanche feedback caused by traffic removal.&lt;/strong&gt; After Amplitude froze the event loop that one time, the first batch of containers got removed and rebuilt because they couldn’t answer health checks. The traffic they had been carrying shifted onto the remaining containers. Those remaining containers suddenly had to handle more requests, so competition for the connection pool naturally became more intense. Then more containers got judged unhealthy because of slower responses and were removed as well. Once this feedback loop starts spinning, it can sustain oscillation for a while even without Amplitude continuing to add fuel, until the number of removed/rebuilt nodes reaches some equilibrium. This also explains why the errors came in waves: each wave roughly corresponds to one cycle of “removal → impact → more removal,” matching the cadence of health checks plus the orchestrator’s node rebuild cycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second layer: restarts themselves were creating new resource pressure.&lt;/strong&gt; When a container gets removed and rebuilt, the first thing the new container has to do is &lt;strong&gt;rebuild the entire connection pool from scratch&lt;/strong&gt;. Those connections don’t magically exist just because a number is written in the config—they have to establish real TCP connections one by one and go through database authentication handshakes. If a batch of containers restarts in a short period, those new containers all compete to establish a large number of new connections at nearly the same time. That “creating connections” action itself takes time and can queue up, which is effectively another round of contention piled on top of already stressed resources, further extending recovery time.&lt;/p&gt;

&lt;p&gt;In other words, that one synchronous Amplitude call merely &lt;strong&gt;knocked over the first domino&lt;/strong&gt;. After it stopped, the chain reaction already had its own inertia. The two mechanisms above had to burn themselves out before things could truly settle down. That’s what makes this kind of avalanche failure so nasty: the root cause may only act up for one minute, but cleaning up the aftermath can take ten times longer.&lt;/p&gt;

&lt;h2 id=&quot;6-break-down-the-mechanics-why-can-one-synchronous-call-drag-so-many-things-down-together&quot;&gt;6. Break down the mechanics: why can one synchronous call drag so many things down together?&lt;/h2&gt;

&lt;p&gt;There are a few details here that I repeatedly verified during the investigation, and they’re worth calling out separately.&lt;/p&gt;

&lt;h3 id=&quot;1-written-inside-an-async-function-does-not-mean-this-line-of-code-is-asynchronous&quot;&gt;1. “Written inside an async function” does not mean “this line of code is asynchronous”&lt;/h3&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await&lt;/code&gt; only suspends a coroutine and yields control at an actual &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await&lt;/code&gt; expression. If you directly call an ordinary synchronous blocking function inside the function body (with no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await&lt;/code&gt;, because it isn’t awaitable at all), Python will not automatically insert a suspension point just because the outer function is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async def&lt;/code&gt;—it just executes normally, waits normally, and returns normally.&lt;/p&gt;

&lt;p&gt;The key point is: the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio&lt;/code&gt; event loop runs on &lt;strong&gt;a single thread&lt;/strong&gt; the whole time, and all coroutines take turns executing on that one thread. A real asynchronous wait gives control back to that thread so it can schedule other coroutines; a naked synchronous call, on the other hand, &lt;strong&gt;occupies that one and only thread in place&lt;/strong&gt; until it returns. During that time, the thread can do nothing else, including scheduling other coroutines or accepting new connections.&lt;/p&gt;

&lt;h3 id=&quot;2-if-we-rewrote-the-health-check-endpoint-as-synchronous-could-that-avoid-frequent-task-removal&quot;&gt;2. If we rewrote the health check endpoint as synchronous, could that avoid frequent task removal?&lt;/h3&gt;

&lt;p&gt;Based on the analysis above, this is really a two-layer issue: &lt;strong&gt;first, the event loop really was blocked for those few seconds&lt;/strong&gt;—that’s just a fact; there’s no way around it. &lt;strong&gt;Second, tasks were judged unhealthy because they couldn’t answer health checks, and that traffic removal in turn amplified the subsequent congestion&lt;/strong&gt;—that’s the nearly ten-minute avalanche described in the previous section.&lt;/p&gt;

&lt;p&gt;So the question needs to be phrased more precisely: &lt;strong&gt;if the health check endpoint were implemented synchronously, could it at least still respond during those few seconds when the event loop was blocked, thereby avoiding the task being judged unhealthy, avoiding traffic removal, and cutting off the avalanche chain at the source?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer is still no, for the same reason as before: regardless of whether an endpoint is ultimately implemented synchronously or asynchronously, &lt;strong&gt;the very first step when a request arrives always has to go through the event loop&lt;/strong&gt;. The event loop is what listens on the port, accepts the new connection, parses the HTTP request, and only then can we even talk about “should this endpoint run directly, or be handed off to a thread pool?” If the event loop can’t get through that first step, it never even reaches the point of “hand it to the thread pool.” So even if the health check endpoint were rewritten as synchronous, during those few seconds when the event loop was truly frozen, it still wouldn’t be able to answer, would still be judged unhealthy, and would still be removed from traffic. This avalanche chain cannot be cut off at the level of “is the endpoint sync or async”—that path simply doesn’t work.&lt;/p&gt;

&lt;h3 id=&quot;3-why-did-the-timeouts-explode-in-a-cluster&quot;&gt;3. Why did the timeouts explode in a cluster?&lt;/h3&gt;

&lt;p&gt;Database connection pool queue timeouts are generally implemented underneath using the event loop’s timer mechanism (“invoke a callback after N seconds to check whether the resource became available”), so they also depend on the event loop continuing to run in order to fire on time. While the event loop is frozen, all the timers that should have expired one after another just sit there. The moment the event loop resumes, they all fire in a batch—that’s the cause of the “clustered explosion” pattern.&lt;/p&gt;

&lt;p&gt;This is also a useful signal for judging whether “the event loop may have been frozen”: &lt;strong&gt;if a large number of the same kind of timeout errors are squeezed into one very short window instead of being evenly distributed, you can strongly suspect a scheduling-layer problem rather than actual resource pressure being that severe.&lt;/strong&gt;&lt;/p&gt;

&lt;h3 id=&quot;4-process-heartbeat--response-time-of-a-single-endpoint&quot;&gt;4. Process heartbeat ≠ response time of a single endpoint&lt;/h3&gt;

&lt;p&gt;Gunicorn decides whether a worker is “alive” based on whether that worker is sending heartbeats on time. That is a &lt;strong&gt;process-level liveness check&lt;/strong&gt;, which is a completely different dimension from how long any specific request has been running. As long as the event loop itself is still scheduling normally, the heartbeat can continue on time. Even if a request runs for several minutes, as long as it is “politely awaiting” rather than “hogging the thread with a naked synchronous call,” the process heartbeat is unaffected, and gunicorn won’t kill the worker.&lt;/p&gt;

&lt;p&gt;Conversely, &lt;strong&gt;“process heartbeat is normal” can never prove that “every endpoint is responding in time”&lt;/strong&gt;. Those are two independent things. Relying on one mechanism to cover for the other is itself an architectural blind spot that’s easy to overlook.&lt;/p&gt;

&lt;h2 id=&quot;7-fix-strategy&quot;&gt;7. Fix strategy&lt;/h2&gt;

&lt;p&gt;What actually needed fixing was that “fake async” call: just honestly move it into a thread pool:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;variants&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_thread&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;fetch_v2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;user&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;With this change, even if the SDK call really blocks for several seconds, it only occupies one thread slot in the thread pool. The event loop itself remains completely unaffected, and other coroutines—including health checks—continue to be scheduled normally.&lt;/p&gt;

&lt;p&gt;While I was at it, I also revisited the relationship between several internal timeout values: the database connection pool queue timeout, the load balancer’s health check threshold, and gunicorn’s own worker heartbeat timeout. Originally they had all been chosen more or less by gut feel, and by coincidence they landed in the same order of magnitude. Once resource pressure actually happened, several mechanisms fired almost simultaneously and reinforced each other into “it really is dead,” which amplified the avalanche effect instead.&lt;/p&gt;

&lt;p&gt;Afterward, I reworked the priority of these values according to the principle that &lt;strong&gt;the mechanism with the smaller blast radius should always fail first&lt;/strong&gt;: local resource wait timeouts should be clearly shorter than process-level watchdog timeouts. That way, under resource pressure, what happens first is always “this one request fails gracefully,” rather than escalating to a much more violent mechanism like “kill the whole process and take a bunch of otherwise healthy requests down with it.”&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;The biggest takeaway from this investigation was finally separating two concepts that are easy to blur together: &lt;strong&gt;process liveness&lt;/strong&gt; and &lt;strong&gt;single-request health&lt;/strong&gt;. They are two independent mechanisms. One only cares whether “the event loop is still turning,” while the other cares “how long this specific request has been running.” Neither can substitute for the other.&lt;/p&gt;

&lt;p&gt;The other, more straightforward lesson is: &lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;async def&lt;/code&gt; is only a syntactic shell; it does not automatically verify whether the calls inside are truly asynchronous.&lt;/strong&gt; When integrating a third-party SDK, even if the docs or docstring confidently claim “this already runs in a thread and won’t block,” it’s still worth taking the time to inspect whether that is actually true underneath—especially for old wrapper code that may not have kept pace with the rest of the project’s async evolution. Those are exactly the kinds of places where this sort of “fake async” likes to hide. And once you really step on it, the cost is often not just “this one endpoint gets a bit slower,” but, like this time, dragging a whole pile of seemingly unrelated mechanisms down together along the event loop’s single scheduling axis.&lt;/p&gt;
</description>
        <pubDate>Sat, 22 Aug 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/cloud/2026/08/22/event-loop-frozen-by-sync-call-cascade.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/cloud/2026/08/22/event-loop-frozen-by-sync-call-cascade.html</guid>
        
        <category>en</category>
        
        <category>cloud</category>
        
      </item>
    
      <item>
        <title>An Emergency Production Incident: Redis Memory Exhaustion, Precise Cleanup, and an Architectural Postmortem</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;Today I received a monitoring alert: memory usage on the primary Redis instance in production had spiked straight to &lt;strong&gt;99.98%&lt;/strong&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Evictions&lt;/code&gt; on the dashboard had started climbing continuously.&lt;/p&gt;

&lt;p&gt;Once this metric fires, it’s actually pretty dangerous, because this instance is configured with the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;allkeys-lru&lt;/code&gt; eviction policy. In other words, once memory is completely full, Redis will start randomly evicting old keys using the LRU strategy so it can keep accepting new writes. But this Redis instance is also serving user sessions, API rate-limit counters, and some foundational business caches. If I just let it keep evicting keys indiscriminately like this, it would only be a matter of time before critical business data got taken out by mistake.&lt;/p&gt;

&lt;p&gt;Since I had confirmed Redis was indeed full, the first step definitely wasn’t to panic-delete things at random. I needed to first understand the situation clearly: &lt;strong&gt;analyze memory usage distribution -&amp;gt; identify the biggest consumer -&amp;gt; design a safe cleanup plan -&amp;gt; eliminate the underlying risk&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So here’s a write-up of this investigation process, the practical script I used to precisely clean up 50% of the data based on hit patterns, and a few architectural issues that surfaced during the postmortem.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260821235633107.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787374585754&quot; style=&quot;zoom:25%;&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: a CloudWatch / Datadog dashboard showing Redis memory hitting 100% and Evictions starting to increase)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;1-no-guessing-in-production-analyze-memory-distribution-across-the-entire-keyspace&quot;&gt;1. No Guessing in Production: Analyze Memory Distribution Across the Entire Keyspace&lt;/h2&gt;

&lt;p&gt;When dealing with a production Redis that’s completely full, the number one taboo is deleting keys based on gut feeling. The number two taboo is taking the “easy” route and running &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;KEYS *&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;FLUSHDB&lt;/code&gt; directly (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;KEYS *&lt;/code&gt; can block the single-threaded server hard, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;FLUSHDB&lt;/code&gt; can take down all online traffic immediately).&lt;/p&gt;

&lt;p&gt;I had to first figure out what was actually consuming the memory. So I wrote a lightweight Python script that uses &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SCAN&lt;/code&gt; cursors to iterate through the entire keyspace in batches, sampling and estimating key counts and memory usage by prefix.&lt;/p&gt;

&lt;p&gt;At the time, the memory distribution across the whole database looked roughly like this (sensitive business prefixes have been masked):&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Key Prefix (masked)&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Key Count&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Share by Count&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Avg per Key&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Estimated Memory Usage&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Business Purpose&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:search_emb:*&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;585k&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;22.8%&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;14.5 KB&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;~8.07 GB (87.5%)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🔴 &lt;strong&gt;Search query vector cache&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:img_url:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;1.063M&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;41.4%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;537 B&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~544 MB (5.9%)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Presigned image URL cache&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;service_go:prod:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;188k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;7.3%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;1.95 KB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~350 MB (3.8%)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Microservice acceleration cache&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:card_cache:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;38k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;1.5%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;2.03 KB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~72.7 MB (0.8%)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Operations card cache&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:session:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;428k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;16.7%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;136 B&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~55.5 MB (0.6%)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;User session state&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:qclass:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;171k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;6.7%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;211 B&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~34.4 MB (0.4%)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Classification tagging cache&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:user:*&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LIMITS:*&lt;/code&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~30k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;1.2%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;100~400 B&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;~15 MB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;User info, rate limiting, and other core business data&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Once I saw this table, the case was basically solved on the spot:
the entire Redis instance only had about &lt;strong&gt;9.2 GB&lt;/strong&gt; of usable memory, and &lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod:search_emb:*&lt;/code&gt; alone was consuming 8.07 GB (87.5%)&lt;/strong&gt;! Sessions, rate limiting, and other core business data were actually taking up only a tiny fraction. They were simply getting squeezed out by this one giant memory hog.&lt;/p&gt;

&lt;p&gt;Here’s the non-blocking script I used to analyze the full keyspace distribution:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;redis&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;redis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;Redis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;your-redis-host&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;6379&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;socket_timeout&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{}&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;while&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;scan&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5000&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;k_str&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;decode&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;utf-8&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;errors&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;replace&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;parts&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k_str&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;parts&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;parts&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;parts&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;parts&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_bytes&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_cnt&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
        
        &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_cnt&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;50&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;memory_usage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_bytes&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prefix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_cnt&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;break&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;total_keys&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;sum&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;s&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;values&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Total keys: &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;total_keys&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;s&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;sorted&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stats&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;items&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;key&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;lambda&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_bytes&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;max&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_cnt&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;reverse&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;avg_m&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_bytes&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;max&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sample_cnt&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;est_mb&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;avg_m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1024&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1024&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;30&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; | &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;8&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; | &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;avg_m&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;mf&quot;&gt;8.1&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;B | &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;est_mb&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;mf&quot;&gt;10.2&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; MB&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;2-check-the-code-can-this-biggest-consumer-be-cleaned-up&quot;&gt;2. Check the Code: Can This Biggest Consumer Be Cleaned Up?&lt;/h2&gt;

&lt;p&gt;After identifying the prefix, I immediately went to the codebase to inspect the corresponding implementation.&lt;/p&gt;

&lt;p&gt;It turned out this was used for semantic search: caching the 1024-dimensional embedding vectors generated by an LLM for user search queries.&lt;/p&gt;

&lt;p&gt;The logic in the code followed a standard &lt;strong&gt;Read-Through&lt;/strong&gt; pattern:&lt;/p&gt;
&lt;ol&gt;
  &lt;li&gt;A user enters a search query, and the system first checks Redis to see whether the query’s embedding is already there;&lt;/li&gt;
  &lt;li&gt;If it’s a hit (Cache Hit), it directly uses the stored vector to query the vector database;&lt;/li&gt;
  &lt;li&gt;If it’s a miss (Cache Miss), it calls the LLM API to generate the 1024-dimensional vector, then writes it back to Redis with a &lt;strong&gt;3-day (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ex=3*86400&lt;/code&gt;)&lt;/strong&gt; TTL.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That explained why there were 585k keys: given the search volume over the past few days, users were entering all kinds of different queries, and the 3-day accumulation had piled up into 8 GB of vector data.&lt;/p&gt;

&lt;h3 id=&quot;can-it-be-cleaned-up-and-how&quot;&gt;Can it be cleaned up? And how?&lt;/h3&gt;
&lt;p&gt;From a business-logic perspective, this is purely an acceleration cache. Deleting it won’t cause any data inconsistency or business errors—because if the cache is missing, the code will automatically fall back to the LLM and regenerate it.&lt;/p&gt;

&lt;p&gt;But &lt;strong&gt;can I just wipe all of it in one shot? No&lt;/strong&gt;.
If I deleted all 585k vectors at once while search traffic was still live, every request would suddenly become a Cache Miss. That would send all concurrent search traffic straight to the downstream LLM API, likely triggering rate limits or timeouts on the model service and causing an even bigger secondary incident.&lt;/p&gt;

&lt;p&gt;So the best solution was: &lt;strong&gt;reduce memory usage without clearing everything, prioritize deleting the coldest data that hasn’t been accessed for a long time, and keep the hot cache intact.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;3-fine-grained-cleanup-precisely-remove-the-coldest-50-by-hit-pattern-idle-time&quot;&gt;3. Fine-Grained Cleanup: Precisely Remove the Coldest 50% by Hit Pattern (Idle Time)&lt;/h2&gt;

&lt;p&gt;So the next question was: how do I determine which vector cache entries in Redis are “cold data” that haven’t been hit in a long time?&lt;/p&gt;

&lt;p&gt;The answer was Redis’s native &lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;OBJECT IDLETIME key&lt;/code&gt;&lt;/strong&gt;.
Under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;allkeys-lru&lt;/code&gt;, Redis maintains an LRU clock for each key. With &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;OBJECT IDLETIME&lt;/code&gt;, I can precisely check how many seconds have passed since the key was last read or written (and this query only reads the clock—it does not alter the key’s access time).&lt;/p&gt;

&lt;p&gt;I first used a Pipeline to collect the idle-time distribution across all 585k vector keys:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Idle Time (since last hit)&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Key Count&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Share&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Decision&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;&amp;gt; 2 days (no access for over 48 hours)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;149k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;25.59%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🔴 &lt;strong&gt;Delete candidate (very cold data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;1 ~ 2 days (no access for 24~48 hours)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;144k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;24.72%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🔴 &lt;strong&gt;Delete candidate (cold data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;12 ~ 24 hours&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;118k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;20.28%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🟢 &lt;strong&gt;Keep (warm data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;6 ~ 12 hours&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;69k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;11.82%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🟢 &lt;strong&gt;Keep (active data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;1 ~ 6 hours&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;80k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;13.73%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🟢 &lt;strong&gt;Keep (high-frequency hot data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;&amp;lt; 1 hour&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;22k&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;3.86%&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;🟢 &lt;strong&gt;Keep (core hot data)&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Once the numbers came out, the picture was very clear: &lt;strong&gt;keys with idle time $\ge 24$ hours (queries nobody had searched for in a full day) accounted for exactly 50.3% (294k keys)!&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those 294k keys were very likely long-tail, low-frequency search queries. Keeping them around brought basically no value other than consuming memory. The remaining 49.7%, on the other hand, had all been hit within the past 24 hours and represented the actually popular queries.&lt;/p&gt;

&lt;h3 id=&quot;execute-a-safe-asynchronous-cleanup&quot;&gt;Execute a Safe Asynchronous Cleanup&lt;/h3&gt;

&lt;p&gt;To guarantee zero impact on production traffic, I used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SCAN&lt;/code&gt; cursors + batched &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Pipeline&lt;/code&gt; checks + &lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UNLINK&lt;/code&gt;&lt;/strong&gt; (which frees memory asynchronously in a background thread and does not block the main event loop), processing 1000 keys per batch with a 10ms sleep between batches:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;redis&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;redis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;Redis&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;host&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;your-redis-host&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;6379&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;socket_timeout&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;# Threshold: cold data not hit for more than 24 hours
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;IDLE_THRESHOLD&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;86400&lt;/span&gt; 

&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;total_scanned&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;total_unlinked&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[]&lt;/span&gt;

&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;清理前内存: &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;info&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;memory&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;used_memory_human&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;while&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;scan&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;match&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;prod:search_emb:*&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2000&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;pipe&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;pipeline&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;transaction&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;pipe&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;object&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;idletime&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;idles&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pipe&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

        &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;zip&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;keys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;idles&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;total_scanned&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;and&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;it&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;IDLE_THRESHOLD&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

            &lt;span class=&quot;c1&quot;&gt;# 每凑齐 1000 个调用一次异步 UNLINK
&lt;/span&gt;            &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1000&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;unlink&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;total_unlinked&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[]&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;sleep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mf&quot;&gt;0.01&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# 短暂停顿，避免主线程抖动
&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cursor&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;break&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;unlink&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;total_unlinked&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;batch_to_delete&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;sleep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;清理完成! 共扫描 &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;total_scanned&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; 个，异步清理冷 Key &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;total_unlinked&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; 个&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;清理后内存: &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;info&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;memory&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;used_memory_human&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;cleanup-results&quot;&gt;Cleanup Results&lt;/h3&gt;
&lt;p&gt;The script ran for about 20 seconds:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;It scanned 585,152 keys and precisely deleted &lt;strong&gt;294,350 cold keys (50.30%)&lt;/strong&gt;;&lt;/li&gt;
  &lt;li&gt;Redis memory dropped immediately from &lt;strong&gt;9.19 GB to 5.22 GB&lt;/strong&gt;, freeing up &lt;strong&gt;4.0 GB&lt;/strong&gt; of precious space;&lt;/li&gt;
  &lt;li&gt;Memory usage fell straight from &lt;strong&gt;99.98% back down to 56.5%&lt;/strong&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Evictions&lt;/code&gt; on the monitoring dashboard dropped to zero, and all alerts affecting Sessions and other core business traffic were fully resolved;&lt;/li&gt;
  &lt;li&gt;The downstream LLM API saw no traffic spike at all, and the system got through the incident smoothly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822000031624.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787374821275&quot; style=&quot;zoom:25%;&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: after cleanup, the monitoring dashboard shows memory usage back down to 56% and business traffic returning to a stable state)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;4-postmortem-what-architectural-problems-did-this-expose&quot;&gt;4. Postmortem: What Architectural Problems Did This Expose?&lt;/h2&gt;

&lt;p&gt;Although the issue was resolved within a little over ten minutes, this memory saturation incident essentially exposed several hard flaws in the earlier architecture design:&lt;/p&gt;

&lt;h3 id=&quot;1-the-serialization-format-was-extremely-wasteful-string-json-vs-binary-pack&quot;&gt;1. The serialization format was extremely wasteful (string JSON vs binary pack)&lt;/h3&gt;
&lt;p&gt;When I checked the code, I found that this 1024-dimensional float vector was being written into Redis using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;json.dumps(vector)&lt;/code&gt; directly.
A 1024-dimensional float array, once converted into a JSON string, stores every float with a bunch of decimal digits, plus commas and brackets, so each key ended up taking &lt;strong&gt;14.5 KB ~ 21 KB&lt;/strong&gt;!&lt;/p&gt;

&lt;p&gt;But in reality, 1024 &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;float32&lt;/code&gt; values packed into binary using Python’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;struct.pack(&apos;&amp;lt;1024f&apos;, *vector)&lt;/code&gt; only take a fixed &lt;strong&gt;4096 bytes (4 KB)&lt;/strong&gt;. Even if I wrapped it in Base64 for compatibility, it would still only be about 5.5 KB.
If binary packing had been used from the start, memory usage for the same number of keys would have dropped by &lt;strong&gt;62%&lt;/strong&gt; immediately, and the original 8 GB of data would have gone down to under 3 GB.&lt;/p&gt;

&lt;h3 id=&quot;2-large-acceleration-caches-and-core-business-data-were-sharing-the-same-instance-no-physical-isolation&quot;&gt;2. Large acceleration caches and core business data were sharing the same instance (no physical isolation)&lt;/h3&gt;
&lt;p&gt;This Redis instance was currently being used as a giant “everything bucket”:
it was serving both &lt;strong&gt;large-volume, disposable, recomputable&lt;/strong&gt; read-through acceleration caches such as search vectors and presigned URLs, and also &lt;strong&gt;small-volume, low-latency, absolutely must-not-be-lost&lt;/strong&gt; core business data such as Sessions, user info, and API rate limiting.&lt;/p&gt;

&lt;p&gt;As soon as one acceleration-cache feature gets rolled out or traffic spikes, it can immediately consume the entire instance’s memory and cause core business keys to be evicted by LRU.
This design is obviously unreasonable. Going forward, &lt;strong&gt;high-consumption caches for search/algorithms&lt;/strong&gt; must be physically separated from the &lt;strong&gt;core business Redis&lt;/strong&gt;.&lt;/p&gt;

&lt;h3 id=&quot;3-the-instance-size-itself-was-too-small&quot;&gt;3. The instance size itself was too small&lt;/h3&gt;
&lt;p&gt;The current primary instance was a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cache.r4.large&lt;/code&gt; (12 GB total physical memory, with only 9.2 GB actually usable by Redis after reserving 25%).
By comparison, Redis instances used by our recommendation system and other compute-heavy services are commonly sized at 64 GB. As search and AI features are being used more and more frequently, a 9.2 GB capacity ceiling is clearly no longer keeping up with business growth. Upgrading the instance size or migrating it is basically inevitable.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;When troubleshooting incidents in high-risk production components, the rhythm is usually: &lt;strong&gt;use monitoring metrics to quickly classify the problem -&amp;gt; use read-only / low-overhead methods to locate the root cause -&amp;gt; weigh business risk and design the lowest-cost mitigation -&amp;gt; finally close the hole at the code and architecture level&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This time, cutting by LRU idle time let me quickly push memory usage back into the safe zone while preserving cache hit rate as much as possible. But more importantly, the postmortem helped firmly surface the real issues around serialization compression, physical instance isolation, and capacity planning.&lt;/p&gt;
</description>
        <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/middleware/2026/08/21/redis-out-of-memory-cleanup-and-reflection.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/middleware/2026/08/21/redis-out-of-memory-cleanup-and-reflection.html</guid>
        
        <category>en</category>
        
        <category>middleware</category>
        
      </item>
    
      <item>
        <title>How to Elegantly Connect a Local GUI Client to a Private Redis on AWS / GCP</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;Recently, I was troubleshooting a Redis memory issue in production where usage was getting dangerously close to full. I wanted to dig into which keys were growing like crazy, whether there were any big keys, or data without a TTL set. In situations like this, the most intuitive approach is usually to fire up a local GUI tool—something like Another Redis Desktop Manager or RedisInsight—connect to Redis, and search or browse the data visually.&lt;/p&gt;

&lt;p&gt;But if you’ve ever used managed Redis on AWS or GCP, you’ve probably run into the same wall: &lt;strong&gt;cloud providers simply do not let you expose Redis to the public internet&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Quick disclaimer first: in large companies with very strict compliance processes, connecting directly from your local machine to a production Redis instance might get you a lot of attention from the security team. But in many startups, small teams, or when you’re acting as the tech lead / ops person handling an urgent incident, being able to observe and diagnose production data quickly and intuitively is often far more practical than fighting through a pile of approval workflows.&lt;/p&gt;

&lt;p&gt;So today I figured I’d write down why cloud providers absolutely refuse to give Redis a public endpoint, and how I moved away from the painful SSM command-line port forwarding workflow to a much cleaner setup using an SSH tunnel directly inside a GUI client.&lt;/p&gt;

&lt;h2 id=&quot;1-why-can-rds-be-public-but-redis-absolutely-cannot&quot;&gt;1. Why can RDS be public, but Redis absolutely cannot?&lt;/h2&gt;

&lt;p&gt;When I first started using cloud services, this felt a little counterintuitive to me: on AWS or GCP, if you buy an RDS instance (PostgreSQL / MySQL) or Cloud SQL, you can usually just click a few buttons in the console, assign a Public IP, whitelist your own IP in the security group, and connect directly from your laptop.&lt;/p&gt;

&lt;p&gt;But when it comes to ElastiCache or GCP Memorystore for Redis, the console doesn’t even offer a “public access” toggle. It’s hard-locked into a private subnet.&lt;/p&gt;

&lt;p&gt;This really isn’t cloud providers being lazy. It’s mostly a consequence of Redis itself:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Historical baggage and weak brute-force resistance&lt;/strong&gt;: Redis was originally designed with the assumption that it would run inside a &lt;strong&gt;fully trusted internal network&lt;/strong&gt;. Early versions didn’t even have password authentication. Even though modern Redis supports Auth and ACLs, Redis can still handle hundreds of thousands of requests per second on a single instance, which makes password brute-forcing over the public internet extremely cheap for attackers.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Dangerous commands and script escape risks&lt;/strong&gt;: Redis has a fairly broad privilege boundary. Commands like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CONFIG SET&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;EVAL&lt;/code&gt; (Lua scripts), and module loading have all historically been involved in serious exploits where attackers used unauthorized access or weak passwords to escape the Lua sandbox, write files to the host machine, or even spawn a reverse shell.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Its single-threaded model is easy to block&lt;/strong&gt;: Redis core request handling is single-threaded. If you expose it directly to the public internet, even a small amount of malicious traffic—or someone accidentally running an expensive &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;KEYS *&lt;/code&gt;—can freeze the whole instance instantly and trigger a cascading failure across downstream services.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So cloud providers are very consistent on this point: &lt;strong&gt;no matter how you configure it, Redis instances must stay inside the VPC private network&lt;/strong&gt;.&lt;/p&gt;

&lt;h2 id=&quot;2-why-does-aws-ssm-port-forwarding-always-feel-awkward&quot;&gt;2. Why does AWS SSM port forwarding always feel awkward?&lt;/h2&gt;

&lt;p&gt;Since direct access isn’t possible, the traditional approach is to use an EC2 bastion host in the same VPC and rely on AWS SSM for local port forwarding:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;aws ssm start-session &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--target&lt;/span&gt; i-0123456789abcdef0 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--document-name&lt;/span&gt; AWS-StartPortForwardingSessionToRemoteHost &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--parameters&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{&quot;host&quot;:[&quot;my-redis-prod.xxxxxx.ng.0001.use1.cache.amazonaws.com&quot;],&quot;portNumber&quot;:[&quot;6379&quot;],&quot;localPortNumber&quot;:[&quot;6379&quot;]}&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Once this command is running, your local &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;127.0.0.1:6379&lt;/code&gt; is effectively mapped to the Redis instance in the cloud, and your GUI client can just connect to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;localhost:6379&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you only need to run an occasional local test from code, this is actually pretty convenient. But if you frequently inspect data in a management tool, a few annoying pain points show up quickly:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Switching between environments is extremely tedious&lt;/strong&gt;: In real life you usually have &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dev&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stage&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prod&lt;/code&gt;, and sometimes a dedicated Redis for recommendation systems or search. Every time you want to switch environments, you have to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Ctrl+C&lt;/code&gt; the current session in the terminal, change the host parameter, and rerun the command. Or you end up opening a bunch of local ports like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;6379&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;6380&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;6381&lt;/code&gt;, and after a while you can’t even remember which port maps to which environment.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Idle sessions disconnect all the time&lt;/strong&gt;: SSM sessions have heartbeat and timeout behavior. If you step away from your computer for a bit without sending requests, or your laptop goes to sleep, the connection in the terminal quietly dies—often stuck at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Waiting for connections...&lt;/code&gt;. Then you go back to the GUI tool, hit refresh, get a timeout, and have to restart everything from the terminal again.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;You have to keep a terminal window open all the time&lt;/strong&gt;: Every day you end up wasting a terminal tab just to keep SSM running, constantly worried you’ll close it by accident.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;3-a-better-approach-use-a-local-key-to-connect-through-the-bastion-host-directly&quot;&gt;3. A better approach: use a local key to connect through the bastion host directly&lt;/h2&gt;

&lt;p&gt;In fact, most modern Redis desktop clients—like Another Redis Desktop Manager and RedisInsight—already support &lt;strong&gt;SSH Tunnel&lt;/strong&gt; natively.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260821234439597.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787373870235&quot; /&gt;&lt;/p&gt;
&lt;center&gt;(Image placeholder: the SSH Tunnel configuration screen in a desktop client; can be replaced during review)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;If you let the desktop client manage the SSH tunnel itself, save each Redis environment as a separate connection card, and just click whichever one you need, it can automatically bring up the tunnel and put it to sleep when idle. The experience becomes much smoother.&lt;/p&gt;

&lt;p&gt;A lot of people get stuck here because &lt;strong&gt;they can’t find the old PEM key for the bastion host, or their access has already expired&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But if you already have permission to run AWS SSM commands locally, then you don’t need to ask anyone for a key again. You can simply generate a new public key on your own machine and use SSM to inject it into the bastion host in one shot.&lt;/p&gt;

&lt;h3 id=&quot;1-generate-a-dedicated-ssh-key-pair-locally&quot;&gt;1. Generate a dedicated SSH key pair locally&lt;/h3&gt;
&lt;p&gt;Run a single command in your local terminal to generate a standard RSA key pair (without a passphrase, so GUI tools can load it directly):&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# Generate a key dedicated to the Redis tunnel&lt;/span&gt;
ssh-keygen &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; rsa &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 2048 &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; PEM &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; ~/.ssh/my_redis_tunnel.pem &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;redis_tunnel_key&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;2-use-aws-ssm-to-inject-the-public-key-into-the-bastion-host&quot;&gt;2. Use AWS SSM to inject the public key into the bastion host&lt;/h3&gt;
&lt;p&gt;Find the EC2 instance ID of the bastion host in the same VPC that has a public IP, then send a CLI command that appends your public key to the bastion host’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;authorized_keys&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;PUB_KEY&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; ~/.ssh/my_redis_tunnel.pem.pub&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;

aws ssm send-command &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--instance-ids&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;i-0123456789abcdef0&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--document-name&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;AWS-RunShellScript&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--parameters&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;commands=[
    &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;mkdir -p /home/ubuntu/.ssh&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;,
    &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;chmod 700 /home/ubuntu/.ssh&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;,
    &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;echo &apos;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$PUB_KEY&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&apos; &amp;gt;&amp;gt; /home/ubuntu/.ssh/authorized_keys&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;,
    &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;chmod 600 /home/ubuntu/.ssh/authorized_keys&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;,
    &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;chown -R ubuntu:ubuntu /home/ubuntu/.ssh&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&quot;&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;
  ]&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The idea is simple&lt;/strong&gt;: the AWS SSM Agent runs inside EC2 as a background daemon with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;root&lt;/code&gt; privileges. Once it receives instructions from the AWS API, it can write your public key directly into the allowlist, completely bypassing the classic deadlock of “you need to log into the server first in order to add your public key.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3 id=&quot;3-configure-the-connection-in-your-redis-client&quot;&gt;3. Configure the connection in your Redis client&lt;/h3&gt;

&lt;p&gt;Now the bastion host already trusts your local private key, so open your Redis desktop client and create a new connection:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;General connection settings&lt;/strong&gt;:
    &lt;ul&gt;
      &lt;li&gt;&lt;strong&gt;Host&lt;/strong&gt;: enter the Redis private endpoint (for example, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;my-redis-prod.xxxxxx.ng.0001.use1.cache.amazonaws.com&lt;/code&gt;)&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Port&lt;/strong&gt;: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;6379&lt;/code&gt;&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Auth&lt;/strong&gt;: leave blank if no password is configured&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;TLS / Cluster&lt;/strong&gt;: choose based on your actual setup (for a standard primary-replica deployment, this is usually left unchecked)&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;SSH Tunnel&lt;/strong&gt;:
    &lt;ul&gt;
      &lt;li&gt;&lt;strong&gt;Check&lt;/strong&gt; ☑️ &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SSH&lt;/code&gt;&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Address&lt;/strong&gt;: enter the bastion host’s public IP (for example, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;54.xxx.xxx.xxx&lt;/code&gt;)&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Port&lt;/strong&gt;: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;22&lt;/code&gt;&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Username&lt;/strong&gt;: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ubuntu&lt;/code&gt; (or the default username for your system image)&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Private Key&lt;/strong&gt;: browse and select the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/.ssh/my_redis_tunnel.pem&lt;/code&gt; you just generated&lt;/li&gt;
      &lt;li&gt;&lt;strong&gt;Password / Passphrase&lt;/strong&gt;: leave blank&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Click test connection, and it should work immediately.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;Once this is set up, the whole development experience improves a lot.&lt;/p&gt;

&lt;p&gt;You can save &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Prod-Redis&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Stage-Redis&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Dev-Redis&lt;/code&gt; as separate connection profiles in your GUI client. All of them can share the same bastion host and private key, while each one points to its own internal Redis host. From then on, whenever you need to inspect a specific environment, just double-click and connect. If the connection drops, the client can usually reconnect automatically, and you can finally stop keeping a pile of SSM terminal sessions running in the background.&lt;/p&gt;

&lt;p&gt;In day-to-day development and operations, a lot of infrastructure security policies—like private network isolation—are non-negotiable. But by making good use of a bastion host and the native capabilities of your client tools, you can absolutely stay within the security boundary while still maximizing your local troubleshooting efficiency.&lt;/p&gt;
</description>
        <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/cloud/2026/08/21/connect-aws-gcp-redis-via-gui.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/cloud/2026/08/21/connect-aws-gcp-redis-via-gui.html</guid>
        
        <category>en</category>
        
        <category>cloud</category>
        
      </item>
    
      <item>
        <title>Investigating a Slow Memory Leak: Creating a New OpenAI Client Every Time, with the Real Issue in the Event Loop</title>
        <description>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;

&lt;p&gt;Recently, I was investigating a slow memory leak in one of our services. After watching the memory curve in production, I noticed the process RSS was steadily creeping upward. It didn’t look like the stepwise growth you’d expect from increasing business data volume. Instead, it was that kind of “slow but constant” diagonal line that just keeps going no matter how long you wait. Given enough runtime, it was clearly going to hit the container memory limit.&lt;/p&gt;

&lt;p&gt;This was actually pretty obvious in monitoring, but it hadn’t turned into a real incident yet—mainly because the business iterates quickly, and we deploy basically every day. Each deployment restarts the machines, and RSS gets reset back to baseline every time, so the leak never had enough time to accumulate all the way to the top before being dragged back to the starting point. But that doesn’t mean the problem wasn’t there. Extend the curve far enough and the conclusion is obvious: as long as the service runs continuously for long enough without restarting, sooner or later it will hit the container memory limit. At that point, either gunicorn will decide memory usage is too high and start killing workers, or some lower-level mechanism will kill it more forcefully. Frequent releases were only delaying the inevitable by accident.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822002614538.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787376366426&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: memory leak curve, RSS keeps climbing over time without converging, eventually approaching the container memory limit during long-running execution)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;After tracing through the code, I found the culprit was a piece of “asset enrichment” logic a teammate had written earlier: when a user uploads an asset, the backend calls an LLM to do tagging and classification, then computes an embedding and writes it into the search index. This logic itself isn’t fast, and we didn’t want it occupying the main thread’s event loop—after all, the main loop still needs to handle incoming requests normally, and nobody wants a single LLM call to stall the whole service. So the original idea was to throw this work into a dedicated background thread pool. That idea itself was fine. The problem was in how the async SDK was being called inside the thread pool. That’s where a subtle trap got buried.&lt;/p&gt;

&lt;p&gt;This post is a record of how I investigated it, why this client couldn’t just be reused directly, and what solution we ended up shipping.&lt;/p&gt;

&lt;h2 id=&quot;1-the-original-approach-start-another-event-loop-inside-the-thread-pool&quot;&gt;1. The original approach: start another event loop inside the thread pool&lt;/h2&gt;

&lt;p&gt;First, here’s roughly what the original code looked like:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;submit_background_job&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;coro_func&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;c1&quot;&gt;# Throw it into the thread pool so the current request isn&apos;t blocked
&lt;/span&gt;    &lt;span class=&quot;n&quot;&gt;executor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;submit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;run&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;coro_func&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;async&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;enrich_asset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;asset_id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;AsyncOpenAI&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;API_KEY&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BASE_URL&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;resp&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;chat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;completions&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;create&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(...)&lt;/span&gt;
    &lt;span class=&quot;bp&quot;&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It’s not hard to guess why someone would write it this way: the main service is built on an async framework, but the enrichment logic needs to call an external LLM API, and we still wanted to keep using the async SDK style of programming (retries, timeouts, streaming, etc. are already built into the SDK, so there’s no need to wrap a separate sync version ourselves). So the implementation picked a very common combination: &lt;strong&gt;inside each thread from the thread pool, use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.run()&lt;/code&gt; to spin up a separate event loop, and then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await&lt;/code&gt; the async SDK inside that loop&lt;/strong&gt;. This way it doesn’t occupy the main service’s event loop, while still letting us keep the async style. On the surface, it looks like the best of both worlds.&lt;/p&gt;

&lt;h2 id=&quot;2-the-cost-you-have-to-create-a-brand-new-client-every-time&quot;&gt;2. The cost: you have to create a brand-new client every time&lt;/h2&gt;

&lt;p&gt;What &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.run(coro)&lt;/code&gt; does is: create a new event loop, run the coroutine, then destroy that loop. In other words, in the code above, every call to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;enrich_asset&lt;/code&gt; goes through a full cycle of “create a brand-new loop -&amp;gt; use it once -&amp;gt; throw it away.”&lt;/p&gt;

&lt;p&gt;The problem starts at this line: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AsyncOpenAI(api_key=..., base_url=...)&lt;/code&gt;. Under the hood, this client wraps an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;httpx.AsyncClient&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;httpx.AsyncClient&lt;/code&gt; internally maintains a real connection pool plus SSL handshake context. Those things are &lt;strong&gt;bound to the event loop they were created on&lt;/strong&gt;. Since every call runs on a brand-new loop, the client is naturally one-shot as well—it can’t really be reused after this request, because the next request will run on a completely different new loop. So the code chose the “seems easiest” option: just create a fresh one every time.&lt;/p&gt;

&lt;p&gt;But as soon as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AsyncOpenAI&lt;/code&gt; is instantiated, it really does allocate the underlying connection pool (even if no request is ever sent). If you create one every time and then leave it to GC after use, in theory that connection pool and its SSL context should be reclaimed together with that event loop. But in practice, cleanup wasn’t happening cleanly. At the time, I pulled a long-window container memory curve from Datadog—about 14 hours—and the fitted numbers looked like this:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Metric&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Value&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Overall memory growth rate over 14 hours&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;about 393 MB/h&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Total net memory increase over 14 hours&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;about 5.35 GB&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Peak memory per instance&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;about 13.4 GB&lt;/strong&gt; (close to the 16 GB container limit)&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;I also ran a simple stress test script that repeatedly hit this enrichment path. The memory curve climbed almost in lockstep with the number of calls and never stopped. It had nothing to do with business data volume. That basically confirmed it: &lt;strong&gt;the leak point was this client that was “created fresh every time and never reused”&lt;/strong&gt;, and at this growth rate, the container would eventually be eaten alive by this logic itself.&lt;/p&gt;

&lt;h2 id=&quot;3-so-why-not-just-reuse-the-client-directly&quot;&gt;3. So why not just reuse the client directly?&lt;/h2&gt;

&lt;p&gt;Once you see “creating a new client every time,” the obvious question is: &lt;strong&gt;why not just cache the client and reuse a single global instance?&lt;/strong&gt;&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;AsyncOpenAI&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;API_KEY&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BASE_URL&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;async&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;enrich_asset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;asset_id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;resp&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;await&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;chat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;completions&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;create&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(...)&lt;/span&gt;
    &lt;span class=&quot;bp&quot;&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The answer is: with the current execution model, &lt;strong&gt;it simply cannot be reused&lt;/strong&gt;. The reason goes back to the previous section: the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;httpx.AsyncClient&lt;/code&gt; inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AsyncOpenAI&lt;/code&gt; is bound to the event loop that existed at the moment it was created. And in this service, background tasks are run via &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.run()&lt;/code&gt;, which means every call to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;enrich_asset&lt;/code&gt; runs on a brand-new event loop, different from the previous one. Even if you store the client in a global variable, it can only be created on &lt;strong&gt;the loop from the first call&lt;/strong&gt;—and that loop is destroyed immediately after use. All later loops would then be trying to send requests through a client that is effectively “attached to a dead loop.” At the mechanism level, that’s simply wrong. Resources from different event loops can’t be shared across them like that.&lt;/p&gt;

&lt;p&gt;In other words, &lt;strong&gt;the prerequisite for “reusing the client” is first having an event loop that itself can be reused&lt;/strong&gt;. As long as every task still creates a fresh loop and throws it away, the client has nowhere stable to “live,” so caching it or not makes no real difference.&lt;/p&gt;

&lt;h2 id=&quot;4-the-fix-dont-touch-business-logic-bind-the-thread-pool-to-long-lived-event-loops-and-reuse-them&quot;&gt;4. The fix: don’t touch business logic, bind the thread pool to long-lived event loops and reuse them&lt;/h2&gt;

&lt;p&gt;The constraint I set for this fix was: &lt;strong&gt;don’t modify the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;enrich_asset&lt;/code&gt; business logic itself&lt;/strong&gt;. How it calls the LLM and how it processes the result should remain unchanged. The only thing to change is the outer execution layer—how this logic gets run. Since the client must be bound to a fixed event loop in order to be safely reused, the solution should be inverted: &lt;strong&gt;instead of creating a throwaway loop for every task, give each thread in the thread pool its own long-lived event loop, and run all tasks dispatched to that thread on the same loop.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The concrete implementation uses thread-local storage (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;threading.local&lt;/code&gt;) to keep a persistent &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.Runner&lt;/code&gt; per worker thread (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.Runner&lt;/code&gt; was introduced in Python 3.11+ as a persistent event loop wrapper; functionally it’s equivalent to manually maintaining an event loop that isn’t destroyed):&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;threading&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;contextvars&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;concurrent.futures&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ThreadPoolExecutor&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;_state&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;threading&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;local&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# each thread has its own .runner
&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_get_runner&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&amp;gt;&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;asyncio.Runner&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;runner&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;getattr&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;_state&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;runner&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;runner&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;runner&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_state&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;runner&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;Runner&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;runner&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_run_on_worker_loop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;coro_fn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_get_runner&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;().&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;run&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;coro_fn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;context&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;executor&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;max_workers&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;thread_name_prefix&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;bg-worker&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;submit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;coro_fn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;contextvars&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;copy_context&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;executor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;submit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;_run_on_worker_loop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ctx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;coro_fn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;kwargs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;With this setup, every task running on the same thread uses &lt;strong&gt;the same event loop&lt;/strong&gt;. Once that prerequisite is in place, client caching finally becomes valid:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;_clients&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{}&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# (id(loop), api_key, base_url) -&amp;gt; client
&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;get_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;str&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;str&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;loop&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;asyncio&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;get_running_loop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;key&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;loop&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_clients&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;key&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;bp&quot;&gt;None&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;AsyncOpenAI&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;api_key&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;base_url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;_clients&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;key&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;client&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Including &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;id(loop)&lt;/code&gt; in the cache key is basically saying: “this client may only be used on the loop it was created on; different loops maintain their own separate instances.” Since each thread in the thread pool now has exactly one long-lived loop, the practical effect is that &lt;strong&gt;the whole process will only ever have as many clients as the thread pool size&lt;/strong&gt;—for example, 4 threads means 4 clients—instead of growing without bound as the number of calls increases.&lt;/p&gt;

&lt;p&gt;When the process exits or the thread pool shuts down, remember to close the client on &lt;strong&gt;that thread’s own loop&lt;/strong&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;await client.close()&lt;/code&gt;). Don’t directly close another thread’s connection pool across threads, or you may run into some very weird errors.&lt;/p&gt;

&lt;p&gt;That leads to the core conclusion of this fix: &lt;strong&gt;“Can an async client be reused?” is fundamentally the same question as “Can the event loop behind it be reused?”&lt;/strong&gt; The two are tightly coupled. Client caching only becomes meaningful once the event loop itself is long-lived and reusable.&lt;/p&gt;

&lt;h2 id=&quot;5-results&quot;&gt;5. Results&lt;/h2&gt;

&lt;p&gt;After deploying the fix, I compared two runtime windows of similar length—both around 14 hours:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Metric dimension&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Before fix&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;After fix&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Change&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Growth rate over 14 hours&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 393 MB/h&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 210 MB/h&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;down about 47%&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Total net memory increase over 14 hours&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 5.35 GB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 2.99 GB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;&lt;strong&gt;net increase reduced by about 44%&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Peak memory per instance (14h)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 13.4 GB (close to 16 GB limit)&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 9.4 GB&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;about 6.5 GB of safety margin left&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822002808695.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787326277575&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Image placeholder: overlay comparison of memory curves before and after the fix over similar runtime windows, with a clearly reduced slope)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The growth rate was cut almost in half, and the peak moved completely out of the danger zone near the container limit. So the biggest part of the problem was effectively contained.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;At the end of the day, this whole issue boils down to one sentence: &lt;strong&gt;“Can an async client be reused?” is really asking “Can the event loop behind it be reused?”&lt;/strong&gt; The two are tightly bound together, and you can’t solve only the surface-level part.&lt;/p&gt;

&lt;p&gt;A few takeaways worth writing down:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncio.run()&lt;/code&gt; creates a brand-new event loop every time. Wrapping each background task with it is a very common convenience pattern, but if the task uses resources that are bound to the event loop (connection pools, SSL contexts), this “one-shot loop” model makes those resources impossible to safely reuse. The result is that you have to recreate them every time, and the cost is connection-pool leakage.&lt;/li&gt;
  &lt;li&gt;On the flip side, simply turning the client into a global singleton doesn’t necessarily help. If the event loop itself is still one-shot, then the global singleton just gets permanently bound to the already-dead loop from the first call, and becomes unusable.&lt;/li&gt;
  &lt;li&gt;The correct order is to &lt;strong&gt;first make the event loop long-lived and reusable&lt;/strong&gt; (for example, by binding a persistent loop to each thread in the thread pool), and only then cache resources that are bound to that loop. That’s the real fix.&lt;/li&gt;
  &lt;li&gt;Even after deploying the fix, I didn’t assume the story was over. The growth rate was cut in half, but the curve didn’t become perfectly flat. There’s still a small residual increase of around 200 MB per hour, which is probably due to a few other smaller issues in other modules and not the same root cause as this one. &lt;strong&gt;This wasn’t a magical “one fix cures all” patch. It just plugged the biggest hole. The remaining small tails still need to be shaved down one by one.&lt;/strong&gt;&lt;/li&gt;
  &lt;li&gt;For this kind of “slow leak,” a stress test script plus the memory curve is the most direct qualitative tool. If the curve rises in sync with the number of calls, you can usually narrow the suspect range down pretty quickly to “what resource isn’t being properly reused or closed.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One last note. Strictly speaking, the way this business logic is written isn’t really best practice. Throwing a slow, heavy LLM call directly into the web service’s own thread pool always comes with trade-offs. But the code had already been written this way and had been running in production for a long time, and there wasn’t a strong business reason to rewrite it. So for this fix, I deliberately set the principle to &lt;strong&gt;zero changes to business logic&lt;/strong&gt;, and only worked on the layer of “how this logic gets executed,” keeping both the scope of change and the risk as small as possible. And in practice, the result did meet expectations.&lt;/p&gt;

&lt;h2 id=&quot;postscript&quot;&gt;Postscript&lt;/h2&gt;

&lt;p&gt;That said, if I were designing this from scratch, the better approach would be &lt;strong&gt;not to run this kind of heavy task inside the web container at all&lt;/strong&gt;. Move it to offline machines, or process it with offline compute resources like Lambda. There are mainly two reasons:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Isolation&lt;/strong&gt;: this kind of heavy task that consumes both CPU and IO will naturally interfere with the web service if they share the same process. And the web service itself has much stricter requirements for stability and low latency, so the cost of interference is much higher. It’s far more cost-effective to let a less latency-sensitive offline task absorb that risk than to make the web service shake along with it.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Cost&lt;/strong&gt;: memory in web services is usually much more expensive than offline compute resources. This kind of heavy logic—loading model requests once, processing large objects—even if everything can eventually be reclaimed by GC, will still create memory spikes during execution. And if you’re even slightly careless (like this time), some of it may stay stuck in the heap forever and turn into a real leak. Putting this kind of high-memory operation on cheaper offline resources is simply the more economical architectural choice.&lt;/li&gt;
&lt;/ul&gt;
</description>
        <pubDate>Wed, 19 Aug 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/middleware/2026/08/19/openai-async-client-event-loop-leak.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/middleware/2026/08/19/openai-async-client-event-loop-leak.html</guid>
        
        <category>en</category>
        
        <category>middleware</category>
        
      </item>
    
      <item>
        <title>A Database Avalanche Incident Investigation</title>
        <description>&lt;h2 id=&quot;preface&quot;&gt;Preface&lt;/h2&gt;

&lt;p&gt;That day I was busy with something else when I suddenly got pulled into an incident chat. The only message was one line: &lt;strong&gt;every single service endpoint is timing out like crazy&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When you see a description like “timeouts across the board,” the first instinct is: is this that old issue happening again? We had investigated a similar avalanche before, and the root cause back then was a call that was “written inside an async function but was actually synchronously blocking,” freezing the event loop and then dragging down health checks, which caused the container orchestration platform to mark instances unhealthy, remove traffic, and restart them. (I wrote up the full investigation of that one in a separate post, so I won’t repeat it here.) I started digging with that assumption in mind, but this time it turned out to be a completely different pit from start to finish. Still, once I kept digging, it pulled out more problems than I expected, and the chain was much longer too.&lt;/p&gt;

&lt;p&gt;This post starts with damage control, then moves on to the retrospective—because when something breaks, what everyone wants to know first is always “how is it now,” and the analysis can come later.&lt;/p&gt;

&lt;h2 id=&quot;1-first-rule-out-the-possibility-of-the-old-problem-happening-again&quot;&gt;1. First, rule out the possibility of “the old problem happening again”&lt;/h2&gt;

&lt;p&gt;Based on last time’s experience, the first step as usual was to check the service’s own basic metrics: CPU usage, memory pressure. Everything looked normal, with no abnormal spikes at all.&lt;/p&gt;

&lt;p&gt;Since what people were reporting was “all endpoints are timing out,” what mattered more was really the latency of the core business endpoints themselves. But since I suspected the old issue, the first thing I still did was take a look at the TP95 of the health check endpoint—if this was another event-loop-freeze situation, the health check would very likely be slowed down too, which makes it a very useful signal. This time, though, the health check TP95 was completely normal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Basic service metrics normal + health check normal&lt;/strong&gt;—once those two were confirmed, I could basically rule out a recurrence of the same old issue. Since the application process itself looked healthy, the problem was probably not in the application layer. Time to look toward the middleware layer.&lt;/p&gt;

&lt;h2 id=&quot;2-shift-to-middleware-database-cpu-was-already-maxed-out&quot;&gt;2. Shift to middleware: database CPU was already maxed out&lt;/h2&gt;

&lt;p&gt;I immediately checked the database—and one look was enough to tell something was seriously wrong. CPU had already &lt;strong&gt;hit the ceiling&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822181335361.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;Database CPU spike screenshot&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Database CPU usage during the incident window. Usually calm and flat, but here it went straight to 100%&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I quickly went to the AWS console hoping to see more detailed execution info, only to find that this database &lt;strong&gt;did not have Performance Insights enabled&lt;/strong&gt;, so I couldn’t directly inspect historical Top SQL, wait event distribution, or other fine-grained data. Fortunately, the console still exposed some basic SQL latency stats. I glanced at the numbers—and &lt;strong&gt;the durations were basically all in the thousands of seconds&lt;/strong&gt;, which honestly startled me. For a moment I even wondered whether I had misread the unit or whether the dashboard was broken. Later I confirmed I hadn’t misread it. At that scale, some statements had already been stuck running in the database for close to an hour, or even longer.&lt;/p&gt;

&lt;p&gt;Without Performance Insights, fine-grained historical analysis was off the table. So I had to fall back to the dumb method: connect directly with a DBA account and manually inspect the live state using system views like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pg_stat_activity&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;3-dig-directly-into-pg_stat_activity-two-deadlocked-offenders-surfaced&quot;&gt;3. Dig directly into pg_stat_activity: two deadlocked offenders surfaced&lt;/h2&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;SELECT&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pid&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;state&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;wait_event_type&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;wait_event&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;now&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query_start&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;AS&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;duration&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
       &lt;span class=&quot;k&quot;&gt;left&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;150&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;AS&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;FROM&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pg_stat_activity&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;WHERE&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;state&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;idle&apos;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;ORDER&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;BY&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;duration&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;DESC&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;LIMIT&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;30&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Sorting by duration and scanning the results, two obvious offenders jumped out immediately:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first one&lt;/strong&gt;: a complex analytical query (cross-table aggregate statistics, something like “count the distribution of dialogue rounds per game from the event table”). Its state was &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;active&lt;/code&gt;, and it had been &lt;strong&gt;running continuously for more than two days&lt;/strong&gt;. It wasn’t just hanging there either—&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;wait_event_type&lt;/code&gt; showed that it was genuinely burning CPU, and it even had two parallel worker processes attached. In other words, this was not some “forgotten idle transaction quietly lying around.” It was a large query that had been actively consuming CPU for two and a half days, with parallelism enabled.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second one&lt;/strong&gt;: a batch of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; statements with identical parameters and only different PIDs, all targeting the same table. Their state was also &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;active&lt;/code&gt;, and the wait events were all over the place—row locks, transaction locks, internal buffer locks. Their durations ranged from several minutes to &lt;strong&gt;nearly 5 hours&lt;/strong&gt;. At a glance, it was obvious this was the same scheduled job being triggered over and over, piling up into a queue layer by layer—the oldest one had already blocked nearly 30 later arrivals.&lt;/p&gt;

&lt;p&gt;At that point the profiles of both offenders were pretty clear: one was continuously bleeding the system dry, and the other was deadlocked in a queue. And both were the kind where letting them continue was completely pointless.&lt;/p&gt;

&lt;h2 id=&quot;4-stop-the-bleeding-first-kill-them-and-let-the-database-breathe&quot;&gt;4. Stop the bleeding first: kill them and let the database breathe&lt;/h2&gt;

&lt;p&gt;At an incident site, the first principle is always stop the bleeding before investigating the cause. Once I confirmed these sessions were effectively stuck and there was no value in letting them continue, I cleared them directly to free up resources. This wasn’t just killing the two “representatives” mentioned above—I cleaned them up in batches by category:&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;-- First cancel the longest-running session in that analytical query chain;&lt;/span&gt;
&lt;span class=&quot;c1&quot;&gt;-- the other parallel workers will terminate along with it&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;select&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pg_cancel_backend&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pid&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;ul&gt;
  &lt;li&gt;The analytical query itself plus its two parallel workers: &lt;strong&gt;3 sessions&lt;/strong&gt; terminated together;&lt;/li&gt;
  &lt;li&gt;The &lt;strong&gt;32&lt;/strong&gt; queued &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; sessions in the lock chain: after confirming their durations and states one by one, I cleared all of them, not just the head of the queue;&lt;/li&gt;
  &lt;li&gt;Sessions stuck on leaderboard count queries due to missing indexes and doing full table scans: there were &lt;strong&gt;738&lt;/strong&gt; of them running at the time, and I batch-canceled those too.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After that, I watched the database metrics closely: CPU dropped immediately, active sessions fell from several thousand down to a few dozen, and connection count dropped along with them. &lt;strong&gt;A few minutes later&lt;/strong&gt;, timeout errors across all online endpoints had basically disappeared, and containers were no longer repeatedly being marked unhealthy by health checks, restarted, and removed from traffic—the avalanche chain had been cut off at the root.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822182733778.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787441244261&quot; /&gt;&lt;/p&gt;
&lt;center&gt;(Placeholder image: after killing the two stuck session groups, database active sessions / CPU usage dropped sharply back to normal levels)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;One thing to make clear: this step was only damage control. The two “lesions” were removed, but it answered none of the “why” questions—why that analytical query could run for two and a half days without anyone noticing, why a scheduled job that looked protected could still pile up nearly 30 concurrent runs, and so on. Those were the things I only had time to investigate after the system was stabilized.&lt;/p&gt;

&lt;h2 id=&quot;5-retrospective-question-one-why-did-it-only-blow-up-after-two-and-a-half-days&quot;&gt;5. Retrospective question one: why did it only blow up after two and a half days?&lt;/h2&gt;

&lt;p&gt;Tracing backward through logs and the incident timeline, the first question was: that analytical query had clearly started two and a half days earlier, so why did nothing happen for so long, and why did it collapse at this exact moment?&lt;/p&gt;

&lt;p&gt;The answer had to do with the machine spec of this database. This instance was quite beefy (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;db.m7i.8xlarge&lt;/code&gt;): 32 vCPUs, 128 GB memory, and high-performance storage. On a machine of that size, even with one heavy query continuously running at full parallelism, it might not be enough to crush the system outright in the short term. More often it just “quietly consumes more resources” and “slows background cleanup tasks a bit” in ways that aren’t very obvious, while the business side barely notices.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822182904008.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1787441341044&quot; /&gt;&lt;/p&gt;

&lt;center&gt;(Placeholder image: database instance spec screenshot, dozens of CPU cores / over 100 GB memory / high-performance storage)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;In other words, the database had actually &lt;strong&gt;been brute-forcing its way through for more than 50 hours&lt;/strong&gt;. That analytical query had been bleeding resources in the background the whole time, garbage collection had been slowed down, and tables had been bloating day by day—but the machine had enough headroom to survive it. It wasn’t until the day’s normal traffic peak arrived that the database, which had already lost a big chunk of its margin, finally couldn’t take it anymore, and the problem exploded all at once. &lt;strong&gt;That’s exactly what makes this kind of “chronic resource drain” issue so nasty: it doesn’t alert immediately. Instead, it suddenly blows up at the moment the remaining headroom is exhausted, making your first reaction “how did this happen with no warning?”—when in reality the warning signs were there the whole time, just masked by the machine’s excess capacity.&lt;/strong&gt; &lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260822182806142.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;image-20260822182806103&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;6-retrospective-question-two-how-did-the-scheduled-jobs-row-lock-avalanche-build-up&quot;&gt;6. Retrospective question two: how did the scheduled job’s row-lock avalanche build up?&lt;/h2&gt;

&lt;p&gt;The second offender I killed during damage control came from a very plain scheduled job: find records stuck in “creating” status for too long without updates, and mark them as failed once they time out. The code explicitly used a Redis distributed lock to prevent multiple instances in the cluster from processing the same batch of data at the same time. (To avoid exposing the exact implementation, field names and variable names have been replaced, but the logic is unchanged.)&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;LOCK_KEY&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;env&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;:item:monitor_creating:lock&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;LOCK_TIMEOUT&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;SCHEDULE_MIN_INTERVAL&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# 300 seconds, same as the minimum scheduling interval
&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;lock&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;redis_client&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;lock&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;LOCK_KEY&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;timeout&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;LOCK_TIMEOUT&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;blocking&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;lock&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;acquire&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# Lock is occupied, skip this round
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;try&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;c1&quot;&gt;# ... execute batch UPDATE ...
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;finally&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;lock&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;release&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;At first glance, all the expected protections seemed to be there. So why did the live incident still end up with nearly 30 concurrent copies of the same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt;, with the longest one stuck for almost 5 hours? There were &lt;strong&gt;two separate paths&lt;/strong&gt; here that could both make the lock expire too early, and the real problem was the combination of both:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;The TTL itself was too short&lt;/strong&gt;: the job ran every 7.5 minutes, but the lock TTL was only 300 seconds (5 minutes). Once the database was already under heavy pressure and a single execution of this &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; got dragged out to several hours, the lock would automatically expire before the job had actually finished—Redis neither knows nor cares whether the previous round is still executing.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;When a container was removed from traffic and restarted, the lock was “gracefully” released early&lt;/strong&gt;: during database overload, lots of containers were marked timed out by health checks because responses got too slow, and were then removed and restarted. If a container holding the lock while executing this &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; got taken down, the cleanup logic in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;finally: lock.release()&lt;/code&gt; would still run, actively releasing the lock—even if the database transaction it was responsible for had not committed yet and was still holding a bunch of row locks. Once the lock was released, a newly started container could immediately acquire it and launch &lt;strong&gt;another round of the same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; with the same conditions&lt;/strong&gt;, competing for locks on &lt;strong&gt;the exact same rows&lt;/strong&gt; with the previous transaction that still hadn’t committed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Both paths lead to the same situation: &lt;strong&gt;the lock looks free, but the underlying transaction is not actually done&lt;/strong&gt;. So new rounds of the job keep pouring in and queueing up, one after another, piling higher and higher. &lt;strong&gt;In this kind of batch-processing scenario, “the lock has been released” should not be treated as equivalent to “it is now safe to start the next round.” The semantics of the distributed lock need to cover whether the transaction it protects has truly ended, not just whether the scheduler layer should launch another invocation.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;7-side-note-i-thought-it-was-an-index-issue-and-then-explain-slapped-me-in-the-face&quot;&gt;7. Side note: I thought it was an index issue, and then EXPLAIN slapped me in the face&lt;/h2&gt;

&lt;p&gt;Since this &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; itself was taking hours to run, my first instinct was of course to suspect a bad or missing index. I checked the table and the size was indeed scary: nearly 800 GB and more than 46 million rows. Then I checked how many rows matched the target status—only a little over a hundred in the whole table. That number actually made me even more convinced it was an index problem. I guessed the existing index didn’t cover the time field, so even after narrowing down to those hundred-odd candidate rows, it still had to fetch rows one by one to evaluate the timeout condition, and on a nearly 800 GB table that random I/O could be very expensive. I was almost ready to recommend adding a composite index on the spot.&lt;/p&gt;

&lt;p&gt;Fortunately, before touching anything, I took one extra step and ran &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;EXPLAIN&lt;/code&gt;. The result showed it was already using an index scan, and the cost was negligible—an existing composite index had already filtered those hundred-odd rows very cleanly. &lt;strong&gt;There was nothing wrong with the execution plan of this statement itself&lt;/strong&gt;. Adding a new index had nothing to do with this incident; my earlier line of reasoning was simply wrong.&lt;/p&gt;

&lt;p&gt;This side note is worth remembering: &lt;strong&gt;when you have “the table is huge” + “very few rows match,” it’s extremely easy to instinctively suspect indexing, but instinct is not a substitute for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;EXPLAIN&lt;/code&gt;.&lt;/strong&gt; Luckily, all I did was run one extra read-only verification step. I didn’t actually go and perform an unnecessary index change on a nearly 800 GB table—which would also have consumed extra I/O. When the system is already under pressure, this kind of “trying to help but making it worse” can cost more than doing nothing.&lt;/p&gt;

&lt;h2 id=&quot;8-retrospective-question-three-where-was-the-real-load-coming-from&quot;&gt;8. Retrospective question three: where was the real load coming from?&lt;/h2&gt;

&lt;p&gt;If you only looked at that batch of lock-queued &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt;s, it still didn’t explain why CPU had hit the ceiling. The reason they were stuck looked more like a consequence of “the whole instance is already under heavy pressure,” not the cause of “these few statements dragged the whole instance down.”&lt;/p&gt;

&lt;p&gt;Only after comparing several load metrics did the picture become clear: active sessions spiked instantly to &lt;strong&gt;2000+&lt;/strong&gt;, and peak load briefly reached &lt;strong&gt;2332&lt;/strong&gt;—while the safe baseline for this 32-core instance under normal conditions was only around 32, meaning it was overloaded by &lt;strong&gt;more than 70x&lt;/strong&gt;. At the same time, total database connections approached &lt;strong&gt;2700&lt;/strong&gt;. The real load hog was a very unremarkable count query on another table:&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;SELECT&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;FROM&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;t_item_result&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;WHERE&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;game_id&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;err&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;AND&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;public&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;IS&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;true&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This query backed a leaderboard-style endpoint. The existing index on the underlying table did not have &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;game_id&lt;/code&gt; as its leading column, so filtering by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;game_id&lt;/code&gt; couldn’t use it effectively. Under normal low traffic, that was tolerable. But when the day’s traffic peak hit, hundreds of concurrent requests all landed on it at once, and every single one became a full table scan, instantly burning CPU to the ground. At the time, there were nearly 500 queries that had been running for more than 60 seconds. Only about one-third of them were related to the earlier &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;version&lt;/code&gt; table; &lt;strong&gt;the remaining two-thirds were innocent queries dragged down by this environment&lt;/strong&gt;—which also confirmed that the main load source was not the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; path. There was also another comment-related table where a join field lacked supporting indexes, which added fuel to the fire too, though not at the same scale as this count query.&lt;/p&gt;

&lt;h2 id=&quot;9-making-things-worse-even-the-last-self-healing-escape-route-was-blocked&quot;&gt;9. Making things worse: even the last self-healing escape route was blocked&lt;/h2&gt;

&lt;p&gt;In theory, when the database hits this kind of bloat and backlog, there is still one last line of defense: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt;, which can slowly clean things up and let the situation recover on its own. But when I checked the live state, even that escape route had been blocked. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt; process responsible for cleaning a large TOAST storage area had been stuck for &lt;strong&gt;3 hours and 07 minutes&lt;/strong&gt; without finishing. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt; processes on two other core tables had also been dragged down by throttling for &lt;strong&gt;12 minutes 49 seconds&lt;/strong&gt; and &lt;strong&gt;6 minutes 01 seconds&lt;/strong&gt; respectively, unable to make normal progress.&lt;/p&gt;

&lt;p&gt;The fact that three &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt; processes were each stuck on different tables showed this was not a localized issue on one table. The whole instance was already so resource-starved that even background cleanup tasks couldn’t get enough CPU time slices to run. &lt;strong&gt;At that point, self-healing was completely impossible. Manual intervention was the only option&lt;/strong&gt;—which was exactly the step described earlier in section 4.&lt;/p&gt;

&lt;h2 id=&quot;10-putting-the-full-chain-together&quot;&gt;10. Putting the full chain together&lt;/h2&gt;

&lt;p&gt;If you connect the answers to the previous questions, the escalation path looks roughly like this:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Two and a half days earlier, an analytical query with parallelism started running and never exited. It not only consumed resources itself, but because its transaction stayed open for so long, it pinned the global MVCC snapshot horizon, preventing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt; from cleaning dead tuples normally, so core tables kept bloating.&lt;/li&gt;
  &lt;li&gt;The database had enough capacity to brute-force through more than 50 hours without obvious symptoms, until a routine traffic peak arrived. A hot count query missing the right index got amplified by concurrency into a large number of full table scans, and CPU was instantly burned through.&lt;/li&gt;
  &lt;li&gt;In that CPU-overloaded environment, a scheduled job that was originally protected by a distributed lock started queueing up round after round of the same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt;, because the TTL was too short and the lock was also being released early when containers were removed and restarted. It piled up to nearly 30 sessions.&lt;/li&gt;
  &lt;li&gt;Nowhere in the whole chain was there any statement-level timeout configured. So a large number of requests could only wait indefinitely, until the containers serving them were themselves marked timed out by health checks and removed/restarted, at which point the coroutines were passively canceled. Containers were replaced in waves—two waves in total, adding up to more than a dozen containers being marked unhealthy and restarted due to probe timeouts—which in turn fed step 3: newly started containers immediately grabbed the just-released lock and launched another round of conflicting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt;s, amplifying the loop.&lt;/li&gt;
  &lt;li&gt;By then, the self-healing mechanism (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;autovacuum&lt;/code&gt;) had already been dragged to a standstill, so only manual intervention could stop the bleeding.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;With all five stacked together, the database and every service depending on it slowed down across the board. From the business side, the direct symptom was simply “all endpoints are timing out.” And once those two root sources were killed, the chain snapped immediately and all metrics quickly returned to normal.&lt;/p&gt;

&lt;h2 id=&quot;11-the-real-holes-that-need-patching&quot;&gt;11. The real holes that need patching&lt;/h2&gt;

&lt;p&gt;What we did on-site was only first aid. It solved none of the root causes. After sorting it all out, there are quite a few things to fix, roughly in this priority order:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Add statement-level timeouts to database connections&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I checked, and right now there is no statement-level timeout configured at all on the database connection layer (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;statement_timeout&lt;/code&gt; in PostgreSQL). That means how long a SQL statement can stay stuck in the database depends entirely on which outer layer times out first. Leaving this unset is basically giving every slow query an unlimited line of credit.&lt;/p&gt;

&lt;p&gt;The setup is straightforward. PostgreSQL supports several levels, so you can choose based on need:&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;-- Option 1: session level, run once after connection is established; only affects the current connection&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;SET&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;statement_timeout&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;15s&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;-- Option 2: per database role; automatically applies to all future connections for this user&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;ALTER&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;ROLE&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;app_user&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;SET&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;statement_timeout&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;15s&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;-- Option 3: instance-level default (configured in postgresql.conf or a cloud provider parameter group)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;statement_timeout&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;15000&lt;/span&gt;   &lt;span class=&quot;c1&quot;&gt;-- unit is milliseconds&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If you’re using a connection pool / ORM, it’s even better to pass it directly when creating connections, for example with Python &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;asyncpg&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nf&quot;&gt;create_async_engine&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;DATABASE_URL&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;connect_args&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;server_settings&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;statement_timeout&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;15000&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}},&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# milliseconds
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The exact timeout value can’t just be guessed—you need to tier it by endpoint type. User-facing core read endpoints usually get single-digit to low double-digit seconds; background batch jobs and offline tasks can be looser, maybe tens of seconds. But &lt;strong&gt;no matter which tier, it must be an explicit number, not “unset.”&lt;/strong&gt; In this incident, the longest batch of queries was stuck for nearly 5 hours. Even if we had set a very relaxed timeout like 60 seconds, it would never have dragged on to the point where manual intervention was required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Change the scheduled job to commit in batches; the distributed lock cannot rely on a fixed TTL alone&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The table is still growing, and the one-shot large-batch &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UPDATE&lt;/code&gt; approach itself does not age well as data volume increases. It should be split into smaller batches with incremental commits. On the distributed lock side, having a TTL that cannot cover “how long the job might take in the worst case” is itself a hidden risk, and the lock release timing needs to be redesigned too—service restarts should not automatically mean “it is safe to release the lock and wake up the next round.” Either add a renewal mechanism, or tie the lock lifecycle much more tightly to the transaction state it is protecting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Add the two indexes that are actually missing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Add an index on the table hit by the hot count query that actually covers the real query conditions, using a non-blocking index build. Also add the corresponding index for the field used by the comment join query. This incident was also a reminder to myself: during troubleshooting, don’t decide by intuition which table “should get an index.” Use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;EXPLAIN&lt;/code&gt; or historical load analysis tools to find the real load source first, then make changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Add caching or periodic pre-aggregation for the hot count endpoint&lt;/strong&gt;, instead of doing a real-time &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;count(*)&lt;/code&gt; every time, to reduce direct pressure on the underlying large table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Add a read replica dedicated to offline analysis&lt;/strong&gt;: even just having one read-only instance that does not serve online business traffic, dedicated to offline analysis and data verification queries, would avoid situations like “someone forgets to close an analysis script connection and drags down the production primary database.” Letting the primary be used everywhere for ad hoc analysis creates too much risk exposure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Fill in the missing monitoring and alerts&lt;/strong&gt;: core database load metrics, slow SQL distribution, and so on. This incident also confirmed one thing: tools like Performance Insights, which let you inspect historical Top SQL and wait events, do cost a bit, but the difference between being able to locate the root cause in a few minutes during an incident versus not being able to is huge. For core databases, it’s worth enabling them by default.&lt;/p&gt;

&lt;h2 id=&quot;summary&quot;&gt;Summary&lt;/h2&gt;

&lt;p&gt;My biggest takeaway from this investigation is: &lt;strong&gt;production incidents are rarely caused by a single factor. Most of the time, they’re several small issues that each look non-fatal on their own, but happen to line up at the same moment.&lt;/strong&gt; A forgotten analytical script, a hot endpoint missing an index, a distributed lock whose TTL was too short and then got released at exactly the wrong time, plus no statement-level timeout anywhere in the chain—looked at separately, each one is the kind of oversight people can easily forgive. Stack them together, and you get a very real avalanche. And because the machine itself had strong enough performance, the problem stayed latent for more than 50 hours before being triggered by a traffic peak. That kind of “delayed explosion” is even more dangerous than “immediate alerting.”&lt;/p&gt;

&lt;p&gt;Another more concrete lesson is this: &lt;strong&gt;during troubleshooting, even conclusions that “look very reasonable” are worth validating with tools before you act.&lt;/strong&gt; At one point I was convinced it was an index issue. If I hadn’t run that extra &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;EXPLAIN&lt;/code&gt;, I might have made a completely unnecessary index change on a nearly 800 GB table and consumed even more I/O in the process. Investigating a blocking chain can tell you who is waiting on whom, but identifying the real load source always comes back to one question: “who is actually burning CPU?” The two are not interchangeable. And when an incident happens, stopping the bleeding and giving the system room to breathe is always more important than trying to analyze everything slowly while still under pressure.&lt;/p&gt;
</description>
        <pubDate>Tue, 18 Aug 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/middleware/2026/08/18/postgres-cascade-failure.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/middleware/2026/08/18/postgres-cascade-failure.html</guid>
        
        <category>en</category>
        
        <category>middleware</category>
        
      </item>
    
      <item>
        <title>I Built an IntelliJ IDEA Plugin That Flattens All the REST APIs in a Project into a Single Panel</title>
        <description>&lt;h2 id=&quot;preface&quot;&gt;Preface&lt;/h2&gt;

&lt;p&gt;Back when I was writing in other languages, I’d already used this kind of “API browsing” plugin, and I always found it pretty handy. Once a project gets big, you &lt;em&gt;really&lt;/em&gt; end up with a ton of endpoints. Old and new APIs get mixed together, and trying to dig them out of the codebase purely by memory is basically impossible—especially in a huge monorepo. You might vaguely remember “this one is user-related,” but what it’s called and which class it lives in? Total guesswork, and it’s hard to pinpoint precisely.&lt;/p&gt;

&lt;p&gt;The IDE’s built-in global search doesn’t help that much either, because the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/&lt;/code&gt; in an endpoint path is easily treated as a separator, so you can’t do a continuous match on the whole string. The most typical scenario: you copy a full path straight from Chrome DevTools’ Network panel, like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/api/v1/user/detail/list&lt;/code&gt;, paste it into global search, and… nothing. You’re forced to split the path and try word by word, which is super annoying. This is especially deadly for backend folks—an API endpoint is the entry point of the entire data chain. From there you trace down through Service, Repository, and finally to the DB. You’re following that &lt;em&gt;full&lt;/em&gt; chain; if you can’t even locate the entry point, everything after that is pointless.&lt;/p&gt;

&lt;p&gt;What really pushed me to write my own was a few other things. First: as IDEA keeps updating, the plugins I used to rely on gradually became incompatible. Every so often I had to go hunt for replacements, and that hassle alone is irritating. Second: while looking for alternatives, I tried a bunch of similar plugins and found they were stuffed with bloated features I didn’t need—not that those features are bad; I get and respect other people’s product thinking. It’s just not what I want. Third, and most directly: lately I’ve been writing more and more Python. I searched the market for this category of plugin and almost all of them only recognize Java-framework Controller annotations. FastAPI? Nobody cares. But the amount of backend APIs written in Python is absolutely not less than Java.&lt;/p&gt;

&lt;p&gt;So I just wrote one myself: cover both the Java ecosystem frameworks and FastAPI first, and stick as hard as possible to a single responsibility—“API browsing + navigation”—without piling on features I don’t use. If you need other languages/frameworks, feel free to open a PR. Let’s test together and fill in this puzzle piece by piece.&lt;/p&gt;

&lt;h2 id=&quot;main&quot;&gt;Main&lt;/h2&gt;

&lt;h3 id=&quot;first-draw-a-clear-line-what-this-version-does-and-doesnt-do&quot;&gt;First, draw a clear line: what this version does and doesn’t do&lt;/h3&gt;

&lt;p&gt;I took a look at RestfulBox screenshots and realized it’s a bit more ambitious: it even has panels for sending requests and viewing responses—kind of like embedding Postman into the IDE. I’m not planning to follow that path—first, that feature set isn’t trivial in complexity; second, there are already dedicated tools for sending requests, so there’s no need to reinvent it inside an IDE plugin. It also aligns with the “don’t be bloated” idea I mentioned earlier. This version is going all-in on one scenario: &lt;strong&gt;browse + navigate&lt;/strong&gt;. Scan all endpoints, select one, jump straight to the code. That’s enough.&lt;/p&gt;

&lt;h3 id=&quot;architecture-one-extension-point-to-cover-four-frameworks&quot;&gt;Architecture: one extension point to cover four frameworks&lt;/h3&gt;

&lt;p&gt;The core is an extracted extension point interface, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ApiParserContributor&lt;/code&gt;, with one implementation per framework:&lt;/p&gt;

&lt;div class=&quot;language-kotlin highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kd&quot;&gt;interface&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;ApiParserContributor&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;fun&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;collect&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;project&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;Project&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;scope&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;GlobalSearchScope&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;List&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;ApiEndpoint&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;&amp;gt;&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;For the annotation recognition logic of Java frameworks like Spring, JAX-RS, and Micronaut, I didn’t reinvent the wheel—this is directly &lt;strong&gt;ported and adapted from RestfulHelper&lt;/strong&gt; (MIT-licensed open source), including its robust value-extraction logic for handling string concatenation and constant references. After the adaptation, everything is unified to fit this project’s own &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ApiEndpoint&lt;/code&gt; model, and the license and attribution are properly preserved in NOTICE. The FastAPI part is newly written, based on Python PSI to parse path composition across &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;@app.get(...)&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;APIRouter(prefix=...)&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;include_router&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For Python support I used a small trick: in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plugin.xml&lt;/code&gt;, declare an optional dependency with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;depends optional=&quot;true&quot;&amp;gt;com.intellij.modules.python&amp;lt;/depends&amp;gt;&lt;/code&gt;. That way, users who don’t have JetBrains’ official Python plugin installed can still install this plugin without errors—they just won’t see FastAPI endpoints. Everything else works as usual.&lt;/p&gt;

&lt;p&gt;The UI layer doesn’t care which framework an endpoint comes from; it only consumes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;List&amp;lt;ApiEndpoint&amp;gt;&lt;/code&gt;. In the tool window, you can switch grouping across four dimensions: path prefix, method, file/class, and module. There’s also a search box for real-time filtering.&lt;/p&gt;

&lt;h3 id=&quot;one-pitfall-full-scans-blocking-the-ui-thread&quot;&gt;One pitfall: full scans blocking the UI thread&lt;/h3&gt;

&lt;p&gt;During development I ran into a pretty classic IDEA plugin pitfall: at first, every keystroke in the filter box would synchronously call all parser contributors to rescan the entire project, and then wait for results on the EDT (the UI main thread). On small projects you don’t notice it, but once the number of endpoints grows, the UI freezes on input and the experience is awful. Fundamentally, I was doing heavy index access on the wrong thread.&lt;/p&gt;

&lt;p&gt;Later I switched to non-blocking async refresh: scan results are cached in a project-level service; when PSI changes, only the affected files are rescanned incrementally instead of doing a full rescan; searching and switching groupings operate only on an in-memory snapshot of the cache and no longer touch the index. Along the way I also found and fixed a related issue—when refresh failed or got canceled, the cache state used to be incorrectly marked as “clean,” which caused it to &lt;em&gt;not&lt;/em&gt; rescan when it actually should have.&lt;/p&gt;

&lt;h3 id=&quot;global-navigation-and-the-settings-page&quot;&gt;Global navigation and the settings page&lt;/h3&gt;

&lt;p&gt;Besides the tree list in the tool window, I also added a global hotkey (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Cmd+Option+\&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Ctrl+\&lt;/code&gt;) to bring up a fuzzy-search popup. Type a path fragment or method name and you can locate and jump directly—no need to manually open the tool window first and then hunt around.&lt;/p&gt;

&lt;p&gt;There’s a small pitfall here too: across different OSes and user-customized keymaps, the actual effective shortcut can vary. If you hardcode a hint string, you’ll mislead people. So I built a dedicated settings page (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Settings → Tools → RestfulController&lt;/code&gt;) that reads the currently active keymap and displays the shortcut the user can actually press, rather than a hardcoded default. On that page you can also change the shortcut directly, and adjust whether the popup appears centered on the screen or follows the mouse. After changes, it writes back into the IDE’s own keymap—the same data you see under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Settings → Keymap&lt;/code&gt;—so there’s no “plugin stores its own copy and gets out of sync with the system” situation.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260730224122359.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1785469273949&quot; /&gt;&lt;/p&gt;
&lt;center&gt;The settings page reads the currently active Keymap and shows the shortcut the user can actually press; you can also adjust the popup position here&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260730224437712.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1785469467685&quot; /&gt;&lt;/p&gt;
&lt;center&gt;All endpoints in the tool window are grouped by the selected dimension, with real-time search filtering&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://mypicgogo.oss-cn-hangzhou.aliyuncs.com/tuchuang20260730224506053.png?x-oss-process=image/auto-orient,1/resize,w_1200,limit_0/format,webp/quality,Q_80&quot; alt=&quot;QQ_1785469500004&quot; /&gt;&lt;/p&gt;
&lt;center&gt;Fuzzy search triggered by the global shortcut—enter a path fragment and press Enter to jump to the corresponding handler method&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h3 id=&quot;development-story-i-kept-writing-then-someone-else-took-over-and-kept-writing&quot;&gt;Development story: I kept writing, then “someone else” took over and kept writing&lt;/h3&gt;

&lt;p&gt;This plugin was completed entirely with AI assistance, and in the middle I even swapped “people” a few times in a relay. At the very beginning I used another AI coding tool to kick things off, set the design, and build the first implementation. Halfway through, that tool’s session context was basically used up, so I dumped the handoff chat logs to Claude and had it continue. Claude pushed it forward a lot—the main features like the four framework parsers, the tool window, and global search were largely done in that phase. Then, as I kept going, Claude’s token budget expired too, so I switched to Codex to finish it off: the settings page, async refresh, shortcuts, and other UX details. I also did a round of performance and architecture self-check (the EDT-freeze pitfall above was found during this round).&lt;/p&gt;

&lt;p&gt;At first I worried that having multiple different AIs take turns writing the same project would produce stylistically fragmented code. In practice it wasn’t a big problem—as long as the design docs and implementation plan are clear enough, and you hand those docs over during the transition, the new AI can align context quickly and keep moving. The human in the loop mainly steers direction, runs tests, and catches the obviously wrong bits.&lt;/p&gt;

&lt;h2 id=&quot;afterword&quot;&gt;Afterword&lt;/h2&gt;

&lt;p&gt;The plugin is now up and running, and the code is open-sourced at &lt;a href=&quot;https://github.com/oreoft/restful-controller&quot;&gt;github.com/oreoft/restful-controller&lt;/a&gt;. Overall, my takeaway is: for a multi-framework IDE utility like this, the hard part isn’t the parsing details of any single framework (a lot of that can be learned from open source), but designing the extension point cleanly enough—adding a new framework shouldn’t require touching the UI or other parsers. This time, I think I nailed that.&lt;/p&gt;

&lt;p&gt;Going forward I’ll probably keep polishing it based on the little annoyances I hit in real use. For now I don’t plan to expand toward “sending requests”—for me, browsing plus navigation already solves the most painful problem. As for support for other languages: same line as before, I’m waiting for your PR.&lt;/p&gt;
</description>
        <pubDate>Thu, 30 Jul 2026 00:00:00 +0800</pubDate>
        <link>https://www.someget.cn/en/tools/2026/07/30/restful-controller-idea-plugin.html</link>
        <guid isPermaLink="true">https://www.someget.cn/en/tools/2026/07/30/restful-controller-idea-plugin.html</guid>
        
        <category>en</category>
        
        <category>tools</category>
        
      </item>
    
  </channel>
</rss>
