We had 30+ Telegram bots, and each one was its own deployment: its own container, its own Python runtime, its own copy of every library, its own liveness probes to babysit. Most of them were idle most of the time — a bot that answers a few messages an hour still holds its full interpreter and dependency tree in memory around the clock.
That zoo cost us roughly 15 GB of idle memory and a long tail of operational noise: 30 deployments to upgrade, 30 sets of alerts, 30 places for config to drift.
The fix was structural, not heroic: one service that speaks for all of them.
- Every bot’s webhook points at the same deployment; requests are routed by bot token to the right handler.
- Handlers are plain Python modules registered in one place — adding a bot became writing a handler, not provisioning infrastructure.
- The service autoscales on Kubernetes by actual load. Thirty idle bots cost what one idle process costs; a traffic spike on any bot scales the shared pool, not a single starved pod.
The ~15 GB came back immediately, but the quieter win was operational: one deployment to upgrade, one dashboard to watch, one incident surface. Rollouts that used to mean touching 30 manifests became a single deploy.
The general lesson: when N things are identical except for their config, N deployments is not isolation — it’s just N copies of the same overhead. Isolation belongs at the routing layer, not the process boundary, until load proves otherwise.