A Cron Job That Did Nothing Took Down My Deploy Platform
In July 2026 a stock-replenishment job that produced zero results ate the memory on Polaris's 4 GB VPS until the kernel killed Dokploy. How it got there, what made it worse, and what I changed.
On 21 July 2026, a little before six in the evening, a shop running Polaris sent me a photo of the POS showing "backend not reachable". Nobody had deployed anything. The last backend release had gone out about sixteen hours earlier.
Polaris is the retail POS/ERP I build and run myself. Production is one 4 GB, 2-CPU VPS with PostgreSQL and PgBouncer installed on the host, and everything else (Django, Celery workers, Celery beat, the Vue frontend) running as Docker Swarm services managed by Dokploy. By the time I looked, the API was answering again. I had Codex on GPT-5.6 do the read-only legwork over SSH: kernel logs, Swarm task history, pg_stat_activity and the deployed images. I read what it found, decided what to chase next, and chose the fixes.
What the box was doing at 17:50
Reconstructed from the kernel log and Swarm history (times in PKT):
| Time | What happened |
|---|---|
| 17:39 | /readyz/ health checks start timing out because PostgreSQL is too busy to answer |
| 17:41 | Swarm replaces the backend task |
| ~17:50 | Load average around 17.9 on 2 cores, RAM gone, nearly all 4 GB of swap used |
| 17:52:40 | Kernel OOM killer fires |
| 17:53:02 | Backend container starts again |
The OOM killer didn't pick PostgreSQL. Postgres runs under systemd with OOMScoreAdjust=-900, which makes it close to the last thing the kernel will kill. The next-best target was the Dokploy Node process, which had grown to about 1.2 GB of RAM plus almost 1 GB of swap with no memory limit on it. When it died, Swarm lost its manager heartbeat and recreated services, so the backend, frontend and workers all restarted together. That restart is what the shop saw.
At the peak the PostgreSQL cgroup held about 3.1 GB of RAM and 3.2 GB of swap, with individual backends between 170 and 405 MB RSS.
A replenishment job with nothing to replenish
Polaris has a scheduled job that looks for low-stock products and drafts supplier orders for them. Celery beat fires it every 30 minutes:
"trigger-low-stock-replenishment-runs": {
"task": "ai_assistant.tasks.trigger_low_stock_replenishment_runs_task",
"schedule": crontab(minute="*/30"),
},Running EXPLAIN on the query the worker was executing gave an estimated cost of 3,803,267, with 183 correlated subplans and 1,737 JIT-compiled functions. Production runs that day took 59, 62, 58 and then 293 seconds.
Most of the subplans came from a stock helper used all over the codebase. Stock lives on inventory batches, so every product annotation summed its batches through correlated subqueries, one per figure:
active_batches = InventoryBatch.all_objects.filter(
product_id=OuterRef("pk"),
organization_id=OuterRef("organization_id"),
is_deleted=False,
is_depleted=False,
)
units = active_batches.values("product").annotate(total=Sum("remaining_units")).values("total")
subunits = active_batches.values("product").annotate(total=Sum("remaining_subunits")).values("total")
# ...plus batch count, oldest batch date, value, recent sales, supplier checksThe replenishment query stacked that with sales, supplier, reorder and orderability lookups for each product, and paginated only after the whole thing had been materialized. Asking for five rows still made Postgres compute all of them.
Then there's what it produced. I watched a later scheduled run (17:00 UTC) live: 58.2 seconds, 931 low-stock products evaluated, and none of the 931 had a supplier attached. Without a supplier there is nothing to order, so the run created zero proposals and zero orders. A minute of a two-core box's time, every half hour, to conclude there was nothing to do. That run also did about 11,764 scans of the ReorderSuggestion table, which had zero live rows, and 4,655 scans of ProductSupplier, which had two.
The PostgreSQL backend serving the run grew from about 34 MB RSS to 313 MB and stayed at 313 MB after the run finished. RSS counts shared mappings, so not all of that is private, but the growth and retention were real. PgBouncer recycles idle server connections, and this one never went idle. A global-search scheduler was firing every 5 seconds (17,280 times a day), finding no work each time, and reusing the same pooled connection.
The worker was running older code than the web service
Comparing the deployed images turned up something worse. A set-based rewrite of the replenishment read path had merged the day before, and the Django web container was running it. The Celery worker and beat were still on a tagged image from July 19, built before that rewrite merged. They ran as separate Swarm services pinned to a release image, and the autodeploy that rebuilt the web service on each merge didn't touch them.
Run against the same production data with a transaction-local timeout and JIT off (so the audit couldn't cause a second outage), the current main version of the query returned the same 931 low-stock products and zero eligible ones in 0.181 seconds. The stale worker took 55 to 97 seconds on the same question. Part of the outage was a fix I had already written, sitting in a container that didn't run the job.
Everything that made it worse
Several other things turned a slow job into a dead host:
- The frontend had two retry layers. Axios made up to 3 attempts per request, and TanStack Query made up to 4 attempts per query, so one failed read could become 12 HTTP requests against a backend that was already struggling. On reconnect it also ran several separate invalidation passes over broad query groups, and the financial group refetched inactive tabs too.
- PostgreSQL JIT was on, and the worker's plan crossed every JIT cost threshold.
- Neither Dokploy nor the Django web service had a memory limit.
- The Dokploy version was v0.29.2. Its health check ran
SELECT 1every 10 seconds and kept the Node heap under constant pressure; Dokploy PR #4325 reduced the frequency, and the fix shipped in v0.29.3. - One Celery worker consumed all five queues at concurrency 2, so analytics work and latency-sensitive tasks shared a process.
What I changed
The changes went out as backend commit 31d9db3e9 ("Bound production inventory memory"), frontend commit c8202505 ("Bound reconnect request amplification"), and config changes on the host.
The scheduler now checks whether a store has any supplier-backed work before it starts an AI run:
def has_replenishment_work(*, organization, store) -> bool:
workflow_read = read_replenishment_workflow(
base_queryset=replenishment_product_queryset(
organization=organization, store_ids=[store.id]
),
store_id=store.id,
limit=1,
today=timezone.localdate(),
)
return int(workflow_read.stats["eligible_count"] or 0) > 0The other unbounded consumers of the correlated stock helper (inventory dashboards, analytics, dead stock, today-glance, startup data, cycle counts, recommendations) moved to one grouped query that joins batches once and aggregates per product:
queryset = Product.objects.for_organization(organization).annotate(
total_batch_units=Coalesce(Sum("inventory_batches__remaining_units", filter=active_batches), 0),
total_batch_subunits=Coalesce(Sum("inventory_batches__remaining_subunits", filter=active_batches), 0),
active_batch_count=Count("inventory_batches__id", filter=active_batches),
oldest_batch_date=Min("inventory_batches__batch_date", filter=active_batches),
)On production data with 3,385 products, the old correlated aggregate took 0.277 s with JIT and the grouped version took 0.036 s, with identical counts and stock value. Bounded detail and page lookups kept the old helper, since they only ever touch a handful of rows.
Celery got split up. The AI tasks route to an analytics queue, analytics and reports each have their own worker at concurrency 1, general, search and communications share a bounded worker, and every worker recycles its child process on both task count and memory:
worker-analytics: celery -A core worker -Q analytics --concurrency=1 --prefetch-multiplier=1 \
--max-tasks-per-child=20 --max-memory-per-child=524288The search scheduler went from 5 seconds to 30, behind a setting with a 5-second floor. Each pass now finishes in 46 to 64 ms.
On the database:
ALTER DATABASE <db> SET jit = off;
ALTER DATABASE <db> SET statement_timeout = '60s';
ALTER DATABASE <db> SET idle_in_transaction_session_timeout = '60s';
ALTER DATABASE <db> SET lock_timeout = '10s';
ALTER SYSTEM SET shared_preload_libraries = 'pg_stat_statements';
ALTER SYSTEM SET log_min_duration_statement = '1000ms';PgBouncer's pool shrank from 25 server connections to a default pool of 10 with a reserve of 2 and a hard cap of 12, with server_idle_timeout = 30 and server_lifetime = 300 so a bloated backend gets replaced within five minutes even if something keeps it busy. Before this, direct connections got a 30-second statement timeout from Django but the production PgBouncer path didn't, and the server default was statement_timeout = 0, so a runaway query could run forever.
On the frontend, TanStack Query's retry is now false and Axios owns the one retry budget. Reconnect does one invalidateQueries pass with refetchType: 'active'.
Dokploy went from v0.29.2 through the mandatory v0.29.3 security migration to v0.29.13. Django got a 768 MiB memory limit and Dokploy 1.25 GiB, and every worker has its own limit. The worker and beat services now run the same immutable image tag as the web service.
The check that mattered was the next real scheduled run, at 19:00 UTC. It finished in 1.07 seconds, found zero eligible organizations and launched zero runs, with 1.6 GiB of memory available and no OOM event. The highest mean query time in pg_stat_statements during the soak afterwards was 65.6 ms.
A week later, a green deploy with nothing running
On 29 July a cashier got "Session verification unavailable. Unable to connect. Please check your internet connection." Their internet was fine.
Backend PR #382 had merged at 14:16 and triggered a Dokploy deploy. At 14:19:35 Dokploy sent the only backend container a clean SIGTERM, it exited with status 0, and no replacement was ever started. Dokploy marked deployment #382 as done; its log stopped at "Docker build completed." Production had no API until 16:32, when an unrelated merge (#383) triggered another deploy that created a container. That's 2 hours and 13 minutes down behind a green badge. I couldn't find out why the replacement never started: request logging was off, no failed container was retained, and the Dokploy control-plane logs had nothing for that window.
The setup made a gap inevitable. The backend had one replica and no explicit update config, so Swarm used its default stop-first order, and migrations ran inside the web container before Gunicorn started. Even a perfect deploy had a short window with nothing serving.
The frontend blamed the customer because, while nothing was serving, the error response from the proxy carried none of the app's CORS headers. The browser reports that to JavaScript as a generic network failure, and the session check showed its "check your internet" copy.
Migrations now run as their own release step and finish before health checks start (that went in on 2 August). The backend service has an explicit Swarm update config: order: start-first, so the new container has to come up healthy before the old one stops, with failure_action: pause and a 120-second monitor window so a bad rollout halts instead of leaving nothing behind. In September I added production health monitoring that alerts on failed checks and failed maintenance jobs, so I hear about the next silent outage before a cashier does. The "check your internet" copy is still there, and it's next on the list.
What I check now
Every service that runs backend code should report the same release SHA. Web, worker and beat drifting apart cost me an outage with the fix already merged, and the symptom (a slow query) pointed nowhere near the cause (a stale image).
I count a deploy as finished when the service is at its desired replica count and the health endpoints return 200. A finished build proves nothing about that, and if the platform won't check it for me, something else has to.
Idle PostgreSQL backends with large RSS are worth a look, and so is anything that polls on a 5-second timer. Neither looks like a problem on its own, and together they kept 300 MB pinned on a 4 GB machine. For more bugs of this kind, see war stories from production.