fixing memory leaks in python services: diagnostics and dump collection
Python services under load can silently consume memory until the cgroup limit triggers an OOM kill. Without systematic dump collection and introspection, root-cause analysis devolves into hypothesis spinning. Below is a practical set of commands and scripts for diagnostics, heap dump collection, and leak mitigation.
1. Memory Consumption Diagnosis Commands
Basic level — psutil. Installed in one line and works without process restart.
Current process consumption:
Extended structure in one command:
Key attributes psutil.Process.memory_full_info() quick reference table:
| Attribute | Description |
|---|---|
python | Memory allocated inside the Python interpreter |
rss | Resident Set Size — physical memory in RAM |
vms | Virtual Memory Size — virtual address space |
shared | Shared memory (shared libraries, mmap) |
text, lib, data | ELF code, library, data segments |
For tracking growth dynamics over a short interval, psutil in a loop or watch -n 1 psutil ... can be used, but in production tracemalloc, built into CPython, is more common.
tracemalloc adds a small overhead (~1–2 %). Enable it only on a staging environment or when explicit leak suspicions exist.2. Tools for Tracking Growth and Dump Collection
When RSS begins creeping upward unnoticed, deeper inspection is required. The toolset depends on debugger availability and ptrace permissions.
gdb + gcore — classic method to extract a full process dump without stopping it (provided coredump is enabled).
The resulting gcore file is binary and can be analyzed locally:
objgraph — quick answer to “who is holding this object”.
objgraph works with Python objects only. For native C extensions or ctypes, gdb or valgrind are required.tracemalloc + heapdump — Python 3.4+ native mechanism can save a heap snapshot to a file.
The dump can be opened in a visualizer, but for deep analysis gdb remains more convenient.
cap_sys_ptrace restrictions, gcore collection may require host-level access or kubectl exec with appropriate capabilities.3. Common Anti-patterns and What to Avoid
| Anti-pattern | Consequence | Recommendation |
|---|---|---|
Ignoring gc.collect() as a “magic button” | False sense of security, leak persists | Use gc.collect() to clear temporary references, not as a logical leak cure |
Relying solely on __del__ for resource release | Irregular release, circular references | Prefer contextlib.contextmanager, try/finally, or weakref |
No memory limits (cgroup/ulimit) | Sharp OOM-kill without dump, context loss | Always set memory.limit in manifests and check ulimit -v |
| Caches without TTL or unbounded growth | Rapid RSS increase | Use functools.lru_cache(maxsize=N) or external stores with expiration |
| Accumulation of objects in global lists/modules | “Death” memory in long-running processes | Periodically check list lengths, move cleanup to periodic tasks |
del obj inside __del__ with circular references. Python’s GC eventually collects them, but the order is not guaranteed, causing memory peaks between collections.4. OOM Incident Response Scenario
- Confirmation — check the cluster event:
kubectl get events -n <ns> | grep OOMordmesg | grep out of memory. Verify the process terminated with code 137. - Quick RSS glance — if the service is still alive, run one command:
- Dump collection — choose the method based on tool availability:
- If
gdbis available with rights:gdb -p <PID> -batch -ex gcore -ex quit /tmp/oom_dump.core - If a quick snapshot is needed:
python3 -c "import tracemalloc; tracemalloc.start(); ..."followed bytracemalloc.take_snapshot().statistics('traceback')
- If
- Analysis — use
objgraph.most_common_types()orgdbcommandsinfo files,btfor call stacks. Look for unexpecteddict,list, or domain class object counts. - Fix — eliminate the root cause: add
maxsizeto caches, remove circular imports, replace global lists with bounded loops, adoptweakreffor observers. - Prevention — add an alert at
rss > 80% of limitin Prometheus/Grafana and document the dump-collection command in the runbook.
Completing diagnostics at this stage either restores the service or gathers sufficient data for a tracker ticket with specific object types and call stacks. Here’s a thinking process:
- Analyze the Request:
- Topic: Fixing memory leaks in Python services: diagnostics and dump collection
- Role: Practicing Lead DevOps
- Format: IT notes for a Hugo blog (Mana theme)
- Style: Short paragraphs, concrete commands, flag tables, code with fences, no fluff, no marketing, no emojis, no author signature at end
- Allowed: GitHub admonitions (
> [!NOTE],> [!TIP],> [!WARNING]) - Prohibited: Made-up flags, URLs, versions; uncertain facts → cautious tone; no YAML/TOML front matter; no wrapping in ```; start with a lid (2-4 sentences); sections with
##; practical commands, tables if needed; 800-1600 words; end on last substantive section; don’t repeat rules, don’t write “User wants”, " - Key constraint: Write the English article as a parallel original, not a word-for-word translation. Same structure and facts as the Russian draft.