AI Coding Agents
Reliability
Memory leak tests for agent-generated services
September 28, 2026 - 2 minute read
AI Coding Agents
Reliability
September 28, 2026 - 2 minute read
Memory leak testing measures whether a service retains memory after repeated work should have released it. Agent-generated code can introduce long-lived listeners, timers, caches, closures, or request objects while still passing functional tests. A repeatable workload and an explicit memory budget make that risk visible before a change reaches production.
Heap snapshots are useful evidence, but they can pause a process and may require about twice the current heap while being created. The official Node.js heap snapshot guidance recommends care in production because snapshot creation can block the event loop or exhaust available memory. Capture them in an isolated test process whenever possible.
Choose one bounded workload that exercises the changed path. Warm the service first so module loading, connection setup, and lazy caches do not look like leaks. Run the workload in several batches, release references, allow garbage collection under controlled test settings, and record memory after each batch.
Resident set size can reveal process growth, while heap measurements narrow the problem to managed memory. Track both when the runtime exposes them. A fixed threshold alone can be brittle across machines, so also compare the slope after warm-up. Continued growth across equivalent batches is stronger evidence than one high sample.
Keep the environment stable. Use fixed fixture sizes, bounded concurrency, and local dependencies. Disable unrelated background jobs that add noise.
When the test detects growth, compare heap snapshots from after two equivalent workload batches. Look for object types whose retained count or size rises and never returns toward the baseline. Trace retaining paths back to changed code rather than treating the largest object as the cause.
Common causes include event listeners without cleanup, maps keyed by request data, promises retained by queues, unbounded telemetry attributes, and response bodies held after parsing. Add a regression test around the smallest behavior that reproduces the retention.
Native allocations and subprocesses may increase resident memory without appearing in the JavaScript heap. Match the diagnostic tool to the runtime and allocation source.
Provide the exact workload, batch count, warm-up behavior, metric, and acceptable growth. State whether the agent may expose garbage collection in the test process. Include a timeout so a stalled diagnostic run fails clearly.
Factory’s Local Code Review supports custom review instructions for domain-specific concerns. A team can ask the reviewer to inspect changed listeners, caches, resource ownership, and cleanup paths. The agent should return the test command and measurements, not a general statement that memory looks stable.
Store compact measurements in CI output and keep full snapshots as access-controlled artifacts when policy permits. Heap snapshots can contain application data, so treat them as sensitive. Do not attach them to public issues or pull requests.
Repeat a suspected failure before accepting it. A leak produces a consistent retention pattern. One noisy run calls for better isolation, not an arbitrary increase in the budget.
Start building