Use kmemleak to identify allocations that are no longer reachable from live references after cleanup or error handling. Heavy allocation is not a leak if the memory is still referenced and eventually released, but a logical leak leaves objects allocated with no remaining path to them. The signal is unreached allocation, not raw allocation volume.
How do you tell a leak from heavy allocation in a kernel module?
The practical distinction is whether allocations remain reachable from live kernel references after the code path should have released them. Heavy allocation can be normal if ownership is still valid and the memory is eventually freed; a leak is the allocation that survives cleanup or error handling with no remaining reference path. The question is about reachability and lifetime, not peak allocation volume.
kmemleak is useful because it looks for objects that are still allocated but can no longer be reached from known roots, which is the signal that matters in this diagnosis. That makes it a better fit than raw allocator counters when you need to separate legitimate buffering or batching from a true lifecycle failure.
In practice, the strongest clue is consistency across a code path. If repeated runs show the same object type accumulating after teardown, module unload, or failed initialization, the problem is more likely a missing free, lost pointer, or ownership transfer bug than benign load-driven allocation pressure.
What makes the memory a leak rather than an intentional hold?
A module can allocate heavily for caching, queues, staging buffers, per-CPU data, or deferred work, and none of that is a leak by itself. The deciding factor is whether the module still has a valid ownership path to the allocation and a planned release point. If the object is intentionally retained until a later event, it should still be traceable through live references.
For kernel debugging, that means you should compare allocation timing with ownership transitions. If memory appears during normal operation but remains referenced by a live object, list, or work item, it is probably part of the design. If it appears after error unwinding, object destruction, or module exit and never becomes reachable again, it is behaving like a leak.
That is why the analysis needs a lifecycle view. A module often allocates as part of setup, retries, or temporary buffering, but cleanup logic must still be able to walk back every successful allocation. Leaks usually come from one missed free path, one forgotten reference drop, or a partial initialization case that returns early.
How should teams validate the diagnosis in a real module?
Start with the code paths that create and destroy the object, then verify that every successful allocation has a matching release on all exit branches. The useful test is not whether the code allocates a lot, but whether every allocation can still be explained by a reachable owner after init, error handling, and teardown. If that owner disappears, kmemleak should flag the object as unreachable.
It also helps to test the suspected path under repeatable setup and shutdown cycles. A genuine leak tends to accumulate across identical cycles, while intentional allocation usually stabilises once the cache, queue, or pool reaches its expected size. If the footprint grows without plateau and the references do not persist, the suspicion strengthens.
When the module uses deferred work, asynchronous callbacks, or shared objects, verify that the last live reference really is the last one. A common failure mode is freeing the obvious pointer while a secondary reference, list entry, or callback context keeps the object alive, or the reverse, where the programmer assumes another path will free it but no path actually does.
Risk and Threat Considerations
Kernel leaks matter because they degrade reliability before they become obviously visible. A slow leak can consume unreclaimed memory, trigger allocator pressure, and eventually affect unrelated workloads, so the operational risk is often instability rather than immediate failure.
Failure mechanism: A module drops the last real reference incorrectly, skips a cleanup branch, or loses ownership during error handling, leaving allocated objects unreachable but still resident. Over time, the leaked objects accumulate and the kernel cannot reclaim that memory.
Impact: Memory pressure rises, performance can degrade, and in severe cases the system can hit allocation failures, OOM conditions, or service disruption. In long-running systems, that makes leak detection a reliability control, not just a debugging convenience.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | SI-2 — Flaw Remediation | Kernel leaks are software flaws that require detection and correction. |
| SI-4 — System Monitoring | kmemleak is a monitoring signal for unreachable kernel allocations. | |
| CM-6 — Configuration Settings | Module behavior and test instrumentation depend on controlled kernel configuration. | |
| Recommendation — Track and remediate the leak as a software defect with verified fix and regression testing. Use memory monitoring to detect unreachable allocations during runtime testing. Enable the required diagnostic settings and keep them consistent across test runs. | ||
| CIS Controls v8 | CIS-7 — Continuous Vulnerability Management | Memory leaks are defects that should be identified, tracked, and fixed in testing. |
| CIS-10 — Data Recovery | Leak-related instability can require recovery planning when kernel memory pressure escalates. | |
| Recommendation — Prioritise repeated leak findings for remediation and retest the module after changes. Validate recovery procedures for memory pressure and service degradation scenarios. | ||
Practitioner Guidance
What to verify: Check that every successful allocation has one clearly owned release path across normal exit, error unwind, and module unload. If you cannot point to the exact code path that will free the object, treat it as suspect until proven otherwise.
Decision rule: If memory is still reachable from a live owner, treat the growth as legitimate allocation first. If the object is unreachable after cleanup or repeated teardown cycles, treat it as a leak even if the total volume is still modest.
Practitioner takeaway: The right question is not “how much memory was allocated?” but “does any live reference still justify that allocation?” That reachability test is what separates normal kernel behaviour from a real leak.