Finding Node.js Memory Leaks in Production
Signals, heap snapshots, and PM2 restarts when Node.js memory grows on long-running Linux servers.
Symptoms
Gradual RSS growth, increasing GC pause times, and PM2 max_memory_restart loops indicate leaks or unbounded caches. Distinguish leaks from normal heap growth after warm-up by plotting memory over days.
Vcom Web Tech correlates memory graphs with deploy times and traffic patterns before profiling.
OOM kills at the kernel level manifest as sudden 502 errors without graceful Node stack traces.
Safe profiling
Capture heap snapshots during elevated usage on staging clones first. Production profiling requires brief maintenance windows or sampled diagnostics.
Use clinic.js or built-in inspector with caution on live traffic. Prefer reproducing load tests locally when possible.
Document which routes were hit during snapshot capture to narrow suspect code paths.
Common culprits
Global arrays accumulating webhook payloads, timers never cleared, and closures holding large request contexts appear frequently in API servers.
Third-party SDKs with internal caches may need configuration limits.
Misconfigured logging that retains entire response bodies in memory hurts more under high RPS.
Mitigation
Set max_memory_restart as a safety valve while fixing root cause. Rotate workers during low traffic with pm2 reload.
Fix forward with tests that assert stable memory over scripted load when regressions are likely.
Share findings in postmortems so similar patterns are caught in code review.
Additional operational notes
Operational excellence on Linux hosting requires documenting every change to Nginx, systemd, PM2, Docker, and DNS in a runbook your team shares. Vcom Web Tech clients benefit when staging environments mirror production firewall rules, TLS versions, and mail authentication so surprises appear before customers notice. Schedule quarterly reviews of backups, certificate expiry, DMARC reports, and monitoring alerts even when traffic feels stable.
When incidents occur, capture timelines and root causes in blameless postmortems. Patterns from past 502 errors, failed renewals, or bounce spikes inform checklists for the next deployment. Training new team members on SSH access, log locations, and escalation paths reduces dependency on single maintainers.
Security patches, dependency upgrades, and framework migrations should ride the same CI pipelines that deploy application code. Automate smoke tests that hit health endpoints and send test mail through staging SMTP relays. Small consistent investments beat heroic firefighting during launch weekends.
Capacity planning matters on VPS hosts where vertical scaling has limits. Watch disk inode usage, connection counts, and database connection pools as traffic grows. Proactive upgrades cost less than emergency migrations during peak sales or campaign sends.
Finally, communicate with stakeholders using plain language about risk, downtime windows, and deliverability metrics. Technical depth supports trust when email authentication or deployment strategy changes affect revenue-facing systems.
Additional operational notes
Operational excellence on Linux hosting requires documenting every change to Nginx, systemd, PM2, Docker, and DNS in a runbook your team shares. Vcom Web Tech clients benefit when staging environments mirror production firewall rules, TLS versions, and mail authentication so surprises appear before customers notice. Schedule quarterly reviews of backups, certificate expiry, DMARC reports, and monitoring alerts even when traffic feels stable.
When incidents occur, capture timelines and root causes in blameless postmortems. Patterns from past 502 errors, failed renewals, or bounce spikes inform checklists for the next deployment. Training new team members on SSH access, log locations, and escalation paths reduces dependency on single maintainers.
Security patches, dependency upgrades, and framework migrations should ride the same CI pipelines that deploy application code. Automate smoke tests that hit health endpoints and send test mail through staging SMTP relays. Small consistent investments beat heroic firefighting during launch weekends.
Capacity planning matters on VPS hosts where vertical scaling has limits. Watch disk inode usage, connection counts, and database connection pools as traffic grows. Proactive upgrades cost less than emergency migrations during peak sales or campaign sends.
Finally, communicate with stakeholders using plain language about risk, downtime windows, and deliverability metrics. Technical depth supports trust when email authentication or deployment strategy changes affect revenue-facing systems.