Back to BlogApplication Deployment

PM2 Cluster Mode for Next.js in Production

When to use PM2 cluster mode with Next.js, how many instances to run, and pitfalls with in-memory state.

Why process managers matter for Node

Next.js production servers are long-lived Node processes. Without a supervisor, a single uncaught exception removes your site until someone restarts it manually. PM2 provides restart policies, log aggregation, and cluster mode that spreads HTTP connections across CPU cores.

Vcom Web Tech deploys PM2 on Ubuntu and Debian hosts where Kubernetes is unnecessary overhead. Cluster mode uses Node cluster module under the hood to fork workers that share the same port through SO_REUSEPORT or PM2's built-in balancing.

Cluster mode helps CPU-bound rendering and API routes more than it helps a single slow database query. Measure before assuming more instances fix all latency.

Configuring instances safely

Start with instances set to the number of vCPUs or one less if you reserve capacity for PostgreSQL on the same host. Set max_memory_restart to avoid one leaking worker consuming all RAM.

Use pm2 reload for zero-downtime restarts during deploys when your app boots quickly. Slow builds should happen before reload, not during.

If you use sticky sessions because of in-memory caches, document that choice. Prefer external Redis for session storage so all workers share state.

Logging and monitoring

Centralize PM2 logs with logrotate or ship JSON logs to your observability stack. Tag each line with instance id when debugging race conditions.

Combine PM2 metrics with Nginx access logs to see whether one worker receives disproportionate traffic due to long-lived connections.

Alert on restart counts. Frequent restarts often precede visible 502 errors at Nginx.

Integration with Nginx

Nginx upstream blocks can list multiple backend ports if you run one process per port, but PM2 cluster on a single port is simpler. Keep proxy headers consistent so Next.js middleware reads correct hostnames.

Health endpoints should respond before heavy initialization completes, or Nginx will mark the node down during deploy.

Document pm2 startup systemd unit changes after OS upgrades. Package updates sometimes reset systemd drop-ins.

Additional operational notes

Operational excellence on Linux hosting requires documenting every change to Nginx, systemd, PM2, Docker, and DNS in a runbook your team shares. Vcom Web Tech clients benefit when staging environments mirror production firewall rules, TLS versions, and mail authentication so surprises appear before customers notice. Schedule quarterly reviews of backups, certificate expiry, DMARC reports, and monitoring alerts even when traffic feels stable.

When incidents occur, capture timelines and root causes in blameless postmortems. Patterns from past 502 errors, failed renewals, or bounce spikes inform checklists for the next deployment. Training new team members on SSH access, log locations, and escalation paths reduces dependency on single maintainers.

Security patches, dependency upgrades, and framework migrations should ride the same CI pipelines that deploy application code. Automate smoke tests that hit health endpoints and send test mail through staging SMTP relays. Small consistent investments beat heroic firefighting during launch weekends.

Capacity planning matters on VPS hosts where vertical scaling has limits. Watch disk inode usage, connection counts, and database connection pools as traffic grows. Proactive upgrades cost less than emergency migrations during peak sales or campaign sends.

Finally, communicate with stakeholders using plain language about risk, downtime windows, and deliverability metrics. Technical depth supports trust when email authentication or deployment strategy changes affect revenue-facing systems.

Additional operational notes

Operational excellence on Linux hosting requires documenting every change to Nginx, systemd, PM2, Docker, and DNS in a runbook your team shares. Vcom Web Tech clients benefit when staging environments mirror production firewall rules, TLS versions, and mail authentication so surprises appear before customers notice. Schedule quarterly reviews of backups, certificate expiry, DMARC reports, and monitoring alerts even when traffic feels stable.

When incidents occur, capture timelines and root causes in blameless postmortems. Patterns from past 502 errors, failed renewals, or bounce spikes inform checklists for the next deployment. Training new team members on SSH access, log locations, and escalation paths reduces dependency on single maintainers.

Security patches, dependency upgrades, and framework migrations should ride the same CI pipelines that deploy application code. Automate smoke tests that hit health endpoints and send test mail through staging SMTP relays. Small consistent investments beat heroic firefighting during launch weekends.

Capacity planning matters on VPS hosts where vertical scaling has limits. Watch disk inode usage, connection counts, and database connection pools as traffic grows. Proactive upgrades cost less than emergency migrations during peak sales or campaign sends.

Finally, communicate with stakeholders using plain language about risk, downtime windows, and deliverability metrics. Technical depth supports trust when email authentication or deployment strategy changes affect revenue-facing systems.