
“The site is slow” usually isn’t a web-server bug. On a hosting VPS, the usual culprits are CPU contention, RAM pressure (swap thrash), or storage I/O stalls.
This VPS performance troubleshooting tutorial gives you a repeatable workflow. Identify the bottleneck, prove it with numbers, then apply a specific fix instead of guessing.
The commands below assume Ubuntu 24.04/26.04 LTS or Debian 12/13. AlmaLinux/Rocky has the same tools, but package names and service paths can differ.
What you’ll collect first (so you don’t troubleshoot blind)
Before you change anything, capture a 10–15 minute baseline while the site is actually slow. Treat it as your “before” snapshot for every fix you test.
- Top-level resource view: CPU, load average, memory, swap, disk I/O wait
- Process-level view: which PIDs and users are consuming CPU/RAM
- Web-level view: slow requests, 5xx errors, upstream time, PHP-FPM saturation
- Disk-level view: is storage the limiter (await/util)?
Install the basic toolbelt
On Ubuntu/Debian:
sudo apt update
sudo apt install -y htop sysstat iotop iftop nginx-common apache2-utils
sysstat provides iostat and sar. They stay quiet most days. On bad days, they’re indispensable.
Step 1 — Confirm the bottleneck class (CPU vs RAM vs Disk vs Network)
Start with a quick snapshot to orient yourself. Then switch to time-series so you can spot patterns.
Quick snapshot
uptime
free -h
vmstat 1 10
- High load + low CPU usage often means disk I/O wait. It’s not “a slow CPU.”
- Free RAM is not the goal. Linux uses spare RAM for cache. Watch swap use and paging instead.
- In
vmstat, track wa (I/O wait) and si/so (swap in/out). Persistent non-zero so is a problem.
CPU: is it real compute or blocked time?
mpstat -P ALL 1 5
If CPU sits at 90–100% in usr/sys, you’re compute-bound. If it piles up in iowait, the CPU is mostly waiting on storage.
Disk: is your VPS stuck waiting on storage?
iostat -xz 1 10
Focus on:
- %util near 100% consistently: the device is saturated.
- await climbing (tens of ms on NVMe is suspicious): latency is hurting you.
- r/s, w/s spikes that line up with slow page loads: you’ll correlate this with logs later.
Network: less common, but easy to rule out
iftop
If one IP is pushing a lot of bandwidth, you may be looking at bot traffic or hotlinking. That’s usually an application/WAF/CDN issue, not a reason to buy more CPU.
If you prefer to start from logs and work upward, pair this workflow with VPS log analysis tutorial.
You can then match resource spikes to specific endpoints and user agents.
Step 2 — Identify the exact processes causing pain
Once you know the bottleneck type, pinpoint what’s driving it. On hosting VPS setups, “which user/site” matters as much as “which PID.”
Find top CPU consumers
htop
In htop, press:
- F6 to sort by CPU%
- F4 to filter (try:
php-fpm,mysqld,apache2,nginx)
If PHP-FPM is at the top, the next question is simple: which pool (which site) is doing the work?
Map PHP-FPM processes to pools/sites
On many servers, PHP-FPM runs per-site pools. Start here:
ps -eo pid,user,cmd --sort=-%cpu | head -n 25
Look for pool clues in the command line or user (for example, site1, customer123).
Common config paths:
- Ubuntu/Debian:
/etc/php/8.3/fpm/pool.d/(or8.2,8.4) - Service:
systemctl status php8.3-fpm
Find memory hogs and OOM risk
free -h
swapon --show
ps -eo pid,user,rss,cmd --sort=-rss | head -n 20
If swap is active and keeps growing during slow periods, performance can collapse fast.
Your goal is to stop swapping. Reduce concurrency, reduce per-process memory, or add RAM.
On production systems, adding RAM is often the cleanest fix.
Step 3 — Web stack diagnostics: prove whether it’s PHP, Apache/Nginx, or the app
System metrics tell you where the pressure is. Logs and service status tell you what created it.
Nginx: check active connections and upstream timing
If you run Nginx, enable a minimal stub_status endpoint (local-only). This shows whether workers are backing up.
Example snippet:
sudo tee /etc/nginx/conf.d/status.conf >/dev/null <<'EOF'
server {
listen 127.0.0.1:8080;
location /nginx_status {
stub_status;
allow 127.0.0.1;
deny all;
}
}
EOF
sudo nginx -t && sudo systemctl reload nginx
curl -s http://127.0.0.1:8080/nginx_status
If Active connections rises and doesn’t drop, you’re usually waiting on upstream responses (PHP/app).
You may also have worker limits set too low.
Apache: check scoreboard and busy workers
On Apache, enable mod_status (if it isn’t already). Then inspect the scoreboard and busy workers.
In many monitoring setups, this shows up as “MaxRequestWorkers reached” during traffic peaks.
Resist the urge to tune from muscle memory. Capture the evidence first. Then change one variable.
PHP-FPM: check pool saturation (the most common root cause)
PHP-FPM tends to fail in predictable ways. Usually it’s too many concurrent requests, not enough workers, or workers so memory-heavy that the kernel starts swapping.
Check logs:
sudo journalctl -u php8.3-fpm --since "-2h" | tail -n 200
Look for lines like:
server reached pm.max_childrenslowlogentries (if enabled)
For a production-safe way to tune pools per site (and stop one WordPress install from dragging others down), follow VPS PHP-FPM pool tuning tutorial.
Step 4 — Fix the common causes (in the right order)
Once you’ve pinned down the bottleneck, apply one change. Then re-measure.
This discipline prevents “fixing” one symptom and creating a new problem elsewhere.
Fix A: swap thrashing and random freezes
If swap usage is heavy during normal traffic, the VPS is undersized for the concurrency you’re serving. Two moves usually help right away:
- Reduce concurrency (web/PHP worker limits) so the kernel stops swapping.
- Add swap safely if you have none (swap is a safety net, not a performance plan).
If you need to add swap, keep it conservative (often 1–4 GB on small VPS plans). Then confirm it’s enabled correctly.
Use this guide: VPS swap configuration tutorial.
Fix B: bot spikes that eat CPU and PHP workers
Typical signs: CPU jumps, PHP-FPM hits max children, and access logs fill with wp-login.php, xmlrpc.php, or random probing paths.
The goal is simple: cut the number of dynamic hits that reach PHP.
- Add rate limiting at the web server for obvious abuse endpoints.
- Enable caching for anonymous traffic (WordPress sites often benefit immediately).
- Use a WAF rule set if you’re seeing exploit scans.
If you run WordPress and you can use Nginx FastCGI cache, it’s one of the few changes that can materially reduce CPU load on anonymous traffic.
Use WordPress full-page caching tutorial and test before/after against the same page.
Fix C: “disk is 100% utilized” (slow admin, slow checkout, slow everything)
If iostat shows saturated storage, you have three practical levers:
- Stop doing expensive writes: verbose logging, runaway backups, cache plugins writing constantly, malware creating files.
- Move hot paths: put caches on tmpfs (carefully), or relocate logs to a separate volume if available.
- Upgrade storage tier: on busy WordPress hosting, NVMe and higher IOPS often matter more than another CPU core.
Also confirm you’re not simply running out of space.
A nearly full filesystem increases write amplification. It can also slow routine operations. If you suspect it, run:
df -h
sudo du -xhd1 /var | sort -h
Then follow VPS disk space troubleshooting tutorial to reclaim space without deleting the wrong things.
Fix D: “CPU is maxed” but only for one or two sites
This shows up constantly on multi-tenant VPS hosting and reseller boxes. The fix is isolation and per-site limits, not a server-wide tweak.
- Set per-site PHP-FPM pools with sane
pm.max_children. - Cap execution time and memory per site.
- Make sure one account can’t read or exhaust another account’s resources.
If you’re running cPanel on a hosting VPS, account isolation features (CageFS + controlled PHP-FPM) are designed for exactly this scenario.
See cPanel account isolation tutorial.
Step 5 — Run a controlled load test (so you know the fix really worked)
Don’t punish production with unrealistic benchmarks. You’re answering one question: “At what concurrency does the server start to saturate and slow down?”
Simple test with ApacheBench (AB)
From a separate machine (or at least outside the request path), run:
ab -n 500 -c 20 https://example.com/
-cis concurrency. Step it up gradually (10 → 20 → 40).- Test a cached page and an uncached dynamic page (like a search result) separately.
Watch server metrics while the test runs:
vmstat 1
iostat -xz 1
mpstat -P ALL 1
If latency rises while wa spikes, disk is your ceiling. If CPU hits 100% in user time, you’re compute-bound.
If swap churn begins, you’re memory-bound.
Step 6 — Quick hardening moves that also improve performance stability
Not every security change makes a site faster. Several do help it stay responsive under pressure, which customers experience as “performance.”
- Firewall basics: only expose what you need (80/443, SSH, panel ports). See VPS firewall setup guide tutorial.
- Keep time in sync: time drift causes SSL errors and confusing cache behavior. If you’ve seen symptoms like that, use VPS time sync troubleshooting tutorial.
- WAF for noisy WordPress targets: ModSecurity + OWASP CRS can reduce exploit traffic that would otherwise trigger expensive PHP execution. Roll it out gradually and monitor false positives.
Troubleshooting checklist (printable, no fluff)
- Capture baseline:
uptime,free -h,vmstat,mpstat,iostat - Confirm the bottleneck class (CPU vs RAM vs disk I/O vs network)
- Identify top processes and map to sites/users (
ps, pool configs) - Check web logs for patterns (same path, same IPs, same UA)
- Fix one thing at a time: concurrency limits, caching, bot control, storage tier, RAM sizing
- Re-test with controlled concurrency and compare numbers
- Set alerts so you catch regression before customers do
Where HostMyCode fits (and when to stop tuning)
If the data says your workload has outgrown the box, the best “tweak” is right-sizing. Chasing config changes on an undersized VPS burns time and increases risk.
If you want full control, a HostMyCode VPS gives you predictable resources for PHP, mail, and control panels.
If you’d rather offload patching, monitoring, and safe change management, managed VPS hosting is a better fit for production.
If you’re troubleshooting slow sites every week, you’re paying the “performance tax” in time. Move workloads that need stable CPU/RAM and fast storage to a HostMyCode VPS, or hand the operational burden to managed VPS hosting so you can focus on the sites instead of the server.
FAQ
How do I know if my VPS is CPU-bound or disk-bound?
Run mpstat -P ALL 1 and iostat -xz 1 during the slowdown. High CPU user/system time points to CPU-bound work. High wa and high disk %util/await points to storage limits.
Is swap always bad on a hosting VPS?
No. Swap is a safety net. The problem is sustained swap-in/swap-out during normal traffic, which causes request latency spikes. If you see constant swap activity, reduce concurrency or add RAM.
Why does load average look high when CPU isn’t maxed?
Load includes processes waiting on I/O. A high load with moderate CPU often means the server is stuck waiting on disk. Confirm with vmstat (I/O wait) and iostat (await/util).
What’s the quickest win for WordPress performance under traffic?
Full-page caching for anonymous visitors typically yields the biggest reduction in PHP work. On VPS setups using Nginx, FastCGI cache is a practical option as long as you exclude login/cart/checkout endpoints.
Should I increase PHP-FPM workers to fix slowness?
Only if you have spare RAM. Increasing workers on a memory-limited VPS often triggers swapping and makes performance worse. Tune per-site pools and keep an eye on RSS memory per worker.
Summary: your repeatable workflow
Classify the bottleneck with metrics, identify the exact process and site responsible, then apply one targeted change and re-measure.
If the numbers show the VPS is undersized, stop tuning and move to a plan with headroom.
HostMyCode’s VPS hosting and managed VPS hosting are built for day-to-day hosting realities: spiky traffic, PHP concurrency, and customers who notice every slowdown.