
A VPS can have plenty of free disk space and still feel “slow.” Pages hang, wp-admin crawls, SSH commands pause, and PHP workers pile up. In many cases, the culprit is storage latency and high iowait, not a full disk.
This VPS disk I/O troubleshooting tutorial shows how to confirm the bottleneck, find the noisy process, and apply fixes that help common web hosting workloads.
The steps below assume Ubuntu 24.04/26.04 LTS or Debian 12/13 on a VPS or dedicated server. The commands are safe on production systems. Still, take a snapshot or backup before you change filesystem mounts or database settings.
What you’re trying to diagnose (and what “iowait” really means)
CPU iowait is time the CPU spends waiting on storage. You’ll see it as wa in top, or as “iowait” in monitoring charts.
High iowait usually points to one of these situations:
- One or more processes are doing heavy reads/writes (database activity, backups, antivirus scans, log bursts).
- Your storage is slow or saturated (IOPS/throughput limits, contended node storage, failing disks on dedicated servers).
- Your application pattern is inefficient (too many small writes, missing caching, slow queries forcing disk reads).
One detail trips people up: I/O pressure often looks like a CPU issue.
Load average climbs while CPU usage stays moderate. That mismatch is a strong early signal.
Prerequisites: baseline checks and tools to install
SSH in as a sudo-capable user. Capture a quick 60-second “what’s happening right now” baseline.
uptime
free -h
df -hT
lsblk -f
Install the usual I/O tools:
sudo apt update
sudo apt install -y sysstat iotop nmon linux-tools-common linux-tools-generic
On Debian, linux-tools-generic may not exist. Use:
sudo apt install -y linux-perf
If you host multiple customer sites or you regularly see traffic spikes, consistent NVMe performance matters as much as CPU.
For more predictable latency, a HostMyCode VPS gives you isolation and resources that shared hosting can’t reliably offer for disk-heavy workloads.
Step 1: Confirm it’s an I/O bottleneck (not RAM pressure or CPU)
Open two SSH sessions. In the first session, run:
top
Focus on two things:
- Load average rising while CPU
idisn’t near 0% wa(iowait) consistently above ~5–10% during slow periods
In the second session, use vmstat to catch short spikes and correlate symptoms:
vmstat 1 60
Quick reading guide:
waclimbing supports the “blocked on storage” theory.- High
si/so(swap in/out) usually means RAM pressure; address memory before tuning storage. - High
b(blocked processes) often rises during I/O stalls.
If you’re swapping heavily, fix that first. Swap can prevent OOM crashes.
It can also turn minor I/O latency into a full-on slowdown.
Step 2: Identify which disk and which process is causing the I/O
Start with per-device stats. iostat gives a clean view of saturation and latency:
iostat -x 1 60
Pay attention to:
%utilnear 100%: the device is saturated.awaitstaying high (for example, 20–50ms+ on NVMe workloads): latency is the issue, not just volume.r/s,w/s,rkB/s,wkB/s: whether the pressure is read-heavy or write-heavy.
Next, map that device activity to a process. iotop is usually the fastest way to catch top offenders.
Root is required for full visibility.
sudo iotop -oPa
On hosting servers, the usual suspects are predictable:
- mysqld/mariadbd flushing dirty pages or writing redo logs
- php-fpm or lsphp writing sessions/caches or doing lots of small reads
- tar/rsync backups running at the wrong time
- clamd or malware scanners reading the entire filesystem
- log bursts from bots hammering endpoints and inflating access logs
If logs are part of the story, avoid the “just delete the file” impulse.
Rotate them, compress old logs, and set sane retention.
Our VPS log rotation tutorial covers a safe setup that reduces surprise I/O spikes from multi-GB log files.
Step 3: Quick isolate tests for web hosting stacks (WordPress, cPanel, Nginx/Apache)
Now you’re narrowing the source. It’s usually real traffic, background jobs, or routine maintenance.
Test A: Is traffic triggering I/O?
Check top URL paths and look for obvious abuse. For Nginx:
sudo awk '{print $7}' /var/log/nginx/access.log | sort | uniq -c | sort -nr | head
For Apache (Debian/Ubuntu paths vary):
sudo awk '{print $7}' /var/log/apache2/access.log | sort | uniq -c | sort -nr | head
If you see repeated hits to /wp-login.php, /xmlrpc.php, or ugly query-string endpoints, the disk can pay twice.
You get log writes plus PHP session churn. Add rate limits and a WAF rule set so the server isn’t “logging itself to death.”
For fast bot-spike triage, use our VPS log analysis tutorial.
Test B: Are backups or cron jobs hammering storage?
List system timers and recent runs:
systemctl list-timers --all | head -n 50
On WordPress, watch WP-Cron behavior.
If wp-cron.php runs on every request, you’ll create a steady stream of small writes.
On busy sites, a real cron plus disabling WP-Cron is often the cleaner setup.
Test C: Is it PHP session I/O?
Many stacks still write PHP sessions to disk by default. Confirm where your sessions land:
php -i | grep -E 'session.save_handler|session.save_path'
If session writes show up in iotop, moving sessions to Redis can cut random disk I/O sharply. This is especially true on logged-in sites.
Fewer small writes usually means fewer fsync stalls and better tail latency.
Step 4: Fix the most common hosting I/O problems (practical changes)
Resist the urge to tweak everything. Make one change, measure it, then decide what to do next.
Fix 1: Stop daytime backups from saturating disk
Backups should be invisible. If tar, rsync, or a control panel backup task dominates iotop, move it to a quiet window.
Then throttle it.
Example: throttle an rsync job to ~25MB/s to reduce impact:
rsync -aHAX --delete --info=progress2 --bwlimit=25000 /source/ user@backup:/dest/
If you run cPanel/WHM, tune backup scheduling and destinations carefully (remote is better than local).
For a WHM-first workflow, follow our cPanel backup configuration tutorial.
Fix 2: Reduce log write pressure
During bot spikes, access/error logs can become a write-amplifier. You get constant appends, plus whatever parses those logs later.
Make rotation predictable (daily or size-based) and compress older files.
For Nginx, you can buffer access logs. The trade-off is real: you may lose the last few seconds of logs after a crash. Example snippet:
access_log /var/log/nginx/access.log combined buffer=256k flush=5s;
Fix 3: Mount options that reduce needless writes (with care)
On ext4, relatime is usually the default and works well.
If you somehow have atime enabled (uncommon on current distros), switching to relatime reduces metadata writes.
Check mount options:
mount | grep ' / '
If you need to adjust, update /etc/fstab and remount. Example (do not copy blindly):
UUID=xxxx / ext4 defaults,relatime 0 1
If the workload is database-heavy, don’t start with risky filesystem tweaks.
You’ll usually get better results by reducing writes at the source first (queries, caching, logging).
Fix 4: Tune PHP-FPM to avoid disk thrash from process churn
Too much PHP worker churn can amplify I/O.
Common multipliers include opcode cache warmups, autoloader reads, session writes, and cache rebuilds.
Set realistic limits per site and keep pools from ballooning under bursty traffic.
On Nginx + PHP-FPM stacks, per-site pool tuning is one of the highest ROI changes you can make. Use our VPS PHP-FPM pool tuning tutorial, then re-check iowait during your normal peak window.
Fix 5: Add full-page caching so PHP and the database stop hitting disk
On WordPress, missing caching often turns into repeated DB reads and cache misses that spill to disk.
If most of your pages are cacheable, full-page caching can cut backend work dramatically.
If you run Nginx, FastCGI cache is a solid option. It serves anonymous traffic without waking PHP, and it uses the filesystem efficiently.
Start here: WordPress full-page caching tutorial.
Step 5: Advanced diagnostics when iowait persists
If you have a device with high await and %util, but no single obvious offender, you need more signal.
Check for filesystem-level stalls and flush pressure
Dirty page flushing can cause periodic “I/O storms.”
Check whether writeback builds up and then dumps in bursts:
cat /proc/meminfo | egrep 'Dirty|Writeback'
If Dirty grows steadily and then drops sharply, you may be watching burst flushes.
The safest fixes are usually workload-driven. Reduce write volume (logs, backups), or add RAM so the kernel can cache more.
Measure latency directly with a small, safe test
You can run a small read test without beating up the server:
sudo dd if=/var/log/syslog of=/dev/null bs=1M status=progress
For writes, be careful on production.
If you must test, write to a temp file and remove it right after:
sudo dd if=/dev/zero of=/tmp/io-test.bin bs=64M count=4 oflag=direct status=progress
sudo rm -f /tmp/io-test.bin
If throughput swings wildly between runs, you’re likely hitting contention or storage limits.
On dedicated servers: rule out disk health issues
Drives often fail as latency first, not obvious errors. Check SMART quickly:
sudo apt install -y smartmontools
sudo smartctl -a /dev/sda | egrep -i 'reallocated|pending|uncorrect|error|fail'
If you run RAID, monitor it instead of waiting for a rebuild surprise.
Use our VPS RAID monitoring tutorial as a baseline.
Step 6: Remediation playbook (choose based on what you found)
This is the decision tree you’ll use during real incidents.
- If one process dominates writes (backups, scans, rsync): reschedule, throttle, or move it to a separate backup node.
- If logs dominate writes: rotate aggressively, buffer where acceptable, and block abusive traffic (WAF/rate limits).
- If database dominates reads: add caching, fix slow queries, and raise memory allocation so hot data stays in RAM.
- If everything looks “a bit busy”: you’re saturated. Upgrade to faster storage and/or move high-traffic sites to their own VPS.
Once you’re consistently hitting saturation, tuning stops being the best tool.
Moving a busy WooCommerce store off a mixed-use VPS to a larger NVMe VPS—or a dedicated plan—often drops tail latency fast. You remove I/O contention.
If you want predictable performance for production sites, consider managed VPS hosting where the platform and tuning are handled with hosting workloads in mind.
Verification: prove the fix worked (numbers you can screenshot)
After each change, collect a before/after set during a known busy period:
iostat -x 1 30(device utilization andawait)vmstat 1 30(iowait and swap activity)- Web response times from an external check (TTFB and 95th percentile, not just “it feels faster”)
Watch the secondary signals too.
You want fewer 502/504 errors, shorter PHP-FPM queues, and a lower load average under the same traffic.
Summary: keep disk latency from becoming your next outage
Disk I/O trouble rarely announces itself cleanly. It shows up as “random slowness,” slow admin actions, and timeouts across multiple sites at once.
The workflow that holds up in 2026 stays simple.
Confirm iowait, identify the device, map the I/O to a process, then apply one measured fix at a time.
If your sites have outgrown best-effort storage performance, switching to a plan built for predictable I/O is often the cleanest path forward.
Start with a HostMyCode VPS for isolation, or step up to dedicated servers when sustained writes and traffic spikes demand it.
High iowait during peak traffic usually means two things: your storage performance isn’t consistent, and the stack needs workload-aware tuning. HostMyCode offers managed VPS hosting if you want changes applied safely and verified, plus flexible HostMyCode VPS plans if you prefer to run the server yourself.
FAQ
How much iowait is “too high” on a web hosting VPS?
If iowait is consistently above ~10% during user-facing slowdowns, treat it as an incident.
Brief spikes can be normal, but sustained iowait usually lines up with timeouts and poor response times.
My load average is high, but CPU usage is low. Is that I/O?
Often, yes. High load with moderate CPU typically means processes are blocked on disk or network I/O.
Confirm with top (wa%) and iostat -x (await/%util).
Will adding swap fix disk I/O problems?
Swap can prevent OOM crashes, but it can also add disk reads/writes and make latency worse.
If the root issue is slow storage or saturated IOPS, swap won’t help.
What’s the fastest way to find the process causing heavy disk writes?
Run sudo iotop -oPa and watch the top writer during the slowdown.
Cross-check with iostat -x 1 to confirm the disk saturates at the same time.
When should I upgrade vs keep tuning?
If the device hits ~100% utilization regularly under normal traffic, you’re at the limit.
Scheduling and caching help, but a bigger/faster VPS or a dedicated server is usually the durable fix.