Back to tutorials
Tutorial

VPS RAID Monitoring Tutorial (2026): Detect Failing Disks Early on Dedicated Servers with mdadm + SMART + Alerts

VPS RAID monitoring tutorial for 2026: configure mdadm, SMART checks, and email alerts to catch disk failure early.

By Anurag Singh
Updated on Oct 02, 2026
Category: Tutorial
Share article
VPS RAID Monitoring Tutorial (2026): Detect Failing Disks Early on Dedicated Servers with mdadm + SMART + Alerts

A degraded RAID can look normal—until a reboot, a rebuild, or a burst of writes pushes it over the edge. This VPS RAID monitoring tutorial focuses on early warnings: mdadm array state, SMART drive health, and alerts that actually reach you.

The same approach works on dedicated servers and on VPS plans where you manage storage inside the OS.

Examples include software RAID on your own hardware and storage-heavy VPS setups.

If you host client sites, email, or WooCommerce stores, a disk that starts reallocating sectors isn’t a “later” problem.

Catch it early to avoid rushed migrations, broken backups, and painful restore windows.

The goal is simple.

You find out before the array drops a member or the filesystem flips read-only.

What you’ll build (and what you need)

  • RAID state monitoring with mdadm (degraded arrays, rebuilds, mismatch counts).
  • SMART monitoring with smartmontools (pending sectors, media errors, wear indicators for SSD/NVMe).
  • Email alerts via a local MTA (lightweight) so notifications still work during partial outages.
  • Optional webhook-style alerts using a simple script (for teams using chat/incident tools).

Assumptions: Ubuntu 24.04/26.04 LTS or Debian 12/13, root access, and an array managed by mdadm (/dev/md*).

If you’re on hardware RAID (LSI/Adaptec/HP Smart Array), SMART may still be possible.

RAID health checks, however, move to vendor tooling.

You’ll get the most predictable results on a server where you control the storage stack end to end.

That’s usually a HostMyCode dedicated server or a storage-optimized HostMyCode VPS where you run your own monitoring and alerting.

Step 1 — Confirm your RAID layout and name the arrays

First, confirm you’re using mdadm software RAID.

Then identify the devices involved.

cat /proc/mdstat
lsblk -o NAME,SIZE,TYPE,FSTYPE,MOUNTPOINT,MODEL,SERIAL
sudo mdadm --detail --scan

Typical healthy output looks like:

md0 : active raid1 sda1[0] sdb1[1]
      976630336 blocks super 1.2 [2/2] [UU]
  • [UU] means both members are present. Anything like [U_] is degraded.
  • If you see a resync or recovery line, the array is rebuilding.

Pull full detail for each array you care about:

sudo mdadm --detail /dev/md0

Pitfall: older arrays often have inconsistent mdadm config.

Common issues include missing MAILADDR and no monitoring integration. You’ll fix that next.

Step 2 — Install mdadm + SMART tools (and a small MTA)

Install the packages you’ll use. On Ubuntu/Debian:

sudo apt update
sudo apt install -y mdadm smartmontools mailutils bsd-mailx

About email delivery: you don’t need a full mail server just to send alerts.

A lightweight relay (Postfix in “satellite” mode) or provider SMTP is usually enough.

If you already run customer mail on the same machine, keep monitoring alerts on a separate path.

Otherwise, “mail is down” can also mean “alerts are down.”

If you need a clean SMTP path on a VPS, see: Postfix mail server setup tutorial (2026).

Step 3 — Enable mdadm monitoring and set alert recipients

On current Ubuntu/Debian releases, mdadm monitoring runs under systemd.

Start by setting an alert recipient. You can also set a friendly sender.

Edit /etc/mdadm/mdadm.conf:

sudo nano /etc/mdadm/mdadm.conf

Add or update these lines:

MAILADDR ops@example.com
MAILFROM raid-alerts@yourdomain.com

Rebuild initramfs so the array config stays consistent across boots.

This matters on some setups.

sudo update-initramfs -u

Enable and start the monitor service:

sudo systemctl enable --now mdadm --quiet || true
sudo systemctl enable --now mdmonitor
sudo systemctl status mdmonitor --no-pager

Quick diagnostic: if mdmonitor doesn’t exist on your build, list what’s available:

systemctl list-unit-files | grep -E 'mdadm|mdmonitor'

On some distributions, the service is mdadm-monitor:

sudo systemctl enable --now mdadm-monitor

Step 4 — Trigger a safe test alert (without breaking production)

Don’t wait for a real failure to learn your outbound mail is blocked or misrouted.

Start with a plain email test:

echo "RAID alert email test from $(hostname)" | mail -s "RAID test" ops@example.com

If you want mdadm to generate an event, you can simulate a failure.

Mark one member as failed, then re-add it. Only do this if you understand the impact.

This is generally safe on RAID1/RAID10 with proper redundancy. Do not do this on RAID0.

If you’re unsure, skip this and move on to SMART tests.

# Example: mark a member as failed
sudo mdadm /dev/md0 --fail /dev/sdb1

# Verify degraded state
cat /proc/mdstat

# Remove and re-add
sudo mdadm /dev/md0 --remove /dev/sdb1
sudo mdadm /dev/md0 --add /dev/sdb1

Watch rebuild progress:

watch -n 2 cat /proc/mdstat

Operational note: rebuilds can punish I/O and drag down WordPress and databases.

Run tests in low-traffic windows. Keep an eye on load and latency.

If your sites already struggle under disk pressure, read: VPS performance troubleshooting tutorial (2026).

Step 5 — Configure SMART monitoring with smartd

mdadm tells you about the array.

SMART tells you about the drives inside it. You want both signals.

Identify the physical disks behind your array:

lsblk -d -o NAME,MODEL,SERIAL,TYPE,SIZE

Grab a baseline SMART report for each SATA/SAS disk:

sudo smartctl -a /dev/sda
sudo smartctl -a /dev/sdb

For NVMe:

sudo smartctl -a /dev/nvme0n1

Now configure smartd. Edit /etc/smartd.conf:

sudo nano /etc/smartd.conf

A practical, readable starting point for two SATA drives:

/dev/sda -a -o on -S on -n standby,q -s (S/../.././02|L/../../6/03) -m ops@example.com
/dev/sdb -a -o on -S on -n standby,q -s (S/../.././02|L/../../6/03) -m ops@example.com
  • -a: enables the common checks (attributes, error logs, and more).
  • -n standby,q: avoids waking sleeping drives just to run checks (useful for some storage profiles).
  • -s: schedules a short test daily at 02:00 and a long test weekly (Saturday 03:00 here).
  • -m: sets the alert email recipient.

For NVMe, use the NVMe device path:

/dev/nvme0n1 -a -o on -S on -s (S/../.././02|L/../../6/03) -m ops@example.com

Enable and start the service:

sudo systemctl enable --now smartd
sudo systemctl status smartd --no-pager

Quick diagnostic: confirm smartd accepted your config.

Also confirm it’s watching the right devices:

sudo journalctl -u smartd -b --no-pager | tail -n 80

Step 6 — Know which SMART attributes actually matter

SMART output includes a lot of noise.

For hosting, focus on values that change over time and correlate with read/write trouble.

  • Reallocated_Sector_Ct (HDD/SSD SATA): any upward trend is a problem.
  • Current_Pending_Sector (HDD): non-zero often lines up with read errors and timeouts.
  • Offline_Uncorrectable: treat increases as urgent.
  • UDMA_CRC_Error_Count: often points to cabling/backplane issues, not the disk itself.
  • NVMe “Media and Data Integrity Errors”: should stay at 0.
  • NVMe “Percentage Used”: useful for endurance planning; replace before wear becomes uncomfortable.

Save a baseline report so you have something to compare later:

sudo smartctl -a /dev/sda > /root/smart-baseline-sda.txt
sudo smartctl -a /dev/sdb > /root/smart-baseline-sdb.txt

If you pack a lot of accounts onto one box, plan replacements instead of debating them.

A “working” disk that retries reads can still cause slow queries, PHP timeouts, and backups that hang.

Step 7 — Add a simple daily RAID + SMART status email (cron)

mdadm and smartd will alert on events.

A daily “everything looks normal” email serves a different purpose.

It helps you spot drift. It also gives you a clean report to attach to a ticket.

Create /usr/local/sbin/storage-health-report.sh:

sudo nano /usr/local/sbin/storage-health-report.sh
#!/bin/bash
set -euo pipefail

HOST=$(hostname -f 2>/dev/null || hostname)
DATE=$(date -Is)

{
  echo "Storage health report for: $HOST"
  echo "Generated: $DATE"
  echo
  echo "=== /proc/mdstat ==="
  cat /proc/mdstat || true
  echo

  echo "=== mdadm details ==="
  for md in /dev/md*; do
    [ -e "$md" ] || continue
    echo "--- $md ---"
    mdadm --detail "$md" || true
    echo
  done

  echo "=== SMART summary ==="
  for d in /dev/sd[a-z] /dev/nvme*n1; do
    [ -e "$d" ] || continue
    echo "--- $d ---"
    smartctl -H "$d" || true
    smartctl -A "$d" | egrep -i 'Reallocated_Sector_Ct|Current_Pending_Sector|Offline_Uncorrectable|UDMA_CRC_Error_Count|Media and Data Integrity Errors|Percentage Used' || true
    echo
  done
} | mail -s "[$HOST] Daily RAID/SMART report" ops@example.com

Make it executable:

sudo chmod 750 /usr/local/sbin/storage-health-report.sh

Add a cron entry (daily at 07:15):

sudo crontab -e
15 7 * * * /usr/local/sbin/storage-health-report.sh

Log growth pitfall: if you start keeping lots of reports or verbose logs, make sure rotation is in place.

This guide helps keep /var from filling up: VPS log rotation tutorial (2026).

Step 8 — Alert on RAID degradation in monitoring (optional but practical)

If you already have a monitoring server, add a remote check.

Flag anything except [UU] (or the expected pattern for RAID10/RAID6).

The simplest version is a single SSH command. You can plug it into almost any monitoring system:

cat /proc/mdstat | grep -E '\[[U_]+\]' -n

If you want a stricter “fail on any underscore” check:

if grep -q '_' /proc/mdstat; then
  echo "CRITICAL: RAID degraded"
  exit 2
else
  echo "OK: RAID healthy"
  exit 0
fi

If you allow SSH for monitoring, lock it down.

A jump host is the cleanest way to reduce exposure: SSH Jump Host Setup Guide Tutorial (2026).

Step 9 — What to do when you get an alert (your “don’t panic” runbook)

Alerts only help if your response stays consistent.

This sequence works well for hosting servers.

  1. Confirm current state:
    cat /proc/mdstat
    sudo mdadm --detail /dev/md0
    
  2. Check kernel logs for timeouts:
    sudo dmesg -T | egrep -i 'md|raid|blk_update_request|I/O error|reset|timeout' | tail -n 80
    
  3. Pull SMART data for the suspected drive:
    sudo smartctl -a /dev/sdb
    
  4. If the array is rebuilding, monitor progress and postpone heavy disk work (big backups, full malware scans) until it finishes.
  5. Replace the drive as soon as you can schedule it. Once SMART trends go bad, they rarely reverse for long.

If you host WordPress for clients: consider pausing heavy cron jobs and large media processing while a rebuild runs.

Rebuild I/O pressure is a common cause of “everything feels randomly slow” complaints.

Step 10 — Tighten the basics that make alerts reliable

Most “why didn’t I get the email?” incidents come down to two boring failures.

One is wrong time. The other is broken mail/DNS.

  • Time sync: bad time breaks TLS and can shift scheduled tests.
    timedatectl
    systemctl status systemd-timesyncd --no-pager
    

    If time is drifting, follow: VPS time sync troubleshooting tutorial (2026).

  • Outbound mail path: if you relay alerts through a domain, make sure DNS is correct (SPF helps). If you also run mail services, keep MX records clean: MX record setup tutorial (2026).

Step 11 — A practical checklist for hosting admins

  • Array health: /proc/mdstat shows no _, no unexpected resync loops.
  • mdadm alerts: MAILADDR set and test email confirmed.
  • SMART monitoring: smartd enabled, scheduled tests configured.
  • Baseline SMART captured and compared monthly (trends matter more than a single value).
  • Runbook exists: who replaces disks, what’s the maintenance window, how you notify customers.
  • Backups are offsite and restore-tested (RAID is not a backup).

Summary: storage monitoring that actually reduces downtime

mdadm tells you when redundancy is at risk.

SMART tells you when a drive is degrading before it drops out.

Combine both, route alerts through a dependable mail path, and you’ll catch most failures early enough to replace hardware on your terms.

If you want predictable performance and full control over storage monitoring, run this on a HostMyCode dedicated server or choose managed VPS hosting when you’d rather have help with the operational details.

HostMyCode keeps the pitch simple: Affordable & Reliable Hosting you can actually administer.

If you’re consolidating multiple sites onto one box, storage problems become customer-visible fast. A HostMyCode VPS fits hands-on admins who want control. If you’d rather have guided maintenance and monitoring, managed VPS hosting takes the edge off day-to-day operations.

Need to move sites before you touch hardware? HostMyCode can help you plan safer cutovers and migrations with fewer surprises.

FAQ

Does RAID mean I don’t need backups?

No. RAID covers single-disk failure (sometimes two, depending on level). It doesn’t protect you from accidental deletes, ransomware, filesystem corruption, or a bad update.

Can I use this on a VPS?

You can use the alerting approach on any VPS, but mdadm RAID monitoring only applies if the OS actually manages RAID devices.

On many VPS platforms, storage redundancy is handled underneath the VM.

What’s the earliest SMART warning I should treat as urgent?

Any increase in reallocated sectors or pending sectors is worth attention.

One bad value isn’t always fatal, but an upward trend usually predicts a future failure.

Will SMART tests slow my websites?

Short tests are usually low impact. Long tests can increase disk activity, especially on HDD arrays.

Schedule them during off-peak hours and avoid running them during RAID rebuilds.

What if I’m on hardware RAID?

You’ll use the controller’s CLI/tools for RAID status (vendor-specific).

SMART may still work, depending on passthrough support. If SMART data isn’t visible, prioritize controller event alerts and regular patrol reads.

VPS RAID Monitoring Tutorial (2026): Detect Failing Disks Early on Dedicated Servers with mdadm + SMART + Alerts | HostMyCode