ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

RAID is not backup: what backup actually means

RAID is not backup: what backup actually means Every homelab forum has the same conversation about once a week: someone built their first NAS, set up RAIDZ2 across six drives, and wants to know “do I need backup, since RAID gives me redundancy?” The answer is yes, and the reason involves understanding what RAID actually protects against - and what it doesn’t.This guide is the long-form version of “RAID is not backup.” Not as a slogan; as a concept. RAID protects againsta specific type of failure during normal operation(one or more drives dying mechanically). That is one failure mode out of dozens. The other failure modes - accidental deletion, ransomware, filesystem corruption, fire, theft, software bugs, your own typos - are unaffected by RAID, and several of them happily destroy a perfectly redundant array.If you take one thing away:redundancy is not backup, and backup is not redundancy. The two solve different problems. You need both.What RAID actually doesRAID (in any flavor - hardware RAID, software RAID with mdadm, ZFS RAIDZ, btrfs RAID levels) does one specific thing well: it lets your storage survive the failure of one or more disks while staying online and serving data. With RAIDZ1, you lose any single drive; the data stays available; replace the failed drive; the array rebuilds. With RAIDZ2, you survive any two drives. With mirror pairs, you survive one drive per pair.RAID is auptime feature. It keeps the storage system available during a specific class of hardware failure. That’s a real benefit, especially for systems that need to keep running while a drive replacement happens.Hardware RAID / ZFS / btrfs / mdadm covers the comparison among RAID approaches; RAIDZ1 vs RAIDZ2 vs Mirror covers the specific ZFS layout decisions.What RAID does NOT doThe list of failure modes that RAID doesnotprotect against is long. Every one of them has destroyed somebody’s RAID array.Accidental deletionYou ranrm -rfin the wrong directory. RAID dutifully and instantly deleted the files from all redundant copies. The data is gone. Recovery requires a backup that has a previous version of the file before you deleted it.This is the #1 cause of data loss in homelabs. Not drive failures.Operators making mistakes.RansomwareMalicious code encrypts files on the NAS. RAID dutifully and instantly encrypts the files on all redundant copies. The data is intact (cryptographically) but unreadable to you. Recovery requires a backup made before the encryption happened.If your backup is on the same NAS - even on a separate dataset - ransomware can encrypt that too. If your backup is on a network share that the infected machine has write access to, ransomware can encrypt that too. Real backup must beinaccessible to the infected machine.Filesystem corruptionZFS and btrfs are checksummed; mdadm ext4 / xfs are not. With ZFS, a corrupted block is detected during read or scrub, and if you have parity (RAIDZ1), the bad block is reconstructed from parity.Usually. Edge cases exist:Two corrupted blocks in different drives in a RAIDZ1 stripe unrecoverable.A bug in ZFS itself (rare but real) corrupts data on write, replicated to all redundant copies.A bad RAM corrupted the data before ZFS computed checksums; the checksum is consistent but wrong.For non-checksummed filesystems (mdadm ext4 typical Linux RAID), corruption is silent - you find out months later when a file won’t open. Recovery requires a backup that has the file before the corruption.Fire / flood / theftThe NAS is in your basement; the basement floods. The NAS is in your office; thief takes it. The NAS is in the rack; the lithium-ion UPS battery vents and ignites. RAID’s redundancy is pointless when the entire array is destroyed simultaneously.This is whyoffsite backupis non-negotiable.Power surgeLightning strike near the building; the surge protector doesn’t fully absorb the spike; the NAS’s PSU and several drives die simultaneously. RAID can’t help if the failure is correlated across drives.Buggy controller / bad cableA bad SATA cable corrupts writes intermittently. ZFS catches some; mdadm ext4 doesn’t. Hardware RAID controllers fail in creative ways - some hardware RAIDs have lost arrays because the controller’s firmware buggily reordered writes during a power-loss event.Software bugsThe application writing to your NAS has a bug that overwrites the wrong file. Or your backup tool has a bug thatdeletesthe wrong file from the source. The NAS dutifully deletes; RAID dutifully replicates the deletion.This is whyversioned backupsmatter.Your own typossh nas01 rm -rf /tank/importantinstead of/tank/temp. The data is gone in the time it took to press Enter. RAID is irrelevant.Lost decryption keysYou set up encrypted ZFS datasets. You forgot to back up the encryption key (or you did back it up to a Vaultwarden instance that lives on the same NAS). The encryption key is gone; the encrypted data is mathematically unrecoverable.What backup actually meansBackup is aseparate, independent copy of the data, made at a specific point in time,inaccessible from the source system except via deliberate restore action.The four properties matter:Separate: Different storage; not just a different folder on the same array.Independent: Not dependent on the source system continuing to work; you can restore even if the source is gone.Point in time: Has a known snapshot date; you can restore to “yesterday” or “last month.”Inaccessible from source: A compromise of the source machine does not corrupt or delete the backup.A backup that fails any of these four properties is not really a backup.Common things that aren’t really backups:Acp -rto a different folder on the same drive - fails #1 and #2.Anrsyncto a second drive on the same machine - fails #2 in the fire/theft case.A snapshot of a ZFS dataset - fails #2 in the disk-pool-loss case.A backup mounted as a writable share that the source can write to - fails #4 (ransomware encrypts both).A cloud sync (Dropbox, OneDrive) - fails #4 (ransomware on local syncs to cloud, deletions sync).Things thatarereal backups:An encrypted snapshot pushed to a cloud bucket the source has write-only access to.A pull-style backup to a separate machine the source can’t reach (the backup machine reaches in to pull data).An offline drive that’s only mounted during backup runs.A versioned archive at a service like Backblaze B2 with object lock enabled.The 3-2-1 ruleThe simplest framework for “real backup” is 3-2-1:3 copiesof the data.2 different storage media(e.g., two different drives, or drive cloud).1 copy offsite(different physical location).Worked example for a homelab:Copy 1: Production data on the 16TB NAS. RAIDZ1 protects it from a single drive failure during normal operation.Copy 2: Nightly snapshot pushed via Borg / Restic / Duplicacy to a second drive on a separate machine (a Raspberry Pi 5 with a USB drive in a cabinet).Copy 3: Same Borg/Restic snapshot pushed to Backblaze B2 / Wasabi / S3 Glacier - different physical location.This satisfies 3-2-1: three copies, two media types (NAS HDD B2), one offsite (B2).The backup-stack scorer walks through this concretely; it scores each rule independently (3 copies, 2 media, 1 offsite, plus versioning, restore-tested, encryption-at-rest).What 3-2-1 specifically protects againstMap the failure modes from earlier to the layer that catches them:Failure modeRAID catches3-2-1 backup catchesSingle drive dies during normal operationYesYes (also restorable from backup)Two drives die simultaneously (RAIDZ1)NoYesAccidentalrm -rfNoYes (restore from previous version)Ransomware encrypts filesNoYes (if backup is inaccessible from source)ZFS bug corrupts data on writeNoYes (older backup pre-corruption)Bad RAM corrupts before checksumsNoYes (older backup pre-corruption)Fire / flood destroys the NASNoYes (offsite copy survives)Theft of the NASNoYes (offsite copy survives)Power surge destroys all drivesNoYes (offsite copy survives)Software bug deletes wrong filesNoYes (versioned backup)User typo with destructive commandNoYes (versioned backup)Lost ZFS encryption keyNo (data unrecoverable from RAID)Yes (older unencrypted backup, or backup of the key)RAID catches one failure mode. Backup catches the other twelve.This is why “I have RAID so I’m fine” is the single most dangerous belief in homelab storage.What versioning meansA backup that overwrites the previous backup with the new state isnot versioned. If you ranrsyncto your backup destination yesterday, and ransomware encrypted everything today, today’srsyncfaithfully copies the encrypted versions over your good backup. Now both copies are encrypted.Aversionedbackup keeps multiple historical states. Borg / Restic / Duplicacy all do this natively - each backup is a snapshot, distinct from previous snapshots, and old snapshots can be retained for any policy you want (last 7 daily, last 4 weekly, last 12 monthly, last 7 yearly is a common pattern).When ransomware hits, you restore yesterday’s snapshot, before the encryption. When you accidentally delete a file, you restore the version from before the deletion. When ZFS bug corrupts data, you restore from before the corruption.Versioning is the difference between “I have a backup” and “I have a useful backup.”What “restore-tested” meansA backup that has never been restored istheoretical. It might restore correctly. It might fail with corruption you didn’t notice. It might be missing the encryption key. It might restore to a directory you can’t access. It might take 14 hours to restore when you needed it in 30 minutes.AdvertisementYou don’t know until you actually restore.The restoring from Restic when everything broke walkthrough covers the practical disaster-recovery case. The discipline:Quarterly restore tests: pick a snapshot; restore to a test directory; validate the data.Annual full-restore drill: spin up a clean VM; restore everything you need to rebuild from scratch; confirm services come up.Document the procedure: in a place you can access without your homelab.Most homelab users skip this. Most also discover during a real disaster that their backup wasn’t restoring what they thought.What “encryption at rest” meansBackups have your data. If the backup is in the cloud or on a friend’s NAS, the storing party can theoretically read it unless it’s encrypted. Encryption-at-rest means the backup is encrypted before it leaves your control; the storage holder sees ciphertext only.Borg, Restic, and Duplicacy all do this natively - encryption is part of the protocol, the encryption key is yours. The storage backend (Backblaze B2, Wasabi, etc.) sees encrypted blobs.The catch:lose the encryption key, lose the backup. The encryption key needs the same backup discipline as the data - store in Vaultwarden / Bitwarden / Passbolt (which itself needs backing up - break the circular dependency by storing the password manager’s master password on paper).The “RAID is backup” antipatternsSpecific things people convince themselves count as backup but don’t:“I have RAIDZ2; I can lose two drives.”RAID handles two-drive failures. It doesn’t handle ransomware, deletion, or fire. Half a backup at best.“I have a snapshot.”A ZFS snapshot is point-in-time copy on the same pool. It’s amazing for “oh no, I deleted the wrong file 5 minutes ago” - fast recovery. It’s not backup. Lose the pool, lose the snapshots.“I have replication to another pool.”Better. But the replication target is reachable from the source; ransomware can encrypt both. Replication ! backup. Replication airgap (the target is read-only or only accessible during pull windows) closer to backup.“I have everything on Dropbox.”Cloud sync. Deletions sync. Encryptions sync. Not backup. Some cloud sync products have version history (Dropbox Rewind, OneDrive version history) - those count as backup, but for a limited time window. Don’t rely on a 30-day cloud history if you might miss the corruption for 60 days.“I have an external USB drive I plug in occasionally.”Closer. If you actually plug it in, copy data, and unplug it, that’s a real backup. The discipline matters; most people forget. Automate where possible.“I have RAID at home and at the colo.”Two RAIDs. Now you’re protected against a single drive failure in two locations. Still not protected against ransomware, deletion, or fire. RAID isn’t the right primitive for backup.What the minimum viable real backup looks likeIf you take everything above seriously, the minimum is:Production dataon whatever storage (RAID is fine; RAID isn’t required).Nightly versioned encrypted backupto a destination the source can write but not read or modify previous versions of. Restic with--read-onlyafter first use is one approach; B2 with object-lock enabled is another.Offsite encrypted backupto a different physical location. Cloud bucket usually; trusted friend’s NAS via Tailscale also works.Quarterly restore test.Encryption key stored separately from the data- your password manager paper backup.That’s the floor. Pair with Borg / Restic / Duplicacy, Backblaze B2 / Wasabi / Glacier, the 3-2-1 backup validator, and the restic recovery walkthrough.Common gotchas in backup designThe backup machine is on the same network.If ransomware spreads through SMB, both the production NAS and the “separate backup machine” get hit. Mitigation: pull-style backup (the backup machine reaches in to grab data, never the other way), or air-gapped backup (the destination is only reachable during scheduled windows).The backup credentials are stored on the source.If the source machine is compromised, the attacker has the backup credentials and can delete the backups. Mitigation: append-only credentials (B2 application keys can be append-only), or pull-style where source has no creds for the destination.Backup runs faster than you check it.Backup completes successfully in 30 minutes; you don’t notice it’s been failing for two months. Mitigation: alerting via Uptime Kuma / Healthchecks / Statping - backup script POSTs to Healthchecks on success; if it stops POSTing, Gotify / ntfy / Apprise / Pushover wakes you up.Bandwidth-bound initial seed.First backup of a 10TB library over a 100Mbps upload link takes 10 days continuous. Plan for this. Some providers offer “ship a hard drive” import to seed; or do the initial offsite to a friend’s NAS over your LAN, take it to their house, then pick up incrementals over the cloud.Cold-storage retrieval costs.Backblaze B2 / Wasabi / Glacier compares costs. S3 Glacier Deep Archive is cheap to store, very expensive to retrieve in a hurry. If your disaster scenario needs fast restores, don’t pick the cheapest tier.Backup of dynamic state without quiescing.Backing up a running database file produces a snapshot of inconsistent state - the file looks fine but the DB won’t open. Mitigation: ZFS or btrfs snapshots that are atomic; or usepg_dump/mysqldumpto export consistent dumps before backup.Forgetting to back up the boot drive.The NAS data is backed up; the OS configuration is not. Rebuilding TrueNAS / Unraid / Proxmox from scratch is a few hours of work. Mitigation: include/etc,/var, and config-file directories in the backup set, or back up the boot drive separately.The “do I really need this?” gateHonest framing: how much would it actually cost you if you lost everything in your homelab right now? Family photos, paperwork, media library, configurations, years of work?For most homelab users, the answer isa lot. Photos especially are irreplaceable.The cost of doing 3-2-1 properly: - B2 storage at 5TB (typical homelab): ~$30/year. - Pi USB drive for second-machine local backup: ~$80 one-time a 4TB drive. - Time to set up: 1-2 weekends. - Time to maintain (after setup): 30 min/quarter.The cost of not doing it: years of irreplaceable data, gone.This isn’t a marginal-value decision; it’s a basic-hygiene one. If you’re already running a homelab, you can run a backup system. The friction is mostly perceived.TL;DRRAID is not backup.RAID protects against drive failure; not against the dozen other failure modes.Backup is: separate, independent, point-in-time, inaccessible-from-source.3-2-1: 3 copies, 2 media, 1 offsite. Plus versioning, encryption at rest, restore-tested.The minimum viable: nightly encrypted versioned backup → offsite cloud bucket. Restore-test quarterly.The cost: ~$30/year a weekend to set up. The cost of not doing it: catastrophic.Most homelab data loss isn’t drive failure. It’s user error, ransomware, or fire. RAID solves none of those. Real backup solves all of them.RelatedBorg / Restic / Duplicacy - backup tool comparison.Backblaze B2 / Wasabi / Glacier - offsite destinations.check your backup architecture - score your backup policy.Restoring from Restic when everything broke - disaster recovery walkthrough.Hardware RAID / ZFS / btrfs / mdadm - RAID approaches comparison.RAIDZ1 vs RAIDZ2 vs Mirror - pool layout choices.RAIDZ rebuild calculator - failure-during-rebuild math.ECC vs non-ECC for ZFS - corruption layer.ZFS-aware drive health monitoring - early-warning layer.Predictive SMART attributes - drive failure prediction.Reading smartctl output - interpret SMART numbers.Smartctl interpreter - automated reading.Reallocated sectors when to replace - replacement threshold.Building a quiet 16TB NAS for under $800 - companion build guide.TrueNAS / Unraid / Proxmox - OS for the storage system.Uptime Kuma / Healthchecks / Statping - monitor backup-job success.Gotify / ntfy / Apprise / Pushover - alert on failures.Vaultwarden / Bitwarden / Passbolt - store encryption keys.Hardening a homelab Debian/Ubuntu box - reduce ransomware surface area.WireGuard / Tailscale / ZeroTier - air-gap pattern for backup machines.pfSense / OPNsense / VyOS - VLAN isolation between source and backup.Hetzner / OVH / Contabo / DigitalOcean - cheap VPS for offsite backup destination.CMR vs SMR - drives matter for backup pools too.WD Red Plus / IronWolf / Toshiba N300 - drives for backup NAS.USB-attached drives lying about SMART - caveats for offline-USB backups.Drive-power calculator - budget for the backup machine too.Big Iron’s NAS planner - sizing for the backup target.
返回列表