RAID 5: One Dead Drive Is Fine, But Two? Data Recovery Guide

RAID 5: One Dead Drive Is Fine, But Two? Data Recovery Guide
Late last year, a printing factory in Lai Chi Kok called: their Dell server ran RAID 5, four 2TB drives, holding the entire factory's ERP data and client files. Morning shift found the server wouldn't boot — the RAID card reported two drives offline. The manager's voice shook: "One drive died yesterday, we planned to replace it this weekend, then the second died this morning." I told him straight: "RAID 5 with two dead drives is one of the toughest situations — but not hopeless." Here's how RAID 5 works, why two dead drives is so serious, and what can still be saved.
RAID 5 Basics: Trading One Drive's Space for Safety
Simple analogy: RAID 5 is like four people sharing accounts — each handles a portion, but everyone holds a "check sheet" recording the others' totals. If one person disappears, the remaining three can reconstruct their share from the check sheets.
Technically:
- Minimum 3 drives, typically 4-6
- Parity distributed across all drives (not concentrated on one — that's RAID 4, obsolete)
- Usable capacity = (N-1) × single drive: four 2TB drives give 6TB usable; one drive's worth goes to parity
- Survives one dead drive: any single failure is reconstructible from parity
- Two dead = theoretically unrecoverable: each stripe has only one parity copy; losing two data chunks can't be mathematically reconstructed
Why Two Dead Drives Isn't Always the End
Theory says RAID 5 with two dead drives is hopeless, but reality has variables:
- Not both fully dead: Usually one is completely unreadable while the other has partial bad sectors — maybe 80% still readable. We use the readable portions plus parity, reconstructing stripe by stripe, saving the majority.
- Different failure times: The first may have died three months ago, the array running degraded since, with old parity data. If the second died recently, we can reference pre-first-failure parity states.
- Not all stripes contain data: Drives aren't full everywhere. If bad sectors land on empty space, actual loss may be minimal. We analyze which stripes hold real data and focus there.
The Lai Chi Kok Case: How We Saved 70%
Four 2TB drives, numbered 0-3. Findings:
- Disk 0: Completely dead, motor not spinning (physical — needs cleanroom)
- Disk 1: Partial bad sectors, ~15% unreadable
- Disks 2, 3: Healthy
Our strategy:
- Disk 0 first (cleanroom): Opened in dust-free cleanroom, swapped heads, read platter by platter. Most time and cost, but it held unique data. Recovered ~70% of sectors.
- Disk 1 imaged, skipping bad sectors: Professional equipment auto-skips bad sectors without force-reading. Recovered 85%.
- Disks 2, 3 imaged directly: Healthy, fully read.
- Analyze RAID parameters: Dell PERC card, RAID 5, 64KB stripe, left-symmetric parity rotation (Dell default). Took half a day to confirm — wrong parity direction means garbage output.
- Stripe-by-stripe reassembly: Each stripe has 4 blocks (3 data + 1 parity). All four readable? Use directly. One unreadable? Calculate via parity. Two unreadable? Mark damaged, check if the filesystem can skip it.
- Mount NTFS, extract files: Windows server, NTFS filesystem. Two days checking MFT entries one by one. Saved ~70% of files. The ERP database main file opened; some old client images were damaged.
Eight days, $28,000 HKD (including cleanroom). The manager said, "70% beats zero — at least we're not closing down." They immediately bought two NAS units: one local RAID 6, one offsite at another factory.
RAID 5 vs RAID 6 vs RAID 10
After this case, many clients ask which RAID to use:
- RAID 5: One dead drive fine, two is disaster. Good for 4-6 drives balancing capacity and speed. Not for 8+ drives (rebuilds take too long, too risky).
- RAID 6: Survives two dead drives (dual parity). Good for 6+ drives, important data. Costs one extra drive's capacity and slower writes. We now recommend RAID 6 for all important data.
- RAID 10 (1+0): Mirroring plus striping, fastest, instant mirror takeover on single failure. But only half capacity (four 2TB = 4TB usable). Best for databases needing speed.
- RAID 0: No protection — one dead drive loses everything. Only for temp/cache or with solid backups. Never for important data.
My recommendation: drives are cheap now — go straight to RAID 6. One extra drive's cost buys peace of mind during rebuilds. That's our standard answer when designing storage.
Hardware RAID Card vs Software RAID
- Hardware RAID (Dell PERC, HP Smart Array): Config lives on the card; a dead card needs the same or compatible model to read. We keep various cards for testing, but very old models may be unfindable. Always note the model and firmware version.
- Software RAID (Linux mdadm, Windows Storage Spaces): Config in drive superblocks; readable on any machine. More straightforward recovery, no card dependency.
- Motherboard RAID (Intel RST): Half-hardware, half-software — the most troublesome. A motherboard swap may not recognize it; manual metadata analysis needed. We've seen many such cases requiring per-drive examination.
Honest Pricing and Success Rates
- RAID 5 one dead drive (standard rebuild failure): $5,000-$10,000 HKD, 90%+ success
- RAID 5 two dead (one fully, one partial): $15,000-$30,000, 60-80% success depending on bad sector distribution
- RAID 5 two dead (both need cleanroom): from $30,000, 40-60% success
- Dead RAID card (drives fine): $3,000-$6,000, 95%+ success (easiest)
Free diagnosis and firm quotes on every job. Cleanroom work is done in our own dust-free lab — few in Hong Kong can do this.
FAQ
1. Q: Can I keep using the server during RAID 5 rebuild?
A: You can, but shouldn't. Rebuild already stresses the drives; adding normal workload slows rebuild and increases second-drive failure risk. For production servers, schedule rebuilds at night or weekends and warn users it'll be slow.
2. Q: Can I convert from RAID 5 to RAID 6?
A: Yes, but it needs extra drives and migration. Most RAID cards support online migration, but it takes dozens of hours — no power cuts allowed. Safest: back up everything, rebuild as RAID 6 fresh, restore. We do these migrations; pricing depends on capacity.
3. Q: Any differences with SSD RAID 5?
A: Yes. SSDs have no mechanical parts — no spreading bad sectors — but have write endurance (TBW) and sudden death (unlike HDDs' warning noises). SSD RAID 5 rebuilds much faster, but firmware-bug deaths read as completely dead, same trouble as physical HDD death. SSDs need TRIM, but many RAID cards don't support it, affecting lifespan.
4. Q: Can cloud replace RAID?
A: Different things. RAID protects against hardware death; cloud backup protects against site-wide disaster (fire, theft, ransomware). Best to have both: local RAID 6 for daily operations, cloud for offsite backup. We always recommend both, never either/or.
Summary
RAID 5 isn't bad, but understand its limit: it only protects against one drive. Drives keep getting bigger, rebuilds take longer, and second-drive-during-rebuild deaths get more likely. If you're still running RAID 5 for important data, seriously consider RAID 6 — or at minimum ensure solid offline backups. When disaster strikes, power off immediately: WhatsApp +852 6558 6806 for a free assessment.
Further reading: NAS buying guide, data backup services.
Found this article helpful?
Feel free to share it with your friends or colleagues.