NAS RAID Rebuild Failed? Stop Trying Immediately!

NAS RAID Rebuild Failed? Stop Trying Immediately!
Last month, an ad agency in Kwun Tong had a NAS with one failed drive. Their IT guy swapped in a new drive and hit rebuild. It stuck at 67%, then "Rebuild failed." He thought "let me try again" — second attempt failed at 41%. Third try, the entire RAID vanished. Four drives held 8TB of client footage and project files. When they brought it in, I told him: "Every attempt was flogging the three dying drives even harder." Here's why rebuilds fail and how to recover after failure.
Why Do Rebuilds Fail?
Many assume rebuilding is simple: one drive dies, swap in a new one, the system recalculates. Reality is less kind:
- The surviving drives have bad sectors too: The most common cause. Four drives bought together, spinning for four years — when the first dies, the other three are close behind. Rebuild reads all three end-to-end; hitting bad sectors stalls or fails it. What started as "one dead drive" can become "two dead drives."
- The replacement drive is faulty: New doesn't mean good. We've seen brand-new drives with bad sectors out of the box, or slightly smaller actual capacity (same 4TB label, a few hundred fewer sectors) — rebuild fails writing at the tail end.
- Power loss / reboot mid-rebuild: During rebuild, the array is at its most vulnerable. A power cut, reboot, or fiddling at this moment interrupts the rebuild; worst case, the entire RAID metadata gets scrambled.
- RAID card / NAS motherboard issues: Sometimes it's not the drives — the controller is unstable. Rebuild is compute-intensive; a flaky controller reveals itself exactly then.
Why You Must Not Retry After Failure
This is what most clients don't understand: "How will I know if I don't try?" Let me explain:
- Every rebuild is a full-disk scan: A 4TB rebuild reads the entire drive — 8-12 hours. If the drive has physical issues, each full read is torture; bad sectors spread.
- The failed attempt already wrote partial data: Failing at 67% means 67% of rebuilt data is on the new drive. Retry with a different order or drive, and the system may use this "half-baked" data, corrupting parity calculations.
- RAID metadata may already be damaged: After repeated failures, the NAS may mark the array "damaged" or "offline." Further attempts only make the state messier.
Professional Recovery: How We Saved the Kwun Tong Case
Four original drives plus the replacement — all five brought in. Our process:
- Image everything: All five imaged with dd. We found that besides the first completely dead drive, another had 200+ bad sectors — the real cause of rebuild failure. Lucky they stopped early; a few more attempts could have killed that drive too.
- Map bad sector distribution: Marked bad sector locations — were they clustered or scattered? Scattered but not too many; parity could still reconstruct.
- Virtual reassembly, skipping bad sectors: Reconstructed stripe by stripe using parity from three healthy drives plus readable portions of the bad drive. Never force-read bad sectors (the head crashing into them can scratch the platter).
- Verify file integrity: After reassembly, we don't just hand it over. We sample photos, videos, documents — confirm they open, aren't corrupted. A few video tails were damaged, but the bulk was saved.
Recovered ~92%; losses were scattered files on bad-sector locations. Five days, $9,500 HKD. The IT guy learned: "Rebuild isn't just pressing a button."
Preventing Rebuild Failure
- Check S.M.A.R.T. regularly: Monthly drive health checks in the NAS. Replace drives showing bad-sector growth while rebuilds still work — don't wait for death. Both Synology and QNAP have scheduled checks; enable them.
- Don't batch-replace same-batch drives: Four drives bought together die together. Stagger purchases, or replace one at a time, waiting for rebuild completion before the next.
- Back up before rebuilding: If possible, copy critical data out before rebuilding. Rebuild is high-risk; a second copy removes the fear.
- Use NAS-rated drives: WD Red, Seagate IronWolf — firmware optimized for RAID with TLER (time-limited error recovery) so one bad sector doesn't hang the entire rebuild. Desktop drives in a NAS fail rebuilds far more often.
- Don't touch during rebuild: No heavy reads/writes, no reboots, no power cuts, no fiddling. Let it run — typically 8-24 hours depending on capacity.
Already Failed Multiple Times — Still Recoverable?
Depends. If it's just failed rebuilds with nothing else done (no initialize, no RAID recreation, no format), recovery odds are good. We:
- Image all drives (including the rebuild-attempt new drive)
- Analyze each drive's RAID superblock, confirming disk order and last good array state
- Ignore partial rebuild data on the new drive; reassemble from the original four
- Reconstruct stripe by stripe with parity for bad sectors
But if after failed rebuilds someone clicked "recreate RAID" or "initialize," it's much harder — new RAID metadata overwrites the old, requiring manual reverse-engineering of previous parameters. Not impossible, but more time and cost.
Pricing Reference
- Single rebuild failure, intact RAID structure: $5,000-$8,000 HKD, 3-4 days
- Bad sectors needing stripe-by-stripe calculation: $8,000-$15,000, 5-7 days
- Multiple physical drive failures: from $15,000, cleanroom work extra
FAQ
1. Q: Rebuild stopped halfway due to power cut. Now it says RAID damaged. What now?
A: Don't power on again. A mid-rebuild array is in a "half-written" state; booting may trigger auto-repair treating half-baked data as real. Bring it in — we'll analyze on images and find the last consistent state before interruption.
2. Q: Can I use a larger drive for rebuild?
A: Yes, with conditions: the new drive must equal or exceed the original's actual sector count (not label capacity). Same brand/series is preferable — similar firmware behavior means higher rebuild success.
3. Q: RAID 6 (dual parity) means no rebuild worries, right?
A: RAID 6 survives two dead drives, so rebuilds have an extra parity layer — relatively safer. But not immune: three simultaneous problem drives, or a second death during rebuild, still kills RAID 6. And RAID 6 rebuilds are slower (dual parity calculation).
4. Q: Any way to completely avoid rebuild risk?
A: No such thing as zero risk, but minimize it: (1) use good drives, (2) check S.M.A.R.T. regularly, (3) back up before rebuilding, (4) most important — replace at the first sign of bad-sector growth, don't wait for death. Most clients are "no coffin, no tears" — then they pay recovery fees.
Summary
Remember: one rebuild failure = stop. Don't retry, don't reboot, don't click randomly. Bring all drives in as-is for a free check — we'll tell you if it's recoverable and what it costs. WhatsApp +852 6558 6806; we handle several of these cases monthly.
Want to understand RAID basics? See our NAS buying guide. For backup questions, see data backup services.
Found this article helpful?
Feel free to share it with your friends or colleagues.