Unlike a clicking drive, there is no clock running on the hardware. The clock that is running is the one on well-intentioned interventions — rebuilds, re-creations, restores — and each of them writes to an array whose description of itself is already unreliable.
The arithmetic differs by level but the sequence never does: stop the array, preserve the bay order, image every member read-only, and reconstruct from the copies. Only the last step depends on which level you have.
Redundancy works by absorbing failure silently. When the first disk goes, the array carries on serving data exactly as before — no interruption, no visible error, usually nothing but an alert email nobody read or a light on a chassis in a cupboard.
Weeks or months later a second disk fails and the volume disappears. That is the failure people report, and it is why so many arrays arrive here having run unprotected for far longer than anyone realised.
The practical consequence matters: by the time you know there is a problem, the margin is already gone. And the remaining disks are the same age, from the same batch, with the same hours on them as the ones that already died.
A rebuild reads every sector of every surviving member and writes across the replacement. On a healthy array that is routine. On one that has been degraded for weeks, it is hours of sustained full-surface load on disks that are demonstrably at the end of their life.
That load is precisely what causes the second failure. The array survives its first disk loss and dies during the rebuild that was supposed to fix it, and this sequence accounts for a large share of the total losses we see.
If the data matters more than the downtime, image first and rebuild afterwards. If the downtime genuinely matters more, that is a legitimate business decision — but make it knowingly rather than by default.
Every disk, including the failed one. Especially the failed one. On RAID 5 with two members out, the first failure is frequently the more readable of the two and often supplies the missing parity.
Bay order, written down. Photograph the front of the chassis before anything comes out. Order is part of the array geometry and while it can be derived, having it saves real time.
The controller, if you have it. Not essential — the configuration is derivable from the disks — but knowing the make and generation shortens the work considerably.
An honest account of what was tried. If somebody attempted a rebuild, cleared a foreign configuration or initialised a disk, say so. It changes where we look first and nobody here will think less of you.
Usually not. The second disk has often been ejected for a timeout rather than failing outright, and its content may be almost entirely intact. Both are imaged and the array assembled from the images.
Common after a chassis move or backplane fault. Do not accept the prompt to import or clear — that writes new metadata over the old. Send the disks as they are.
Expected, and it is most of the work. Disk order, stripe size, parity rotation and start offset are all derived by finding the arrangement where the parity mathematics holds.
Please do not. It is the highest-risk operation available on a degraded array. Power it down and decide with the disks safe rather than under load.
Five to ten working days. Each member is imaged individually before reconstruction begins, and on large modern disks that imaging accounts for most of the elapsed time.
From £500 +VAT regardless of level, member count or controller, fixed in writing after the free 48-hour diagnostic.