Call us — 01223 655015
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · More Failures Than the Design Allows

Simultaneous Failures After a Power Event Are Usually Drops, Not Deaths

This enquiry reports a situation that should not be survivable and frequently is. A four-disk array using a scheme that tolerates two failures, where after a weekend power cut "three disks have failed and the fourth is in a warning state", with around 2TB held. Three genuine simultaneous failures would be extraordinary — and after an abrupt power loss, what usually happens is that the controller stopped trusting members that are perfectly readable.

MediaFour-disk array on a network appliance using dual-parity redundancy, approximately 2TB held — three members reported failed and one in a warning state following an uncontrolled power loss
Reported situationPower failure occurring over a weekend · appliance losing power without controlled shutdown · three of four array members subsequently reported as failed · remaining member reported in a warning state · array not presenting · image and application content sought
Fault classMembers marked failed by the controller following abrupt power loss — genuine simultaneous failure of three members improbable; member readability to be established independently
Equipment usedReported failure counts treated as controller state rather than as confirmed member condition · no rebuild, initialisation or member replacement permitted · each member imaged individually write-blocked under strict per-sector timeouts · genuine condition established per member from the imaging results · array geometry reconstructed from member metadata and assembled offline from the images

The decode: why the count is probably wrong, and the action that would make it true

Why three simultaneous genuine failures is implausible: drives fail independently. Three separate mechanical or electronic failures occurring in the same moment is vanishingly unlikely, and a common cause is far more probable.

What that common cause usually is: the event itself. An abrupt power loss can leave members with inconsistent metadata, and a controller that cannot reconcile them marks them failed rather than risk using them.

Why marking them failed is reasonable behaviour: caution. A controller unable to establish which members are current is right to refuse rather than to guess, and refusing looks identical to failure.

What else produces this after a power event: the power itself. A supply that browned out rather than cutting cleanly can leave drives in an unresponsive state that a full power cycle clears.

Why the warning state on the fourth is consistent with all this: it is a different threshold. Warning generally reflects error counters rather than an inability to communicate.

Why the redundancy scheme matters to the reading: it tolerates two. The array survives losing any two members, so if even two of the three are readable the volume reconstructs completely.

What determines that: reading them individually. Each member is imaged on its own, outside the appliance, and its genuine condition established from what it returns.

Now the action that would convert a recoverable situation into a real loss: replacing a member and rebuilding. A rebuild writes parity across the surviving members, and with the array already beyond its stated tolerance the calculation is performed from an incorrect picture.

Why initialising is worse still: it discards the arrangement. An appliance offering to set up a new array will write fresh metadata over the descriptors a reconstruction depends on.

What must happen now: the appliance stays off and the members are labelled by bay. Order matters to the reconstruction, and it is easily lost once drives are out.

On the bench

Reported failure counts were treated as controller state rather than as confirmed member condition — drives failing independently, so three simultaneous genuine failures is vanishingly unlikely against a common cause, while abrupt power loss leaves members with inconsistent metadata and a controller unable to establish which are current marks them failed rather than guessing. Dual-parity redundancy tolerates two absent members. Each member was imaged individually and its condition established from the results.

The outcome

Reported counts treated as controller state, no rebuild or initialisation permitted, and each member imaged individually before offline assembly. Free assessment, one fixed written figure including VAT, charged per drive, with 50% of parts and labour upfront where a drive has to be opened. The decode: three drives do not fail in the same instant. The power event left the controller unable to establish which members were current, and marking them failed is what a cautious controller does when it cannot tell.

An array reporting more failures than it can tolerate

Switch the appliance off, label every drive with the bay it came from, and refuse any offer to replace a member, rebuild or set up a new array. Read the count sceptically: drives fail independently, so three at the same moment is vanishingly unlikely next to a common cause, and an abrupt power loss leaves members with inconsistent metadata that a cautious controller responds to by marking them failed. Refusing to use a member looks identical to that member having died.

More disks reported failed than your array allows?
Power it down — call Cambridge Data Recovery on 01223 655015; reported counts treated as controller state, each member imaged individually, geometry reconstructed and assembled offline.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.