An HP ProLiant ML370 running a finance business on five 300 GB SAS drives. A member failed, the hot spare rebuilt the array automatically, and four months later a second disk failed into an array whose spare had never been replaced.
← All case files · from £500 + VAT
An HP ProLiant ML370 acting as the primary storage server for a finance business, with five 300 GB SAS drives in a Smart Array set and offices across Europe working off it daily. Everyday business data — not archives, not backups, the live working store.
Four months earlier a member had failed. The hot spare did exactly what a hot spare is for: it came online automatically, the array rebuilt onto it, and nobody had to do anything. That is the system working as designed, and it is also the moment the problem started.
Because the spare had been consumed, and nobody replaced it. The array carried on for four months with full redundancy on paper and none in reality. When a second drive failed, there was nothing left to absorb it.
The failure mode here is organisational rather than technical, and it is common enough to be worth naming.
A hot spare is a single-use resource. It sits idle until a member fails, then it becomes a member, and at that instant the array's spare capacity is gone. Everything still reports healthy because the array is healthy — it has simply used up the margin that made it safe.
Nothing about the day-to-day behaviour changes to signal that. The volume mounts, performance is normal, monitoring shows a rebuilt and optimal array. The only indication is a spare count of zero on a screen nobody looks at once the crisis has passed.
Four months is a very typical interval. It is long enough for the original failure to stop being fresh in anyone's mind and short enough that the drives — same batch, same age, same duty cycle — are still well inside the window where a second failure is likely.
All five drives arrived. The four working members were imaged first and returned to their carriers, which is standard practice: get clean copies of everything readable before touching the difficult one.
The failed drive turned out not to be mechanically damaged at all. It carried a well-documented Seagate SAS firmware fault — a defect in the drive's own controller code that leaves it unable to present itself correctly, while the platters and heads underneath remain in good order.
That is a considerably better position than a head crash, but it is not one that any imaging hardware can work around directly. The drive has to be persuaded to talk before it can be read, and that means addressing the firmware itself.
The array reconstructed completely and every file was returned. Because the failed member's problem was firmware rather than physical, there was no partial image and no gaps to explain — an outcome that would not have been available had the second failure been mechanical.
The more useful outcome was procedural. A hot spare that has been consumed is an open item, not a closed one, and the monitoring that reports an array as optimal will not tell you that. The client's own alerting was reporting exactly what it was designed to report.
Worth checking on your own arrays. If a hot spare has ever activated, confirm it has been replaced. An array running with a spare count of zero is a RAID 5 with no redundancy at all, and it will report itself as healthy right up until the moment it is not.
The moment a hot spare comes online it stops being a spare. Put replacement on the ticket that closes the original incident, not on a list for later.
Members bought together, installed together and worked identically tend to fail within months of each other. The window after a first failure is the riskiest period an array has.
A drive that will not present itself may be mechanically perfect. On SAS in particular, known controller-code defects account for a meaningful share of what looks like total failure.