Servers rarely fail as servers. They fail as arrays, usually having lost more than one member behind redundancy that was still reporting healthy. The pressure of downtime creates an urge to act immediately, and most of the damage we see was done in that first hour by somebody trying to help. Power it down, change nothing, record the disk order, and ring. The diagnostic costs nothing and the figure that follows is fixed in writing.
$ cdr diagnose /dev/server → Server: Dell PowerEdge · RAID 5 · 6 × 2 TB → Status: ARRAY OFFLINE — 2 disks failed → Client: confidential · Cambridge CB2 $ cdr engineer-working → Member disks: all 6 imaged read-only → RAID 5: rebuilt off the controller → Windows Server: volumes mounted $ cdr verify → ✓ file shares — 6.1 TB → ✓ SQL databases — restored → ✓ server recovered — data back
Downtime creates pressure to act immediately. Almost every irreversible mistake we see on server jobs was made in that first hour by somebody trying to help.
A server with a failing array left running is spending what protection remains. Shut it down properly if you can, pull the power if you cannot, and stop there.
Bay order is part of the array structure. Before anything comes out, photograph the chassis and label each caddy. Two minutes here saves hours later.
Rebuilding onto a replacement disk loads every surviving member at full surface for hours. On disks of the same age and batch as the ones that failed, that is when the next one goes.
Reinstalling the operating system, running a file system repair or letting the controller initialise a foreign configuration all write over the structures we need.
The chassis is mostly irrelevant. What matters is the array inside it and how long it has been running without protection.
A server is a box around some disks and a controller. When it stops, the fault is nearly always in that array rather than in the machine, and the array has usually been degraded for considerably longer than anyone realised — the first disk failed silently, redundancy absorbed it, and nobody saw the alert because the alert went to a mailbox that no longer exists.
There is a second pattern particular to small estates, and it accounts for more total loss than disk failure does: the backup was running to a share on the same array. When the array went, the backups went with it. It is worth checking where yours actually writes to before you need to know.
Cache modules and battery packs. On HP ProLiant and similar, a failed cache battery leaves writes stranded in volatile memory and the controller refuses to bring the array online. It looks catastrophic and frequently is not — the disks are intact and the array is reconstructable.
Peterborough and Milton Keynes send corporate and back-office servers — claims and insurance systems, finance and payroll, document management under statutory retention. Chelmsford and Bury St Edmunds send professional practice servers holding case files and land records. These are conventional machines, conventionally maintained, and they usually fail through disk age rather than any event.
The third source is different in character. Process and production servers out of food manufacturing and packing across the fens and along the A14, running batch, recipe and traceability systems that a supplier audit will suddenly require. Those arrive with a deadline belonging to a contract rather than a project, and frequently on an operating system that went out of support years ago.
Server and array recovery starts at £500 +VAT regardless of member count, RAID level or controller. Database and mail store extraction from the recovered volume is included rather than charged separately.
Send the disks labelled by bay where you can, and the controller if it is available. Sending the whole chassis is fine and sometimes easier — tell us on the call and we will advise which suits your situation.
Leave it powered down. Power events commonly leave a cache module holding writes that never reached the disks, and the controller refuses to assemble the array as a result. The disks are usually intact and the volume reconstructable.
Common after a chassis move or backplane fault, and it does not mean the data is gone. Do not accept the prompt to import or clear the configuration — that writes new metadata over the old. Send the disks as they are.
More common than anyone would like, and it is why total loss usually involves a backup decision rather than a hardware one. The array is still recoverable; it just means the recovery is now the only route rather than a convenience.
Yes, and often that is the fastest route back to working. Once the volume is reconstructed, SQL, Exchange and file shares can be extracted individually so the critical system comes back before the rest.
Not usually. The disks and their bay order matter; the chassis and motherboard rarely do. Send the controller if you have it, since knowing the generation shortens the reconstruction.
The diagnostic is 48 hours and reconstruction typically five to ten working days, driven mostly by imaging time on large disks. Tell us if there is an operational deadline and we will be honest about whether it is achievable.
From £500 +VAT, with enterprise SAN environments quoted separately from £1,250. Out-of-hours collection across Cambridgeshire where a room full of people is waiting on it, and a named contact from diagnostic to handover.