IBM TotalStorage Enterprise Storage Server Data Recovery
The IBM TotalStorage Enterprise Storage Server, widely known by its internal codename "Shark", was IBM's mainframe-and-open-systems enterprise array through the late 1990s and 2000s, later succeeded by the DS6000 and DS8000 families. A recovery evaluation on ESS-era hardware has to work with SSA loop-attached disk, logical subsystem structures, and — on mainframe-attached systems — CKD volume formats that differ fundamentally from open-systems FBA layouts.
Platform Lineage and Naming
IBM's TotalStorage Enterprise Storage Server (machine type 2105), models F20 and 800, was the flagship enterprise array bridging IBM's mainframe DASD heritage and open-systems SAN storage, attaching over ESCON and later FICON for System z hosts and Fibre Channel/SCSI for open systems. Internally, the array became known by the engineering codename "Shark".
IBM's enterprise array line continued from ESS through the TotalStorage/System Storage DS8000 family, which replaced Shark's SSA-loop design with switched Fibre Channel back-end connectivity while retaining CKD volume support for mainframe customers. The mid-range DS4000/FAStT line (originating from IBM's relationship with the former NetGuard/LSI FAStT technology) and the DS6000 filled complementary roles alongside ESS and its DS8000 successor rather than replacing it outright.
- IBM TotalStorage Enterprise Storage Server (ESS), machine type 2105
- Codename "Shark" — models F20, 800
- Predecessor to IBM TotalStorage/System Storage DS8000 (see the dedicated DS8000 page)
- Related mid-range lines: DS6000, DS4000/FAStT
Generations and Models We Evaluate
| Generation / family | Models |
|---|---|
| ESS 2105 | Model F20, Model 800 — Shark generation, SSA-attached back-end disk |
| Successor | IBM TotalStorage/System Storage DS8000 — covered on the dedicated DS8000 page |
| Related mid-range | DS6000, DS4000/FAStT |
Architecture and Data Layout
ESS uses redundant cluster processors, each with its own cache and non-volatile storage, connected to back-end disk over Serial Storage Architecture (SSA) loops rather than a switched fabric — drives are addressed by their position on a loop, and a loop-level fault can affect every drive behind it.
Physical disks are organised into RAID ranks (predominantly RAID 5, with RAID 10 options on later code), and ranks are carved into logical subsystems (LSSs) that group logical volumes for management and mainframe addressing purposes.
On mainframe-attached configurations, ESS presents Count Key Data (CKD) volumes emulating traditional IBM DASD geometry, addressed over ESCON or FICON channels; open-systems hosts instead see Fixed Block Architecture (FBA) LUNs over Fibre Channel or SCSI. CKD and FBA layouts are structurally different and require separate handling during reconstruction.
Protocols and formats: ESCON, FICON, Fibre Channel, SCSI, CKD (mainframe), FBA (open systems)
- RAID ranks (typically RAID 5, later RAID 10) provide the underlying protection; a rank's SSA loop topology means a single severe loop fault can present as multiple simultaneous drive failures.
- Logical subsystems (LSSs) group logical volumes for addressing and mainframe device-number purposes, and must be correctly mapped when reconstructing which volumes belonged to which host.
- CKD volumes emulate mainframe DASD track/record geometry and are unrelated in structure to open-systems FBA LUNs, even where they sit on the same physical ranks.
Failure Scenarios
Logical failures
- Deleted or reformatted logical volumes within a logical subsystem
- CKD volume or VTOC corruption on mainframe-attached storage
- Failed microcode update affecting rank or LSS configuration
- Corrupted FBA volume tables on open-systems LUNs
- Mismanaged PPRC/FlashCopy relationship overwriting a volume
Hardware failures
- Multiple disk failures within a RAID rank beyond its parity tolerance
- SSA loop failure affecting all drives behind the break point
- Cluster processor or cache/NVS failure, including failed failover between clusters
- ESCON/FICON or Fibre Channel adapter failures affecting host connectivity
- Power or environmental failure affecting a full frame
What Not To Do Before an Evaluation
- Do not run rebuilds, reconstructions or re-initialisations against an array that has already lost more drives than its protection level allows.
- Do not recreate pools, aggregates, disk groups, storage pools or clusters — these operations write new metadata over the structures a recovery needs.
- Do not swap drives between slots, and do not reorder shelves. Record the original slot and shelf positions before removing anything.
- Do not run file-system repair tools against production volumes before the underlying storage layer has been evaluated.
- Do not restore a backup or replication set over the affected volumes until the recovery scope has been assessed.
- Do not eradicate deleted volumes or empty recycle/destroyed states on platforms that hold deleted data for a retention window.
Our Evaluation and Recovery Process
- Intake and platform identification — array model, generation, firmware, protection layout and the sequence of events that led to the failure.
- Read-only evaluation of the media and array structures, including assessment of drive health and the extent of any physical damage.
- Forensic imaging of all contributing media, with cleanroom work where drives require it. Originals are preserved unaltered.
- Reconstruction of the storage layer — pools, aggregates, parity groups, chunklets, extent groups or objects — from the images.
- Extraction of the layers above: file systems, virtual machines, databases, mailboxes and shares.
- Verification against a file list and customer-nominated critical data, followed by secure return on encrypted media.
Frequently Asked Questions
What is the difference between recovering CKD and FBA volumes?
CKD volumes emulate mainframe DASD track and record geometry, while FBA volumes use fixed-size blocks as on standard open-systems storage. They require different tooling and interpretation even when hosted on the same array.
Is ESS/Shark the same as DS8000?
No, though they are related in lineage. DS8000 succeeded ESS with a switched Fibre Channel back end in place of SSA loops. See the dedicated DS8000 page for that platform's specifics.
Does an SSA loop fault mean multiple drives have actually failed?
Not necessarily. A loop-level fault can make many drives appear inaccessible at once even though the drives themselves may be undamaged — an evaluation distinguishes loop faults from genuine multi-drive failure.