RAID, ZFS, DRBD or Ceph:
what each one gives you, what it costs
These four technologies are often compared as if they answered the same question. They do not: each one absorbs a different failure, is paid for in a different currency — disks, servers, network, operating hours — and some of them stack while others work against each other.
Contents
1. Three failure levels, not four competitors
The right question is not "which one is best?" but "what can fail without the service stopping?". Asked that way, the comparison sorts itself out.
| Technology | What can fail | Minimum servers |
|---|---|---|
| RAID, ZFS mirror or RAIDZ | One disk (two with RAID 6 / RAIDZ2) | 1 |
| ZFS replication | One server, losing the writes made since the last copy | 2, plus a witness |
| Synchronous DRBD | One server, with no acknowledged write lost | 2, plus a witness |
| Ceph | One disk, one server, even a rack depending on where the copies are placed | 3, preferably 4 or 5 |
None of the four is a backup
All of them faithfully copy whatever is written to them, including a deletion, a table emptied by mistake or files encrypted by ransomware. Redundancy protects against hardware that breaks, not against a write that does damage. That takes a separate, versioned copy kept out of reach: the job of an offsite backup of your Proxmox VMs, a different subject from this one.
2. RAID: protecting the disks of one server
RAID spreads data across several disks of the same machine so that a dead disk does not stop it, either through a hardware controller or in software (mdadm on Linux). It is the foundation of almost every server, and also the one most often overestimated: it protects the disks, not the server. To compare the usable capacity of each level, our RAID calculator works it out for your number and size of disks.
What it gives you
- • Losing a disk (two with RAID 6) without downtime or urgent intervention
- • No dependency on the network: everything happens inside the machine
- • Transparent to the operating system and applications
- • Good performance with RAID 10, disks as close to the CPU as they get
- • Mature, well-tooled technology that every administrator knows
What it costs
- • Half the capacity with RAID 1 and 10, two disks with RAID 6
- • Motherboard, power supply, controller and the server itself remain single points of failure
- • Rebuilds lasting hours on large disks: a RAID 5 is exposed to a second failure during that window, hence RAID 6 beyond a few terabytes
- • With hardware RAID, a dependency on the controller: reading the disks back needs a compatible model, and the write cache depends on its battery
3. ZFS: checking what it reads, and replicating
ZFS plays two roles that are worth separating. Inside a server, it replaces RAID: mirror or RAIDZ, managed by the file system itself, with no controller. Between two servers, it can send a disk's changes to another node, and Proxmox VE storage replication is built on that mechanism. To size a pool, our ZFS RAIDZ capacity and PBS storage calculator gives the usable capacity for the layout you choose.
Locally: a RAID that knows what it reads
What it gives you
- • Every block carries a checksum: data corrupted on one disk is detected on read and repaired from the other copy, where a classic RAID sees nothing
- • Near-free snapshots, handy before an upgrade
- • On-the-fly compression, which often wins back part of the space
- • No proprietary RAID controller: the disks can be read on any server
What it costs
- • RAM for its cache: at install time, Proxmox VE gives it 10% of memory, capped at 16 GiB, to be balanced against the VMs
- • Direct disk access: no RAID controller with its own cache underneath
- • Performance that degrades as the pool gets close to full
- • A RAIDZ that can grow one disk at a time only since OpenZFS 2.3; with mirrors, you add pairs
Between two servers: scheduled replication
Proxmox VE copies the virtual machines' disks to another node at a regular interval: every fifteen minutes by default on Proxmox VE, and the cron-style schedule goes down to once a minute or up to once a week. If the node fails, the machine restarts on the other one from the last copy. It is asynchronous: writes made since that copy are lost, up to a whole interval — a quarter of an hour with the default setting, a day or a week if replication runs daily or weekly. Writes themselves stay local and pay no network latency.
Two servers and a small witness are enough, with no dedicated storage network. For many SMBs it is the best ratio between what you protect and what you pay, provided you accept that loss window; we go through that trade-off in the case without shared storage.
4. DRBD: a disk mirror between two servers
DRBD does between two machines what RAID 1 does between two disks: every write is sent to the other server. In synchronous mode ("protocol C"), it is only acknowledged to the application once written on both sides. If the active server fails, the other one holds every acknowledged write, and a cluster manager such as Pacemaker restarts the service there.
What it gives you
- • No acknowledged write lost when a server fails, in synchronous mode
- • Two servers are enough, plus a lightweight witness for quorum
- • Reads from local disks, as fast as local
- • Standard servers, no shared storage array
- • An asynchronous mode for replicating to a remote site
What it costs
- • A split-brain risk: without fencing and quorum, both nodes may each write on their own
- • Every write waits for the network round trip: a dedicated, fast, short link is essential
- • Active/passive in practice; dual-primary requires a cluster file system
- • A failover that restarts the service: the outage is short, but real
- • On Proxmox VE, no native integration: it goes through LINBIT's LINSTOR plugin
The latency cost is what decides. It quickly becomes prohibitive once the two servers move apart; we put a number on it, write by write, in the latency constraint of synchronous replication.
5. Ceph: distributed storage that heals itself
Ceph pools the disks of several servers into a single storage system and places each piece of data in several copies on different servers — three by default. When a disk or a node disappears, the cluster recreates the missing copies elsewhere on its own. Proxmox VE integrates it directly, user interface included, which makes it the natural shared storage for a hypervisor cluster.
What it gives you
- • No single point of failure: disk, server, and even rack or room if copies are spread across them
- • Automatic healing, spread across the whole cluster
- • Shared storage: VMs live-migrate without copying their disks, and restart on any node
- • Gradual growth, by adding disks or servers
- • Block, file (CephFS) and S3-compatible object storage from one cluster
What it costs
- • Three servers as a bare minimum; with three, losing a node leaves the cluster no room to heal, hence 4 or 5 in practice
- • A third of raw capacity with three copies, and less once you keep room to absorb a lost node; the fill warning triggers at 85%
- • A dedicated network of at least 10 Gbit/s, more with NVMe disks
- • Higher write latency than a local disk, noticeable for very transactional databases
- • Carefully chosen hardware: SSDs with power-loss protection, 4 GiB of RAM per disk by default and 8 recommended, an HBA rather than a RAID controller
- • Demanding operations: data placement, recoveries, version upgrades
6. What each one gives, what it costs
| What it gives | Usable capacity | Hardware | Writes | Operations | |
|---|---|---|---|---|---|
| RAID 1 / 10 | Losing one disk | 50% | 1 server | Local | Light |
| RAID 6 | Losing two disks | (n−2)/n | 1 server | Penalised | Light, long rebuilds |
| ZFS mirror / RAIDZ2 | Losing one or two disks, and detecting corrupted data | 50% / (n−2)/n | 1 server, RAM for the cache | Local | Light |
| ZFS replication | Losing a server, minus the writes since the last copy (1 minute to 1 week depending on the interval) | 25% (mirror on two nodes) | 2 servers + witness | Local | Light, built into Proxmox |
| DRBD on RAID 1 | Losing a server, with no acknowledged write lost | 25% | 2 servers + witness, dedicated link | Wait for the network | Medium: fencing and quorum |
| Ceph, three copies | Losing a disk or a server, automatic healing, live migration | 33%, minus headroom | 4 to 5 servers, 10/25 Gbit/s network | Wait for three copies | Heavy |
Capacity percentages are arithmetic on total raw capacity, not measurements. Each row costs more than the one before, in hardware as in operations: the only way to know which one is justified is to set it against the price of an hour of downtime and the amount of data you are prepared to lose.
7. Which ones stack
| Combination | Verdict | Why |
|---|---|---|
| ZFS mirror + ZFS replication | Yes | The typical SMB setup: each node survives a disk, the whole survives a node |
| RAID 1 + DRBD | Yes | The classic use: RAID covers the disk, DRBD covers the server, and a dead disk does not trigger a failover |
| RAID 1 for the Ceph nodes' system disks | Yes | System disks hold no Ceph data: no objection |
| Ceph on RAID sets | No | See below |
| ZFS on hardware RAID | No | ZFS sees a single disk: it detects corrupted data but has no copy left to repair it, and the controller cache undermines its write guarantees |
| DRBD + Ceph | Pointless | Both replicate between servers: you would pay twice for the same protection |
Ceph on RAID 1 sets: a false good idea
It is the most common mistake in first deployments, and the Proxmox VE documentation advises against it in so many words: Ceph handles redundancy itself, and a RAID controller improves neither performance nor availability.
- • You pay for redundancy twice. Three Ceph copies on RAID 1 disks means six copies of every piece of data: one sixth of raw capacity is left usable, minus headroom. The same disk budget with direct access stores twice as much.
- • Ceph no longer sees the disks. It is Ceph that spots a failing disk, takes it out and recreates the copies elsewhere. Behind a controller, it sees a "healthy" volume while a disk degrades, and a repair spread across the whole cluster becomes a slow RAID rebuild inside a single machine.
- • The controller cache undermines the guarantees. Ceph assumes an acknowledged write is on disk. With a write cache whose battery is worn, a power cut can lose writes that Ceph has already acknowledged to the virtual machines.
- • Two repair mechanisms get in each other's way. During a RAID rebuild, the Ceph disk concerned stays in service but slow, and slows down the whole cluster without Ceph knowing why.
- • The "RAID 1 underneath, two Ceph copies on top" variant is worse. With two copies, losing a node forces a choice between blocking writes or running on a single copy, which the Proxmox VE documentation explicitly rules out: one more failure and the data is gone.
Good practice: an HBA, or a controller switched to HBA ("IT") mode, with the disks presented as they are. On a controller that cannot do that, the usual workaround is one RAID 0 per disk with the cache disabled — for lack of anything better, not by choice.
8. What we recommend, case by case
One server, no high availability
A ZFS mirror rather than hardware RAID: the disks stay readable elsewhere and corrupted data is detected. Offsite backup does the rest.
Two servers, a few minutes of loss acceptable
A ZFS mirror on each node, ZFS replication between the two, a lightweight witness for quorum. It is the cheapest setup that survives losing a server.
Two servers, no acknowledged write may be lost
Synchronous DRBD on RAID 1, driven by Pacemaker, over a dedicated link between two nearby servers. For a database on its own, replication at the engine level is often a better fit: it only pays the latency when transactions commit.
A growing virtualisation cluster
Ceph, from four or five nodes, with a dedicated storage network and disks in direct access. It is the option that takes the most operating effort, and the one that gives the most back: live migrations, automatic healing, growth without a redesign.
The right level depends on what downtime costs you, not on the most complete technology. Start with the order of magnitude from our downtime cost calculator: it tells you whether the third server, or the 25 Gbit/s network, pays for itself.
Frequently asked questions
Is RAID enough for high availability?
No. RAID protects the disks of a server, not the server itself: a power supply, motherboard or controller that fails stops the machine, whatever the RAID level. High availability starts when the data also exists on a second server, through ZFS replication, DRBD or Ceph.
Does ZFS replace a RAID controller?
Yes, and it brings more: every block carries a checksum, which makes it possible to detect and repair corrupted data that a classic RAID would read back without noticing. In return, ZFS wants direct access to the disks, with no RAID controller and its own cache underneath, and RAM for its cache.
How many servers does Ceph need?
Three as a bare minimum, which is what Proxmox VE requires. But with three servers and three copies, losing a node leaves the cluster nowhere to recreate the missing copy: it keeps running, degraded, until the node comes back. From four or five servers, it heals itself. Below three, ZFS replication or DRBD is a better fit.
Can Ceph run on RAID disks?
Technically yes, but it is a mistake. Ceph handles redundancy itself: on RAID 1, every piece of data exists in six copies, Ceph no longer sees the real state of the disks, and the controller cache can lose writes that Ceph believes are already on disk. The Proxmox VE documentation explicitly advises against it. You need an HBA and disks in direct access.
DRBD or ZFS replication: which one on two servers?
It depends on how much data you are prepared to lose. ZFS replication is asynchronous: simple, no write penalty, but a failure loses what was written since the last copy: up to fifteen minutes with the Proxmox VE default, more if replication runs daily or weekly. Synchronous DRBD loses no acknowledged write, at the cost of network latency on every write and a more demanding setup, with fencing and quorum.
Does redundant storage remove the need for backups?
No. RAID, ZFS, DRBD and Ceph faithfully copy everything written to them, including a deletion or files encrypted by ransomware. They protect against hardware that breaks. Against mistakes and attacks, only a separate, versioned copy kept out of reach lets you go back.
Would your storage survive losing a server?
We audit what you have, we identify what is truly redundant and what is not, and we quote the level that matches your cost of downtime.
Get a quoteFurther reading
The 4 high-availability architectures
What we deploy, from 4 hours to 10 seconds of failover
Hypervisor failure
Quorum, shared storage, fencing: restarting VMs in 5 minutes
Service cluster: keepalived or Pacemaker
Failing the service over, not just the machine
Disaster recovery with Proxmox
Ceph stretched across two sites, recovery objectives
Sources
The hardware recommendations and default values quoted on this page come from the following primary documentation. Links checked on 29 September 2026.
- [1] Proxmox VE — Deploy Hyper-Converged Ceph Cluster (node count, network, memory per OSD, avoid RAID, min_size)
- [2] Proxmox VE — ZFS on Linux (direct disk access, ARC cache)
- [3] Proxmox VE — Storage Replication (replication intervals)
- [4] OpenZFS 2.3.0 — release notes (RAIDZ expansion)
- [5] LINBIT — DRBD 9 User Guide (protocol C, quorum tiebreaker)
- [6] LINBIT — LINSTOR User Guide (Proxmox VE plugin)