In one sentence: PBS splits each backup into chunks, identifies each chunk by the SHA-256 fingerprint of its content, and stores each identical chunk only once in a datastore. A backup remains a complete restore point, even if almost all of its chunks already existed.
The documentation puts it this way: “multiple indexes can reference the same chunks, reducing the amount of space needed to contain the data (even across backup snapshots)” (Technical Overview).
1. Two ways to split: fixed and dynamic chunks
VM disks: fixed 4 MiB chunks
For a block-based backup (VM disk, .img image), PBS splits the content into chunks of the same length: “typically 4 MiB”, says the documentation. In the code it is more than typical: the backup API has never created a fixed index with another size, and since PBS 4.2.8 (October 2, 2026) a fixed index of any other size is rejected on read and on write (pbs-datastore/src/fixed_index.rs). The last chunk of a disk can be smaller.
In practice, the --chunk-size option of proxmox-backup-client is useless for an image. Our bench found it before we read it in the code: below 4096 KB, every .img archive fails.
Files and containers: dynamic chunks from 1 to 16 MiB
For a file backup (pxar archive: LXC containers, servers backed up with the client), PBS places chunk boundaries according to the content, with a rolling hash: “We use a variant of Buzhash” (Technical Overview). An insertion in the middle of a file therefore does not shift all the following chunks.
- Default target average size: 4 MiB.
- Actual chunk size: from a quarter to four times the average, so 1 to 16 MiB by default (
pbs-datastore/src/chunker.rs). --chunk-size(64 KiB to 4 MiB, power of 2) sets this average, not a maximum.
Each chunk, fixed or dynamic, is then compressed with zstd, when compression makes it smaller, before being encrypted if needed and written. For what follows, remember that compression adds to deduplication.
2. Scope: the whole datastore, but only one
A datastore's chunks live in a single .chunks/ directory, where a chunk's path depends only on its fingerprint. A chunk that is already there is not written again, whatever machine or namespace sends it. The documentation even presents namespaces as a way to reuse “a single chunk store deduplication domain for multiple sources” (Terminology).
- Storage: deduplication across the whole datastore, all machines and namespaces combined.
- Network: the client downloads the chunk list of the previous snapshot of the same group and does not upload the chunks it finds there. The first backup of a new VM therefore transfers its chunks, even though the server keeps a single copy.
- Across datastores: nothing. Each one has its own chunk store.
3. Encryption and deduplication
An encrypted chunk cannot be identified by the fingerprint of its encrypted content, otherwise nothing would deduplicate. PBS computes the fingerprint on the plain data, concatenated with a key derived from your encryption key (in the code: SHA-256(data ‖ id_key), with an id_key derived by PBKDF2, pbs-tools/src/crypt_config.rs). The documentation gives the reason: two identical chunks encrypted with different keys produce two different fingerprints (Encrypted Chunks).
- Same data, same key, several machines: normal deduplication.
- Two different keys: no deduplication between them.
- Encrypted and unencrypted backups: no deduplication. Turning encryption on later therefore sends all the data once more.
The client only encrypts the chunks it actually has to upload: deduplication is decided on your machine, and the server only receives encrypted chunks. At NimbusBackup: key not requested for backup hosting only; it stays with you. To create and keep it, see the PBS encryption key.
4. Reading the deduplication factor
In the PBS interface, on the datastore's Summary tab, the Stats from last Garbage Collection panel shows a Deduplication Factor. Three things to know to read it correctly:
- It is only computed at the end of a garbage collection, and kept until the next one. Before the first one, it is 1.00.
- It is a volume ratio: the cumulative logical size of all indexes of all snapshots (
index-data-bytes), divided by the real space of the chunks on disk (disk-bytes). It therefore includes compression. - Each snapshot counts as a full copy. For a VM, that is the whole virtual disk size, empty blocks included. The more snapshots a datastore holds, the higher the factor, without deduplication being any “better”.
On the command line, the status of the last garbage collection gives both raw values; the factor is their quotient.
proxmox-backup-manager garbage-collection status <datastore> --output-format json
# factor = index-data-bytes / disk-bytes5. Our measured figures
We do not publish a “typical rate”: it depends on your data. Here are two measurements we made ourselves, with their dates. Note that neither is the PBS factor described above: they are shares of data (or chunks) already present at the time of a backup.
MariaDB to PBS (SQL dumps, dynamic chunks averaging 1 MB)
Six databases, about 30 GiB of uncompressed dumps, one backup per night. Share of the dump's data already present on the datastore, computed as 1 − (new / total). Full details in our MariaDB to Proxmox Backup Server bench.
| Backup | Date | Already present | Note |
|---|---|---|---|
| 2nd, 7 h after the 1st | 29/09/2026 | 99.57% | 132.2 MiB of new data out of 29.77 GiB |
| 3rd, 24 h later | 30/09/2026 | 97.82% | 668.9 MiB new out of 29.93 GiB |
| After dropping a monthly partition | 02/10/2026 | 54.4% | the boundaries of extended INSERTs shift |
| Following night | 03/10/2026 | 98.69% | 36.0 MiB sent |
The 54.4% row is the most useful one: a routine database operation is enough to halve one night's deduplication. An average would have hidden it.
A 1 TB Windows server
On an 877 GB Windows server backed up with the Windows client, the 2nd run (18→19/04/2026) found 99.97% of its chunks already present, and only sent 1.73 GB. Here the percentage counts chunks, not bytes. Details in the 1 TB Windows case study.
6. Garbage collection: why space does not come back right away
A chunk shared by several backups cannot be erased as soon as one of them no longer needs it. PBS leaves this to garbage collection, in two phases (Maintenance):
- Mark: all indexes are read, and the access time (
atime) of every referenced chunk is updated. - Sweep: only chunks whose access time is older than the cutoff are erased. By default the cutoff is 24 h 5 min before the start of the garbage collection, or earlier if a backup is still running. Chunks within this grace period are reported at the end of the task as Pending removals.
The delay comes from the relatime mount option, which only updates a file's access time at most once every 24 h. It can be tuned per datastore with the gc-atime-cutoff option, in minutes (1 to 2880). In practice, a chunk that is no longer needed still takes up space until the garbage collection that runs more than 24 h after its last marking.
7. What it means for pricing at NimbusBackup
At NimbusBackup, you reserve a capacity in TB. What fills that reservation is the chunks, stored once and compressed, not the sum of your backups: deduplication decides how many restore points fit in the reserved TB. To size it, start from your data volume, then adjust the reservation after the first garbage collections, once the space actually used is known.
To compare plans and estimate your needs, see choose my backup.
Frequently asked questions
What deduplication rate should I expect with Proxmox Backup Server?
It depends entirely on your data, and we do not publish an average. On our MariaDB bench, 97.82% of a dump's data was already present 24 h after the previous one (30/09/2026), but only 54.4% the day after a partition was dropped (02/10/2026). The "Deduplication Factor" shown by PBS is a different indicator: it includes compression and counts each snapshot as a full copy.
Are two identical VMs stored twice in PBS?
No, if they are backed up to the same datastore: each identical chunk is stored only once there, across all namespaces and machines. However, the client only skips uploading chunks already present in the previous snapshot of the same group: the first backup of a new VM transfers its chunks, even though the server keeps a single copy. Two different datastores share nothing.
Does client-side encryption prevent deduplication?
No, as long as the key is the same. The fingerprint of an encrypted chunk is computed on the plain data and a key derived from the encryption key: the same data with the same key gives the same fingerprint, and deduplicates. With two different keys, or between an encrypted and an unencrypted backup, there is no deduplication at all.
A hosted PBS, sized on the real space
NimbusBackup hosts managed Proxmox Backup Servers: you reserve TB, deduplication and compression make your backups fit in them, and the encryption key stays with you. From €12 excl. VAT per TB per month.
