# Difference between the available splitters

**URL:** <https://kopia.discourse.group/t/difference-between-the-available-splitters/894>\
**Category:** General\
**Created:** [January 31, 2022, 4:01pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894 "2022-01-31T16:01:26Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![lordraiden](https://yyz2.discourse-cdn.com/free1/user_avatar/kopia.discourse.group/lordraiden/32/309_2.png) [@lordraiden](https://kopia.discourse.group/u/lordraiden)\
**Post date:** [January 31, 2022, 4:01pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894/1 "2022-01-31T16:01:26Z")

</div>

Can some explain, or where I can find and explanation of between the different splitters available?

And how can I interpret this results?

 ![imagen](https://global.discourse-cdn.com/free1/uploads/kopia/original/1X/e1bceb2edf7f1dfee3a40d6085f168950c8143ed.png)

---

<div class="post-metadata">

**Author:** ![josef](https://avatars.discourse-cdn.com/v4/letter/j/ec9cab/32.png) [@josef](https://kopia.discourse.group/u/josef)\
**Post date:** [January 31, 2022, 9:24pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894/2 "2022-01-31T21:24:04Z")

</div>

Maybe you find some information in this thread:

> [@Low read performance per distinct hashing process](https://kopia.discourse.group/t/low-read-performance-per-distinct-hashing-process/90/18):
>
> FWIW fixed splitters are not ideal if you have files that are partially modified as they don’t perform content-based splitting. So if you have a large video file (say 10 GB) and add one byte somewhere at the beginning (for example by changing embedded text metadata), fixed splitter will need to reupload 10 GB. content-based splitters (buzhash and rabin karp) will detect the change and typically only need to upload one or two chunks (\<10MB total). Fixed splitter will be much faster, though.

Best seems to be the one with the lowest timespan.  
Not sure why you see 0ms on all FIXED splitters though.

---

<div class="post-metadata">

**Author:** ![dimejo](https://yyz2.discourse-cdn.com/free1/user_avatar/kopia.discourse.group/dimejo/32/287_2.png) [@dimejo](https://kopia.discourse.group/u/dimejo)\
**Post date:** [February 1, 2022, 1:30pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894/3 "2022-02-01T13:30:08Z")

</div>

> [@lordraiden](#):
>
> Can some explain, or where I can find and explanation of between the different splitters available?

AFAIU the difference is the algorithm used and the size of the chunks (1M, 2M, 4M and 8M indicating the size).

I’m unaware of any official documentation for the splitter functions in Kopia, but these links may provide some basic understanding of how it works.

> **[Rolling hash](https://en.wikipedia.org/wiki/Rolling_hash)**
>
> A rolling hash (also known as recursive hashing or rolling checksum) is a hash function where the input is hashed in a window that moves through the input.
> A few hash functions allow a rolling hash to be computed very quickly—the new hash value is rapidly calculated given only the old hash value, the old value removed from the window, and the new value added to the window—similar to the way a moving average function can be computed much more quickly than other low-pass filters; and similar to th...

[https://restic.net/blog/2015-09-12/restic-foundation1-cdc/](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)

> [@lordraiden](#):
>
> And how can I interpret this results?

My guess is:  
algorithm name, compute time, number of chunks, minimum chunk size, …, maximum chunk size

---

<div class="post-metadata">

**Author:** ![jkowalski](https://yyz2.discourse-cdn.com/free1/user_avatar/kopia.discourse.group/jkowalski/32/3_2.png) [@jkowalski](https://kopia.discourse.group/u/jkowalski)\
**Post date:** [February 1, 2022, 4:17pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894/4 "2022-02-01T16:17:28Z")

</div>

FIXED is super quick because given an input data it is trivial to determine where it needs to be split since the output size is, well, fixed, other splitters rely on rolling hash, which is harder to compute (and Rabin/Karp tends to slower than buzhash).

---

<div class="post-metadata">

**Author:** ![flomp](https://avatars.discourse-cdn.com/v4/letter/f/7ba0ec/32.png) [@flomp](https://kopia.discourse.group/u/flomp)\
**Post date:** [February 2, 2022, 7:09pm UTC](https://kopia.discourse.group/t/difference-between-the-available-splitters/894/5 "2022-02-02T19:09:51Z")

</div>

I did some experiments with Veeam Backup Files as input (`.vbk` and `.vbi`). Regarding the resulting size, I did not find a big difference between Buzhash and Rabin/Karp if the block size is the same.

Here is part of my results:

| splitter | du -d0 | find | wc | time real/usr/sys |
| --- | --- | --- | --- |
| DYNAMIC-4M-BUZHASH | 1188 / 88% | 116290 | 307 / 153 / 20 |
| DYNAMIC-1M-BUZHASH | 968 / 72% | 97057 | 323 / 158 / 20 |
| DYNAMIC-1M-RABINKARP | 975 / 72% | 97541 | 305 / 179 / 20 |
| restic | 975 / 72% | 231521 | 382 / 399 / 42 |

Input here was 41 VIBs (367 GiB) and 2 VBKs: 487 GiB + 492 GiB  
Total size: 1346 GiB  
I was especially interested in the deduplication of the two full backup files (VBK).  
In this run, the data was already compressed but not encrypted.
