Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios)
github.com164 pointsby nonadhocproblem49 comments
codec decode ratio encode
misa77 -0 5302 MB/s 42.64% 58.7 MB/s
selkie -3 3078 MB/s 39.58% 97.0 MB/s
lzo1x -1 421 MB/s 47.48% 314 MB/s
memcpy 12135 MB/s 100.00% 11952 MB/s codec decode ratio encode
misa77 -0 5309 MB/s 42.64% 58.2 MB/s
misa77 -1 4332 MB/s 39.65% 54.5 MB/s
selkie -1 3095 MB/s 43.48% 197 MB/s
selkie -2 3013 MB/s 41.80% 163 MB/s
selkie -3 3081 MB/s 39.58% 95.8 MB/s
selkie -4 3498 MB/s 36.95% 43.3 MB/s
selkie -5 2607 MB/s 33.58% 2.97 MB/s
lz4 2530 MB/s 47.60% 373 MB/s
Note: selkie -4 corresponds to the "Normal" level here. - 1 token byte + 2 distance bytes
- a somewhat unpredictable number of literal bytes
If these streams are interleaved, then it's harder for the out-of-order core to process the token bytes of several blocks in advance (as the position of the next token byte depends on the unpredictable number of literal bytes in the current block). However, if you separate them (all token bytes in a prefix, and literal bytes in a suffix), the CPU can speculatively parse token bytes of a lot of blocks in advance (as most of the time, it just has to step forward by three bytes to go to the next block's token byte). - match length per block is capped to 32
- distance to a match must be >= a fixed constant
- unlike lz4, tokens and literals have separate streams
- format of the token byte has been changed
Now, this format allows our decompressor's hot loop to be very simple (in terms of the number of branches it has). This simplicity in turn allows our compressor to create a compressed stream that is friendly to the (small number of) branches in the decompressor. codec decode ratio encode
misa77 -0 4061 MB/s 61.52% 40.4 MB/s
misa77 -1 2851 MB/s 59.04% 36.2 MB/s
lz4 2561 MB/s 62.66% 488 MB/s
lz4hc -12 2428 MB/s 55.53% 6.23 MB/s
Results on equipment asset (WoTR): codec decode ratio encode
misa77 -0 4675 MB/s 54.33% 48.6 MB/s
misa77 -1 3752 MB/s 51.96% 40.9 MB/s
lz4 3101 MB/s 55.21% 497 MB/s
lz4hc -12 3036 MB/s 47.62% 6.25 MB/s
Results on texture asset (DOS2): codec decode ratio encode
misa77 -0 5546 MB/s 67.99% 47.0 MB/s
misa77 -1 2991 MB/s 63.79% 31.5 MB/s
lz4 3602 MB/s 68.53% 623 MB/s
lz4hc -12 2689 MB/s 59.30% 9.01 MB/s
Note: the benchmarking setup is identical to the intel x86-64 one described in the readme.