I've been using one of these for about a year. It's been great having a device that I can actually reach all of, and the small size also helps to not be a distraction.
About the only thing I miss from the Pixel series is camera skills and unnatural photo enhancements. It takes some nudging to get my jelly star to focus properly.
Each beat is probably better considered a symbol, so by your analysis the human audio channel is around 20 baud not 20 bits/s. The symbol size obviously has a large impact on the information bit rate.
We can estimate the size of a symbol though:
13b for primary tones(20khz range at about 3Hz resolution) + 4b for rising/falling/etc tonal patterns + 7b for volume (1dB detection over 120db range) + (42 phonemes in english so round up to 256 for convenience) 8b for phoneme patterns + 2 ears preprocessed for 2d directions is maybe 4b direction and 4b distance gives a total of around 40b per symbol.
So with that wild approximation, it's at 800 bps, or 8.64 MB / 24h
ZFS is also very far away from the state of the art in online dedup. For instance, http://users.soe.ucsc.edu/~avani/wildani-icde13dedup.pdf has a theoretical dedup regime that needs only 1% of the RAM for 90% of the benefit.
About the only thing I miss from the Pixel series is camera skills and unnatural photo enhancements. It takes some nudging to get my jelly star to focus properly.