the size is given because one can understand how many machines it would take to hold the data in the memory, how much bandwidth is required to move data around and latency.
Its hard to understand if you are just a programmer, a computer science graduate can easily identify with size in bytes since they represent accurate amount rather than vague billion entries.
In later case you need to specify two variables size of entry and number of entries.
the larger entries become the easier is to sort them, if they get smaller then you can have an O(n) sorting ability.
The 100 byte records is fixed in literature and well understood by practitioners, what you are experiencing is your naivete.
As some one who actively works in this area. the problem with queries is that most of the algorithms that work require deep learning and not simple sql like syntax. I would love the data in simple csv format hosted on a cloud such as Amazon EC2 [for safety] rather than some beautiful interface by a latte sipping ruby hippy.