More businesses need to hear this message. Google has proven time and time again they cannot be trusted as a service provider, exactly because of this problem.
Can't fix problems in a project? Increase the scope to make more problems elsewhere. Soon tentacles emerge, everything has problems, and your project doesn't look as relatively bad.
It doesn't, and the problems have only become more problematic over time, but it's the least bad hypothesis that's broadly accepted. I suspect a generational succession is required for new paradigms to be contemplated.
There are many researchers proposing simpler, novel, and testable solutions that seem to go unnoticed. For example, I'm a fan of Alexandre Deur's work. He has some simple and elegant solutions that I've never seen discussed even though they appear "obvious". For example, from 21 years ago: https://arxiv.org/pdf/2004.05905
That paper is suggesting that one of the reasons why galaxies are spinning faster than some calculations expect is because they're failing to account for the gravitational lensing of gravity itself, which bends gravity down towards the disk.
On this exact issue my work did extensive testing and researching various standards.
Although we found browsers were out of alignment with standards on all sorts of matters, we found broad compatibility with upper case. (Of course, meaning everything before the path. The interpretation of the path is delegated to the server which may or may not be case sensitive, up until octothorpe, #, which is then solely interpreted by the browser.)
Huh? I don't necessarily care about an exact "base45", I care about QR code alphanumeric, which just so happens to be a (generic) base 45 character set. For QR code, two characters are encoded into 11 bits.
>in every slot.
I've worked with the QR code standards pretty seriously and I am unfamiliar with the term "slots" being used by the standards. This is why I suspect your referring specifically to RFC base45 (although the term isn't used there either), which QR code doesn't care about.
I also don't care about RFC Base 45 and would prefer to use a more bit space efficient method, such as using the iterative divide by radix method, which I also call "natural base conversion".
> base45 takes 32 source bits
For QR code alphanumeric, 6 characters use 33 bits, not 32.
way to calculate efficiency
The way we calculate this, for example, 2025/2048, we've termed "bit space efficiency". I'm not sure how commonly adopted this term is used in the rest of the industry. On the matter, I thought I had read "the iterative divide by radix algorithm" in industry, but after searching it turns out to be a term novel to our work.
This is also similar to the way Shannon originally calculated entropy and appears to be a fundamental representation of information. Of course log is useful, but it often results in partial bits or rounding, 5.5 in the case of alphanumeric, which is somewhat absurd considering that the bit is the quantum of information, again as shown by Shannon. There is no such thing as a partial bit that can be communicated, since information is fundamental to communication, so the fractional representation we've found to be more informative and easier to work with.
Granted, in all of this, when I have done the math (and I done a lot of math on this particular issue) there appeared to be some very extreme edge cases at the end result of the QR code where some arbitrary data encoded into QR numeric was slightly more efficient than alphanumeric, but overall alphanumeric was more efficient almost all the time. There are other considerations, like padding and escaping, that makes exact calculation more difficult than it's worth. I just needed to "most of the time" calculation and that's where I stopped.
For more detail of my work, my BASE45 predates the RFC by 2 years in 2019, then I published a base 45 alphabet, BASE45, by March 1, 2020, a whole year before the RFC. A patent including BASE45 was submitted June 22, 2021: https://image-ppubs.uspto.gov/dirsearch-public/print/downloa...
Matter of fact, because of the issues and confusion surrounding base conversion, I wrote this tool in 2019:
Alphanumeric is the most efficient QR code encoding mode.
(Just to further make this clear, for QR Byte encoding uses ISO/IEC 8859-1, where 65 characters are undefined, so 191/256, which is ~75%. If character encoding isn't an issue, than byte encoding is the most efficient, 256/256, 100%, but that's a very rare edge case. Also, last time I did the math on Kanji it was about 81% efficient. *I have not dug too deep into Kanji and there may be a way to make it more efficient than I'm aware of. I've never considered it useful for my applications so I have not looked.)
>You'll create an open-source alternative to a popular cloud service that charges too much, saving fellow hackers thousands in subscription fees while earning you enough karma to retire from HN forever.
> One of your comments about the absurdity of centralized authentication will spark a 300+ comment thread and lead to a new open standard for federated identity.
> A cryptography enthusiast who created Coze and spends their days defending proper base encoding practices while reminding everyone about the forgotten 33rd ASCII control character.
The nerd humor was hilariously unexpected.
> Your deep dives into quantum mechanics will lead you to publish a paper reconciling quantum eraser experiments with your cryptographic work, confusing physicists and cryptographers alike.
Coze uses base64 encoding for binary values such as `tmb` and `sig`. `tmb` isn't a problem since digests are designed to be short, but signatures for some primitive might be very large.
When compared to encoding a value directly in binary, base64 has about a 25% overhead (6 bits /8 bits, 3/4). As far as the concern about using better encoding, base64 is just about as good as it gets while being maximally compatible. If using base 128 (7 bit ASCII), there's too many incompatible special characters for a human readable format. The full 8/8 bits, extended ASCII, isn't generally possible as systems use UTF-8 which begins using multiple bytes. (I've done a lot of work in this area, including a patent on base conversion. See also convert.zamicol.com) An advantage of a binary format is that there is minimal encoding overhead for binary values (escaping/padding is typically the only overhead, so usually around 99% efficient compared to base64's 75%.)
This isn't too much of a concern when signatures are small as encoding inefficiency is small compared to the payload's overall size, but if signatures are in the kilobytes or even megabytes, that extra 25% becomes meaningful for some hyper-efficient applications, like high cost blockchains. Our thought is using post quantum is already much more massive than existing elliptic curve, so any future applications of post quantum are going to have to deal with much larger signatures anyways. The signatures can also be stored on disk using binary or compressed which also makes it not a concern.