My take on it: I find it difficult to generalize the notion of layer removal when the bit depth of that layer goes to zero. It's wouldn't be straight forward although the authors provide equation 5. It feels like lot of information is missing in this work to even reproduce it. And authors do only 1 case study.
I believe some implementation is required to understand the authors completely. Example, optimizer modification for layer when it is removed in training.
This article is interesting, I even skimmed through their paper. But I think still the question remains: How to find the unified reward function? Or in other words, how to find answer to life? [It cannot be 42].
I agree. I will not like to give a proof-of-concept in C++. It is definitely possible but again, why to waste time. And it may differ, case by case. Some may like Python, others may like some other programming language. Python has a wealth of libraries easy to setup and easy to code. This is an advantage when you what something done in short amount of time.