So they can train on everyone's copyrighted works to create their model, but when someone trains a model off their model it's not okay? Seems kind of hypocritical.
I would think modern optimized binaries would be too complicated for current LLMs to handle. LLMs currently get overwhelmed with large code bases, imagine turning that all into binary.