One can do ZFS snapshots so one does not need do insanely huge backups all the time. Just transfer off the diffs as needed. If an attack happens it's pretty easy to roll-back to a known good state. It's also not that complex to set some process in place that does random checksum verification of some files to trigger an alarm that such an attack has taken place. It is really perplexing me that very large institutes don't do this
Many thanks for sharing your ideas and views on the world, they have planted a seed in many a mind. Your influence is far greater than many could imagine.
Reading these comments it makes me think that what is being show here in many of these responses is one of core issues with the software world at large.
The issue is that of "It's not the way I do it therefore no-one should do it like that."
If a person wants to hold on to and horde masses of data then that is their prerogative. To give suggestions on how one would do it from their own view is acceptable but to outright dismiss another persons wants and needs is very myopic.
One could even view the building of the data storage systems as a hobby in itself and the act of doing so and documenting it will be of use to others, even in other industries.
I know of some professional photographers who have really poor data setups as they are very not that tech savvy so linking them to an article like this is very helpful.
Thought they used Parsons Code as it is space efficient as a fingerprinting technique and less across the wire too for a partial fingerprint and it handles tempo drift. In addition I know they where becoming CPU bound and then moved to GPU to do matching, that greatly helped them.
I do highly suggest that a quick intro demo video and/or screen shots of a tool like this would be beneficial to the project.