I am working on a MicroPHP to shove into something like a RP2040 since PHP is really just a shell script for a whole bunch of C functions... but PHP eats too much ram currently on the micros
Full Archival with the standards required by the Internet Archive require that full unmodified headers are required, and unmodified content. This tends not to work well with modern browsers. Chrome and Firefox both fail at this currently. Someone is looking into a kind of modified Firefox to help with this. but its just not that how this system works. Now the Archive.org does have a API of sorts to say hay archive this URL, and a little working on the backend goes and does it..
What the Archive Team does is on a much more massive scale. Like SETI at home scale of scraping data across the internet. At almost every point we have had to make custom tools to ensure it meets our needs in our archival efforts.
I really think this is what Gabe did with Valve. To this day they are a privately owned company. It may not be the best place to work, but I do know they get paid well.
https://jrwr.io -- It's my personal website with a fun "fake" terminal with a custom command set with a set of challenges to solve to gain more access levels in the machine. Its been fun to watch the commands come in.
Over at a University we run, we like to run like a ISP and only have a /16 to work with, its very tight even now, we have thousands of students using the Wifi, Dorm Networks and such. I do wish we had more.
Here is the API you are looking for: -- https://archive.org/wayback/available?url=example.com×t... -- this will show the oldest timestamp of a archive that the wayback has. The trick is to set the timestamp to be /really/ old and it will show the first snapshot it has.