The holds are not "extremely" taxing, I would say. The tremors can be induced through a simple 7 step sequence of holds like wall sits or calf raises. There are walkthroughs on YouTube.
On the TRE subreddit [0], people report being able to tremor at will once they have enough experience with the technique.
Not trying to be pedantic, just clarifying that it is very accessible! I've personally been experimenting with it and find it to be helpful.
Your info might be outdated. Apple Maps is actually better than Google Maps in many ways (e.g. it says "pass this light and at the next one, turn left" instead of "in 300 feet, turn left")
Yes, it wouldn't know what to do with the picture unless you fine-tune the model (which is why they are permissively releasing it).
The embeddings form the vocabulary of the model. The vocabulary "namespace" has 70k empty slots so you could introduce your own tokens and train on top of that, where token = some patch of multimodal data.