A simple cue like asking the model to 'see' or 'hear' can push a purely text-trained language model towards the representations of purely image-trained or purely-audio trained encoders.
Has anyone found these deep research tools useful? In my experience, they generate really bland reports don't go much further than summarization of what a search engine would return.
23X revenue multiple is not an expensive valuation for something growing well over 100% YoY, even by public market standards let alone venture capital valuation.
Housing prices remaining high because people can't move and people can't buy because of interest rates just shows how stupidly hard it is to build new housing in the US.
I would really hope you could get decent utilization on ops as fundamental as GEMM/memcpy on a single device. Translating that to MFU is a completely different story.