jumpCastle·قبل سنتين·discussAttention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle·قبل سنتين·discussAlso the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle·قبل 3 سنوات·discussUse model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle·قبل 3 سنوات·discussWithout open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.