jumpCastle·2 năm trước·discussAttention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle·2 năm trước·discussAlso the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle·3 năm trước·discussUse model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle·3 năm trước·discussWithout open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.