jumpCastle·2년 전·discussAttention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle·2년 전·discussAlso the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle·3년 전·discussUse model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle·3년 전·discussWithout open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.