jumpCastle·hace 2 años·discussAttention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle·hace 2 años·discussAlso the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle·hace 3 años·discussUse model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle·hace 3 años·discussWithout open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.