jumpCastle·2 jaar geleden·discussAttention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle·2 jaar geleden·discussAlso the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle·3 jaar geleden·discussUse model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle·3 jaar geleden·discussWithout open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.
jumpCastle·3 jaar geleden·discussThe future is the AI also writes the high level descriptions of stuff.
jumpCastle·3 jaar geleden·discussBut you can fine tune gpt 3.5 turbo, so your comparison is not clear.