Arrows of Time for Large Language Models(arxiv.org)6 ポイント·投稿者 tianlong·2 年前·3 コメントarxiv.orgArrows of Time for Large Language Modelshttps://arxiv.org/abs/2401.175052 コメントコメントを投稿[–]nyoncore·2 年前返信Isn't it obvious that since LLM are trained to predict the next word they do better than to predict the previous one?[–]frotaur·2 年前返信In the paper it is mentioned that the LLMs predicting the previous token are indeed pre-trained in this way, so it is not true that the difference is obvious.[+][deleted]·2 年前[–]tianlong·2 年前返信There is a link with entropy creation?