i think people say that thinking that only training to produce the next likely word would end up producing some local minimum word that generally fits but doesn't actually lead to intelligent thought.
that feels like a misunderstanding of how the loss function behaves when used within a sequence
it’s hard to isolate the effect size of policy, covid happened, car weights changed, policing may have decreased, US drivers may have driven differently, population size, etc.
that feels like a misunderstanding of how the loss function behaves when used within a sequence