dylanbyte·anno scorso·discussThese are high school level only in the sense of assumed background knowledge, they are extremely difficult.Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves.This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math.The answers are not in the training data.This is not a model specialized to IMO problems.
dylanbyte·5 anni fa·discussMy experience is that replicating papers is actually nontrivial. For example someone announced they had replicated gpt2 some time back but when evals were run it turned about to be the equivalent of a much smaller model.
dylanbyte·5 anni fa·discussCurious to see what parameter size of gpt3 this will end up being equivalent to. Obviously we won't know until they evaluate their models.
Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves.
This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math.
The answers are not in the training data.
This is not a model specialized to IMO problems.