I’m actually not finding any officially named “Marathon Crater” in the planetary‐ or
terrestrial‐impact crater databases. Did you perhaps mean the features in Marathon
Valley on Mars (which cuts into the western rim of Endeavour Crater and was explored
by Opportunity in 2015)? Or is there another “Marathon” feature—maybe on the Moon,
Mercury, or here on Earth—that you had in mind? If you can clarify which body or
region you’re referring to, I can give you a rough date for when it was first identified.
First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement
Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.
Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.
Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.
If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.