>> In the proposal, OpenAI also said the U.S. needs “a copyright strategy that promotes the freedom to learn” and on “preserving American AI models’ ability to learn from copyrighted material.”
Perhaps also symmetric "freedom to learn" from OpenAI models, with some provisions / naming convention? U.S. labs are limited in this way, while labs in China are not.
And no, I don't think the knowledge of language is necessary. To give a concrete example, tokens from TinyStories dataset (the dataset size is ~1GB) are known to be sufficient to bootstrap basic language.
For long context sizes AGI is not useless without vast knowledge. You could always put a bootstrap sequence into the context (think Arecibo Message), followed by your prompt. A general enough reasoner with enough compute should be able to establish the context and reason about your prompt.
I agree, they are only starting the data flywheel there. And at the same time making users pay $200/month for it, while the competition is only charging $20/month.
And note, the system is now directly competing with "interns". Once the accuracy is competitive (is it already?) with an average "intern", there'd be fewer reasons to hire paid "interns" (more expensive than $200/month). Which is maybe a good thing? Fewer kids wasting their time/eyes looking at the computer screens?
The approach of "cutting funding and then observing whether anything critical fails or is impacted" only works if outcomes follow a normal distribution.
This is far from the case — many areas are characterized by heavy-tailed loss distributions, where extreme negative consequences could really ruin the day and erase any efficiency gains.
If you look at the benchmarks of the DeepSeek-V3-Base, it is quite capable, even in 0-shot: https://huggingface.co/deepseek-ai/DeepSeek-V3-Base#base-mod... This is not from scratch. These benchmark numbers are an indication that the base model already had a large number of reasoning/LLM tokens in the pre-training set.
On the other hand, my take on it, the ability to do reasoning in a long context is a general capability. And my guess is that it can be bootstrapped from scratch, without having to do training on all of the internet or having to distill models trained on the internet.
MMMU is not particularly high. Janus-Pro-7B is 41.0, which is only 14 points better than random/frequent choice. I'm pretty sure, their base DeepSeek 7B LLM will get around 41.0 MMMU without access to images, this is a normal number for a roughly GPT4-level LLM base with no access to images.
Sorry, but this was ChatGPT/o1 with access to code execution (Python) and it used almost 4 minutes to do reasoning. It had done a few checks with smaller numbers, all of which had failed. And it proceeded to make a wrong conclusion (with high confidence).
> Thought about large prime check for 3m 52s: "Despite its interesting pattern of digits, 12,345,678,910,987,654,321
is definitely not prime. It is a large composite number with no small prime factors."
Feels like this Online Encyclopedia of Integer Sequences (OEIS) would be a good candidate for a hallucination benchmark...
I understand that it is mostly regulated at the state level. I'm not sure about other states, but The Computer Science Standards for California Public Schools (Kindergarten through Grade Twelve) also tend to be followed by private schools. So they can claim their programs meet state requirements.
This brings computers into the classroom, and once they’re available, it is a slippery slope. It is easier for teachers to have students use semi-gamified "educational" apps rather than engage themselves.
K-2.CS.1 Select and operate computing devices that perform a variety of tasks accurately and quickly based on user needs and preferences.
K-2.CS.2 Explain the functions of common hardware and software components of computing systems.
K-2.CS.3 Describe basic hardware and software problems using accurate terminology.
K-2.NI.4 Model and describe how people connect to other people, places, information and ideas through a network.
...
K–2 K-2.AP.12 Create programs with sequences of commands and simple loops, to express ideas or address a problem
K-2.IC.20 Describe approaches and rationales for keeping login information private, and for logging off of devices appropriately
Another Gorilla is the schools, teachers and state-approved recommendations, that extend their reach even into private schools.
Imagine my frustration one day, when I've discovered that my kindergartner has full access to a brand-new, shiny iPad during class. Despite complaints from parents, the teacher refused to reduce iPad usage (or even activate Screen Distance and Screen Time controls on the iPad, or share usage statistics).
The only thing that I've learned, this is all in line with California’s state-approved computer literacy recommendations.
I wish that "Online Coupon Price Tags" in stores would also be banned. I'm talking about these yellow price tags that show lower than "Club" prices, which are only valid if you collect a coupon online.
Like FTC, I estimate that banning these would save U.S. consumers millions of hours they currently spend searching and clicking on pointless coupons on their phones before making purchases. It would also increase happiness, as it's extremely annoying to pay $20 extra, knowing that a lower price is available if only you spent ten minutes struggling with a store's website on your phone.
Whoever invented this is evil and is destroying happiness.