How do you think the bots are trained? Reddit has really good data since its comments and tags and specific subreddits basically have ranked questions and answers by points driven by the title and the comments are ranked by points which have quality answers.
The problem was that people who know python don't really expect this specific DSL and it is easier for someone that doesn't know python and just learn their DSL and then read the rest of python.
The best description about how it became a problem is one of the paragraphs.
"New starters take an exceptionally long time to get up to speed - and that's if they don't resign in fit of pique as soon as they see the special, mandatory, in-house IDE (as I nearly did). Even months in, new starters are still learning quite fundamental new things: there is a lot that is different."
I think it took me til I was there around two and a half years to fully comprehend it when I was working on it. Not much modern training til they figured out they had to teach it again that was better. The worst part is to make an UI around it coding it and it wasn't approved for new projects.
I was studying Erdos problems by only taking ChatGPT 5.5 outputs and just asking it to keep on attempting to solve it by asking it to go further. I haven't started doing this with chatgpt 5.6 I have some partial results here https://chatgpt.com/g/g-p-69f03400f420819192418b18ca90ffee-d...
What was really interesting is that during the process it was able to find lemmas or theorems that might be related or relevant to be published.
While I was doing that I was also trying to use Aristotle to do the Lean formalization and I have a WIP system to do that at https://github.com/aconsapart/thesisus/
I don't understand the code in the examples after simply reading it. That's a problem. You need to have comments to clarify or pick a better example of code.
You are in competition with Python and Javascript which has lots of code for training data. I wouldn't touch esoteric features unless it improves readability or is under the hood. The bigger problem with esoteric features is that its hard for humans in general to understand unless you cover it with enough syntactic sugar for people to write it down.
Xcode 27 which is in beta ships a much better codex integration. There are problems in Xcode that are frustrating . Also, git integrations are a nice touch that allow you to force push too. Would be interesting to find out how well this works in a team based setting.
I am working on a Jupyter notebook client for VisionOS. It allows for 3D data to be visualized on visionOS. Right now it supports point cloud data, USDZ models and Gaussian Splats. I am working on it to launch on the App Store. Sign up for more information at http://www.pulto.org
I love Claude Code and I don't use the rest of the models since I use ChatGPT for productivity work it. Fable is pretty great and the UI / UX is much better than codex
There is one simple thing you have to realize why Python is the optimal choice. You have so much training data. Python is the second most popular language on GitHub and is easy to read.
I agree with you on every point but it is interesting to see real world benchmarks like this. Showing the standard benchmarks that all LLMs use is not only boring but at this point likely gamed or even has issues (according to OpenAI) by every LLM.
The only thing I can possibly think of is that they could use it internally at possibly a lower cost and offer it to people who have a Tesla cheaply. Owning Cursor might help for integration or data collection.
Xbox around 2021 had around a 12% profit margin and the gaming industry as a whole was around 17-22% . In 2023 the target for the division was put to 30% . We see this new restructuring because the target was put this high. Microsoft really wanted Game Pass to be a steam competitor which is pretty much what everyone in the industry tries to do and fails. The push for Game Pass prices to be higher was to get the 30% margin and that didn't work out. They aren't operating at a loss they are operating at a goal and they failed the goal. From other child comments many studios they bought probably were below average. We can see this restructuring basically is that they failed the target, the old guard went out as the new guard came in.
Sort of interesting license not sure if anyone will do it long term.
The training data and the Apertus LLM may contain or generate information that directly or indirectly refers to an identifiable individual (Personal Data). You process Personal Data as independent controller in accordance with applicable data protection law. SNAI will regularly provide a file with hash values for download which you can apply as an output filter to your use of our Apertus LLM. The file reflects data protection deletion requests which have been addressed to SNAI as the developer of the Apertus LLM. It allows you to remove Personal Data contained in the model output. We strongly advise downloading and applying this output filter from SNAI every six months following the release of the model.
UIC Alum Bachelors in Computer Science 2012.
Website : http://joshuajherman.com
Email: zitterbewegung at gmail dot com
Check out my natural language network scanner: http://www.securday.com