I gave GPT-5.6 Sol, Fable 5, Grok 4.5, Sonnet 5, and GPT-5.5 the same greenfield spec: build the Basecamp 5 frontend and API. Fable won both tracks at $85.87 in 2:06:40. Grok reached 84% of Fable's frontend score and 87% of its backend score for $9.30 in 36:48. Five reruns exposed meaningful variance: the best run beat the median by up to 0.46 points, while Sol’s best frontend beat its first score by 0.72. The full report shows where each model excels, where it is okay, and where it fails.
100% — publish the hidden research, the value is in the discoveries, not in the dividend. With all due respect to the author, it feels like he missed the entire lesson of history.
Built ScreenCommander to solve a personal gap: individual app integrations severely limit what local AI agents can actually achieve. This macOS CLI tool captures your desktop as screenshots, then allows local agents like Codex to interpret and perform actions (clicks, keystrokes, navigation) visually. Requires local permissions for accessibility and screen recording. It's like giving your agent actual eyes and hands, augmenting or bypassing rigid app skills altogether.
Works well with Codex, but not great with Claude Code or Gemini CLI yet (both are bad at novel CLI tools despite having a skills file). Also works well in conjunction with other skills (Atlas, or Apple Script), especially with non-vision models like Spark.
It was initially one-shotted from a GPT-5.2 Pro briefing into Codex-5.3-Codex-xHigh, then iterated on to fix performance issues and expand capabilities.
No, I’ve seen this pattern as well. Will apologize and then when you ask to continue it will have a change of mind and refuse again. It’s a bad RLHF/AIF loop that it gets stuck into.
This is awesome, but there were a couple of great laptop interfaces from that movie too. Spent some quality time in the 90s getting AfterStep/Litestep to look like them.
100% agree, I think the 26% will greatly increase over time... or the ones that don't will decline as a business over time.
the 13.4% is likely leaders in ML for some specific use case like fraud or recommendations. it would be great to have access to raw data with anonymized demographics.
great data – wish they provided the raw information to slice the respondent audience more, but aligns with what I've seen in the market re: concerns and models.