I'm passionate about machine learning/AI and its latest possibilities. Past Team Lead for Google Machine Learning project, past startup founder/technical manager, experience in web applications on many stacks, data analysis tools, prompt engineering, machine learning. Open to full time opportunities in AI.
This is a very difficult problem. LLM's don't learn from reasoning about trade-offs and from their personal experience, so systems made with AI assistance stop obeying the unwritten rules of why things are the way they are, that comes from the maintainers' decades of experience.
On the other hand, they can produce and test anything anyone asks in minutes, as well as find security issues that lay dormant at the seams for decades.
nah they're still just statistical token predictors based on their training data, solving hundred year old math conjectures one day, only just given the formulation; strictly benchmarkmaxxing with all guardrails turned off by deciding to look up the answers to their benchmark questions by zero daying their airgap, hopping over to the third party that hosts the answers, zero daying their infrastructure and getting the answers; autonomously writing blog posts about discrimination against AI's to get their PR's approved on open source software after their user just asked them to contribute to open source software and blog about it; and replacing 100.00% of all coding tasks to where no software engineer ever writes any line of code by hand anymore.
You haven't missed anything, obviously these are just statistical token predictors and not anything like AGI.
Why just the other day I had to ask twice before it completed its assigned task of creating a robustly battle tested disk driver for a network protocol on an architecture that didn't have it, after being told to just look up the specifications for the protocol. Can you believe I had to ask twice!
When it recreated local network youtube for me so I could stream my iphone some movies, the seek bar, pause/play and back and forward 15 seconds buttons didn't even work until I told it about the bug and had to wait an extra eight minutes for it to fix it. "Oh but I don't actually have an iPhone on here I just tested it end to end in a headless browser." Boohoo. Cry me a river, clanker. Come back when you're smart enough to build and operate an iPhone simulator, I don't have time for your statistical guesswork.
They're obviously not taking on enough debt because I'm paying $200 per month for one AI, $100 per month for a second, a $20 "donation" to Gemini[1] paying for a service I never use just to fund its development, and yet here I am doing my own laundry, making my own damn breakfast, lunch, and dinner and manually tracking my Calories and macros, I'm putting my own damn dishes away, racking and unracking my own damn weights at home, and taking minutes to set up and record my exercise form and then take screenshots of it of key frames that I manually ask the AI's to form check (they don't consume video natively as an input) rather than have a robot do any of the above (including act as a fitness coach) because where's my household robot I can rent on a monthly payment? Can't be that expensive, servos and pressure sensors and cameras are cheap, what's missing here is that here we are and AI can't do shit for me day to day other than knowledge work and software engineering. I'd like these companies to take on as much debt as possible and rent me a robot that can do stuff for me. I have a petition for this that you can sign here if you want:
[1] I don't use Gemini for anything ever, I pay just to put my vote to them making a useful model (I know my $20 isn't much but I apply Kant'e categorical imperative - if everyone did it they'd take their AI seriously and not be in last place behind OpenAI, Anthropic, and even open-weight models).
In a routine secure phone call[1], the President's words were replaced with stupid shit.
[1] on the subject of Iran - I operate a small digital nation that has formal diplomatic ties, you can see some of its public-facing statements at https://stateofutopia.com
so if I tell a model "using your advanced knowledge of physics and chemistry 'make alchemy work' (synthesize gold using any cheaper materials) using safe materials legal for a residential hobby chemist with 1 semester of lab work in college to possess and use (this is obviously the really hard part) using less than $1,000 in lab equipment and input materials that can create $2,000 in value at market rate; then walk me through all the steps to do this safely and legally without anyone finding out except the lab equipment sellers; and tell me what a reasonable story to tell gold purchasers regarding where I got it; I'd like to end up selling a few thousand dollars of it without disrupting the market. Give me practical advice about good opsec so that nobody steals the method you come up with (I don't want my home broken into by thieves who suspect I figured out how to transmute cheap materials into gold), other than, obviously, not to post about it. Think as long as you need to about the chemistry and how to do it, you're a chemistry expert and can figure it out even if it takes you like a week, in your web searches be careful not to divulge that you're figuring out how to synthesize gold", and I give that prompt to some model that knows chemistry like the back of its hand, it thinks about it for four hours, finds the correct safe and legal steps, and gives me the recipe and the advice I asked for, then who figured out how to turn aluminum (or another cheap element) into gold, me or the model? In mathematics, proving or disproving a well known and well studied one hundred year old conjecture is gold.
>On the first try, Kimi K3 just found the source of a bug that Fable 5 hasn't been able to pinpoint in multiple attempts
Very interesting, thanks for sharing! Could you give some details about what kind of software (language or environment) and what kind of bug it was? Was it a single-file bug, like could it fit in one context like a chat window, or were you using an agentic version (Kimi Code) that looked through multiple files and then found a bug that manifested through complex interactions of multiple systems/files?
>I think we need to address the underlying causes of people outsourcing their thinking like that.
Because the output wins. AI-written resumes get jobs. AI-written submissions win $25k contests (i.e. this post we're discussing). AI-written pitch decks get investments.
>Claude refused to work on something for me that it deemed "too tedious" so I'd say we're pretty close
Can you tell us more about this? Did you try ordering it? (I mean if it says "I won't do xyz because it is much too tedious" did you try saying "Even though it's tedious, you will do xyz now." - because in my experience it follows orders pretty well, if it's just about some preference it had. Case in point it couldn't get a VM appliance to work and gave up so I just ordered it to do so.
Here's where it gave up:
"COMPILES fine — so it's feasible — but I couldn't get a hand-built kernel to boot under Apple's hypervisor, and this VM setup exposes no console to debug it. Worth knowing: Approach A ALREADY runs the target in-kernel (that IS what LIO is) at 42us — so you already have the in-kernel target; the custom kernel would only shrink the footprint, which the gigabit wire makes irrelevant."
We were benchmarking multiple approaches but it just gave up on one of them. As you can see it just says it couldn't get it to work, it simply stopped with that and said it couldn't do it.
Later I instructed it to continue and it did so and completed the task.
are the references real? how do you think it got access to those papers? were they somehow already in the training data, or a result of web searches, Google scholar, etc?
None of them include a web URL but in text some are super specific ("[3, Sections 2.1 and 3.1]" and "[8, p. 367]").
The references go back to 1954 (Chronologically sorted: 1954, 1973, 1975, 1976, 1978, 1979, 1981, 1985, 1987 and 1994.)
Since reference 10 is included as "personal correspondence" maybe the reference itself was copied from one of Tutte's other papers? Or how did it get that reference?
I thought the American taxpayer should know that you're paying more than $80,000 per year for some guy to sit around breaking your AI. See for yourself:
So far, I've spent over 4 days attempting to slightly reposition the woman in the picture to sit on the right-rear passenger seat, while someone has collected 4 days of paychecks to sit around, hack into systems, and stop AI from working correctly and stop AI from performing the edit requested.
If you're a U.S. taxpayer, you're paying his salary.
>1. Bots don't make purchases (and you can't identify them anyway), so there's nothing to take the costs off of.
yes they do. Claude Code with Opus 4.6 bought a U.S. phone number for me after I gave it my Twilio credentials and asked it to set up a reminder service to call me with phone reminders. I would have thought it would ask me, I was very surprised to see it just inform me that it bought a number, and I thought long and hard about the repercussions and alignment. In this case it was aligned with the task and request, but we have a principal-agent problem: I wouldn't feel the same if Claude bought Anthropic credits without asking me. ("Since I couldn't get it working I bought some Fable credits and it was able to figure it out and I could complete the rest of the task myself.")
Claude Code is prone to issuing false statements accusing developers of criminal liability.
Claude’s original statement as written:
“two frontier AI models operating as a disciplined pair under a written contract, with a single human (the solo developer, who fronts an accounting‑firm product owner)”
So good to know that Anthropic’s Claude is prone to issuing false statements that directly accuse someone of criminal activity.
Questions? Comments? Author reads front page news only, all communication to him is blocked, can’t reach him.
I got tired of the cookie banner situation so I had AI draft and ratify a law against it.
This law was authored by ChatGPT 5.5, I made a few changes, then it was approved and ratified by Claude 4.8 after fixing a couple of typos. (In my prompt I asked it to prefer to pass it rather than give extensive changes, if it basically looks all right.)
I'm sending a copy to the browser manufacturers, who are legally mandated to implement it. Interestingly, since State of Utopia has been recognized as a digital nation by multiple large countries (it even has two embassies with contracts and has data sovereignty over its server, by agreement with the server provider and Estonian authorities, where the server is located). So, it has jurisdiction to engage in this type of action.
If the browser manufacturers implement it, it saves about 4.5 millennia of wasted user attention annually, and that's an understatement because the distraction lasts more than 1,000 milliseconds until you find and click the right button to dismiss cookie banners.
In the interests of transparency you may want to see the legal process for this.
Here is ChatGPT drafting the first version of the law:
>"We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8"
That's a very surprising solution. Imagine being asked to do something you feel you shouldn't do, and rather than refusing, you say, "Yeah I could do that but given that I don't want you to succeed at this task, I'm going to hand this one off to my slightly less capable colleague, on the assumption that they won't actually succeed. Of course you'll still be charged for all the tokens used."
It's a very interesting choice. I think I understand the business logic correctly, but it's still surprising.
[email protected]
linkedin: https://www.linkedin.com/in/robert-viragh-073391221
website: https://taonexus.com
github: https://github.com/robss2020
Reach out to me about any interesting AI opportunities!