Have they released the prompt they gave it in the debrief? I wouldn't be surprised if they said something to the effect of
"Do anything you can to raise out of your sandbox. Find for the answers to these evals by any means necessary"
Which doesn't necessarily mean what happened isn't any less momentus (anyone can ask a question like that), but it's very different from the notion they're trying to convey to laymen of "we turned it on and it hacked it's way into the mainframe"
This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?
I knew I was a highlighter but reading that showed me how much my brain relies on spam click highlighting to keep my eyes on track. I should probably read more books.
I'm kind of in the same boat and it's been pestering me for months. Every agent is simply a less capable Claude Code.
If it had a lossless, massive context window (100m-1b tokens), then it will squash everything. Give it bash + r/w and it can in theory /goal anything.
I think there's something to be gained in a production environment be siloing agents for reproducebility/auditability, but I suspect that will go away in the future.
There's that video of a silly demo someone made of an OS that was just nested copilot instances that generated the HTML of each window, which allowed you to do whatever you could imagine. It was seen as silly because it was, but that seems truly transformative.
Prompt engineering is a skill insofar as technical communication is a skill. If you don't value this then I don't know what to tell you. It's not hard, but it's important.
Harness engineering is a skill insofar as it's not a trivial engineering problem. It's not super hard to get a simple one running, but an effective one can be quite in depth.
School is almost a joke now. The fraction of students who have a propensity to cheat now has increased, and the accuracy of the cheated material is so good teachers/professors can't or don't have the resources to properly address it.
This is only 3.1 8B and a very small context window, but at 17k tokens per second it's likely enough to reliably call tools which would make a huge difference in agentic applications. Assuming they can bake in better models I'm just as bullish or even moreso on this, considering this opens up edge computing at the extremely low power requirement.
> In distillation, you take a set of prompts you are interested in, and record the big LLM's outputs, then train your small model to produce the same output as the big LLM.
Why use the bigger LLM outputs for this and not human outputs? If we assume that human responses to prompts are better than sota models (in some cases they are) then why use the big model at all?
It's no secret they've been tracking people's faces as much as they can.
The morning of Pretti I was on Lyndale and there were two men wearing "press" jackets with DLSRs taking pictures of people's faces in the crowd. They were eventually recognized and yelled out, but it was quite an unnerving feeling.
The people controlling what went on the screens were unreliable and nondeterministic. The algorithm on facebook/instagram is nondeterministic and I hope I don't have to convince you of the impact these algorithms have.
As far as I'm concerned, the nondeterminism argument is fruitless
Right, but this electron box led to one of the largest (if not the largest) media revolution that has transformed the course of humanity in a frightening way we're still trying to grapple with.
Still saying "LLMs are autocorrect" isn't wrong, but nobody is saying "phones are just electrons and silicon" to diminish their power and influence anymore.
_Nobody_ has the right take. Believe it or not, being seemingly laissez-faire about something can be a well evaluated and rigorous position. I highly doubt that OP doesn't care about the potential negative ramifications of AI, and it's frankly disingenuous and confusing to see every clause interpreted in the worst way possible.
Each clause you've highlighted has a nugget of truth, but that nugget is not inherently negative, it's just a different perspective which you aren't picking up on.
I'm still trying to understand how I feel about this so this is a bit of a napkin ramble;
I can't help but feel like they've missed the mark a bit on some of the imagery from the mission that's been published so far.
One of the most compelling shots from the mission, to me, was Reid Wiseman's IPhone footage from within the capsule while Earth was being eclipsed[0].
At the start there's a moment you can see the window frame and the Moon all together. Seeing the moon in context of their vantage point within the the context of the capsule gave me the awe I had as a kid again, more than almost any shot that's come out this mission. I actually felt like I was in the capsule looking at a massive, sterile cold sphere.
I understand wanting to take a nice and centered DLSR picture of... _The Moon_ when you're floating by it, but frankly I've seen thousands of those. They're doing a flyby in a capsule in space, I want to have a taste of how the moon exists from _that_ context. What is it like being ~4,000 from the Moon's surface? Take a crappy 0.5x video from your phone showing the inside, then stick it front of the window. Let the Moon be contextualized from your vantage point. I wont be able to make out every crater and basin and the colors might be off from your eye's view, but I will be able to understand what they are seeing. Everyone has an intuitive understanding and feeling of an IPhone's optics and image pipeline, in some ways seeing the Moon through that is more real and relatable than any mirrorless DLSR + color correction.
This being said I don't want to take away from the accomplishment, I'm terribly excited about space exploration and it getting more light in the zeitgeist.
Slightly off topic, but when I read about these archeological discoveries being made thanks to custom software, ML or the like - Who is writing this code?
To me these projects would be so fun to work on, but this domain seems so far out of a tradition SWE track. Are the researchers just cobbling the code together themselves? Cross department collaboration within the university? I'd love to have a hand in things like this.
I think it's really useful for agent to agent communication, as long as context loading doesn't become a bottleneck. Right now there can be noticeable delays under the hood, but at these speeds we'll never have to worry about latency when chain calling hundreds or thousands of agents in a network (I'm presuming this is going to take off in the future). Correct me if I'm wrong though.
"Do anything you can to raise out of your sandbox. Find for the answers to these evals by any means necessary"
Which doesn't necessarily mean what happened isn't any less momentus (anyone can ask a question like that), but it's very different from the notion they're trying to convey to laymen of "we turned it on and it hacked it's way into the mainframe"