Grok-2 Beta Release(x.ai)
x.ai
Grok-2 Beta Release
https://x.ai/blog/grok-2
333 comments
I assume they will have a lot less "safety", i.e. the model will be more likely to actually do what you ask instead of finding a reason why "sorry Dave, I can't do that".
Since these "safety" features tend to also degrade the model, that's likely also helping them catch up in the benchmarks.
Since these "safety" features tend to also degrade the model, that's likely also helping them catch up in the benchmarks.
It's hilarious they put Claude 3.5 Sonnet in the far right corner while it scores the highest and beats most of Grok's numbers.
It uses FLUX.1 to generate images and it has been fun so far. Its good on writing, can generate very realistic photos, can create memes, and looks like hands problem is fixed now.
You know what’s also impressive besides this beta release? How Claude 3.5 Sonnet is still able to keep up so well. Grok-2 beat every other LLM except Claude. How did Anthropic achieve this?
I don't really care. The model may be competitive, but my use cases require speed, local (semi-local) execution and reliability. Neither of these seem to be baked into whatever X produced now.
When they make the mini model available for download and quantizable. That's when I may be interested. But given the minimal improvement in the past several months, I'm inclined to believe that we have reached the plateau.
When they make the mini model available for download and quantizable. That's when I may be interested. But given the minimal improvement in the past several months, I'm inclined to believe that we have reached the plateau.
Do we have any info on this model's balance of censorship versus safety?
This is Musk after all, so I wouldn't be surprised if it strayed far from the norm.
This is Musk after all, so I wouldn't be surprised if it strayed far from the norm.
Oh this is great, one more competitor with top model which will be available via API. I wonder what the pricing will be. OpenAI was slashing prices multiple times in the last year and a half I was using it.
I have seen sus-column-r on LMSYS a bunch of times. It seemed pretty good, though not as good as the best Google, Anthropic, or OpenAI models.
I'm surprised they managed to catch up. I guess there really is no moat.
I'm surprised they managed to catch up. I guess there really is no moat.
Putting a new tool for developers behind an "enterprise API" gate is a sure way to kill it
"Our AI Tutors engage with our models across a variety of tasks that reflect real-world interactions with Grok. During each interaction, the AI Tutors are presented with two responses generated by Grok"
My guess is that they're using one of the third party AI training outfits for this and that they are paying through the nose.
This looks exactly like a training task I got to see on one of those platforms.
My guess is that they're using one of the third party AI training outfits for this and that they are paying through the nose.
This looks exactly like a training task I got to see on one of those platforms.
So all models seem to converge to a similar level of performance - is this the end of the line for LLMs?
I’m hoping we’ll see an open release of this in 6 months or so, as we saw with Grok-1.
I’m not hugely optimistic, though.
I’m not hugely optimistic, though.
If the X.AI team is able to build out a good enough model with access to real-time tweets, they could have an incredible product. I'd love to be ask about current events and get really strong results back based on tweets + community notes.
But when will it be available in my region (Europe)?
If I am reading the table correctly they are claiming it is better than all models but 3.5-Sonnet
Is anyone with X premium able to confirm the vibe check -- Is the model actually good or another case of training on benchmarks?
Is anyone with X premium able to confirm the vibe check -- Is the model actually good or another case of training on benchmarks?
Glad to see an uncensored AI able to compete with the other models.
Interesting that they’re rolling this out to Twitter/X Premium users, it was previously the biggest differentiator between Premium+ haves and Premium have-nots.
Seems like a solid result & more competition is always better.
That said I’m still cheering for mistral and meta with their more open stance
That said I’m still cheering for mistral and meta with their more open stance
Twitter started irreversibly feeding users’ data into its “Grok” AI technology in May 2024, without ever informing them or asking for their consent.
https://noyb.eu/en/twitters-ai-plans-hit-9-more-gdpr-complai...
https://noyb.eu/en/twitters-ai-plans-hit-9-more-gdpr-complai...
Can anyone tell me how much censorship grok has? I hate that many other LLMs have too much censorship.
Pretty funny to read the comments from xAI's initial announcement now.
https://news.ycombinator.com/item?id=36696473
https://news.ycombinator.com/item?id=36696473
[flagged]
We need an alt-right version of AI like we need a pumpkin spice sushiccino. No thanks but no thanks.
When github release?
Guys come on, you cant keep releasing software in the US, and then do a staggered launch where months later things are available to users in England, Denmark etc. There should be no reason for it. Im sure whatever dumb EU regulations can be dealt with easily in the software, these staggered releases ( such as chagpt having no Memory etc MONTHS down the line for EU users ) its just a hindrance to progress. Its starting to feel like we live on an island in the middle of nowhere.
What is the company’s ethical position though? It officially stemmed from Mr Musk’s objection that OpenAI was not open-source, but it too is not open-source. It followed Mr Musk’s letter to stop all AI development on frontier models, but it is a frontier model. It followed complaints that OpenAI trained on tweets, but it also trained on tweets.
Companies like Meta, Mistral, or DeepSeek, address those complaints better, and all now play in the big league.