Absolutely! We need new and better benchmarks like this.
I have a question: why not use the maximum available reasoning on each LLM? For example, I see that Opus 4.7 at `max` reasoning but Sonnet 4.6 at `high`. Wouldn't it be a fairer comparison if all were at max?
- Opus 4.7 writes the code
- I make GPT-5.5 in Codex to review it (given context)
- I provide the review back to Opus and ask it to verify the review findings
- Make Opus plan the fixes then execute them
- Ask GPT-5.5 to review the fixes and check if they solve the problems
My "trick" was to divide things into batches (which can be big with LLMs with larger context sizes) and classify the items in each batch, then take the resulting categories from each batch and feed them into an LLM to group semantically similar categories into groups with a representative category for each group. The representative category can be chosen from the group or created by the LLM. This is an over-simplification of the process but that's the gist of it.
Language support is not mentioned in the repo.
But from the paper, it offers extensive multilingual support (nearly 100 languages) which is good, but I need to test it to see how it compares to Gemini and Mistral OCR.
Claude Skills seem to be the option that offers highest flexibility to add more capabilities at most simplicity. Better than MCP in my opinion. Hope it becomes a standard and get adopted by OpenAI and the rest of labs.
Good question! I selected the edition with the smallest Goodreads ID¹ that has the publication date and cover photo available. If all editions don't have publication date nor cover photo, then we get the one with the smallest ID.
And you're right, in a few cases, this resulted in getting less widely read editions for some books.
1: Assuming smaller ID means earlier addition to Goodreads' database.
I have Raycast extensions for GPT and Claude models. Whenever I have a question, the most powerful LLMs in the world are two key strokes away.
This way is easier than going to the browser then ChatGPT tab for example then creating a new chat.
I found myself using LLMs more and getting more out of them because of this frictionless interaction. They've become more of actual "helpful assistants."
Can you explain more? Like which tool do you use for this wiki page? Or is it an internal tool? And do you use it to write meeting notes and then discuss on the same page?
Like many people here have noticed, it's definitely less quality now than before. It's annoying to be honest to reduce the quality significantly without a notice while we are paying the same amount. I'm willing to pay $40 for the original GPT-4, though.
GPT-4 was remarkable 2 months ago. It could handle complex coding tasks and the reasoning was great. Now, it feels like GPT-3. It's sad. I had many things in mind that could have been done with the original GPT-4.
I hope we see real competitors to GPT-4 in terms of coding abilities and reasoning. Absence of real competitors made it a reasonable option for "Open"AI to lobotomize GPT-4 without a notice.