How well can language models like Claude Opus and GPT-5.2 write music?
With boogiebench, I ask models make strudel compositions (https://strudel.cc/) in response to music prompts ('hyperpop', 'spaghetti-western theme', etc) and generate ELO rankings based on user votes.
Unlike Suno, LLMs haven't been trained explicitly on this task, making it a nice generalization test (coding, aesthetics, temporal reasoning), akin to the pelican-riding-a-bike with SVGs.
Models often struggle but are rapidly improving, judging by the performance gap between the strongest and weakest models. (The Anthropic models seem to underperform other model families relative what we'd expect, for whatever reason).
It is far from given that what the author predicts will happen will happen. And even if it does, this argument could have been made about any labor-saving technology in the past. “I wish Photoshop had never happened”.
"Hagen told me that her favorite white-space flavor—the one she wished she had created—was Red Bull, because it succeeded in getting consumers to embrace the surreal."
With boogiebench, I ask models make strudel compositions (https://strudel.cc/) in response to music prompts ('hyperpop', 'spaghetti-western theme', etc) and generate ELO rankings based on user votes.
Unlike Suno, LLMs haven't been trained explicitly on this task, making it a nice generalization test (coding, aesthetics, temporal reasoning), akin to the pelican-riding-a-bike with SVGs.
Models often struggle but are rapidly improving, judging by the performance gap between the strongest and weakest models. (The Anthropic models seem to underperform other model families relative what we'd expect, for whatever reason).