Exactly right. You design with the diffusion model, then hand those designs off to an agent to implement.
It's a lot like having an architect create plans for you before handing it off to a builder. In my (obv biased) experience, you end up getting better/more creative results with this approach. You're using the best model for the job at each specific task, ie a diffusion model as the designer, and a LLM as the engineer.
Thank you! That effect is easier than you'd think. Normal maps/depth maps are actually fairly easy to generate via diffusion models. Once I had the designs for the tarot site in diffui, I just copied the build plan, pasted it into claude, and asked it to "generate normal/depth/roughness maps for each of the card designs and dynamically light and displace them based on the mouse position."
The build plan has tooling built in to generate these. Under the hood the model I'm routing to is Fal.ai's Patina model, which does a fantastic job at creating maps.
The voice activity detection alone here is compelling - very useful for doing things like highlighting a speaker who's transmitting in realtime. At that rate the impact on perf will be so minimal that you could easily run it in the browser across devices.
> taking that to a design system is a great way to make something more personal and original
Indeed, that's why the "brands" feature exists. Once you have 5 screens on canvas you like, select all of them, right click and select "Create brand". That will extract and create an internal representation of a design system so that all future screens are in that look/feel/style.
I ran Design Systems for 5 years over at Figma - this part is very important to me.
They've unfortunately been radio silent since their release, and aren't active on socials at all. I've tried to contact them for updates / api usage, but haven't heard anything back.
It's been really interesting seeing how LLMs perceive things differently than humans. I'm working on image->html conversion pipelines right now, and there are glaring issues LLMs run into that are obvious for humans. Any subtle gradients get lost, 75 degree angles get converted to 90 degree angles, etc.
This tracks towards what you're seeing with this font - the high frequency details get picked up, but the low frequency ones dont.
This is really cool. What learnings would you say you could extract on the design side from the Watergate scandal, and how do you think StyleSeed applies to those learnings?
I’d recommend a different approach if you’re looking for unique, non-LLM looking styles. Consider trying diffusion as a starting point, then feeding that into an LLM to build.
All of these started with diffusion renders as a starting point.
LLMs have intrinsic problems when it comes to design. They write code to represent design, but are trained to write consistent code. Consistent code is great if you’re writing a backend function that handles financial, but for design you end up with similar looks. You can use LLMs to police themselves, but it’s still like using backend engineers to review other backend engineers’ designs.
I'm not sure I fully agree with this being a major vuln. There's a lot of up front scary text which was raising a lot of red flags until it actually discussed the "what".
An actor has to place a malicious .exe in the user's code folder, named git.exe, for this to take place.
I see this akin to something like saying "replacing their .bashrc with an alias that says `ls` instead executes `/tmp/mega-big-virus.sh` is a vuln".
Yes it's a vector, but if they've placed something in your filesystem like that already, you've already been compromised.
Yes you do. The teams that typically win the Molokai'i Hoe (the race between the islands) are the ones that ride the waves the longest. You'll stay on the same wave for 20-30min if you're good.
The last thing you'd want is a balloon or an inflatable hampering your movements. Getting into the canoe as it's moving is a tricky, practiced and precise motion.
A hyperbolic example would be imagine a pole vaulter with a safety sausage attached to them.
1. that 15m one was the most fun I've ever had on a wave, ever. Think of it like going down a ski slope in a canoe, except the ski slope moves with you for miles.
2. we had one of the motorized lead boats sink that year due to them. Interestingly, you're more safe in the canoes on them than you are a standard boat.
> Don't think of these waves like the ones you encounter at shore. Open ocean waves are moving mountains.
> It isn't this: /(
> It's this:
.,-~^^~-,.
___.-/ \-.__
The canoes are outrigger canoes specifically designed for open ocean wave surfing. They're made to ride these mountains. There are air bladders in the front and the back, and the canoes are easily recoverable when (not if) you flip.
One fun thing you get to do in long distance outrigger canoe races in hawaii is crew changes.
Generally, outrigger races have 6 people in the boat and a 9 person team. An escort boat will hold your reserve people, and then drop them in front of the canoe when you need to swap people out.
The problem is that you need to drop people around 200m in front of the canoe so the canoe can have enough time to prep for the crew change, but with that distance, the wave height can obscure the crew from the person steering.
The solution? If you're the one being dropped, you're expected to splash violently. Create as much splash as possible so the canoe can see you, even behind a wave.
The fun part is what gives signal to the canoe is the same thing that gives signal to sharks. Our coach used to say the adrenaline helps us in the race.
This is no joke. I've done the crossing from Moloka‘i to Oahu (~45 miles) in a canoe several times, and those open ocean waves can get very nasty (largest I've dealt with were around 15m tall). I can't imagine the mental endurance required here, let alone the physical. My longest crossing took 9 hours, and I was completely drained by the time I touched shore. 44 days is absolutely insane.
Not quite. These cost-per-task benchmarks report the cost of the task after the model gives its initial answer. The total cost is irrelevant, and isn't factored into the model's decisions - a run of the full benchmark for something like Fable might cost $10k.
What I'm looking for is the inverse. I want to give the model a budget of $100, and see how much it can accomplish with that $100. For smaller models, this means they can do more than just choose thinking amount, they can do something like a /loop to keep iterating on a problem until they get it right.
Can something like Deepseek V4 Flash get more answers correct than Fable, when given equal budgets?
Think of it as answering this question: How much intelligence can you get out of a model given a budget of $100? A cost-per-task dash correlates, but it doesn't give you an answer to that question.
I want a new bench - given $100 of api spend, how much can a model accomplish for a suite of benchmark tests?
Give us something that measures a combination of efficiency and intelligence.
I think this would allow for some interesting tactics for smaller models - eg they could do things like computer use to test their results and grind on problems for longer to verify the outputs, whereas larger models may not have budget to self-test.
Working on diffui.ai - diffusion for UI design.
Formerly Figma, Atlassian, and Microsoft.
AMA about design tokens, webcomponents, and design systems!
+1 808 366 1708 [email protected]