Since no one else posted it... I have open-webui pointed at a linux box with 128 gig of ram and an RTX Pro 6000, and after a couple of runs on trivia, had it do one of Open WebUI's conversation starters: "Show me a code snippet of a website's sticky header in CSS and JavaScript."
72.06 t/s. That's the full Qwen 3.6 27B model BF16, using MTP, running on Ollama. Yes I know I should bite the bullet and get vllm running on that box.
That was, also, at a 570 watt limit: I normally run a little less, but when I first tried this I actually forgot I had set the limit to 300 (it's a hot day, I figured why fight the A/C?), and at 300 watts the same question came back at 69.38 t/s. (The extra power matters more for compute bound things, the difference in generating LTX2.3 videos is considerably higher... but still not linear.)
Not a shareholder, but on first try, it won't do it because it recognizes Iger's name. And clearly the deal is fresh because it balked at Mickey Mouse too. But it has no trouble with just, "mouse": https://sora.chatgpt.com/p/s_693ae0d25bbc819188f6758fce3f90c...
It's likely that I'm seeing this from my deep into ComfyUI bubble. My impression was that AUTOMATIC1111 and Forge and the like, were fading as ComfyUI was the "what people ended up on" no matter which AI generation framework they started with. But I don't know that there are any real stats on usage of these programs, so it's entirely possible that AUTOMATIC1111/Forge/InvokeAI are being used by more people than ComfyUI.
Did you test some local image gen software in that you installed the Python code on the github page for a local model, which is clearly a LOT for a normal user... or did you look at ComfyUI, which is how most people are running local video and image models? There are "just install this" versions, which eases the path for users (but it's still, admittedly, chaos beneath the surface).
Well, they sort of do: they keep referring to the 4090, on their Github and primary promotional pages (https://wan.video/).
But really all the various video models really want an 80+ gig vram card, to run comfortably. The contortions the ComfyUI community goes through to get things running at a reasonable speed on the current, dinky-sized vram consumer cards, are impressive.
As jasonjmcghee says, they're available... but if you go to ollama.com and set models to "newest" you'll see Mistral (specifically mistral-small3.2 at this writing) because they seem to not sort the models based on newest update: only newest "group" or however you'd phrase it. So you need to scroll down to "qwen3" to see it's been updated.
I use Markdown Viewer, in Chrome: I'd bet there are multiple equivalents in Firefox and Safari. Well. I don't know what Safari's extension universe is like but it seems likely.
128G of unified memory. $3K. Throw ollama and ComfyUI on that sucker and things could get interesting. The question is how much slower than a 5090, is this gonna be? The memory bandwidth isn't going to match a 512 bit bus.
That's too simple a comparison, against something that's a little more complicated. A 720p stream without adequate bandwidth can have terrible artifacts, mostly showing up during fast motion/panning/action. They can look objectively worse than DVDs.
I don't get it: the new models still need to be used somewhere, unless you're coding everything yourself, and that means they'll mostly get used in the updated ComfyUI nodes that work with the new checkpoints and LoRAs (and are available today).
Well, MS isn't fatal, but it certainly doesn't HELP... people with it have a generally shorter life expectancy (by like 7 years) from complications with something else, like heart disease or cancer.
None of which takes away from it being terrible that we've lost Terri Garr.
Worked with a lot of directors on their first or nearly first movies, like James Cameron, Martin Scorsese, Francis Ford Coppola, Peter Bogdonovich, Ron Howard. Even the Fantastic Four movie he produced had some pluses (the Red Letter Media review is worth watching).
72.06 t/s. That's the full Qwen 3.6 27B model BF16, using MTP, running on Ollama. Yes I know I should bite the bullet and get vllm running on that box.
That was, also, at a 570 watt limit: I normally run a little less, but when I first tried this I actually forgot I had set the limit to 300 (it's a hot day, I figured why fight the A/C?), and at 300 watts the same question came back at 69.38 t/s. (The extra power matters more for compute bound things, the difference in generating LTX2.3 videos is considerably higher... but still not linear.)