Does this require a given hardware platform, or for example, could I run this model with a cheap mini robot with a single camera (i.e. good enough without depth), tracked treads, and an arm?
Oh and also side mounted a 4GB 750Ti for graphics, and/or whisper use, parallel to the mainboard.
Airflow from the front fans to back is unimpeded by anything in this way.
It was AliExpress, during sales, earlier this year, with discounts from playing games. I typically used to see much higher eBay prices.
A bit of hyperbole tbh, not everyone has the patience to do this but I was honest that this is what I got them for. The DDR shortage has pushed the current price up to $130 (?) or >£85 after all discounts now.
It typically runs 120W idling, and 450W peak.
Not great power efficiency wise, but I have a Cron midnight shutdown and and S5 bios event powering it up on a morning (I could use wake on Lan but didn't see the point).
It's minimal kind of speed but for the price I feel it's quite good. Smaller models fly as you'd expect, but the fact it can load these mid range models into something that is dirt cheap is generally a win imho.
No mention of the venerable Tesla P4. 75W peak, 8GB VRAM, about $80 (£60).
I have 6x P4s, a Xeon E5 2696v3 (36 threads, 3.8ghz peak but all core turbo unlocked, so 6 cores at 3.8Ghz - about 8 cores at 3.5ghz, or all cores at 3.1ghz), 48GB DDR4, all fit into a micro atx case running on a 650W MSI psu.
This gives me a virtual 48GB GPU (llama.cpp ftw) to backup that 48GB of RAM.
I typically see scores of at least 7-12t/s on 20-30B Q4KM size dense models, on a 32K/48K/64K context, adequate for modern inference.
The pain point is the prompt loading, it is far far slower, minutes not seconds, than modern tensor core 8GB 5060s (my other machine's 2x GPUs) but is quite similar in regular inference speed once it has loaded.
I don't know tbh, I run Qwen 3.6 27B, Q6, on two 5060TIs with 128K context (2 parallel, 256K total) and get better fairly good performance. It's not perfect, but that's why I also use a low quant Gemma 4 validation model to confirm outputs. It's probably comparable to early chatgpt in terms of ability.
We used to flip display upside down in display options, which also reverses the mouse. We'd then lock the PC and disconnect the keyboard.
After they figured out the keyboard had been pulled they often couldn't work out why their screen was upside down...
I've only an A-Level in Further Maths from 1997, but understand complex numbers and have come across complex inverse trig functions before.
My takeaway for other people like me from this is "computer is correct" because the proof shows that we can't define arccosh using a single proof across the entire complex plane (specifically imaginary, including infinity).
The representation of this means we have both complex functions that are defined as having coverage of infinity, and arccosh, that a proof exists in only one direction at a time during evaluation.
This distinction is a quirk in mathematics but means that the equation won't be simplified because although it looks like it can, the underlying proof is "one sided" (-ve or +ve) which means the variables are fundamentally not the same at evaluation time unless 2 approaches to the range definition are combined.
The QED is that this distinction won't be shown in the result's representation, leading to the confusion that it should have been simplified.
Also, cheaper... X99 + 8x DDR4 + 2696V4 + 4x Tesla P4s running on llama.cpp.
Total cost about $500 including case and a 650W PSU, excluding RAM.
Running TDP about 200W non peak 550W peak (everything slammed, but I've never seen it and I've an AC monitor on the socket).
GLM 4.5 Air (60GB Q3-XL) when properly tuned runs at 8.5 to 10 tokens / second, with context size of 8K.
Throw in a P100 too and you'll see 11-12.5 t/s (still tuning this one).
Performance doesn't drop as much for larger model sizes as the internode communication and DDR4 2400 is the limiter, not the GPUs.
I've been using this with 4 channel 96GB ram, recently updated to 128GB.
The irony of this is that Gen-Z have been mollycoddled with praise by their parents and modern life, we give medals for participation, or runners up prizes for losing. We tell people when they've failed at something they did their best and that's what matters.
We validate their upset feelings if they're insulted by free speech that goes against their beliefs.
This is exactly what is happening with sycophantic LLMs, to a greater extent, but now it's affecting other generations, not just Gen-Z.
Perhaps it's time to rollback this behaviour in the human population too, and no I'm not talking reinstating discipline and old Boomer/Gen-X practices, I'm meaning that we need to allow more failure and criticism without comfort and positive reinforcement.