from my understanding, you can run the inference server (llama.cpp/vllm/whatever) and the agent/harness in different contexts, event different machines.
The risky part is in the agent/harness and what tools it has access to.
You don't need to give GPU passthrough to the VM running the agent/harness.
There is still a risk of a prompt messing with the inference server, but I think that's a much lower risk compared to an agent doing whatever on its own.
Works with any mechanism to turn on and off nodes(IPMI, WoL...)
I have some nodes that I turn on and off via a curl to homeassistant to the power plug.
it's a transparent proxy that automatically launches your selected model with your preferred inference server so that you don't need to manually start/stop the server when you want to switch model
so, let's say I have configured roo code to use qwen3 30ba3b as the orchestrator and glm4.5 air as coder, roo code would call the proxy server with model "qwen3" when using orchestrator mode and then kill llama.cpp with qwen3 and restart it with "glm4.5air"
well, I tried it and it works for me. llm output is hard to properly evaluate without actually using it.
I read a lot of good comments on r/localllama, with most people suggesting qwen3 coder 30ba3b, but I never got it to work as well as GLM 4.5 air Q1.
As for using Q2, it will fit in vram, but with very small context or spill over to RAM, but with quite an impact on speed depending on your setup. I have slow ddr4 ram and going for Q1 has been a good compromise for me, but YMMV.
On my 2x 3090s I am running glm4.5 air q1 and it runs at ~300pp and 20/30 tk/s
works pretty well with roo code on vscode, rarely misses tool calls and produces decent quality code.
I also tried to use it with claude code with claude code router and it's pretty fast.
Roo code uses bigger contexts, so it's quite slower than claude code in general, but I like the workflow better.
I have a OnePlus 8t since last october and have not found any issue on the battery side. On the contrary I'm quite pleased with battery management since i can charge it in 30 minutes from 0 to almost 100%, so usually i just charge it for 5-10 minutes and I can use it all day easily.
With full charge on average I get around a day or two without recharging with mostly firefox, youtube, messaging apps open and some light gaming.
Battery saving mode gets me almost 1 hour more of usage when I'm at 15%
Haven't really had any signal problems and 5g isn't that much of an improvement around where I live so it's not really a concern for me but ymmv
usually you may catch a general exception and throw another one that's caught up the call stack.
In this case i think it may be useful for logging additional causes that are not going to be obvious with just the stacktrace(?)
Not my area either but I think dns queries are cached locally for a certain amount of time so dns lookup for 1000s of requests shouldn't give too much overhead
Edit: yeah and also what the other guys said about not relying on the ip being the same forever
>The grief factor of learning to code is on a different scale to every other major. One missing semicolon will take your whole tower down, and you realise this in the first day of practical exercises
>In fact, every exercise in CS has this problem. You add a new thing (eg inheritance), and it breaks. But not only that, it might be broken because of a banal little syntax problem.
That can be said for math as well, switch a + for a - at any point and the whole solution crumbles down, or use the wrong method of resolution and you are in for a world of pain.
>Oh and learn git, too. And Linux. Just so you can hand in the homework.
Again, comparing that to math, learn all the different notations used just to understand what the problem asks of you.
Point being that any scientific course require a part of self study to get what you are doing.
Although I can certainly agree that the way CS is taught now interwines many basic things that should be on their own course and could do with some reorganization, even though it's hard for universities to keep up with the way most of current day tech advances.
I do the same without using 33mail.
I have my mail hosted on zoho mail which gives me infinite aliases that get redirected to my main address and in case I ever need to forward a mail from an alias I can create a new address with that alias, use it and then delete it.
So when I register to a new site I usually input <sitename>@mydomain.com and then if I want I can create a filter to sort them automatically
Hijacking dns means that when you connected to the bank's website you would connect to their servers first and then they could have just proxied your connection to the real servers, that image->username check wouldn't have saved you from it since the bank's servers still operated normally
The risky part is in the agent/harness and what tools it has access to.
You don't need to give GPU passthrough to the VM running the agent/harness.
There is still a risk of a prompt messing with the inference server, but I think that's a much lower risk compared to an agent doing whatever on its own.