Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight(github.com)
github.com
Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight
https://github.com/logxio/picchio
1 comments
Does this catch intermittent fallback between requests?
Yes, picchio monitors a running llama-server on a timer and flags when the prefill/decode ratio goes CPU-shaped. I hot-swapped one mid-run. probe 4 caught ENGAGED -> NOT ENGAGED.