775
Generative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.
(www.theatlantic.com)
This is a most excellent place for technology news and articles.
Yeah I'm just beginning my local AI journey on a 5080, tried Qwen3.6 27b Q4 and was getting like 1tps because of the vram overflow. Ran it over night at it was still chewing on generating a prompt for a sub agent when I got up in the middle of the night until it simply ended in some kind of "fetch failure" lol. I think I gave it something too large to tackle, but either way 1tps is kinda garbage.
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It's the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well