102
I gave Qwen 3.8 27B a reverse-engineering job I assumed needed a frontier model, and it finished in 30 minutes
(www.xda-developers.com)
This is a most excellent place for technology news and articles.
I thought I was asking a lot when work asked me to pick a computer (with no guidelines) and I asked for a 36GB MacBook and justified it by saying I needed to run local models to save money. I didn't ask for nearly enough. To its credit it does run local models fast, but it's very limited in context window. It starts slowing down long before hitting context sizes I hit in frontier models.
Gorgon Halo, IIRC, goes up to a unified 192GB.
And while it depends on application, I generally agree that amount of memory is the most important factor. I started out with a 24GB RX 7950 XTX and then picked up a 128GB Framework Desktop. The larger amount of memory on the Framework is just a lot more useful than the greater bandwidth on the 7950, gives a lot more flexibility. I was always able to find useful things to do with more memory and could use more. For LLMs, more context, larger models, less quantitization. For image diffusion models, larger models, higher native resolutions without tradeoffs like having an upscaling pass, batch passes.
You can see why the cloud AI companies are hell-bent on getting all the memory that they can get their paws on.