this post was submitted on 26 Jul 2026
283 points (99.3% liked)

Technology

87575 readers
3779 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] brucethemoose@lemmy.world 24 points 1 month ago (20 children)

And Llama and Mistral are ancient history at this point.

The cutting edge of local is lightyears better now. It's basically where ChatGPT/Anthropic were not that long ago, with a bit less world knowledge because of the size.

[–] DJKJuicy@sh.itjust.works 11 points 1 month ago (16 children)

What's the cutting edge now? Skool me...I want to try it. Can I grab one using ollama?

[–] brucethemoose@lemmy.world 15 points 1 month ago* (last edited 1 month ago) (14 children)

https://sleepingrobots.com/dreams/stop-using-ollama/

And this is just the tip of the iceberg for ollama. They're the same kind of scammy tech bros as OpenAI.

The best setup depends on your hardware. There is no "easy button" unfortunately, quantized LLMs are just too intense and finicky to run without making some informed choices.

It also depends on what you want to do with the LLM. For example, some are too slow or bad at long context for agenic use, some quantizations are great at scripts but terrible outside that, or vice versa.

But LM Studio and Qwen 3.5 35B Q4 is probably the "easiest" flat recommendation I can make.

Or... honestly, just pay $40 for basically unlimited usage for a year from an API, then roll your own frontend.

[–] DJKJuicy@sh.itjust.works 2 points 1 month ago (1 children)

Well that sucks. I was really impressed as a novice to open weight LLMs with the ease of use for Ollama on Bazzite.

[–] OnfireNFS@lemmy.world 4 points 1 month ago (1 children)

I've been running LM Studio on Bazzite and I had to do nothing to get it working. Just go to the LM Studio website and download the .appimage for Linux. If you open it with Gear Lever it will install like an app from the app store and show up in your launcher with an icon.

From there I have just been able to download models and use them from in the app. In fact I setup a local server to connect to my IDE and have been trying out local models for coding. It's pretty cool

[–] DJKJuicy@sh.itjust.works 2 points 1 month ago (1 children)
[–] Asafum@lemmy.world 4 points 1 month ago (1 children)

I can also vouch for lmstudio. If you can get Hermes running on Linux I would suggest trying that as well. It connects to lm studio and you use Hermes to communicate with the model. Iook into it as there's a lot to it, I've really been enjoying using it so far it even learns how I like to create tasks and I've stopped having to ask it to delegate certain tasks, it just knows to do it and to break down the tasks so my fairly context starved local model can handle it.

As for a model, the Qwen 3.6 family of models do really well. I'd suggest the Qwen 3.6 35B a3b probably Q4 depending on your hardware. It's large, but because it's a mixture of experts model only 3b of experts are kept on vram at any one time so it stays fast. Qwen 3.6 27b is the smarter "dense" model, but trying to stay with Q4 for quality it becomes too large for 16GB vram and for me runs at like 2 tokens per second lol

[–] brucethemoose@lemmy.world 2 points 1 month ago* (last edited 1 month ago) (1 children)

+1 for Hermes.

If you have a newer Nvidia GPU, you can run Qwen 27B via exllamav3 and get good quality/speed in 16GB. And it's worth the trouble, as 27B is an amazing model.

If it's AMD, yeah, a3B is a good bet, depending on how much spare CPU RAM you have.

[–] Asafum@lemmy.world 1 points 1 month ago (1 children)

Thanks for the tip! I do have a 5080 so I'll have to look up exllamav3

[–] brucethemoose@lemmy.world 2 points 1 month ago* (last edited 1 month ago) (1 children)

You want this one:

https://huggingface.co/turboderp/Qwen3.6-27B-exl3_3.30bpw/tree/main

Or maybe the 3.5bpw one if you don’t mind less context, or 3bpw if you need more:

https://huggingface.co/turboderp/Qwen3.6-27B-exl3

For faster inference at the cost of a little more VRAM usage:

https://huggingface.co/turboderp/Qwen3.6-27B-DFlash-exl3

And you run those in:

https://github.com/theroyallab/tabbyAPI

And FYI, if you have 64GB of RAM or more, you might consider hybrid inference instead.

[–] Asafum@lemmy.world 2 points 1 month ago

Thank you for all the links! Much appreciated :)

load more comments (12 replies)
load more comments (13 replies)
load more comments (16 replies)