isVeryLoud

joined 3 years ago
[–] isVeryLoud@lemmy.ca 0 points 1 day ago (1 children)

Higher efficiency = less hardware demand

[–] isVeryLoud@lemmy.ca 1 points 1 day ago (3 children)

Local models tend to be more efficient since people will be more likely to run compressed and MoE models.

Also, it's basically 1 GPU >= 1 request for the most part in data centres, each request is its own LLM. Each time you make a new request after a set timeout, model weights get loaded in VRAM, context gets initialized, the query gets parsed, and it spits out tokens. These frontier models can be 300 GB in size or more, which all needs to be kept in VRAM for best performance, usually distributed across multiple GPUs, and each loaded model can only answer one query at a time.

Compare this to someone like me, trying to cram Qwen 3.6 MoE on the 16 GB RX 6800XT I already have in my own computer, not using up drinking water or prime real estate to cool my PC, powered using hydroelectricity.

It's so much more eco-friendly and economical, I wish frontier models would just die tbh, or at least only be used for distillation. LLMs offer diminishing returns past a certain point, and you can get 90% of the frontier model with a MoE local model.

[–] isVeryLoud@lemmy.ca 25 points 2 days ago

This isn't even the case for those LG monitors, they trigger a Windows Update download for malware.

[–] isVeryLoud@lemmy.ca 1 points 1 week ago

Fash take, just kill it. LLMs should belong to the people, just like information should belong to the people.

[–] isVeryLoud@lemmy.ca 17 points 1 week ago

My work would be impossible without WSL, and yes I am guilty of hiding from the security systems... Because the security systems are so invasive they keep locking files I'm trying to open so IDEs don't even work properly under Windows.

[–] isVeryLoud@lemmy.ca 4 points 1 week ago

Yeah that's called transformation, which is fine. Means a human went over it.

[–] isVeryLoud@lemmy.ca 5 points 2 weeks ago

The LLM has no idea whether you entered text into gemini.google.com or google.com, it's all the same to it. LLMs don't even know what time and date it is unless you feed it to its context.

As far as it is concerned, you are on the Gemini web app. The rest of the page is search results, but who knows what special sauce Google put in there.

[–] isVeryLoud@lemmy.ca 4 points 2 weeks ago

Pot meet kettle

[–] isVeryLoud@lemmy.ca 2 points 2 weeks ago

Sounds like the LLM is limited by its context window lol

[–] isVeryLoud@lemmy.ca 4 points 3 weeks ago (1 children)

I love you please bear my children

[–] isVeryLoud@lemmy.ca 7 points 3 weeks ago (1 children)

Is that a US only thing?

view more: next ›