this site seems to be better than others ive seen : https://llama.garden/
complaints - the torrents are packs of all different quants .. or some are just safetensor files.
leanleft
ok. my mistake.
either way..
i would think you could just use a web browser. but sometimes app is more preferable
i dont think you need google play services for banking.
furthermore, i disabled play services. i dont need it for 99% of apps.
sometimes theres one cool heavily proprietary app that forces you to use it.
sucks. and i just gotta uninstall it.
i cant get any LLMs to consult reddit at this point.
i suspect, shit like this is going to prompt some idiotic supreme court verdict is going to be set in stone for the next 10 years.
example: https://en.wikipedia.org/wiki/City_of_Grants_Pass_v._Johnson
if its not a guarenteed win, then dont push it to supreme court level. the consequences are severe for everyone.
if you start excluding the 1000+ B param models ..
using smaller models, would initially ease hardware demand by 60% .
OFC you cant.. and probably shouldnt, ignore and disrespect SOTA flagship models
i wish we could focus on making small LLMs better.
some companies are doing this. some definitely aren't
maybe youtube wants us to switch to torrents.
piracy? yes.
but they(youtube) honestly cant even afford to serve the content and remain profitable.
so....
low-level compilers can output very ugly-looking assembly. he probably did this and then used LLM to super-optimize it. may be perfomant, but id guess that theres a risk that its unsafe.
personally, i would be cautious with discord too
i distilled this article
Summary of the article “How China gets better bang for its buck than America in AI” (Aug 3 2026)
-
U.S. AI spending is massive – Bloomberg Intelligence estimates U.S. data‑centre capital outlays could exceed $740 billion in 2026, with Nvidia alone negotiating a $250 billion financing deal for a $500 billion data‑centre run by OpenAI. Alphabet announced a $205 billion AI budget.
-
China spends far less – Chinese tech firms are projected to invest less than one‑tenth of the U.S. amount in data centres. Yet their models perform only slightly behind U.S. equivalents. For example:
- K3 (Moonshot AI) scores ≈ 95 % of Anthropic’s Fable 5 on common benchmarks while being 70 % cheaper to run.
- Alibaba’s newly released model ranks among the world’s best on certain metrics.
-
Why Chinese spending is efficient
- Lower input costs – Land, construction, equipment and labour are cheaper in China.
- Model distillation – Chinese labs often train models using outputs from expensive U.S. models, reducing the compute needed.
- Hidden spending – Some expenditures on high‑end chips are masked as “cost‑saving” techniques that make inferior hardware achieve higher performance (e.g., DeepSeek’s efficiency tricks).
-
Export restrictions limit Chinese capital use – U.S. bans on advanced AI chips (Nvidia designs, TSMC manufacturing) prevent China from buying the most powerful hardware.
- Chinese firms are pushed toward domestic alternatives (Huawei, SMIC).
- Sanctions also block access to cutting‑edge chip‑making equipment, forcing costly work‑arounds and capping production capacity.
-
Domestic demand constraints – Chinese enterprises spend < 10 % of what U.S. firms spend on IT, despite China’s GDP being two‑thirds of the U.S. (or a third larger in PPP terms). This throttles revenue prospects for AI providers, curbing their willingness to invest heavily.
-
Strategic focus differs – The Chinese Communist Party emphasizes diffusing AI across the economy, not pursuing a race toward artificial general intelligence (AGI). Fewer than ten Chinese firms target AGI, compared with dozens of U.S. players.
-
Investor attitudes – Chinese investors have historically punished over‑spending on AI, whereas U.S. investors once rewarded aggressive budgeting. This cultural difference keeps Chinese AI budgets modest.
-
Potential bottlenecks for China – Despite restraint, China may face compute shortages:
- ByteDance experiences ten‑hour processing times for some videos.
- Alibaba Cloud, Zhipu AI, and Moonshot’s K3 have long waiting lists or quickly sell out capacity.
- Over‑restriction could stifle growth if AI services cannot meet user demand.
Overall takeaway: China achieves comparable AI performance to the U.S. while spending a fraction of the capital by leveraging cheaper resources, model‑distillation techniques, and a strategic focus on wide‑scale diffusion rather than raw computational power. However, export bans, limited domestic chip capacity, modest corporate demand, and cautious investors together create both an efficiency advantage and a risk of under‑provisioned infrastructure.
i just read this: https://reddit.com/comments/1w01y1f
With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the
llama.cppproject, together with the entire team behind it.In February 2026 the llama.cpp team was employed by HF in order to continue working on llama.cpp and the ggml library.
This includes:
Now with the acquisition, llama.cpp's future looks a lot less certain given Nvidia's poor track record with open-source.
This is still rather speculative at this stage, but it's definitely possible for the llama.cpp project to change in the future: either by switching to a different license, or by having staff redirected to other projects within the larger company.
Even when a project is open-source the copyright owner has complete control over it, and they can change licensing as they wish.
This has happened before with projects like Redis, Minio, and others.
Source:
https://huggingface.co/blog/ggml-joins-hf
Edit:
The original announcement from Feb 2026 from Gerganov gives a few more details:
https://github.com/ggml-org/llama.cpp/discussions/19759