this post was submitted on 01 Jul 2026
108 points (94.3% liked)

Technology

87627 readers
3838 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] Dran_Arcana@lemmy.world 9 points 1 month ago (8 children)

Can you cite your source on the claim that "inference is currently insanely profitable"? Everything I read suggests that openai and anthropic lose money on their plans.

[–] jbloggs777@discuss.tchncs.de 1 points 1 month ago (5 children)

My caveats were clearly stated... After capital expenditure, it's just operational costs, where electricity & cooling are the big ones.

At that point, it is insanely profitable to serve. The cheap API prices on open weights models hints at the profit margins involved in the US (the frontier labs and hyperscalers don't open their books for us), unsurprisingly)

Therefore, the longer they can serve existing and lower cost models at the current rates, the better for their bottom line. It's just common sense in business.

It doesn't mean the company as a whole is profitable. I expect we'll see turmoil in the coming months and years, and the prize will be compute capacity, with electricity & cooling options.

[–] MangoCats@feddit.it 1 points 1 month ago (1 children)

I just asked Gemini to estimate run costs for a local GLM-5.2 instance, something that a team of a few software engineers might use the way they are using Cursor today... power budget is 6KW, which around here - after facility cooling costs - works out around $1000 per month. Our Cursor subscriptions have $100 per month price tags on them for the developers who use them most extensively, and this $100K to buy in $1K per month to run local instance isn't likely to serve more than a dozen engineers efficiently. Even if you can lease it out at full utilization 24 hours a day, it doesn't sound like much of a money printing machine to me, yet.

My $20/month home subscription to Claude? Even less so.

[–] jbloggs777@discuss.tchncs.de 1 points 1 month ago

Economies of scale... And you said instance, which is AWS terminology... If you have the scale and the expertise to run a DC efficiently, expect significant savings. We pay a premium for opex over capex.

load more comments (3 replies)
load more comments (5 replies)