this post was submitted on 01 Jul 2026
108 points (94.3% liked)

Technology

87598 readers
3638 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] jbloggs777@discuss.tchncs.de -1 points 1 month ago (9 children)

There is also a commercial aspect...

Bigger models are more expensive to train and serve..

Inference is currently insanely profitable if you have the hardware and the automation in place to support and serve it. At that point, it's a money printing machine, and you want to squeeze as much out of it as you can.

While training new models is extremely expensive, and serving them probably makes less profit (at least initially).

Having an external brake applied to the frontier labs is likely good for their bottom line, while increasing hype and directing customers' annoyance away from them.

It's likely only a temporary benefit, though. The dragon will catch up and apply more pressure, both on inference price and capabilities.

[–] Dran_Arcana@lemmy.world 9 points 1 month ago (8 children)

Can you cite your source on the claim that "inference is currently insanely profitable"? Everything I read suggests that openai and anthropic lose money on their plans.

[–] stsquad@lemmy.ml 1 points 1 month ago (1 children)

I suspect it's profitable in the abstract - and their accountants would be bad at their jobs if they couldn't work out what utilisation rate you need to pay for the server runtime.

However how aggressively you amortise the cost of the training is the key, especially if you keep releasing new models every 6 months.

[–] MangoCats@feddit.it 1 points 1 month ago* (last edited 1 month ago)

20 years ago, after 20 years of watching computers get faster and cheaper, I felt like they were "fast enough" - I mean, sure, more faster is more better, but for everything I had used computers for up to that point, they were fast enough - hell, they were already streaming DVD quality video by then on "normal" laptops. Certainly computers today are much faster still, but so much of that performance feels wasted on bloat rather than enhancing actual user experience.

LLM models seem to be evolving faster. A year ago, they were nowhere near good enough, but you could see the potential, much like desktop computers in the mid 1980s. Just make them faster, more powerful, more storage, higher resolution, you'll really have something. Today, I feel about the LLMs (for code) almost like I felt about computers in 2006 - they're good enough. Of course they could always get better, but if I were stuck with what we've got today for the next 5 years, I wouldn't be too disappointed. The interesting question (that nobody seems to have a real answer for) is: how much better will they get. A year ago there were obvious rough edges that have quickly been smoothed off... how smooth can they actually get?

LLMs for graphic arts? Yeah, that feels like MS paint levels of performance at the moment, they definitely have room for improvement.

load more comments (6 replies)
load more comments (6 replies)