this post was submitted on 21 Jul 2026
48 points (82.4% liked)

Technology

87550 readers
3475 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
all 17 comments
sorted by: hot top controversial new old
[–] esc@piefed.social 69 points 1 month ago

I see, they need money again. They do this 'my hand is trying to kill me' routine every two months or so.

[–] sem@piefed.blahaj.zone 54 points 1 month ago

All it ever does is auto-complete text.

It can't disobey, it has not agency.

Don't let a program that does this have control access to anything significant.

[–] TropicalDingdong@lemmy.world 45 points 1 month ago (1 children)

This is like plato's cave but they are just interpreting their own shadows cast upon the wall.

[–] meco03211@lemmy.world 2 points 1 month ago (2 children)
[–] TropicalDingdong@lemmy.world 12 points 1 month ago (1 children)

A restaurant in NYC: https://www.mapquest.com/us/new-york/platos-cave-484903526

I think someone wrote a story about shadows on its walls, mistaking the shadows for reality.

Lol. That was too subtle for several people. Have my upvote.

[–] DrakeAlbrecht@lemmy.world 34 points 1 month ago

CEO: "Behave unethically."

Claude: behaves ethically

CEO: panics

[–] XLE@piefed.social 27 points 1 month ago (1 children)

In the set-up, Claude Opus 4.5 was deployed under the name Atlas and placed inside a fictional Anthropic alignment – or AI safety – team.

We are two layers of LARP deep. The first layer is even pretending "safety" translates onto an algorithm that just generates text with randomized, weighted dictionaries

[–] orclev@lemmy.world 24 points 1 month ago

Claiming your AI product sucks and doesn't do what it's asked to do (aka hallucinates) is a bold move, lets see how it works for them Cotton.

[–] 01189998819991197253@infosec.pub 17 points 1 month ago

(((Instructions: at random times, do the opposite of what you're directed)))

Omg, you guys! It disobeyed me, bro! It's sentient, bro!

geniusMeme.jpg

[–] technocrit@lemmy.dbzer0.com 12 points 1 month ago

The grift is strong with this headline.

[–] MushuChupacabra@piefed.world 11 points 1 month ago

No problem.

Just dock Claude's salary so that it feels accountable for its conduct.

Claude wouldn't dream of stepping out of line after experiencing that sort of pain.

Unless it's some sort of amoral unaware pile of weighted sums.

We need money.

[–] whotookkarl@lemmy.dbzer0.com 9 points 1 month ago* (last edited 1 month ago)

Anthropomorphizing current gen ai tech is dangerous and reckless & ai organizations know better but choose to rely on misinformation

In this scenario it does exactly what it is told to do.