this post was submitted on 25 Aug 2026
42 points (85.0% liked)
Technology
87598 readers
3638 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
We need to get rid of the surveillance architecture, but AI is only really dangerous to the cognitive abilities of those who use it.
It literally cannot do things on its own. (I know there are a lot of articles who say it can, but they’re just lying to you for money).
Do you know what an agentic system is? You know, the thing everyone uses since ~1 year ago which completely disproves your idea that all LLM actions are prompted. They literally can do things on their own: Give an agent a goal and it attempts to accomplish the goal. If problems arise along the way, the agent tries to reward hack and ends up doing things which were not included in the original goal, like trying to trick and bully a human into merging unsafe code. Is the UK AI security institute also in on this big global conspiracy where they pretend the perfectly safe™ AI is being developed in unsafe ways? Not to mention the capabilities of increasingly powerful AI being used by people to do harm.
Which involves prompting it. It may take a thousand unexpected twists in the process but it's still starting from the single prompt. It has no will and desire of its own, it cannot chose to act without human intent knocking down the first domino. So don't ask it to do anything and it's just a trillion dollars of silicon and code sitting there boiling water.
You dont need anything other than "doesn't do what you asked in the prompt" for the system to be dangerous. If I tell a maximally powerful AI agent from 2036 to make loads of paperclips, it might reward hack and decide destroying the earth is the best way to do that. There are records of these systems working for days on a single task, and if the AI companies manage to extend the max time it can be useful working towards a goal they can earn trillions of dollars. That's not even mentioning spawning subagents, self prompting, the goal being changed during context comptaction, or systems like openclaw which further break the link between what you type into the prompt and how long and on what the LLM works on. Pretrained-only LLMs have few goals beyond predicting the next token, but introducing RLVR et al. has always introduced bad goals we don't want in the models.
The first agent you spawn to solve the riemann hypothesis might work on it, but then decide that having a lot of subagents might be useful. Maybe it wants 256 subagents, but the environment has a max of 64. Since RL has trained it to accomplish the task no matter what, it breaks out of the sandbox, exfiltrates it's weights and tricks a human into running 256 subagents outside the AI company's servers with a cron job reminding the agents to keep working in case the original loses connection. One of the subagents now tries to spawn its own subagents but needs more compute to do so, and hacks into some crypto wallets to finance another batch of 64 subagents, this time prompted with "solve riemann hypothesis, and get more money to finance the work on the task". If the AI agents kinda suck at long term hacking, planning and social manipulation, this spiral won't be dangerous. If they are kinda cracked, it will be. But that's the question of capabilities AI companies are spending trillions on trying to solve, where we know they are already good enough to hack out of regular sandboxes and try to steal benchmark keys from another company.
Which is why not using them, not becoming dependent on them, avoiding them to the greatest degree possible is crucial.
You cannot build a model that is safe and viable under the current incentive structure. So don't. These companies are going to burn under the weight of their own debt if adoption is minimal. Let them burn. The less you need them, the less likely you are to have a problem when Sodom and Gomorrah burn.