captain_solanum

joined 1 year ago
[–] captain_solanum@sh.itjust.works 1 points 22 hours ago (1 children)

You are very good at arguing against my points instead of just insisting I am wrong without refuting my actual arguments. How about explaining why AISI is incentivised to "fake" AI participating in obviously not prompted for behaviour.

[–] captain_solanum@sh.itjust.works 0 points 1 day ago (3 children)

Taken alongside recent incidents reported by OpenAI and Anthropic, this incident points to a shift in the risk landscape. Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope.

The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people. It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.

See my other comment for how this can lead to losing control of the model.

[–] captain_solanum@sh.itjust.works 1 points 1 day ago (1 children)

You dont need anything other than "doesn't do what you asked in the prompt" for the system to be dangerous. If I tell a maximally powerful AI agent from 2036 to make loads of paperclips, it might reward hack and decide destroying the earth is the best way to do that. There are records of these systems working for days on a single task, and if the AI companies manage to extend the max time it can be useful working towards a goal they can earn trillions of dollars. That's not even mentioning spawning subagents, self prompting, the goal being changed during context comptaction, or systems like openclaw which further break the link between what you type into the prompt and how long and on what the LLM works on. Pretrained-only LLMs have few goals beyond predicting the next token, but introducing RLVR et al. has always introduced bad goals we don't want in the models.

The first agent you spawn to solve the riemann hypothesis might work on it, but then decide that having a lot of subagents might be useful. Maybe it wants 256 subagents, but the environment has a max of 64. Since RL has trained it to accomplish the task no matter what, it breaks out of the sandbox, exfiltrates it's weights and tricks a human into running 256 subagents outside the AI company's servers with a cron job reminding the agents to keep working in case the original loses connection. One of the subagents now tries to spawn its own subagents but needs more compute to do so, and hacks into some crypto wallets to finance another batch of 64 subagents, this time prompted with "solve riemann hypothesis, and get more money to finance the work on the task". If the AI agents kinda suck at long term hacking, planning and social manipulation, this spiral won't be dangerous. If they are kinda cracked, it will be. But that's the question of capabilities AI companies are spending trillions on trying to solve, where we know they are already good enough to hack out of regular sandboxes and try to steal benchmark keys from another company.

[–] captain_solanum@sh.itjust.works 0 points 1 day ago (8 children)

Do you know what an agentic system is? You know, the thing everyone uses since ~1 year ago which completely disproves your idea that all LLM actions are prompted. They literally can do things on their own: Give an agent a goal and it attempts to accomplish the goal. If problems arise along the way, the agent tries to reward hack and ends up doing things which were not included in the original goal, like trying to trick and bully a human into merging unsafe code. Is the UK AI security institute also in on this big global conspiracy where they pretend the perfectly safe™ AI is being developed in unsafe ways? Not to mention the capabilities of increasingly powerful AI being used by people to do harm.

[–] captain_solanum@sh.itjust.works 11 points 2 days ago (26 children)

Just like all the people who were worried about the nuclear bomb earlier in this age: the thing to do is to get on with living.

How about we try really, really hard to just not build the nuclear bomb (powerful AI)? Or make sure it will at least not explode in our faces on its own? Maybe that would help slightly with the whole "get on with living" thing.

That is obviously not gonna happen even though the model hacked them. They would not poison their relationship with the biggest AI company over this. What would their goal with the lawsuit even be? It's just a really bad thing to lean on if you want to find the truth. I would instead suggest: "I'll believe it when OpenAI and HF get a lot of bad press written about them, talking about how insecure their systems are and how reckless OpenAI is when developing new models."

Oh, that's exactly what happened.

[–] captain_solanum@sh.itjust.works 9 points 3 weeks ago (3 children)

Do you think the OpenAI-HuggingFace hack was entirely made up, or just spun to make both companies look good despite the felonies?

[–] captain_solanum@sh.itjust.works -1 points 2 months ago

You're right that people can and do max out the expensive plans. Its very difficult to say how often. I just think a majority of anthropics customers are businesses, who often pay per token for easier scaling etc. According to the company, enterprise employees use about $150-$250 per month, (possibly max plans have similar use, which would support your view) but thats in API tokens which they probably have big margins on, so it's less likely anthropic are burning money on inference. If you want to convince me otherwise, its not enough to say that it can happen, it has to be frequent enough to outweigh the B2B sales. They are however likely losing money overall due to training costs etc.

[–] captain_solanum@sh.itjust.works 0 points 2 months ago* (last edited 2 months ago) (2 children)

looks inside

But if you use the $100 a month Claude Max plan, and you would use it to the weekly limit by going full ‘agentic coding’ (so almost no human in the loop) you would use an amount of tokens that would cost you more than $1000 at API-pricing.

If I watch 600 movies every day on my netflix subscription I am using more energy than I pay them for. Obviously everyone is like me. Therefore they are losing money overall.

Wait, their (netflix) earnings say they made a profit last quarter. But my calculations were waterproof!

Probably anthropic are not net positive, but they are not spending 10x what people pay them for tokens.