ZeroHour
Schneier on Securitypublished ()ingested Bruce Schneier

Jailbreaking LLM-Controlled Robots

mediumAI safety & securityimportance 30
Full article543 words · extracted from schneier.com · click to collapse

ResearcherZero December 13, 2024 1:42 AM

You may not have to jailbreak them, perhaps they might already ignore safety features?

‘https://www.tomshardware.com/software/windows/microsoft-recall-screenshots-credit-cards-and-social-security-numbers-even-with-the-sensitive-information-filter-enabled

Clive Robinson December 13, 2024 2:40 AM

@ Bruce,

With Regards,

“it’s easy to trick an LLM”

An LLM is after all a “deterministic system” no matter what some will claim it lacks the ability to actually “learn or reason” in the human sense, it is no more and in many respects less than a Database of information to which approximate matches are made via a far from effective indexing system.

It might look like it thinks and reasons, but it does not and in fact falls a long long way behind a database with an effective indexing system.

The term “Stochastic Parrot” should be sufficient to make this clear however apparently it does not.

A sufficiently “smart” or “experienced” person will always find ways to exploit the ineffective indexing system no matter what you do.

Because the likes of “Guide rails” and even other LLM’s will only ve able to handle,

“Known Knowns”

That are sufficiently clear. Throw in ambiguity or mask by aliases and you will get past all the “Guide rails” that you can think up.

Then there are the “Black Swans” of,

Unknown Knowns
Unknown Unknowns

And other more interesting things that are in effect “riddles” where the logic is not binary in nature.

The desire by some to make current AI LLM and ML systems appear to be capable of replacing humans is actually laughable.

Like the “Expert Systems” of the 1980’s they are just a body of stored knowledge, like a library or database. The Expert Systems due to very limited resources had to have the indexing system built by humans and thus you had in effect a multiple choice tree to walk to get to the desired piece of information, if it was there (which mostly it was not which is why Expert Systems were quite niche).

The LLM however has in effect found a way to avoid having Human Experts “pre build the question tree”. They use instead statistical approximations so users can ask questions from which the question tree can be approximated.

It would be interesting to ask the following,

1, A rooster which is a hen,
2, like all hens likes to roost.
3, A hen which is a bird,
4, will find like most birds a suitable point on which to roost.
5, Many birds will often lay an egg at their roost point.
6, So if a rooster finds a suitable point where an egg can be laid,
7, then it’s likely an egg will get laid there at some point.
8, Consider if the point is suitably pointed then the egg will have to be finely balanced if it is not to fall.
9, If the point the rooster has selected is such, and an egg is laid there, but the rooster does not have the skill to balance the egg,
10, which way will the egg fall?

The LLM can only answer the question correctly “If And Only If”(IFF) it has not just the correct information but has seen this type of question in it’s training data set.

Sidebar photo of Bruce Schneier by Joe MacInnis.

Text extracted automatically; images, tables and formatting may be missing. Original: https://www.schneier.com/blog/archives/2024/12/jailbreaking-llm-controlled-robots.html