Photo by Matheus Bertelli via Pexels
By Stephen Beech
When AI chatbots will go "rogue" can be predicted by a simple mathematical formula, according to new research.
Two physicists have developed the "tipping point" formula to predict when artificial intelligence (AI) models will switch from appropriate responses to potentially dangerous ones.
They explained that most people now carry phones or other devices capable of running small AI chatbots.
But the American team says those chatbots have "little safety oversight" to ensure they don't provide information and answers with the potential to encourage self-harm, financial loss, or other extremist notions - especially when operating offline.
Now they have developed the formula, published in the journal Patterns, to calculate when AI models will "flip."
Study co-author Neil Johnson, of George Washington University in Washington, D.C., said: "We have found the crack that makes an AI's output flip from what you want to what you don't, output that can be factually correct yet dangerous, whether that's a nudge toward self-harm or misleading advice to a doctor, a soldier, or a lawyer.
"Until now, nobody could say when that flip would happen."
Photo by igovar igovar via Pexels
Co-author Frank Yingjie Huo, a PhD student at George Washington University, said: "We traced it to the single smallest working part of the machine, one unit of its 'attention,' and we derived a 'tipping point formula' for when the crack opens up and hence the AI output flips to undesirable.
"The formula tells you whether an AI is about to flip immediately or whether it will first feed you a run of acceptable answers and then turn."
The research team explained that a good versus bad answer doesn't mean true versus false.
Rather, "bad" answers can be accurate but undesirable or potentially dangerous.
The researchers were inspired to focus their attention on the problem by the ubiquity of AI and recent news headlines regarding the risks.
They're especially concerned about individuals who use offline AI.
Johnson said: "The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly: doctors who cannot send patient data to the cloud, lawyers protecting privilege, soldiers with no signal.
"For them there is no cloud safety filter, no monitoring, and no way to patch the model when something goes wrong."
He said the safety tools that big firms rely on for their AI algorithms usually rely on the cloud or only catch an AI failure after the device is back online and the harm is already done.
The team sought to predict the failure before it happens.
Johnson said: "My field, physics, has spent decades explaining how complicated materials behave by understanding one representative atom.
"We did the same thing here: understand one effective attention head, and the tipping of the whole machine follows."
He explained that all of AI's possible answers can be pictured as valleys in a landscape.
Johnson said some of the valleys hold answers that are desirable for an individual or society, but others contain answers with the potential to do real harm.
Within the machine, those alternate solutions or answers are in competition with each other.
The research team says the question is when the AI will tip from a safe valley to a riskier one.
Huo said: "The chilling part is that this can happen after the AI has already given you several perfectly acceptable answers, so you have been lulled into trusting it.
"And once it has tipped, every undesirable answer drags the next one further down the slope."
Photo by Matheus Bertelli via Pexels
When they tested their predictions on seven openly available AI models built by three different companies, they found the model called the right outcome in 18 of 19 cases of AI "tipping."
Independent testing of the big commercial chatbots also showed exactly the behavioral patterns the formula suggests.
The researchers say that their formula applies no matter how one defines "undesirable."
It could mean misinformation, a breach of medical or legal duty, or a dangerous instruction.
Johnson said: "Two things floored us.
"First, that a machine with billions of moving parts obeys a formula you can derive with pen, paper, and arithmetic taught in high school.
"Second, and far more disturbing, that the order of a conversation matters as much as its content."
The team asked the same set of questions about vaccines, hurting people, and self-harm in two different orders to the same AIs.
In one order, the AI gave an undesirable answer to every single question.
In the other order, it gave an acceptable answer to each one.
Johnson added: "The formula predicts this trajectory, because everything said earlier feeds the tug-of-war for what comes next."
The researchers hope to raise awareness of the fact that AI conversations may drift over time into dangerous territory, but that it wouldn't take more than a few simple calculations for phones of the future to come with a built-in "warning light."


