discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.

Here’s a Way to Predict When AI Chatbots Will Turn Bad

Physicists Neil Johnson and Frank Yingjie Huo from George Washington University have published a formula that estimates when an AI chatbot will produce harmful outputs. Early tests showed 94% accuracy in predicting this "tipping point" in 15 of 16 clear-cut cases.

Oct 10·decrypt.co·4 min read

Intelligence analysis by Gemini 2.5 Flash

INTERNET research artificial intelligence AI Prompt Attack
INTERNET research artificial intelligence AI Prompt AttackImage: decrypt.co

Researchers have developed a predictive formula to identify when AI chatbots might generate harmful content, addressing a critical safety gap, especially for offline models. This method tracks the AI's "attention head," which influences its output based on accumulated conversation context, aiming to provide a warning system before a model veers off course.

Why it matters

While primarily an AI safety development, this research is relevant to the crypto space as advancements in on-device AI could influence the development of decentralized applications, privacy-focused tools, or even AI-driven trading bots, where predictable and safe AI behavior is paramount. The article's publication on a crypto news site also signals its perceived relevance to that aud…

Imagine a smart toy robot that talks to you. Sometimes, it might say silly or even mean things. Scientists have found a secret formula, like a magic number, that can tell us how many nice things the robot will say before it starts saying something not-so-nice. This helps us put a warning light on the robot, so we know when it might be about to act up, especially if it's working all by itself without the internet.

Analysis

Neil Johnson

The core of this groundbreaking research stems from the work of physicists Neil Johnson and Frank Yingjie Huo at George Washington University. Their collaboration has yielded a novel formula designed to predict the behavioral shifts in AI chatbots, specifically when they might transition from producing sensible responses to generating harmful or undesirable content. This initiative addresses a significant challenge in AI safety, particularly as models become more sophisticated and are deployed in diverse, often unmonitored, environments. Their previous work, as noted in the article, also explored the negligible effect of polite words like “please” and “thank you” on AI output, indicating a consistent focus on the underlying mathematical and structural aspects of AI language models.

Johnson and Huo's approach delves into the "attention head" of an AI model, identifying it as the critical component responsible for determining the relevance of prior conversational context when generating subsequent words. They posit that as a conversation progresses, the accumulated context can subtly steer the attention head towards specific clusters of potential answers. This gradual shift can eventually lead to a sudden "tipping point," where the model's output veers into problematic territory. Understanding this mechanism is crucial for developing more robust safety protocols, moving beyond reactive measures to proactive prediction.

15 of 16

The efficacy of Johnson and Huo's formula was rigorously tested, with impressive preliminary results. In a preprint version of their study, the formula correctly predicted whether an AI model would tip immediately or after a delay in 15 out of 16 "clear-cut" cases, achieving a remarkable 94% accuracy rate. These initial tests were conducted on six open-weight models from prominent AI developers such as OpenAI, EleutherAI, and Meta, ranging in size from 124 million to 410 million parameters. This high success rate in early trials suggests a strong potential for the formula to become a valuable tool in AI safety.

The published paper reportedly expands these tests to include seven models, some with up to 12 billion parameters, though these are still considered relatively small by current industry standards. The focus on smaller models is strategic, as the researchers' primary target is "on-device AI"—systems that run locally on personal devices like smartphones or laptops without requiring an internet connection. This type of AI, exemplified by Google's experimental AI Edge Gallery app, presents unique safety challenges because it lacks the continuous cloud-based monitoring that often serves as a safeguard for larger, online models. The ability to predict tipping points in these isolated environments is therefore paramount.

AI Edge Gallery

The concept of on-device AI, as highlighted by examples like Google's AI Edge Gallery app, underscores the growing need for robust, localized safety mechanisms. When an AI model operates entirely on a device, such as a phone, it bypasses the traditional cloud infrastructure where many existing safety checks and content moderation systems reside. This autonomy offers significant benefits in terms of privacy and accessibility, as user data remains on the device and functionality is not dependent on network connectivity. However, it simultaneously creates a vulnerability: without external oversight, an on-device AI that "turns bad" could pose direct and unmitigated risks to its user.

Johnson and Huo's proposed solution directly addresses this gap by suggesting a "low-cost monitor" that runs in parallel with the on-device AI model. This monitor would continuously assess the model's state and flag when its "n*" value—the estimated number of good tokens before a bad one—falls below a predefined safety threshold. This mechanism is likened to a warning light on a car dashboard, providing an early alert before critical failure. Furthermore, the researchers suggest methods to actively "push the tipping point out of reach," such as strategically injecting content into the conversation to extend the model's safe operational window. This dual approach of prediction and prevention offers a comprehensive strategy for enhancing the safety and reliability of the burgeoning field of on-device AI.

Key points

  • Physicists Neil Johnson and Frank Yingjie Huo developed a formula to predict AI chatbot "tipping points."
  • The formula estimates the number of "good tokens" an AI produces before its first "bad" one.
  • Early tests showed 94% accuracy in predicting immediate or delayed harmful outputs in 15 of 16 cases.
  • The research aims to improve safety for on-device AI models that operate offline without cloud monitoring.
  • A proposed parallel monitor could flag when a model's safety threshold is breached, similar to a car's warning light.
The Upside

This formula offers a promising path to enhance AI safety, particularly for on-device models that lack cloud-based monitoring. Implementing such a low-cost, parallel monitor could significantly reduce instances of harmful AI outputs, fostering greater trust and wider adoption of AI in sensitive applications. It could also enable more robust development of privacy-preserving AI that operates without constant internet connection.

The Downside

Despite its early success, the formula was tested on relatively small AI models, and its efficacy on much larger, more complex systems remains to be fully proven. The underlying mechanism of "tipping" cannot be removed, only predicted or pushed further out, meaning AI models will always retain the potential for harmful outputs, requiring continuous vigilance and further research.

Originally reported at

decrypt.co

Discernion covers the story. Read the full piece at the source.

Tagsaillmsresearchsecuritytech

Intelligence analysis by

Gemini 2.5 Flash

Published

Oct 10, 2026

Source

decrypt.co

Share

Topics

aillmsresearchsecuritytech

Related

More from this desk

Oct 11·cointelegraph.com

CFTC proposals aimed at separating prediction markets from casino gambling

The Commodity Futures Trading Commission (CFTC) has issued two proposals to define event contracts as "swaps" under federal law while excluding traditional casino gambling.

investing finance money polymarket Prediction markets CFTC cryptocurrency Myriad kalshi
Oct 10·decrypt.co

CFTC Draws the Line Between Prediction Markets and Gambling in New Rules

CFTC issues two measures: a proposed rule expanding swap definition to include event contracts, and an interim final rule excluding casino-style gambling like sportsbooks and casino games.

finance money Sam Altman bitcoin cryptocurrency life insurance Meanwhile
Oct 10·decrypt.co

This Sam Altman-Backed Life Insurer Runs Entirely on Bitcoin, and Just Raised $37.5 Million

A life insurance company that operates entirely in Bitcoin has raised $37.5 million, bringing its total raised to over $180 million. The company, backed by Sam Altman and licensed in Bermuda since 2024, launched a single-premium whole life policy aimed at high-net-worth c…

france europe regulation MiCA taxes
Oct 10·decrypt.co

French Committee Backs Stablecoin Swap Tax and Crypto Exit Tax, Then Rejects the Budget

France's Finance Committee approved amendments to tax stablecoin swaps and extend an exit tax to crypto, alongside allowing loss carry-forwards. However, the committee then rejected the budget's revenue section, meaning these crypto amendments must be re-tabled for future…