Here’s a Way to Predict When AI Chatbots Will Turn Bad

Here’s a Way to Predict When AI Chatbots Will Turn Bad

Summary

George Washington University researchers developed a formula to estimate how many safe tokens a chatbot will generate before producing harmful output, based on how conversational context shifts an attention head. In preprint tests, it correctly predicted immediate or delayed failures in 15 of 16 cases across six small open-weight models. The published study reportedly expands testing to seven models, up to 12 billion parameters. The researchers propose a lightweight monitor for offline, on-device AI that flags when a model approaches its predicted safety tipping point. They say alignment training can delay failures for particular prompts but does not eliminate the underlying mechanism. Tests remain limited, and predictions may be off by one token.