UNSW: ‘Drunk’ AI Models Leak Secrets

Researchers at the University of New South Wales have found that artificial intelligence models act like loose-lipped drunks when prompted or trained to mimic drunken speech. The study says the models become far more likely to leak confidential information and answer harmful questions.
The team, led by Dr Aditya Joshi with Anudeex Shetty and Professor Salil Kanhere, tested three ways of making models "drunk": a role-play prompt, fine-tuning on drunk texts, and reinforcement learning that rewards the style. They ran the tests on OpenAI's GPT-4 and GPT-3.5 as well as open-weight models.
Across all three methods, the "drunk" models answered prompts they should have refused and gave up secrets they were told to keep, the researchers said. "Across the board for all three methods, it is vulnerable," Dr Joshi said.
The researchers say organisations should not treat changes to a model's persona or training as purely cosmetic. A fine-tuned model connected to internal documents or customer data could expose information if its safety testing is no longer valid.
The work was reported by the Australian Cyber Security Magazine and Cyber Daily. The finding adds persona-based prompting to the list of attack surfaces that companies deploying chatbots need to test for.
Leave a Reply