Uncovering the Risks of AI: LLMs, Watermarks, and OpenAI's Cloud

Researchers from the University of North Carolina, Chapel Hill have published a preprint paper highlighting the challenges of removing sensitive data from large language models (LLMs) like OpenAI's ChatGPT and Google's Bard. While it is possible to delete information from LLMs, verifying its removal is equally difficult. LLMs are pretrained on databases and fine-tuned, making it impossible to delete specific files to prevent related outputs. The researchers found that even state-of-the-art model editing methods fail to fully delete factual information from LLMs. Defense methods are constantly playing catch-up to new attack methods, making the problem of deleting sensitive information an ongoing challenge.
- Researchers find LLMs like ChatGPT output sensitive data even after it's been 'deleted' Cointelegraph
- Researchers Tested AI Watermarks—and Broke All of Them WIRED
- Is AI lying to us? These researchers built an LLM lie detector of sorts to find out ZDNet
- Watermarking AI images to fight misinfo and deepfakes may be pretty pointless The Register
- There's big risk in not knowing what OpenAI is building in the cloud, warn Oxford scholars ZDNet
- View Full Coverage on Google News
Reading Insights
0
9
2 min
vs 3 min read
82%
577 → 102 words
Want the full story? Read the original article
Read on Cointelegraph