Sharon Lin is an AI security researcher at Google DeepMind. Previously, she was a research fellow at the Max Planck Institute for Security and Privacy and a machine learning engineer at Abnormal Security.
In recent years, security researchers have already demonstrated significant exploits related to large language models, including data exfiltration through bypassing Bard’s content security policy, breaking isolation in ChatGPT code interpreters, buffer overflow of Mixtral’s prediction model, and command injecting Claude with invisible Unicode tags. With AI product marketplaces on the horizon and large language models already used in production, there are ample opportunities for researchers to discover new vulnerabilities in this domain.
This talk will provide an overview of LLM architectures, fine-tuning and serving infrastructures, security mitigations, and tool integrations, and provide strategies for reverse-engineering LLM-integrated systems to discover novel security exploits. We will walk through several of these attacks and discuss techniques for utilizing the emergent properties of LLMs for further vulnerability discovery.