Aha.
正在载入中英对照阅读…

arXiv:2402.00888 · 中英对照阅读

Security and Privacy Challenges of Large Language Models: A Survey

Badhan Chandra Das、M. Hadi Amini、Yanzhao Wu

中文速览

大语言模型在带来强大生成、推理和交互能力的同时,也可能遭遇越狱、提示注入、后门、数据投毒、对抗攻击以及个人信息泄露等安全与隐私风险。研究通过系统梳理模型架构、攻击类型、典型案例和防御方法,进一步分析其在交通、教育、医疗等实际领域中的应用风险,并总结现有研究的不足与未来方向。结果表明,漏洞可能来自训练数据、模型开发者、部署系统和用户交互等多个环节,现有防御虽能缓解部分问题,却仍难以全面应对不断演化的攻击和隐私泄露。厘清这些风险及其防护思路,有助于研究者和企业更可靠地评估、部署大语言模型,尤其能为医疗、金融等对安全和隐私要求较高的场景提供依据。

摘要

Large language models (LLMs) have demonstrated extraordinary capabilities and contributed to multiple fields, such as generating and summarizing text, language translation, and question-answering. Nowadays, LLMs have become very popular tool in natural language processing (NLP) tasks, with the capability to analyze complicated linguistic patterns and provide relevant and appropriate responses depending on the context. While offering significant advantages, these models are also vulnerable to security and privacy attacks, such as jailbreaking attacks, data poisoning attacks, and personally identifiable information (PII) leakage attacks. This survey provides a thorough review of the security and privacy challenges of LLMs, along with the application-based risks in various domains, such as transportation, education, and healthcare. We assess the extent of LLM vulnerabilities, investigate emerging security and privacy attacks for LLMs, and review the potential defense mechanisms. Additionally, the survey outlines existing research gaps in this research area and highlights future research directions.

术语表

Large Language Model (LLM)
大语言模型
Natural Language Processing (NLP)
自然语言处理
Artificial Intelligence (AI)
人工智能
Artificial General Intelligence (AGI)
通用人工智能
Language Model (LM)
语言模型
ChatGPT
ChatGPT
GPT-4
GPT-4
prompt engineering
提示工程
in-context learning
上下文学习
pre-training
预训练
fine-tuning
微调
tokenization
标记化
deep neural networks (DNNs)
深度神经网络
attention mechanism
注意力机制
next-word prediction
下一词预测
jailbreaking attack
越狱攻击
data poisoning attack
数据投毒攻击
personally identifiable information (PII) leakage attack
个人身份信息泄露攻击
privacy attack
隐私攻击
security attack
安全攻击
vulnerability
漏洞
defense mechanism
防御机制
mitigation technique
缓解技术
privacy-preserving
隐私保护
robustness
鲁棒性
reliability
可靠性
Health Insurance Portability and Accountability Act (HIPAA)
《健康保险流通与责任法案》
General Data Protection Regulation (GDPR)
《通用数据保护条例》
California Consumer Privacy Act (CCPA)
《加利福尼亚州消费者隐私法》