arXiv:2402.00888 · 中英对照阅读
Security and Privacy Challenges of Large Language Models: A Survey
中文速览
大语言模型在带来强大生成、推理和交互能力的同时,也可能遭遇越狱、提示注入、后门、数据投毒、对抗攻击以及个人信息泄露等安全与隐私风险。研究通过系统梳理模型架构、攻击类型、典型案例和防御方法,进一步分析其在交通、教育、医疗等实际领域中的应用风险,并总结现有研究的不足与未来方向。结果表明,漏洞可能来自训练数据、模型开发者、部署系统和用户交互等多个环节,现有防御虽能缓解部分问题,却仍难以全面应对不断演化的攻击和隐私泄露。厘清这些风险及其防护思路,有助于研究者和企业更可靠地评估、部署大语言模型,尤其能为医疗、金融等对安全和隐私要求较高的场景提供依据。
摘要
Large language models (LLMs) have demonstrated extraordinary capabilities and contributed to multiple fields, such as generating and summarizing text, language translation, and question-answering. Nowadays, LLMs have become very popular tool in natural language processing (NLP) tasks, with the capability to analyze complicated linguistic patterns and provide relevant and appropriate responses depending on the context. While offering significant advantages, these models are also vulnerable to security and privacy attacks, such as jailbreaking attacks, data poisoning attacks, and personally identifiable information (PII) leakage attacks. This survey provides a thorough review of the security and privacy challenges of LLMs, along with the application-based risks in various domains, such as transportation, education, and healthcare. We assess the extent of LLM vulnerabilities, investigate emerging security and privacy attacks for LLMs, and review the potential defense mechanisms. Additionally, the survey outlines existing research gaps in this research area and highlights future research directions.
术语表
- Large Language Model (LLM)
- 大语言模型
- Natural Language Processing (NLP)
- 自然语言处理
- Artificial Intelligence (AI)
- 人工智能
- Artificial General Intelligence (AGI)
- 通用人工智能
- Language Model (LM)
- 语言模型
- ChatGPT
- ChatGPT
- GPT-4
- GPT-4
- prompt engineering
- 提示工程
- in-context learning
- 上下文学习
- pre-training
- 预训练
- fine-tuning
- 微调
- tokenization
- 标记化
- deep neural networks (DNNs)
- 深度神经网络
- attention mechanism
- 注意力机制
- next-word prediction
- 下一词预测
- jailbreaking attack
- 越狱攻击
- data poisoning attack
- 数据投毒攻击
- personally identifiable information (PII) leakage attack
- 个人身份信息泄露攻击
- privacy attack
- 隐私攻击
- security attack
- 安全攻击
- vulnerability
- 漏洞
- defense mechanism
- 防御机制
- mitigation technique
- 缓解技术
- privacy-preserving
- 隐私保护
- robustness
- 鲁棒性
- reliability
- 可靠性
- Health Insurance Portability and Accountability Act (HIPAA)
- 《健康保险流通与责任法案》
- General Data Protection Regulation (GDPR)
- 《通用数据保护条例》
- California Consumer Privacy Act (CCPA)
- 《加利福尼亚州消费者隐私法》