arXiv:2309.03409 · 中英对照阅读
Large Language Models as Optimizers
中文速览
面对没有梯度、难以用传统算法更新的优化问题,OPRO把大语言模型当作“优化器”,用自然语言说明目标,并让模型根据以往方案及其得分不断生成新方案。研究者在回归和旅行商问题上验证了这种黑盒搜索能力,发现模型能逐步找到较好的参数和路线,有时还超过人工设计的启发式方法。更重要的是,OPRO能自动优化给其他模型使用的提示词,在GSM8K上比人工提示最高提升8%,在Big-Bench Hard上最高提升50%,且效果会随迭代持续改善。它的意义在于,人们不必为每种新任务专门推导优化算法,只需用文字描述目标,就能借助大语言模型探索复杂、离散且难以求导的解空间。
摘要
Optimization is ubiquitous. While derivative-based algorithms have been powerful tools for various problems, the absence of gradient imposes challenges on many real-world applications. In this work, we propose Optimization by PROmpting (OPRO), a simple and effective approach to leverage large language models (LLMs) as optimizers, where the optimization task is described in natural language. In each optimization step, the LLM generates new solutions from the prompt that contains previously generated solutions with their values, then the new solutions are evaluated and added to the prompt for the next optimization step. We first showcase OPRO on linear regression and traveling salesman problems, then move on to our main application in prompt optimization, where the goal is to find instructions that maximize the task accuracy. With a variety of LLMs, we demonstrate that the best prompts optimized by OPRO outperform human-designed prompts by up to $8\%$ on GSM8K, and by up to $50\%$ on Big-Bench Hard tasks. Code at https://github.com/google-deepmind/opro.
术语表
- Optimization by PROmpting (OPRO)
- 通过提示进行优化(OPRO)
- large language model (LLM)
- 大语言模型(LLM)
- derivative-based algorithm
- 基于导数的算法
- derivative-free optimization
- 无导数优化
- optimization problem
- 优化问题
- objective function
- 目标函数
- optimization step
- 优化步骤
- optimization trajectory
- 优化轨迹
- optimization score
- 优化得分
- solution space
- 解空间
- performance landscape
- 性能景观
- exploration-exploitation trade-off
- 探索-利用权衡
- meta-prompt
- 元提示
- prompt optimization
- 提示优化
- prompt engineering
- 提示工程
- meta-instruction
- 元指令
- training accuracy
- 训练准确率
- test accuracy
- 测试准确率
- zero-shot prompting
- 零样本提示
- few-shot chain-of-thought prompting
- 少样本思维链提示
- chain-of-thought prompting
- 思维链提示
- linear regression
- 线性回归
- traveling salesman problem
- 旅行商问题
- heuristic algorithm
- 启发式算法
- GSM8K
- GSM8K 数学推理数据集
- Big-Bench Hard
- Big-Bench Hard 评测集
- PaLM 2
- PaLM 2
- GPT-3.5-turbo
- GPT-3.5-turbo
- GPT-4
- GPT-4