Aha.
正在载入中英对照阅读…

arXiv:2309.03409 · 中英对照阅读

Large Language Models as Optimizers

Chengrun Yang*、Xuezhi Wang、Yifeng Lu、Hanxiao Liu、Quoc V. Le Denny Zhou Xinyun Chen*、* Equal contribution

中文速览

面对没有梯度、难以用传统算法更新的优化问题,OPRO把大语言模型当作“优化器”,用自然语言说明目标,并让模型根据以往方案及其得分不断生成新方案。研究者在回归和旅行商问题上验证了这种黑盒搜索能力,发现模型能逐步找到较好的参数和路线,有时还超过人工设计的启发式方法。更重要的是,OPRO能自动优化给其他模型使用的提示词,在GSM8K上比人工提示最高提升8%,在Big-Bench Hard上最高提升50%,且效果会随迭代持续改善。它的意义在于,人们不必为每种新任务专门推导优化算法,只需用文字描述目标,就能借助大语言模型探索复杂、离散且难以求导的解空间。

摘要

Optimization is ubiquitous. While derivative-based algorithms have been powerful tools for various problems, the absence of gradient imposes challenges on many real-world applications. In this work, we propose Optimization by PROmpting (OPRO), a simple and effective approach to leverage large language models (LLMs) as optimizers, where the optimization task is described in natural language. In each optimization step, the LLM generates new solutions from the prompt that contains previously generated solutions with their values, then the new solutions are evaluated and added to the prompt for the next optimization step. We first showcase OPRO on linear regression and traveling salesman problems, then move on to our main application in prompt optimization, where the goal is to find instructions that maximize the task accuracy. With a variety of LLMs, we demonstrate that the best prompts optimized by OPRO outperform human-designed prompts by up to $8\%$ on GSM8K, and by up to $50\%$ on Big-Bench Hard tasks. Code at https://github.com/google-deepmind/opro.

术语表

Optimization by PROmpting (OPRO)
通过提示进行优化(OPRO)
large language model (LLM)
大语言模型(LLM)
derivative-based algorithm
基于导数的算法
derivative-free optimization
无导数优化
optimization problem
优化问题
objective function
目标函数
optimization step
优化步骤
optimization trajectory
优化轨迹
optimization score
优化得分
solution space
解空间
performance landscape
性能景观
exploration-exploitation trade-off
探索-利用权衡
meta-prompt
元提示
prompt optimization
提示优化
prompt engineering
提示工程
meta-instruction
元指令
training accuracy
训练准确率
test accuracy
测试准确率
zero-shot prompting
零样本提示
few-shot chain-of-thought prompting
少样本思维链提示
chain-of-thought prompting
思维链提示
linear regression
线性回归
traveling salesman problem
旅行商问题
heuristic algorithm
启发式算法
GSM8K
GSM8K 数学推理数据集
Big-Bench Hard
Big-Bench Hard 评测集
PaLM 2
PaLM 2
GPT-3.5-turbo
GPT-3.5-turbo
GPT-4
GPT-4