Is ChatGPT a Good Recommender? A Preliminary Study
Junling Liu Alibaba GroupChina , Chao Liu Ant GroupChina , Peilin Zhou Hong Kong University of Science and Technology(Guangzhou)China , Renjie Lv Ant GroupChina , Kang Zhou Alibaba GroupChina and Yan Zhang Alibaba GroupChina
Abstract
Recommendation systems have witnessed significant advancements and have been widely used over the past decades. However, most traditional recommendation methods are task-specific and therefore lack efficient generalization ability. Recently, the emergence of ChatGPT has significantly advanced NLP tasks by enhancing the capabilities of conversational models. Nonetheless, the application of ChatGPT in the recommendation domain has not been thoroughly investigated. In this paper, we employ ChatGPT as a general-purpose recommendation model to explore its potential for transferring extensive linguistic and world knowledge acquired from large-scale corpora to recommendation scenarios. Specifically, we design a set of prompts and evaluate ChatGPT’s performance on five recommendation scenarios, including rating prediction, sequential recommendation, direct recommendation, explanation generation, and review summarization. Unlike traditional recommendation methods, we do not fine-tune ChatGPT during the entire evaluation process, relying only on the prompts themselves to convert recommendation tasks into natural language tasks. Further, we explore the use of few-shot prompting to inject inte
中文速览
大量传统推荐系统只能针对特定任务训练专用模型,泛化能力有限,而ChatGPT凭借在海量语料上积累的语言与世界知识,理论上有望跨任务迁移到推荐场景。研究者为此设计了一套提示词(prompt),在完全不微调ChatGPT的前提下,将评分预测、序列推荐、直接推荐、解释生成和评论摘要五类任务转化为自然语言任务,并在Amazon Beauty数据集上与传统基线方法进行系统对比;同时引入少样本提示(few-shot prompting)将用户历史交互注入上下文,帮助模型更好地捕捉用户偏好。结果显示,ChatGPT在评分预测上表现出色,在序列推荐和直接推荐的准确性指标上仅与早期基线持平甚至更差,但在解释生成和评论摘要两项任务上,客观指标较低而人工评估却反超当前最优方法,说明标准自动评测指标难以全面衡量ChatGPT的真实能力。这项研究为大语言模型在推荐系统中的应用提供了首个系统性基准,揭示了其潜力与局限,对推动后续相关研究具有重要参考价值。
原文 arXiv:2304.10149;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2304.10149v3