Distilling Reasoning Capabilities into Smaller Language Models
Kumar Shridhar Alessandro Stolfo Mrinmaya SachanDepartment of Computer Science, ETH Zürich{shkumar, Thanks: Equal contribution;
Abstract
Step-by-step reasoning approaches like chain of thought (CoT) have proved to be very effective in inducing reasoning capabilities in large language models. However, the success of the CoT approach is fundamentally tied to the model size, and billion parameter-scale models are often needed to get CoT to work. In this paper, we propose a knowledge distillation approach that leverages the step-by-step CoT reasoning capabilities of larger models and distills these abilities into smaller models.
原文 arXiv:2212.00193;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.00193v2