MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Longhui Yu1,⋆⋆\star⋆ Weisen Jiang2,3,⋆⋆\star⋆ Han Shi4,† Jincheng Yu3,4 Zhengying Liu4 Yu Zhang2 James T. Kwok3 Zhenguo Li4 Adrian Weller1,5 Weiyang Liu1,6,† 1University of Cambridge 2Southern University of Science and Technology 3Hong Kong University of Science and Technology 4Huawei Noah’s Ark Lab 5The Alan Turing Institute 6Max Planck Institute for Intelligent Systems - Tübingen
Abstract
Large language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (e.g., LLaMA-2) are still far away from satisfactory for solving mathematical problems due to the complex reasoning procedures. To bridge this gap, we propose MetaMath, a finetuned language model that specializes in mathematical reasoning. Specifically, we start by bootstrapping mathematical questions by rewriting the question from multiple perspectives, which results in a new dataset called MetaMathQA. Then we finetune the LLaMA-2 models on MetaMathQA. Experimental results on two popular benchmarks (i.e., GSM8K and MATH) for mathematical reasoning demonstrate that MetaMath outperforms a suite of open-source LLMs by a significant margin. Our MetaMath-7B model achieves $66.5\%$ on GSM8K and $19.8\%$ on MATH, exceeding the state-of-the-art models of the same size by $11.5\%$ and $8.7\%$ . Particularly, MetaMath-70B achieves an accuracy of $82.3\%$ on GSM8K, slightly better than GPT-3.5-Turbo. We release the MetaMathQA dataset, the MetaMath models with different model sizes and the training code
原文 arXiv:2309.12284;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2309.12284v4