CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?
Tianwen Wei Jian Luan Wei Liu Shuang Dong Bin Wang Xiaomi AI Lab
Abstract
We present the Chinese Elementary School Math Word Problems (CMATH) dataset, comprising 1.7k elementary school-level math word problems with detailed annotations, source from actual Chinese workbooks and exams. This dataset aims to provide a benchmark tool for assessing the following question: to what grade level of elementary school math do the abilities of popular large language models (LLMs) correspond? We evaluate a variety of popular LLMs, including both commercial and open-source options, and discover that only GPT-4 achieves success (accuracy $\geq$ 60%) across all six elementary school grades, while other models falter at different grade levels. Furthermore, we assess the robustness of several top-performing LLMs by augmenting the original problems in the CMATH dataset with distracting information. Our findings reveal that GPT-4 is able to maintains robustness, while other model fail. We anticipate that our study will expose limitations in LLMs’ arithmetic and reasoning capabilities, and promote their ongoing development and advancement.
中文速览
大语言模型(LLM)到底有没有小学数学水平?为了给这个问题一个直观的答案,研究者从真实的中国小学练习册和考试卷中收集了1700道应用题,按年级标注,构建了 CMATH 数据集,使得评测结果可以像"某模型相当于四年级数学水平"这样被普通人理解。测试结果显示,在所有参评的商业和开源模型中,只有 GPT-4 能在一到六年级全部及格(准确率≥60%),ChatGPT 止步于五年级,其余开源模型甚至在一年级就已失手;进一步在题目中加入无关干扰信息后,GPT-4 依然保持稳健,而其他模型则普遍被误导。这项工作揭示了当前主流大模型在中文算术和逻辑推理上的明显短板,为推动模型能力的持续改进提供了一把接地气的量尺。
原文 arXiv:2306.16636;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.16636v1