Teaching models to express their uncertainty in words
Stephanie Lin University of Oxford Jacob Hilton OpenAI Owain Evans University of Oxford
Abstract
We show that a GPT-3 model can learn to express uncertainty about its own answers in natural language – without use of model logits. When given a question, the model generates both an answer and a level of confidence (e.g. “90% confidence” or “high confidence”). These levels map to probabilities that are well calibrated. The model also remains moderately calibrated under distribution shift, and is sensitive to uncertainty in its own answers, rather than imitating human examples. To our knowledge, this is the first time a model has been shown to express calibrated uncertainty about its own answers in natural language.
原文 arXiv:2205.14334;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.14334v2