Aha.
正在载入中英对照阅读…

arXiv:1810.04805 · 中英对照阅读

BERT:用于语言理解的深度双向 Transformer 预训练

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin、Ming-Wei Chang、Kenton Lee、Kristina Toutanova

中文速览

自然语言模型过去多只能按单向顺序读文本,难以同时利用词语左右两边的信息,尤其影响问答、句子关系判断等任务。BERT用双向Transformer编码器,并在海量无标注文本上通过“遮住词再猜”(Masked Language Model)和判断两句话是否相邻(Next Sentence Prediction)进行预训练,之后只需接上很小的任务输出层再微调。实验显示,BERT在问答、自然语言推断等11项任务上刷新最佳成绩,例如GLUE达到80.5%、SQuAD 1.1的F1达到93.2。它证明了双向预训练表示能让同一个通用模型适配多种语言任务,大幅减少为每项任务单独设计复杂模型的需要。

摘要

We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models Peters et al. (2018a); Radford et al. (2018), BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications.

术语表

BERT
BERT
Bidirectional Encoder Representations from Transformers
基于 Transformer 的双向编码器表示
language representation model
语言表示模型
language model pre-training
语言模型预训练
pre-training
预训练
fine-tuning
微调
downstream task
下游任务
bidirectional representation
双向表示
unidirectional language model
单向语言模型
masked language model
掩码语言模型
MLM
掩码语言模型
next sentence prediction
下一句预测
NSP
下一句预测
Cloze task
完形填空任务
Transformer
Transformer
self-attention
自注意力
contextual representation
上下文表示
contextual word embedding
上下文词嵌入
word embedding
词嵌入
sentence embedding
句子嵌入
feature-based approach
基于特征的方法
fine-tuning approach
基于微调的方法
ELMo
ELMo
OpenAI GPT
OpenAI GPT
GLUE
GLUE
MultiNLI
MultiNLI
SQuAD
SQuAD
natural language inference
自然语言推理
named entity recognition
命名实体识别
question answering
问答