OAG-BERT: Towards A Unified Backbone Language Model For Academic Knowledge Services
Xiao Liu Tsinghua University , Da Yin Tsinghua University , Jingnan Zheng National University of Singapore , Xingjian Zhang Tsinghua University , Peng Zhang Zhipu AI , Hongxia Yang DAMO Academy, Alibaba Group , Yuxiao Dong Tsinghua University and Jie Tang Tsinghua University
Abstract
Academic knowledge services have substantially facilitated the development of the science enterprise by providing a plenitude of efficient research tools. However, many applications highly depend on ad-hoc models and expensive human labeling to understand scientific contents, hindering deployments into real products. To build a unified backbone language model for different knowledge-intensive academic applications, we pre-train an academic language model OAG-BERT that integrates both the heterogeneous entity knowledge and scientific corpora in the Open Academic Graph (OAG)—the largest public academic graph to date. In OAG-BERT, we develop strategies for pre-training text and entity data along with zero-shot inference techniques. OAG-BERT achieves outperformance over baselines on nine academic tasks including two demo applications, demonstrating its potential to serve as one foundation model for academic knowledge services. Its zero-shot capability furthers the path to mitigate the need of expensive annotations. OAG-BERT has been deployed for real-world applications, such as the reviewer recommendation function for National Nature Science Foundation of China (NSFC)—one of the larges
中文速览
学术知识服务平台(如AMiner、Google Scholar)越来越需要能理解科学内容的AI模型,但现有方案往往各立门户、依赖昂贵的人工标注,难以规模化落地。为此,研究团队基于迄今最大的公开学术图谱OAG(包含逾7亿实体、110亿关系及海量论文全文),预训练了一个统一的学术语言模型OAG-BERT,通过异构实体类型嵌入、实体感知二维位置编码和跨度感知实体掩码等技术,将论文文本与作者、机构、期刊、研究领域等结构化实体知识融合进同一模型。在作者名消歧、文献检索、论文推荐、领域标注等九项学术任务上,OAG-BERT全面超越基线模型,并支持无需任何标注数据的零样本(zero-shot)推断。该模型已在AMiner平台和中国国家自然科学基金(NSFC)评审专家推荐系统中实际部署,为构建统一的学术知识服务基础模型提供了切实可行的路径。
原文 arXiv:2103.02410;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.02410v3