Node Feature Extraction by Self-Supervised Multi-scale Neighborhood Prediction
Eli Chien Thanks: This work was done during Eli Chien’s internship at Amazon, USA. Affiliation: University of Illinois Urbana-Champaign, USA Email: Wei-Cheng Chang Affiliation: Amazon, USA Email: Cho-Jui Hsieh Affiliation: University of California, Los Angeles, USA Email: Hsiang-Fu Yu Jiong Zhang Affiliation: Amazon, USA Email: Olgica Milenkovic Affiliation: University of Illinois Urbana-Champaign, USA Email: Inderjit S. Dhillon Affiliation: Amazon, USA Email:
Abstract
Learning on graphs has attracted significant attention in the learning community due to numerous real-world applications. In particular, graph neural networks (GNNs), which take numerical node features and graph structure as inputs, have been shown to achieve state-of-the-art performance on various graph-related learning tasks. Recent works exploring the correlation between numerical node features and graph structure via self-supervised learning have paved the way for further performance improvements of GNNs. However, methods used for extracting numerical node features from raw data are still graph-agnostic within standard GNN pipelines. This practice is sub-optimal as it prevents one from fully utilizing potential correlations between graph topology and node attributes. To mitigate this issue, we propose a new self-supervised learning framework, Graph Information Aided Node feature exTraction (GIANT). GIANT makes use of the eXtreme Multi-label Classification (XMC) formalism, which is crucial for fine-tuning the language model based on graph information, and scales to large datasets. We also provide a theoretical analysis that justifies the use of XMC over link prediction and motiv
原文 arXiv:2111.00064;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2111.00064v3