Long-Term Forecasting using Higher-Order Tensor RNNs
\nameRose Yu \AND\nameStephan Zheng \AND\nameAnima Anandkumar \AND\nameYisong Yue \addrDepartment of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125, USA
Abstract
We present Higher-Order Tensor RNN (HOT-RNN), a novel family of neural sequence architectures for multivariate forecasting in environments with nonlinear dynamics. Long-term forecasting in such systems is highly challenging, since there exist long-term temporal dependencies, higher-order correlations and sensitivity to error propagation. Our proposed recurrent architecture addresses these issues by learning the nonlinear dynamics directly using higher-order moments and higher-order state transition functions. Furthermore, we decompose the higher-order structure using the tensor-train decomposition to reduce the number of parameters while preserving the model performance. We theoretically establish the approximation guarantees and the variance bound for HOT-RNN for general sequence inputs. We also demonstrate $5\sim 12\%$ improvements for long-term prediction over general RNN and LSTM architectures on a range of simulated environments with nonlinear dynamics, as well on real-world time series data.
中文速览
非线性动力学系统(如洛伦兹吸引子、湍流、气候)的长期预测极其困难,根源在于系统同时存在高阶相关性、长程时间依赖和误差累积敏感性,而标准RNN和LSTM只用一步前向状态转移,无法显式捕捉这些特性。为此,作者提出HOT-RNN(高阶张量循环神经网络),通过保留多步历史隐状态并用高阶多项式张量对状态转移建模来直接拟合非线性动力学,同时借助张量列(tensor-train)分解把参数量从指数级压缩到线性级,使模型在不爆炸参数规模的前提下保持高表达能力。理论上,作者严格证明了HOT-RNN对满足一定光滑性条件的动力学函数具有指数级于标准RNN的逼近优势,并给出了方差上界。实验结果表明,在多个模拟非线性动力学场景和真实时间序列数据上,HOT-RNN相比标准RNN和LSTM的长期预测误差降低了5%至12%,为高维多变量时间序列的长程预测提供了一个兼具理论保障和实用效率的新框架。
原文 arXiv:1711.00073;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1711.00073v3