Effective Bayesian Modeling of Groups of Related Count Time Series
Nicolas Chapados ApSTAT Technologies Inc., 408-4200 Boul. St-Laurent, Montréal, QC, H2W 2R2, CANADA
Abstract
Time series of counts arise in a variety of forecasting applications, for which traditional models are generally inappropriate. This paper introduces a hierarchical Bayesian formulation applicable to count time series that can easily account for explanatory variables and share statistical strength across groups of related time series. We derive an efficient approximate inference technique, and illustrate its performance on a number of datasets from supply chain planning.
中文速览
供应链、零售等行业大量存在"计数型时间序列"——比如每周某商品的销售件数——这类数据全是非负整数、零值很多、分布高度偏斜,传统的指数平滑或ARIMA模型根本不适用。这篇论文提出了一个分层贝叶斯负二项状态空间模型(H-NBSS),它以负二项分布刻画计数观测值的过度离散特性,用AR(1)潜在过程捕捉需求随时间的动态变化,同时通过分层结构让一组相关序列共享季节性、促销响应等统计信息,还能自然处理"结构性零值"(如断货导致的零需求)和协变量。由于模型的非共轭性使精确推断无解,作者推导了一种基于拉普拉斯近似(Laplace approximation)的高效近似推断算法,在保持接近最优精度的同时大幅降低了计算成本。在多个供应链真实数据集上的实验表明,该方法在预测精度上显著优于Croston方法等行业常用基准,而其计算效率又远胜于粒子滤波等随机推断方法,为大规模业务场景下的计数序列预测提供了切实可行的解决方案。
原文 arXiv:1405.3738;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1405.3738v1