Dynamical Mass Measurements of Contaminated Galaxy Clusters Using Machine Learning
M. Ntampaka11affiliation: McWilliams Center for Cosmology, Department of Physics, Carnegie Mellon University, Pittsburgh, PA 15213 , H. Trac11affiliation: McWilliams Center for Cosmology, Department of Physics, Carnegie Mellon University, Pittsburgh, PA 15213 , D.J. Sutherland22affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213 , S. Fromenteau11affiliation: McWilliams Center for Cosmology, Department of Physics, Carnegie Mellon University, Pittsburgh, PA 15213 , B. Póczos22affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213 , J. Schneider22affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213
Abstract
We study dynamical mass measurements of galaxy clusters contaminated by interlopers and show that a modern machine learning (ML) algorithm can predict masses by better than a factor of two compared to a standard scaling relation approach. We create two mock catalogs from Multidark’s publicly available $N$ -body MDPL1 simulation, one with perfect galaxy cluster membership information and the other where a simple cylindrical cut around the cluster center allows interlopers to contaminate the clusters. In the standard approach, we use a power-law scaling relation to infer cluster mass from galaxy line-of-sight (LOS) velocity dispersion. Assuming perfect membership knowledge, this unrealistic case produces a wide fractional mass error distribution, with a width of $\Delta\epsilon\approx 0.87$ . Interlopers introduce additional scatter, significantly widening the error distribution further ( $\Delta\epsilon\approx 2.13$ ). We employ the support distribution machine (SDM) class of algorithms to learn from distributions of data to predict single values. Applied to distributions of galaxy observables such as LOS velocity and projected distance from the cluster center, SDM yields better tha
中文速览
准确测量星系团质量是用它们来检验宇宙学模型的关键难题,而现实观测中不可避免地混入"伪成员星系"(interlopers),会让传统方法的质量估计误差大幅增加。研究者从N体数值模拟MDPL1出发,分别构建了"纯净"和"受污染"两种模拟星系团样本,传统幂律定标关系在纯净样本下误差分布宽度已达约0.87,加入伪成员污染后更恶化至约2.13。他们引入一种名为支持分布机(Support Distribution Machine,SDM)的机器学习算法,让它直接从星系沿视线方向速度和投影距离的分布中学习并预测团质量,结果在受污染样本上将误差宽度压缩至约0.67——比传统方法在纯净样本上的表现还要好。这意味着SDM不仅对观测噪声具有更强的鲁棒性,还能更精确地重现星系团质量函数,为未来用大规模星系团巡天约束暗物质与暗能量参数提供了一条切实可行的新路径。
原文 arXiv:1509.05409;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1509.05409v2