Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Soorya Ram Shimgekar Affiliation: Nimblemind Nimblemind Nimblemind Nimblemind Singapore University of Technology and Design University of Illinois Urbana-Champaign Florida International University University of California Los Angeles Department of Health Informatics, Rutgers University Newark Nimblemind Nimblemind Rutgers Institute for Health, Health Care Policy and Aging Research Department of Medicine, Robert Wood Johnson Medical School, Rutgers Health Department of Health Informatics, Rutgers University Newark Michelle Hu Affiliation: Dorisa Shehi Affiliation: Daniel Kang Affiliation: Roy Ka-Wei Lee Affiliation: Koustuv Saha Affiliation: Christian Poellabauer Affiliation: Christopher Lee Affiliation: Sajeev Singh Affiliation: Piyum Zonooz Affiliation: Navin Kumar Affiliation: Zeeshan Ahmed Affiliation: Affiliation: Priyadarshini Kachroo Affiliation:
Abstract
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39–45% of data scientists’ workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demon
原文 arXiv:2608.06366;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2608.06366v1