Computer Science

A general approach for determining applicability domain of machine learning models

L. E. Schultz, Y. Wang, et al.

Discover a new, general method to determine where machine-learning predictions are trustworthy by measuring feature-space distance with kernel density estimation. This approach identifies chemically dissimilar groups, links high dissimilarity to large prediction errors and unreliable uncertainty estimates, and includes automated tools to set dissimilarity thresholds for in-domain versus out-of-domain decisions. Research conducted by Lane E. Schultz, Yiqi Wang, Ryan Jacobs, and Dane Morgan.

00:00

~3 min • Beginner • English

Index

Abstract

Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions. In this work, we develop a new and general approach of assessing model domain and demonstrate that our approach provides accurate and meaningful domain designation across multiple model types and material property data sets. Our approach assesses the distance between data in feature space using kernel density estimation, where this distance provides an effective tool for domain determination. We show that chemical groups considered unrelated based on chemical knowledge exhibit significant dissimilarities by our measure. We also show that high measures of dissimilarity are associated with poor model performance (i.e., high residual magnitudes) and poor estimates of model uncertainty (i.e., unreliable uncertainty estimation). Automated tools are provided to enable researchers to establish acceptable dissimilarity thresholds to identify whether new predictions of their own machine learning models are in-domain versus out-of-domain.

Publisher

npj Computational Materials

Published On

Apr 05, 2025

Authors

Lane E. Schultz, Yiqi Wang, Ryan Jacobs, Dane Morgan

DOI

https://doi.org/10.1038/s41524-025-01573-x

Related Publications

Explore these studies to deepen your understanding of the subject.

Business

Exploring the mechanism of path-creating strategy for latecomers: a combined approach of econometrics and causal machine learning

Y. Teng, Y. Li, et al.

Medicine and Health

Pre-deployment risk factors for PTSD in active-duty personnel deployed to Afghanistan: a machine-learning approach for analyzing multivariate predictors

K. Schultebraucks, M. Qian, et al.

Medicine and Health

A multimodal deep learning approach for the prediction of cognitive decline and its effectiveness in clinical trials for Alzheimer’s disease

C. Wang, H. Tachimori, et al.

Psychology

Building machine learning prediction models for well-being using predictors from the exposome and genome in a population cohort

D. H. M. Pelt, P. C. Habets, et al.

Listen, Learn & Level Up

Over 10,000 hours of research content in 25+ fields, available in 12+ languages.

No more digging through PDFs, just hit play and absorb the world's latest research in your language, on your time.

listen to research audio papers with researchbunny