Faculty, Staff and Student Publications
Language
English
Publication Date
1-1-2025
Journal
IEEE Transactions on Privacy
DOI
10.1109/tp.2025.3628998
PMID
41551373
PMCID
PMC12807549
PubMedCentral® Posted Date
1-16-2026
PubMedCentral® Full Text Version
Author MSS
Abstract
We present a privacy-preserving framework to verify whether a declared data preprocessing pipeline was correctly applied before training a machine learning model on sensitive data. The verifier has only black-box query access to the model and combines three behavior indicators: shift in prediction accuracy, Kullback-Leibler (KL) divergence between output distributions, and explanation vectors from Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). The method requires neither the original training records nor ground-truth labels. It supports two tasks: (i) a binary decision on correctness and (ii) a multi-class diagnosis identifying which step is missing. Experiments on three tabular datasets (Diabetes, Adult-Income, Student-Record) show that the binary detector maintains over 75% F1 even under strong local differential privacy (𝜀=0.1). Machine-learning classifiers consistently outperform simple threshold rules in the binary setting, while the two approaches perform comparably for multi-class diagnosis. A label-free variant that clusters explanation vectors achieves competitive accuracy, enabling verification when no labeled pipelines are available. These results demonstrate a practical and scalable approach for safeguarding preprocessing integrity in privacy-sensitive machine learning workflows.
Keywords
Data preprocessing, differential privacy, explainable AI, local differential privacy, model auditing, tabular data
Published Open-Access
yes
Recommended Citation
Li, Wenbiao; Halimi, Anisa; Vaidya, Jaideep; et al., "Privacy-Preserving Verification of ML Preprocessing via Model Behavior Indicators" (2025). Faculty, Staff and Student Publications. 892.
https://digitalcommons.library.tmc.edu/uthshis_docs/892