Faculty, Staff and Student Publications
Language
English
Publication Date
1-1-2025
Journal
Proceedings of the 16th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics
DOI
10.1145/3765612.3767209
PMID
41960139
PMCID
PMC13059132
PubMedCentral® Posted Date
4-9-2026
PubMedCentral® Full Text Version
Author MSS
Abstract
We present a privacy-preserving selection layer for collaborative population stratification under 𝜖-local differential privacy (LDP). Rather than fixing a single pipeline (e.g., PCA+K-Means with preset 𝐾), our framework lets parties choose among three DP pipelines: PCA→Noise, Noise→PCA, and Noise-Only, according to their resources, and has an honest-but-curious server aggregate only DP shares to automatically select the clustering algorithm (K-Means, GMM, or Hierarchical) and 𝐾 that maximize internal metrics (Silhouette, Calinski–Harabasz, Davies–Bouldin). Because selection operates on DP data, it adds no further privacy loss. On openSNP (942 samples, 28,396 SNPs), the PCA-augmented pipelines yield higher utility and substantially lower communication and runtime than Noise-Only, and the recommended configuration consistently outperforms fixed baselines. Membership-inference attack power remains markedly lower for PCA-based pipelines across privacy budgets 𝜖. In this paper, experiments are limited to two collaborating parties; extensions to multi-site collaboration are left for future work.
Keywords
Population Stratification, Clustering, Principal Component Analysis, Privacy, Differential Privacy, Membership Inference Attack, Data Mining, Machine Learning
Published Open-Access
yes
Recommended Citation
Maryam Ghasemian, Lynette Hammond Gerido, and Erman Ayday, "Privacy-Preserving Collaborative Population Stratification with Dynamic Algorithm and Hyperparameter Selection" (2025). Faculty, Staff and Student Publications. 888.
https://digitalcommons.library.tmc.edu/uthshis_docs/888