Inferring cohort differences in microbiota composition using Latent Dirichlet allocation model

(2026)

Files

Cornez_08131900_2026.pdf
  • Open access
  • Adobe PDF
  • 2.88 MB

Details

Supervisors
Faculty
Degree label
Abstract
Microbiota composition data are inherently characterized by high sparsity and stochastic overdispersion, complicating the reliable detection of cohort-level differences. To address this challenge, this thesis evaluates the capacity of a joint analytical framework, combining Latent Dirichlet Allocation (LDA) and PERMANOVA, to infer and test global differences between microbiota cohorts. An in silico generative process coupled with a full factorial design is implemented to evaluate the pipeline against controlled variations in biological signal amplitude (A), cohort size (Ng), and the Negative Binomial distribution size parameter (ϕ). By exclusively simulating the cohort effect on the latent topic prevalence matrix (Γ), this framework isolates the statistical behavior of the pipeline from the confounding factors of empirical datasets. The systematic evaluation reveals a crucial functional decoupling between latent space reconstruction and statistical testing. First, the structural fidelity of the LDA inference is primarily dictated by overdispersion; extreme variance (ϕ ≤ 1) degrades the accuracy of the inferred matrices, a degradation that increased cohort recruitment alone cannot fully compensate for. Despite this, the downstream statistical testing phase demonstrates consistent reliability. The framework maintains a controlled Type I error rate (oscillating around 5%), ensuring that extreme overdispersion does not translate into spurious cohort separations. Furthermore, the statistical power is primarily governed by the biological effect amplitude and cohort size, exhibiting compensatory dynamics that facilitate accurate detection provided the biological contrast is pronounced or sampling is sufficient. Consequently, this thesis quantitatively maps the operational limits of the LDA-PERMANOVA pipeline, providing data-driven guidelines for the experimental design of future microbiota studies.