**Background:** Genes act in context-specific networks to carry out functions, and variations in these genes can affect disease-relevant biological processes. Transcriptome-wide association studies (TWAS) have helped uncover the role of individual genes in disease mechanisms, but they do not capture gene-gene interactions, which are crucial according to the omnigenic model. The authors introduce PhenoPLIER, a computational approach that integrates gene-trait associations and pharmacological perturbation data into a common latent representation based on modules of genes with similar expression patterns across conditions.
**Methods:** PhenoPLIER uses a latent representation derived from the MultiPLIER models, which were obtained by applying the PLIER matrix factorization method to recount2, a large collection of RNA-seq samples (tens of thousands). This representation consists of 987 latent variables (LVs), each representing a gene module. Gene-trait associations from TWAS (using S-MultiXcan and S-PrediXcan) and drug-induced transcriptional responses from LINCS L1000 (1170 compounds) are projected into this latent space. The approach includes three components: (1) an LV-based regression model to compute associations between LVs and traits, (2) a consensus clustering framework to group traits with shared transcriptomic properties, and (3) an LV-based drug-repurposing approach. TWAS results from PhenomeXcan (4091 traits) were used as discovery, and eMERGE (309 phecodes) as replication. A CRISPR-Cas9 screen in HepG2 cells identified 462 genes associated with lipid regulation, from which high-confidence lipid-decreasing (8 genes) and lipid-increasing (6 genes) gene sets were selected.
**Key Results:** The LV-based regression model found 3450 significant LV-trait associations (FDR<0.05) in PhenomeXcan, with 686 LVs associated with at least one trait and 1176 traits associated with at least one LV. In eMERGE, 196 significant LV-trait associations were found, with 116 LVs associated with at least one trait and 81 traits with at least one LV. For the CRISPR screen, LV246, associated with the lipid-increasing gene set, was expressed in adipose tissue and was significantly associated with plasma lipids, high cholesterol, and Alzheimer's disease in PhenomeXcan (FDR<1e-23). In eMERGE, LV246 was significantly associated with hypercholesterolemia (FDR<4e-9), hyperlipidemia (FDR<4e-7), and disorders of lipoid metabolism (FDR<4e-7). Two high-confidence CRISPR genes, DGAT2 and ACACA, were among the highest-weighted genes in LV246 but were not associated with cardiovascular traits by TWAS alone. The LV-based drug-repurposing approach outperformed the gene-based one, with an area under the curve (AUC) of 0.632 and an average precision of 0.858, compared to the gene-based method. For niacin, the LV-based approach predicted it as a therapeutic drug for atherosclerosis with a score of 0.52, while the gene-based method assigned a negative score of -0.01. The clustering analysis revealed stable trait clusters, including a complex branch with lipids, cardiovascular, autoimmune, and neuropsychiatric disorders.
**Clinical Implications:** PhenoPLIER provides a gene module perspective that improves the detection and interpretation of genetic associations, prioritizing potential therapeutic targets even when single-gene associations are not detected. The approach can identify disease-relevant cell types, predict drug-disease relationships, and infer mechanisms of action, as demonstrated for niacin and cardiovascular traits. This method represents a conceptual shift in interpreting genetic studies, with potential to enhance understanding of complex diseases and their therapeutic modalities.