Zhang Group
About the team
Our group members come from diverse research backgrounds in machine learning, bioinformatics, and wet-lab biology, with a shared interest to integrate multimodal data for a holistic understanding across scales from molecular interactions, cellular states, to tissue-level phenotypes. We work closely with the robotics lab and our experimental collaborators to generate data guided by model predictions, which in turn informs our model design. Every group member is encouraged to lead their own project(s) within a collaborative and supportive environment.
Research Focus and Collaborators
We develop machine learning models to both advance the understanding of biological mechanisms and facilitate the discovery of therapeutic targets, building on theoretical and empirical advances in machine learning and causal inference. We will develop models that integrate multimodal and spatiotemporal data to achieve a holistic understanding of cell states, tissue microenvironments, and perturbation effects. Our goal is to gain mechanistic insights into cellular and tissue regulation across scales – from protein localization and interaction in single cells to cell fate decisions in organoids and tissues. By modeling cellular dynamics and interactions in tissue context, we aim to enable virtual profiling of genetic and chemical perturbations to identify potential therapeutic targets for disease-associated changes in protein localization, cell states, and tissue architecture. We collaborate locally and globally with experts in organoid models, spatial multiomics, live-cell imaging, perturbation screens, cardiovascular diseases, immunology, and cancer.
Candidate’s Profile and Skills
We are looking for 1-2 PhD students with a background in computer science, machine learning, statistics/biostatistics, computational biology, data science, or a related field. We also welcome candidates who would like to explore both computational and experimental projects. We are open to co-supervision arrangements with experimental or computational collaborators.
Project 1: Predicting the impact of gene variants on protein subcellular localization
The subcellular localization of a protein is important for its function, and its mislocalization is linked to numerous diseases. While atlas-scale efforts have profiled thousands of proteins across diverse cell lines, the combinatorial space of proteins and cellular contexts vastly exceeds what has been measured to date. The incorporation of genetic variants further expands the possible combinations, and current experimental studies have been limited to a single cell line in each study and to functional subsets of proteins. We will develop machine learning frameworks that generalize across proteins, variants, and cell states to infer how genetic changes alter protein localization at single-cell resolution. By explicitly modeling both protein-intrinsic features and image-based cell states, we aim to enable “virtual profiling” of variant effects at scales impractical for wet-lab screening and to reveal disease-relevant mechanisms.
We will further develop models to disentangle the contributions of sequence motifs, structure, post-translational modifications (PTMs), interaction partners, and cell state to protein mislocalization. This could reveal actionable mechanisms of mislocalization and enable prediction of interventions that restore correct localization. We will perform imaging-based single-cell perturbation assays and orthogonal validations to test the predicted mechanisms and model attributions.
Project 2: Prediction of perturbation effects on protein localization
We will develop a multi-modal machine-learning framework that predicts how small molecules alter protein subcellular localization at single-cell resolution to enable screening and discovery of compounds that correct mislocalization in silico. While existing single-cell perturbation models focus on genetic perturbations and gene expression readouts, recent work shows that generalization to unseen gene targets is achievable by learning meaningful gene embeddings. Building on this principle, we will develop a model for cellular-context-aware prediction of drug-protein effects, quantification of uncertainty to prioritize experiments, and nomination of candidates predicted to restore correct localization. As multiplexed protein imaging in tissues becomes increasingly available, we will transfer the learned representations to more physiologically relevant settings to aid in therapeutic hypothesis generation.
Project 3: Spatiotemporal models for learning cell fate decisions in cancer and cardiac organoids
How cells decide their fates is shaped not only by their internal programs but also by when and where they sit in tissue. In embryonic organoids, early morphological and signaling cues foreshadow lineage commitment and symmetry breaking. In cancer, such as in pancreatic ductal adenocarcinoma (PDAC), the transition of epithelial progenitors to classical and mesenchymal states along with microenvironmental differences appear to influence prognosis and therapy response. While ML models have been applied to study single-cell trajectories, methods developed in non-biological settings often struggle to yield novel biological insights. This project builds novel machine learning tools based on flow matching, graphical models, and optimal transport to model heterogeneous, asynchronous, and spatially dependent dynamics and to answer challenging biological questions: Which morphological, transcriptional, or microenvironmental features forecast a cell’s lineage commitment and plasticity? How can we distinguish cell-intrinsic and extrinsic programs? What combinations of genetic and chemical perturbations could drive the commitment to a particular lineage? While we will initially develop our models on a cardiac organoid dataset with spatial multiomics and live-cell imaging, our methods will be designed to be broadly applicable to diverse biological questions.
