AITHYRA, the Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences, opens a new call for PhD students from September 1- November 1.

We are an interdisciplinary lab that does both computational and experimental research in the field of enzyme discovery and design. Will is working on flow matching for enzyme mechanisms, Gabriela is generating new enzymes for antibiotics with DNA language models, Luca is designing fitness predictors for proteins using contrastive learning, Sebastian is using DPO to identify better catalysts, Ikumi is designing biosensors, and Manuel is building our automated pipeline to evaluate enzymes. Our flavor of ML methods is to develop models that can translate directly from sequence to function. One of the benefits of this approach is that it scales, so our methods can be applied to the billions of enzyme sequences that have been sampled from nature – enabling new discoveries of distant function across evolutionary scales. By also having an experimental lab, we can generate new data for the models to learn, have lab in the loop for active learning, and ask questions that are of interest and impactful for us, such as, how could we design a better antibiotic that is both potent and less susceptible to resistance. We are 2 postdocs (Gabriela – evolutionary biology, Béla – synthetic biology/ML), 1 technician (Manuel – molecular biology), 2 local PhD students (Luca – biochemistry, Sebastian – chemistry), 2 remote PhD students (Will – bioinformatics, Brisbane, Australia, Ikumi – biosensors – Tokyo, Japan). Lab culture is very important to us, so we strive to be inclusive, friendly, share culture and science.

Enzymes frequently exhibit promiscuous activity beyond their native roles, providing starting points for new functions. Finding these promiscuous enzymes, especially for non-natural chemical transformations, is challenging but highly valuable, as they promise novel, sustainable solutions for chemistry and biotechnology. We take a mechanistic perspective to try to disentangle the complex multidimensional enzyme sequence-function-environment relationship. Our goal is to learn the fundamental processes of evolution, so that rather than doing methods like directed evolution, we can do informed design, where we can a-priori know the impact of mutations on function. Here we do this by developing ML models (generative and predictive) to learn across the sequence and activity space. We collaborate both locally and globally, on new to nature chemistry with the Reisenbauer lab (ISTA, Austria), Peptide based catalysis with the Nguyen lab (EPFL, Switzerland), Causal inference for enzyme mechanisms (Dr. Loomba, Imperial College London, UK), bioinformatics and evolutionary methods with Boden lab (University of Queensland, Australia), and we’re always open to more collaborations as it fits your projects 🙂

In enzymes, evolutionary potential often manifests through catalytic promiscuity, the intrinsic ability to process non-native substrates or catalyze side reactions. Directed evolution (DE) exploits this latent potential by repurposing low-level promiscuous starting points toward target biocatalytic tasks, remaining the primary paradigm to expand biocatalytic scope. However, discovering initial promiscuous starting variants still relies heavily on Edisonian trial-and-error and chemical intuition, as no in silico screening tools currently exist to predict new product formation and novel reactivities.

In your project we will investigate whether generative modeling can be used to screen enzymes for novel catalytic functions. While our baseline model that used conditional flow matching (unpublished) could rank candidates within a substrate class for native functions this has yet to be achieved for “new-to-nature” activity. Hence in your project we’ll ask: (1) Can generative models be used to identify new-to-nature reaction products? (2) Can the architecture be adapted to prioritize lead variants for directed evolution campaigns?

Under this aim, we will determine the optimal representation for an enzyme’s “spheres of influence” on reactivity, balancing the insufficiency of active-site residues against the noise of whole-protein representations. Ultimately, this aim will deliver a generative screening platform capable of learning the latent catalytic potential of enzymes for new biocatalytic transformations.

This is a primarily computational role however you can experience the wet-lab if interested! Looking for ML/computational folks who are excited to develop new models and architectures for enzyme discovery. More important than your skills and background is your curiosity, ability to creatively solve new problems, and ability to work in an interdisciplinary team.