Skip to content

Commit e65d449

Browse files
committed
Expand enzyme tutorial glossary add advanced constraints section
Enhances the advanced enzyme design tutorial with clearer internal cross-links to key terms, adds a new RASA conditioning section with parameter guidance and examples, and fills in/expands glossary definitions (including theozyme, ORI token, catalytic residue, and related terms). Also tightens a few wording and references updates for readability.
1 parent fa3b767 commit e65d449

1 file changed

Lines changed: 64 additions & 14 deletions

File tree

models/rfd3/docs/tutorials/advanced_enzyme_design_tutorial.md

Lines changed: 64 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -5,7 +5,7 @@
55

66
(adv_enzyme_tutorial_intro)=
77
## Introduction
8-
In this tutorial, you will learn how to design a *de novo* enzyme by generating novel protein backbones that scaffold a pre-defined active site using [RFdiffusion3 (RFD3)](https://www.biorxiv.org/content/10.1101/2025.09.18.676967v2). More specifically, you will design a *de novo* metalloprotease for a system comprised of a phosphonamidate transition-state analog, zinc cofactor, and six catalytic residues shown below.
8+
In this tutorial, you will learn how to design a *de novo* enzyme by generating novel protein backbones that scaffold a pre-defined active site using [RFdiffusion3 (RFD3)](https://www.biorxiv.org/content/10.1101/2025.09.18.676967v2). More specifically, you will design a *de novo* metalloprotease for a system comprised of a phosphonamidate [transition-state analog](#adv_enzyme_tutorial_transition_state_analog_def), zinc [cofactor](#adv_enzyme_tutorial_cofactor_def), and six [catalytic residues](#adv_enzyme_tutorial_catalytic_residue_def) shown below.
99

1010
```{figure}
1111
<!-- TODO insert image of initial/final (?) system here -->
@@ -47,11 +47,11 @@ PyMOL is not necessary to complete this tutorial, the steps shown here can be re
4747
### Creating a Theozyme
4848
Structural inputs to RFdiffusion3 typically either come from a structure-prediction model (e.g. [AlfaFold3](https://www.nature.com/articles/s41586-024-07487-w)) or from an experimental structure reported in resources like the [RCSB Protein Data Bank (PDB)](https://www.rcsb.org/)
4949

50-
However, for enzyme design our goal is to stabilize the **transition state** of the reaction involving our ligand and the key catalytic residues it interacts with. If the PDB contains a structure with a bound transition-state analog (a mimic of the transition state), this structure will be most useful for your enzyme design projects.
50+
However, for enzyme design our goal is to stabilize the **transition state** of the reaction involving our ligand and the key [catalytic residues](#adv_enzyme_tutorial_catalytic_residue_def) it interacts with. If the PDB contains a structure with a bound transition-state analog (a mimic of the transition state), this structure will be most useful for your enzyme design projects.
5151

5252
For the metallohydrolase we are designing in this tutorial, we will start with a phosphoester. These are known for being a transition state analog for ester- and amide-cleaving metallohydrolases due to their tetrahedral geometry and localized negative charge. A careful search of the [RCSB Protein Data Bank (RSCB PDB)](https://www.rcsb.org/) leads us to [astacin (1QJI)](https://www.rcsb.org/structure/1QJI), a phosphonamidate transition-state analog. <!-- Not sure if this line is actually relevant, it is the only time zinc is mentioned in the introduction: In the specific case of a zinc protease, peptide substrates bearing a **phosphonamidate** at the cleavage site are especially effective transition-state analogs.--> <!-- TODO: figure out if it is appropriate to include figure1 here, it's for a zinc reaction mechanism, and I'm not sure if it's actually necessary/relevant for what we are trying to accomplish here. -->
5353

54-
Now we need to determine which residues in this protein are important for stabilizing the transition state. For this type of catalytic reaction, it is known that the three histidine residues (H92, H96, and H102 in 1QJI) that surround the zinc ion are crucial for chelation. The glutamic acid residue whose side chain interacts with the zinc ion (E93) is also known to serve as the general base for this hydrolysis reaction. The [article](https://www.nature.com/articles/nsb0896-671) that published the 1QJI structure also reveals that Y149 and M147 may ne necessary for this reaction: Y149 stabilizes the oxyanion that is formed during the reaction and M147 maybe important for conserving the motif that sits below the active site.
54+
Now we need to determine which residues in this protein are important for stabilizing the transition state. For this type of catalytic reaction, it is known that the three histidine residues (H92, H96, and H102 in 1QJI) that surround the zinc ion are crucial for [chelation](#adv_enzyme_tutorial_chelation_def). The glutamic acid residue whose side chain interacts with the zinc ion (E93) is also known to serve as the general base for this hydrolysis reaction. The [article](https://www.nature.com/articles/nsb0896-671) that published the 1QJI structure also reveals that Y149 and M147 may ne necessary for this reaction: Y149 stabilizes the oxyanion that is formed during the reaction and M147 maybe important for conserving the motif that sits below the active site.
5555

5656
So for this example, we will be using 6 catalytic residues to create our theozyme along with the ligand and the zinc ion: H92, E93, H96, H102, M147, and Y149. You can see these residues highlighted in the structure below:
5757

@@ -74,7 +74,7 @@ For how to crop and save structures in PyMOL, see the 'Motif Preparation' sectio
7474
Make sure to save this structure as a CIF or PDB for for use in the next section.
7575

7676
### Adding an ORI token
77-
ORI (origin) tokens allow you to specify where the center of mass of the *designed* portion of your protein should approximately be. It can be used to have greater control over the interactions between the designed and input portions of your final structure. It can be particularly important for enzyme design as it can be used to guide the approximate orientation of how the generated protein should bind the ligand. <!-- TODO: add image here if Seth gives you the files you need. -->
77+
[ORI (origin) tokens](#adv_enzyme_tutorial_ori_token_def) allow you to specify where the center of mass of the *designed* portion of your protein should approximately be. It can be used to have greater control over the interactions between the designed and input portions of your final structure. It can be particularly important for enzyme design as it can be used to guide the approximate orientation of how the generated protein should bind the ligand. <!-- TODO: add image here if Seth gives you the files you need. -->
7878

7979
For our metalloprotease designs, we can start by assuming that the ORI token should be placed near the zinc atom since it should be relatively buried in the enzyme structure. You could just determine the coordinates of the zinc atom and use these for the ORI token input. However, for this tutorial let's say that we know we want the ORI token to actually sit slightly below the zinc atom. We could determine the coordinates of this ourselves, but often times it is helpful to determine the placement of our ORI token visually.
8080

@@ -168,7 +168,7 @@ Here we specify both the zinc atom (ZN) and the astacin molecule (PKF) as ligand
168168

169169
The length setting is telling RFdiffusion3 to design a protein that is between 120 and 150 amino acids in length, inclusive of the input residues. We are not using a `contig` setting here because we are using other ways to specify the designed or input portions of our structure.
170170

171-
We are `unindex`-ing all of the input residues here because their index in the final protein structure is not important, only their orientation with respect to each other, the ligand, and the zinc cofactor matters.
171+
We are `unindex`-ing all of the input residues here because their index in the final protein structure is not important, only their orientation with respect to each other, the ligand, and the zinc [cofactor](#adv_enzyme_tutorial_cofactor_def) matters.
172172

173173
Which brings us to the `select_fixed_atoms` setting. You can see the atom names for each residue in the PDB/CIF file for the input structure and/or view them on PyMOL using its labeling functionality. You'll notice that all of the atoms specified here are from the side chains of the residues. For this design problem, we want to keep the backbone position flexible to avoid overconstraining the designed protein. Any atoms not listed are free to be moved by default. Only the positions of the side chain atoms relative to the ligand and cofactor are important for enzymatic activity.
174174

@@ -209,7 +209,7 @@ Here is a brief description of what each of these additional settings are changi
209209
- `diffusion_batch_size=4` changes the diffusion batch size from 8 to 4, lowering the default batch size can help if you run into GPU memory errors.
210210
- `n_batches=6` run 6 batches of RFD3 instead of 1, for a total of 24 designs. The total number of designs you want to generate will depend on your design needs.
211211
- `inference_sampler.use_classifier_free_guidance=True` turns on classifier free guidance so that the model learns to predict the denoised structure with and without some of the conditioning features. It results in better adherence to your input constraints at the cost of lower diversity in the designed structures.
212-
- `inference_sampler.s_jitter_origin=1.5` adds a small amount of positional noise to the 'motif', the fixed portions of the structure that were given to RFD3 as input. Useful for exploring how the protein scaffold can orient around the active site.
212+
- `inference_sampler.s_jitter_origin=1.5` adds a small amount of positional noise to the ['motif'](#adv_enzyme_tutorial_motif_def), the fixed portions of the structure that were given to RFD3 as input. Useful for exploring how the protein scaffold can orient around the active site.
213213
- `seed=42` sets the seed to increase the reproducibility of results between RFD3 calculations. As of the publication of this tutorial, setting the seed **does not** result in fully deterministic results – they will still be slightly different between runs.
214214

215215
Feel free to reduce the number of batches or batch size if you have limited GPU resources.
@@ -233,7 +233,7 @@ First, let's open the `.cif.gz` file in PyMOL. If you are using the provided tut
233233

234234
Here is what we are looking for:
235235
1. **Overall fold.** Does the protein form a compact, well-folded structure? Is the active site located in a cleft or pocket, as you would expect for an enzyme?
236-
1. **Catalytic geometry.** Are the fixed atoms (H92/E93/H96/H102/Y149/M147) in their expected positions relative to the Zn(II) ion and the astacin ligand? Do the key interactions look reasonable? (You can find the new residue numbers for the catalytic residues in the `diffused_index_map` section of the design's JSON output file.)
236+
1. **Catalytic geometry.** Are the fixed atoms (H92/E93/H96/H102/Y149/M147) in their expected positions relative to the Zn(II) ion and the astacin ligand? Do the key interactions look reasonable? (You can find the new residue numbers for the [catalytic residues](#adv_enzyme_tutorial_catalytic_residue_def) in the `diffused_index_map` section of the design's JSON output file.)
237237
1. **Backbone connectivity.** Does the backbone trace smoothly through the structure without obvious clashes or unnatural loops?
238238
1. **Active site accessibility.** Is the ligand reasonably accessible from the protein surface? A completely buried ligand may be problematic depending on the application.
239239

@@ -311,27 +311,77 @@ You specify the atoms to use for the hydrogen bond donor/acceptor via a dictiona
311311
```json
312312
"select_hbond_acceptor": {
313313
"PKF": "O4,O6,O7"
314-
},
314+
},
315315
"select_hbond_donor": {
316316
"PKF": "N2,N4,N20"
317-
}
317+
}
318318
```
319319

320+
You can find an example output from the addition of these diffusion constarints here. <!-- TODO: add link to files -->
321+
322+
### RASA Conditioning
323+
Relative accesible surface area (RASA) conditioning allows you to tell RFD3 how exposed to the solvent or buried in the designed structure you want portions of your input structure to be in your final design. This is particularly useful in enzyme design as you will want certain parts of the subtrate to be enclused by the protein while others should extend outside of the active site cleft.
324+
325+
There are three input parameters that are used to control RASA conditioning in your design:
326+
- `select_buried` — atoms that should be **surrounded by protein** (low solvent accessibility)
327+
- `select_exposed` — atoms that should be **accessible to solvent** (high solvent accessibility)
328+
- `select_partially_buried` — atoms with **intermediate** solvent accessibility (less commonly used)
329+
330+
The specific atoms you want buried, exposed, or partially buried are specified in the same format as for `select_fixed_atoms`. Here is an example of what these settings could look like for the design example discussed in this tutorial:
331+
```json
332+
"select_buried": {
333+
"ZN": "ZN",
334+
"PKF": "O5,P1,O6,C28,C29,C33,C36"
335+
},
336+
"select_exposed": {
337+
"PKF": "N5,C19,O2,C2,O8,C23,C24,C31,O3"
338+
}
339+
```
340+
The buried were chosen because we want to bury the zinc ions and the atoms near the reactive center of PKF. These are direvely involved in the catalytic mechanism so they should be enclosed by the protein. The atoms on the peptide tail of PKF, meanwhile, should be exposed as they would naturally protrude from the binding cleft in a real enzyme-subtrate complex.
341+
342+
You can find an example output from the addition of these diffusion constarints here. <!-- TODO: add link to files -->
343+
344+
You can find an example output from a design that uses both hydrogen bonding and RASA conditioning here. <!--TODO: add link to files-->
320345

346+
(adv_enzyme_tutorial_conclusion)
347+
## Conclusion
348+
You have now set up an RFD3 calculation and successfully designed enzymes based around a theozyme created from a known structure. While the options discussed here are particularly useful in enzyme design projects, RFD3 has many more that you can explore by looking at {doc}`../input`.
321349

322350
(adv_enzyme_tutorial_glossary)=
323351
## Glossary
324352

353+
(adv_enzyme_tutorial_catalytic_residue_def)=
354+
### Catalytic Residue
355+
Catalytic residues are amino acids known to be cruicial for the enzymatic activity of a given reaction. They can either directly interact with the molecule the enzyme interacts with, or indirectly support the stability of the transition state of the ligand.
356+
357+
(adv_enzyme_tutorial_chelation_def)=
358+
### Chelation
359+
Chelation describes a type of interaction between ligands and metal atoms that form a ring structure.
360+
361+
(adv_enzyme_tutorial_cofactor_def)=
362+
### Cofactor
363+
A cofactor is a non-protein molecule that binds to an enzyme to help it function.
364+
365+
(adv_enzyme_tutorial_motif_def)=
366+
### Motif
367+
The input structure to RFD3 that designs are generated around.
368+
369+
(adv_enzyme_tutorial_ori_token_def)=
370+
### ORI Token
371+
The ORI token is the user-specified center of mass of the struture designed by RFD3. It gives the user some control over the interactions between the designed and input portions of the final structure.
372+
325373
(adv_enzyme_tutorial_theozyme_def)=
326374
### Theozyme
327-
<!-- TODO: Add definition -->
328-
### ORI Token
329-
<!-- TODO: Add definition-->
375+
A theozyme is a small structure formed from the transition state structure of a ligand and any catalytically important residues, atoms, etc. that are necessary to achieve a given catalytic reaction. It is the input structure for enzyme design calculations and can be created from known enzyme structures or via quantum mechanical methods.
376+
377+
(adv_enzyme_tutorial_transition_state_analog_def)=
378+
### Transition State Analog
379+
A transition state analog is a compound that resembles the transition state of a substrate molecule in an enzyme-catalyzed reaction
380+
330381

331382
(adv_enzyme_tutorial_refs)=
332383
## Resources and References
333384
- [RFdiffusion3 preprint](https://www.biorxiv.org/content/10.1101/2025.09.18.676967v2)
334385
- The procedure described here follows the approach discussed in [Kim, D. et al. (2025)](https://www.nature.com/articles/s41586-025-09746-w) and [Chen, A. et al. (2025)](https://www.nature.com/articles/s41586-025-09746-w)
335-
- Astacin structure: Grams, F. et al. (1996). Structure of astacin with a transition-state analogueinhibitor. *Nature Structural Biology* **3**, 671-675. DOI: https://doi.org/10.1038/nsb0896-671
336-
-
386+
- [Astacin structure](https://doi.org/10.1038/nsb0896-671)
337387

0 commit comments

Comments
 (0)