News

Published: August 21, 2026

Blog: OpenBind-0: Advancing Open Molecular Structure Prediction

Where to find the model and data

Quick links

  1. Overview 
  2. OpenBind is powering a new generation of co-folding models
  3. OpenBind-0 is a fully competitive molecular structure prediction model
  4. What does adding four more years of PDB data provide?
  5. The good, the bad, and the ugly: evaluating co-folding methods on newly released fragment-to-lead structure data
  6. Revisiting EV-A71 2A: Fine-tuning is good; pre-training is even better
  7. FatA: co-folding methods struggle to place the ligand
  8. RdRp: A true challenge for co-folding methods
  9. Conclusions
  10. Further Information

Overview

Today, we are releasing OpenBind-0 (OB0), the first in a new series of fully open-source molecular structure prediction co-folding models built on OpenFold3 and specialised for predicting the structures of proteins bound to small-molecule ligands. We are also releasing a dense, discovery-relevant dataset of 717 ligand-bound structures that capture fragment-to-hit progression across three drug targets.  

OB0 is trained on data from the Protein Data Bank (PDB) through June 2025 and uses a nearly final version of the upcoming OpenFold3 model architecture. Its checkpoints were selected to optimise for protein-ligand performance while retaining strong performance on other molecular modalities. We also introduce chemical steering to enhance the physical validity of predicted ligand structures. On protein-ligand benchmarks, OB0 is competitive with the strongest models currently available.

We are additionally demonstrating new best practices for releasing models, by including new structures by which to evaluate model performance, in this case in a drug discovery context. The structures were contributed by antiviral and agrochemical discovery projects running alongside OpenBind and accelerated by its technologies: the RdRp domain of NS5 from both Dengue serotype 2 and Zika viruses (thanks to ASAP Discovery Consortium); and fatty acid thioesterase A (FatA), a target for commercial herbicides with a well-documented resistance risk (thanks to Syngenta-funded student Ekaterina Kot). Each was the subject of a fragment-to-hit campaign, and the release covers 547 fragment binding events and 170 hit binding events, comprising 298 unique fragments and 101 unique hits.

Evaluating OB0 and other co-folding methods on this new data and previously released OpenBind data, we find that performance varies widely by target, from high accuracy on EV-A71 2A protease down to success rates below 10% on both RdRp systems. We also test whether fine-tuning on target-specific fragment structures recovers the gap and find that gains depend heavily on the system. Closing this gap is the subject of ongoing work.

OpenBind is powering a new generation of co-folding models

OB0 is the first in a series of OpenBind models. Each release will incorporate more recent PDB structures and newly available data from OpenBind, with the goal of staying within roughly six months of the latest publicly available structural data. Future versions will also add capabilities tailored to small-molecule applications, with this first release introducing inference-time steering methods to the OpenFold3 family to improve chemical validity.

OB0 establishes a baseline for measuring how much OpenBind data boosts model performance. The training set for OB0 contains no OpenBind data, only public structures, providing a baseline for future models to compare to as new data is added into the training set. Keeping the model current is also important for active learning. OpenBind uses model predictions to help determine which experiments to run next, so that guidance should reflect both the latest public data as well as data generated within OpenBind’s experimental design cycles.

OB0 is fully open source under the Apache 2.0 licence: the training data, code, model weights, and training recipes are publicly available. Users can also retrain or fine-tune the model using mixtures that include their own data. This is particularly important for robust fine-tuning, which often requires retaining examples from the original training distribution.

Below, we walk through OB0’s benchmark results, examine what more recent PDB data contributes, and show how the current crop of co-folding methods perform on the newly released OpenBind targets.

OpenBind-0 is a state-of-the-art molecular structure prediction model

We evaluate OB0 following the Runs N’ Poses (Škrinjar et al., 2025) methodology, applied to a set of complexes released after OB0’s 30 June 2025 training cutoff. The benchmark is designed to measure generalisation beyond the training set. It groups post-cutoff protein–ligand complexes according to their similarity to the nearest training example, considering both binding-pocket geometry and ligand shape. This benchmark was constructed from a preview version of an updated PLINDER (Durairaj et. al. 2024).

The models included in the comparison were trained using data from different periods. AlphaFold3, ESMFold2, and OpenDDE have a cutoff date of September 2021, Boltz-2 uses June 2023, and Protenix-v1-20250630 and OB0 use data through June 2025. For consistency, the similarity bins shown here are calculated relative to the June 2025 training set and therefore describe the difficulty of each target with respect to the data available to OB0 and Protenix-v1.

Across similarity bins, OB0 is comparable to AlphaFold3 and other biomolecular structure prediction models. As expected, performance improves the more closely targets resemble structures in the training set. What was not expected is that models with varying training date cutoffs perform similarly. To understand why, we examine how four additional years of PDB data change the relationship between the training and test sets below.

Figure 1: Performance of co-folding methods on protein–ligand structure prediction, stratified by similarity to the training set. Success rate is the fraction of drug-like ligands with a correct pose (lDDT-PLI > 0.8 and ligand RMSD < 2 Å). Error bars are 95% Wilson score intervals.

We also implemented chemical steering during diffusion sampling, following the approaches taken in Boltz and Protenix. Downstream uses of predicted structures, from free energy calculations to medicinal chemistry design decisions, depend on ligand geometry being physically plausible and not merely correctly positioned. Here, we measure chemical validity using PoseBusters (Buttenschoen et al. 2023). Steering raised the joint success rate (accuracy and PoseBusters validity) from 48% to 61%.

Figure 2: Inference-time steering improves ligand chemical validity for OB0. Bars show the fraction of ligands that have correct poses (lDDT-PLI > 0.8, ligand RMSD < 2 Å) and the fraction with correct poses and valid chemical geometry (lDDT-PLI > 0.8, ligand RMSD < 2 Å, and PoseBusters-valid), with and without steering during diffusion sampling.

What does adding four more years of PDB data provide?

One might expect that extending a model’s training cutoff from 2021 to 2025 would improve performance across all similarity bins. In practice, we observe that the improvements are not uniform.

OpenFold3-preview2 (OF3p2) and OB0 perform similarly across many similarity bins, despite OB0 having been trained on more data (the two share similar, though not identical, architectures). This suggests that simply increasing the number of training examples does not make every test case meaningfully easier. Gains are visible in the highest and lowest similarity bins. Improvements at the high similarity bins may result from test complexes being closer to structures OB0 was trained on, while gains in low similarity bins may reflect broader coverage of chemical and structural space translating into improved generalisation capability.

Figure 3: Success rate for OB0 (2025 cutoff) and OpenFold3-preview2 (2021 cutoff), stratified by similarity to the training set. Despite seeing more data, OB0 and OF3p2 have similar success rates in many similarity bins.

To investigate this, we calculate the similarity of each test complex to two versions of the training set: one ending in 2021 (OF3p2’s training cutoff) and another ending in June 2025. Most complexes remain in the same similarity bin under both cutoffs. In other words, although new protein-ligand entries were added to the PDB during this period, the nearest relevant training example for a large fraction of the benchmark did not change enough to move the target into a different similarity bin.

Figure 4: Similarity to the training set of test entities changes little between the 2021 and 2025 cutoffs.

This helps explain the comparable performance between the two models for many test complexes. Among complexes assigned to the 0-60 similarity bin using the 2021 cutoff, those that remain in the same bin under the 2025 cutoff show little difference in performance between OF3p2 and OB0. In contrast, for complexes that moved into the highest similarity bin under the 2025 training set, OB0 substantially outperforms OF3p2.

Figure 5: Success rates improve substantially for test entries which have more similar training set entries under the updated training cutoff. Performance of OB0 (2025 cutoff) and OF3-p2 (2021 cutoff) on complexes originally in the 0–60 similarity bin, split by whether the target’s bin changed under the 2025 training set.

Our observations suggest that the benefit of four more years of PDB data for these models is in large part determined by whether the training structures make a test structure more in-distribution. The OpenBind Project intends to generate structures selected specifically to close important gaps and improve future models.

The good, the bad, and the ugly: evaluating co-folding methods on newly released fragment-to-lead structure data

Next, we evaluate the performance of OB0 on structures from our previous data release, as well as new benchmark sets spanning three different protein targets. As we will discuss, we find that pre-training on relevant structural data can yield significant performance gains over fine-tuning. Furthermore, we find that co-folding models exhibit particularly poor performance on two highly flexible protein systems that are dissimilar to data in the training set. Below, we highlight strengths and limitations of co-folding models for each benchmark system. An extended analysis of benchmarking performance is included in the Further Information section.

Revisiting EV-A71 2A: Fine-tuning is good; pre-training is even better

Our previous blog post showed that fine-tuning OF3p2 on fragment screening data for the EV-A71 2A protease substantially improved predictions for follow-on compounds. Those fragment structures were deposited in the PDB before the OB0 training cutoff but after the OF3p2 cutoff, providing a reasonably controlled experiment in which OB0 has seen the fragment structures as part of its pre-training mixture but OF3p2 has not, while both are blind to the follow-on compounds.

Evaluating OB0 on follow-on compounds, we find that it outperforms the fine-tuned OF3p2 model by a wide margin, reaching a top-25 success rate of 92.2% and a top-1 success rate of 73.8% (Figure 6). While we cannot solely attribute the gain to the fragment data, it suggests that new fragment data may be utilised in a more effective way during pre-training compared to our current fine-tuning protocol.

Figure 6: Success rates for OF3p2, fine-tuned OF3p2, OB0, AlphaFold3, and Protenix-v1-20250630 on the EV-A71 2A Protease benchmark set.

FatA: co-folding methods struggle to place the ligand

We also benchmarked models on a fragment-to-hit campaign targeting fatty acid thioesterase A (FatA), a commercially relevant herbicide target. Fragments and follow-ons bind almost entirely within a single pocket that shows a moderate degree of flexibility. Our data release covers 311 fragment binding events across 128 unique fragments, along with 91 binding events for 38 unique follow-ons, none of which appeared in the training data of any model tested here.

Success rates are low across the board, although more recent training data clearly helps. OF3p2 reaches a top-25 success rate of 14.5% and OB0 roughly doubles that at 28.2%, with Protenix-v1-20250630 achieving the strongest success rate of 39.4% (Figure 7).

The failure mode differs from what we observe for EV-A71 2A protease (see Further Information section). Models nearly always find the correct pocket and generate a reasonable conformation for the target, but fail to accurately predict the correct ligand pose. Fine-tuning on fragments modestly lifts success rates, but far less than the gains seen for EV-A71 2A protease.

Figure 7: Summary of co-folding performance on FatA benchmark data. A) Structural overview of binding events for FatA. Fragments are displayed in green; follow-ons are displayed in cyan. A zoomed in view shows flexible binding site residues in magenta. B) Success rates for co-folding models

RdRp: A true challenge for co-folding methods

Next, we benchmarked co-folding models using data generated from fragment-to-lead optimisation campaigns targeting the RNA-dependent RNA polymerase (RdRp) domain of the nonstructural protein 5 (NS5) from Dengue virus serotype 2 (RdRp DENV-2) and Zika virus (RdRp ZIKV). The RdRp domain, responsible for replicating the viral RNA, is essential for the viral life cycle and is structurally highly conserved between Flaviviruses, making it an attractive target for the design of broad-spectrum antiviral compounds. We thank the AI-driven Structure-enabled Antiviral Platform (ASAP) Discovery Consortium for providing these structures as part of their open science early-stage drug discovery efforts to develop broad-spectrum antivirals for pandemic preparedness with the goal of globally equitable and affordable access.

For RdRp DENV-2, we release 85 fragment binding events covering 67 unique fragments and 39 follow-on binding events covering 30 unique compounds. For RdRp ZIKV, we release 151 fragment binding events covering 99 unique fragments and 40 follow-on binding events covering 33 unique compounds.

These are the most challenging targets in the release by a wide margin. No model exceeded a top-25 success rate of 10% on either system, topping out at 7.7% for RdRp DENV-2 and 6.7% for RdRp ZIKV (Figure 8).  Fine-tuning models on fragment data improved success rates only slightly (see Further Information).

Figure 8: Summary of co-folding performance on RdRp benchmark sets. A) Structural overview of fragment (green) and follow-on (cyan) compounds binding the RdRp DENV-2. A zoomed view shows flexible residues (magenta) within the binding site. B) Success rates for four co-folding models. Analogous information for RdRp ZIKV is displayed in panels C-D.

What makes this target so challenging? It is a large protein with many pockets, with compounds primarily binding near highly flexible regions of the receptor. Models frequently miss the binding site entirely, and even when the pocket is found, the conformation and pose within are incorrect (see Further Information). Furthermore, all RdRp structures are dissimilar to data in the OF3p2 and OB0 training sets (see Further Information). As a result, we do not observe performance gains when we train on more recently released structures in the PDB.

Conclusions

We hope that OB0 will be a useful model for the community to use and build on. As the OpenBind Project continues to generate more data, OB0 will serve as a baseline to understand the impact of adding that data and will help guide data selection.

The 717 structures released alongside it show how much room remains to improve models. We believe that shared access to difficult, discovery relevant structures helps the field understand the limitations of these methods in real world drug development.

Further Information

Evaluation of co-folding methods on OpenBind generated data

To further investigate the performance of OpenBind-0 on unseen target-specific data, we benchmarked models on structural data collected from four fragment-to-lead optimisation campaigns spanning four protein systems.

For each system, we assessed the performance of two models trained on structures released in the PDB prior to September 30, 2021 (OF3p2 and AlphaFold3) and two models trained with additional data released in the PDB prior to June 30, 2025 (OB0 and Protenix-v1-20250630). For each co-folding model, 5 seeds x 5 diffusion samples were used to generate a total of 25 predicted structures per binding event. Evaluation sets for each protein target comprise binding events for follow-on compounds that were not included in the training data for OB0. Success rates were calculated for top-25 and top-1 ranked predictions for each binding event, where predictions were ranked on the basis of their protein–ligand iptm score. In this benchmark, a co-folded prediction was labelled as successful if it had a ligand RMSD ≤ 2.0 Å, an LDDT-PLI ≥ 0.8, and if it passed all PoseBusters validity checks.

In addition to benchmarking success rates, we perform a failure mode analysis on each benchmarked protein-ligand system. To measure if a ligand was placed within the correct pocket, we calculated pocket recall, defined as the fraction of ground-truth pocket residues within 6.0 Å of the ligand in the predicted structure. When pocket recall < 0.35, we considered the ligand to be placed in the incorrect pocket. As an approximate measure of predicted pocket conformation accuracy, we calculated the ligand pocket LDDT (LDDT-LP) for each co-folded structure, where an LDDT-LP < 0.8 denoted an incorrect pocket conformation. For each binding event, we consider failure mode metrics for the best top-25 prediction from each co-folding model.

For each protein-ligand system, we quantify training set similarity using the SuCOS-pocket similarity metric described by Škrinjar et al. in the Runs N’ Poses publication. To calculate training set similarity, we utilise the custom scoring code deposited in the feat/custom-structure-scoring branch of the PLINDER github repository. All similarity metrics are calculated using structures in an updated version of the PLINDER database, which includes structures released in the PDB prior to July 2026. Similarity metrics for benchmark sets discussed in the following sections are available for download in the GitHub repository that accompanies this blog post.

EV-A71 2A Protease benchmark set

First, we benchmarked OB0 on a subset of 374 binding events for 306 unique follow-on compounds that were included in the first OpenBind data release. This subset was selected to exclude structures which were deposited in the PDB prior to the OB0 training cutoff date, as well as structures used for fine-tuning experiments. All follow-on compounds in the subset bind in or near a single pocket, within which the majority of fragment hits also bind (Figure A1A).

EV-A71 2A Protease Performance

When comparing OB0 to OF3p2, which was trained on data released prior to September 30, 2021, we find that OB0 achieved a large performance boost on the EV-A71 2A Protease benchmark. For this system, the inclusion of additional training data dramatically shifts the training set similarity from the 35–55 range into the 65–100 range (Figure A1B). This shift is reflected in the performance of the OB0 model, which achieved a top-25 success rate of 92.2% and top-1 success rate of 73.8% (Figure A1C). Protenix-v1-20250630 also performs comparably, further indicating that the inclusion of additional training data can significantly boost the accuracy of co-folding models.

Figure A1:  Overview of structural ensemble and co-folding performance on the EV-A71 2A Protease benchmark. A) Structural ensemble of ligand-bound receptor structures with follow-ons (cyan) and fragments (green) displayed as sticks. The zoomed view highlights a flexible receptor motif (magenta) within the primary binding site. B) Distribution of follow-on compound similarities to the training set using 2021-09-30 and 2025-06-30 as training cutoffs. C) Success rates for co-folding models on the benchmark set.

Failure mode analysis for co-folding models on this benchmark set reveals that for the majority of cases, every co-folding model placed the ligand in the correct binding pocket (Figure A2). However, for OF3p2 and AlphaFold3, the majority of failed predictions can be attributed to the generation of an incorrect pocket conformation (Figure A2). In contrast, models that were trained on more recent PDB data typically generated relatively accurate receptor conformations. Given the inclusion of both fragment and follow-on data in the OB0 training set, the observed performance gains are unsurprising. However, the increases in performance relative to models with older training cutoff dates underscore the importance of concerted data collection efforts, which have the potential to expand the domain of applicability for co-folding models to new systems.

Figure A2: Failure modes for co-folding predictions on the EV-A71 2A Protease benchmark set.

FatA benchmark set

Lastly, we benchmarked models from a fragment-to-hit campaign targeting fatty acid thioesterase A (FatA), a commercially relevant herbicide target. FatA belongs to a family of enzymes which terminate de novo fatty acid biosynthesis inside plastids of higher plants and are therefore essential to healthy plant growth and fertility. Currently, FatA is a target of five commercial herbicides designated Group 30 by the Herbicide Resistance Action Committee (HRAC). These herbicides are used as a second line of defence against herbicide-resistant weeds; however, they are limited in their utility due to a similar pharmacophoric profile. Fragment-based discovery has been employed to find chemically novel inhibitors to circumvent this.

The dataset comprises 311 binding events for 128 unique fragment compounds, reported by Kot et al. Additionally, models were benchmarked against a set of 91 binding events, covering 38 unique follow-on compounds. Most fragment hits bind within a single binding site, and almost all follow-on compounds are concentrated within the same pocket (Figure FFA). Within the pocket, a subset of binding site residues exhibit a range of different conformations, indicating that there is some degree of flexibility within the pocket. All fragment data was released in the PDB after 30 June 2025 and was not seen during training for any co-folding model. However, we note that there is a shift in training set similarity between the 2021-09-30 and 2025-06-30 training cutoff dates (Figure A3B). At the time of writing, FatA data is available for download via the fragalysis platform FatA Performance.

When comparing the performance of OF3p2 to OB0, we found that the top-25 success rate roughly doubled from 14.5% to 28.2% (Figure A3C). Furthermore, we found that Protenix-v1-20250630 achieved the strongest performance, with a top-25 success rate of 39.4%.  Notably, prior to the June 30, 2025 training cutoff date, only 7 ligand-bound structures with a sequence identity ≥ 30% to FatA were deposited in the PDB. However, we still observe modest shifts in training set similarity when we change the training cutoff date (Figure FFB). This result highlights the utility of including additional data to retrain co-folding models, even when new training data is not collected for the target of interest. Failure mode analysis for co-folding results on the FatA benchmark set reveals that in most cases, co-folding methods predict the ligand binding to an accurate conformation of the correct binding pocket (Figure A4). However, the relatively low success rate indicates that for this system, accurate ligand placement within the binding site was the primary challenge for co-folding models.

Figure A3: Overview of structural ensemble and co-folding performance on the FatA benchmark. A) Structural ensemble of ligand-bound structure for FatA. Fragments are displayed as green sticks and follow-ons as cyan sticks. The zoomed view highlights a flexible receptor motif (magenta) within the primary binding site. B) Distribution of follow-on compound training set similarities using 2021-09-30 and 2025-06-30 training cutoffs. C) Success rates for co-folding methods on follow-on compounds binding to FatA.
Figure A4: Failure modes for co-folding predictions on the FatA benchmark set.

RdRp benchmark sets

Next, we benchmarked co-folding models using data generated from fragment-to-lead optimisation campaigns which targeted the RNA-dependent RNA polymerase (RdRp) domain of the nonstructural protein 5 (NS5) from Dengue virus serotype 2 (DENV-2) and Zika virus (RdRp ZIKV). The RdRp domain, responsible for replicating the viral RNA, is essential for the viral life cycle and structurally highly conserved between Flaviviruses, making it an attractive target for the design of broad-spectrum antiviral compounds.

The RdRp DENV-2 dataset consists of 85 fragment binding events for 67 unique fragment compounds, and 39 binding events for 30 unique follow-on compounds. For the RdRp DENV-2 dataset, 76/85 fragment binding events covering 60/67 fragment hits are described in a previous publication by Saini et al., and were released in the PDB prior to the training date cutoff for OB0 and Protenix-v1-20250630. The RdRp ZIKV set includes 151 binding events for 99 unique fragments, and 40 binding events for 33 unique follow-ons. Binding events for all follow-on compounds that were used to evaluate model performance are not deposited in the PDB and were therefore not part of the training set for any co-folding model. All crystallographic binding events for fragment hits and follow-on compounds targeting RdRp DENV-2 and RdRp ZIKV are available for download via the fragalysis platform

Across both RdRp systems, we observed that fragments bound in a variety of pockets across the protein surface. Most follow-on compounds are concentrated within a single binding site, characterised by highly flexible receptor motifs (Figure A5A, A5D). This makes RdRp a challenging target for co-folding, in that models must not only generate an accurate conformation of the binding pocket but also place the ligand within the correct pocket. Furthermore, all binding events in the RdRp benchmark sets are highly dissimilar to data in the released PDB prior to both the 2021-09-30 and 2025-06-30 training cutoffs (Figure A5B, A5E).

For both RdRp systems, we find that all co-folding models fail to reliably predict binding events for follow-on compounds (Figure A5C, A5F). For DENV2 and ZIKV, co-folding achieved a maximum top-25 success rates of 7.7% and 6.7%, respectively. This result is reflective of the fact that there is poor coverage of homologous liganded RdRp structures in the training data, causing co-folding methods to consistently fail on these targets. To further investigate how co-folding models fail, we conducted a failure mode analysis across both RdRp targets. In our analysis, we assess the ability of each co-folding method to find the correct binding pocket and generate an accurate pocket conformation.

Figure A5: Overview of structures and results for RdRp benchmark sets. A) Structural ensemble of ligand-bound RdRp DENV-2 receptor structures with follow-ons (cyan) and fragments (green) displayed as sticks. The zoomed view highlights a flexible receptor motif (magenta) within the primary binding site. B) Distribution of follow-on compound training set similarities using 2021-09-30 and 2025-06-30 training cutoffs for RdRp DENV-2. C) Success rates for co-folding methods on the RdRp DENV-2 benchmark set. Analogous information for RdRp ZIKV is displayed in figures D-F.

A summary of the failure modes for each binding event in the RdRp DENV-2 and RdRp ZIKV benchmark sets are displayed in Figure A6. These results indicate that across all models, a large proportion of failed co-folding predictions placed the ligand in an incorrect binding site and generated an incorrect conformation for the main binding pocket. We suspect that these failures primarily result from two factors: 1) fragment hits in the training set bind to a variety of pockets in which no follow-on compounds bind and 2) conformational modelling remains a challenge for co-folding methods, and the primary binding site in RdRp is highly flexible.

Figure A6: Failure modes for co-folding predictions on the A) RdRp DENV-2 benchmark set and B) RdRp ZIKV benchmark set.

Does fine-tuning on fragment data help? It depends!

In our previous blog post, we found that fine-tuning OF3p2 on fragment screening data for the EV-A71 2A Protease significantly improved co-folding performance for predicting follow-on compounds. However, our previous analysis was limited to a single system. To further investigate fine-tuning behaviour, we aimed to conduct similar fine-tuning experiments across additional systems. For each system, we fine-tuned OF3p2 and OB0 on all available fragment binding events. Fine-tuning runs were performed with a batch size of 8 and a learning rate of 10-3, remaining consistent with the protocol used in the previous blog post. All evaluations were conducted on the same set of binding events for follow-on compounds, which were not seen during training or fine-tuning for any of the benchmarked co-folding methods.

In short, we find that performance gains obtained for fragment fine-tuning vary on a target-by-target basis. For the EV-A71 2A Protease benchmark set, we find that OB0 significantly improved upon the success rate obtained by the fine-tuned OF3p2 (Figure A7A). As noted previously, a subset of the fragment and follow-on binding data for EV-A71 2A Protease is included in the training set for OB0. This suggests that while fine-tuning may improve the performance of a base model, retraining the model with new data can have a greater impact on performance.

Figure A7: Success rates for co-folding models fine-tuned on fragment data for the A) EV-A71 2A Protease benchmark set, B) FatA benchmark set, C) RdRp DENV-2 benchmark set, and D) RdRp ZIKV benchmark set.

For FatA, we find that fine-tuning OF3p2 and OB0 results in similar performance gains (Figure HHB). This is likely because fragment screening data for this target was released after the training cutoff date for both models. Although only a moderate overall success rate is observed, this result underscores the fact that the addition of new training data alone can improve co-folding performance. For RdRp DENV-2 and RdRp ZIKV, fine-tuned models still exhibited poor performance, albeit with slightly improved success rates when compared to base models (Figure A7C, A7D).

On the whole these results indicate that fine-tuning does not work out of the box for all systems, and highlight that naïvely fine-tuning on all available structures may not be sufficient for obtaining meaningful improvements in co-folding accuracy. To establish generalisable fine-tuning protocols that are applicable to a variety of systems, further work is needed.

Acknowledgements

We thank everyone who was involved in generating and processing data, and preparing this release, as well as the funders supporting OpenBind

OpenBind received funding from the UK Department for Science, Innovation and Technology under grant number G2-SCH-2025-06-16537.

We acknowledge the use of resources provided by the Isambard-AI National AI Research Resource (AIRR). Isambard-AI is operated by the University of Bristol and is funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) via UK Research and Innovation; and the Science and Technology Facilities Council.

For the FatA dataset, we thank Ekaterina (Kate) Kot, graduate student at the Centre for Medicines Discovery at University of Oxford, and her funders Syngenta and industrial supervisors Nick Mulholland and Mark Montgomery.

For the antiviral datasets, RdRp and 2A protease, we thank the AI-driven Structure-enabled Antiviral Platform (ASAP) Discovery Consortium for funding the discovery work and its close collaboration with OpenBind to generate the structures.

We acknowledge the Diamond Light Source for access to the fragment screening facility XChem, for usage of DSi-Poised library, and beamlines I03 and I04-1 for data collected under proposals lb42888 and lb32627.

We would like to thank Peter Skrinjar, Janani Durairaj, and Vladas Oleinikovas and the Schwede Lab at Biozentrum for providing access to an updated PLINDER database, curating an updated Runs N’ Poses style benchmark set, and aiding with training set similarity calculations.

Related Announcements

August 21, 2026

OpenBind releases its first predictive AI model

OpenBind has released its first open predictive AI model, trained on high-quality experimental protein–ligand data. Released alongside the project’s first public dataset, it provides a foundation for predicting molecular interactions and benchmarking structure-based AI methods for drug discovery.

Our partners

Diamond logo
Columbia University logo
EBI logo
IPD logo
MedChemica logo
MSKCC logo
OMSF logo