Minjae Park, Roman Nett, Brian Petersen, and Arvind Sivasubramanian
bioRxiv preprint, DOI: 10.64898/2026.08.20.745880
Overview
While general protein:protein complex prediction has advanced rapidly, the progress made for antibody-antigen prediction has been slower. Antibody recognition is mediated largely by the complementarity determining region (CDR) loops, and accurately predicting their conformations, particularly that of the long and highly variable CDRH3 loop, remains a challenge. In addition, antibodies evolve independently of their targets, so the evolutionary co-variation captured by multiple sequence alignments, which underpins the success of modern co-folding methods on natural protein complexes, is largely absent at the antibody-antigen interface. Recent advances have been achieved by increased sampling, improved co-folding models (such as AlphaFold3 and its successors), and the incorporation of experimental data (e.g., epitope constraints).
Study rationale and objectives
Even as co-folding methods, sampling strategies, and experimental constraints have each advanced, it remains unclear how much of the recent progression in antibody-antigen prediction is attributable to each factor individually. This study addresses that question directly by benchmarking ten co-folding protocols under matched conditions.
Objectives:
Approach and techniques:
HuMonoAg-Bench, a benchmark data group of over 400 experimentally determined antibody complexes (VHH and Fv) with human monomeric antigens, was assembled. The benchmark included complexes released after a uniform training date cutoff of September 30, 2021, and partitioned accordingly. The benchmark was used to independently evaluate ten co-folding protocols incorporating unconstrained methods and those that are epitope-constrained. In addition, complementary structural features were characterized that could influence benchmark difficulty, similar to the use of conformational change in protein-protein docking benchmarks, namely the antigen length, buried interface area, and CDRH3 length, representing the antigen, interface, and antibody, respectively.
Major findings
Conclusions and impact
The accuracy of antibody-antigen co-folding depends on simultaneously modeling several structural features: the antigen structure, the conformations of the CDR loops, and the formation of the antibody-antigen interface. Prediction accuracy and confidence depend on epitope constraint correctness. Other classes of antigens, including multimeric assemblies and viral proteins, may present additional challenges not captured here. Consequently, many antibody-antigen complexes remain beyond the reach of current co-folding methods, with CDRH3 modeling and sampling of accurate binding modes remaining major limitations. Ranking becomes increasingly important as successful binding modes become more readily accessible.
Implications for therapeutic antibody development
Antibody-antigen complexes remain sparsely and redundantly represented in current structural databases. Increasing the diversity of experimentally determined complexes could support further progress by broadening the range of antibody-antigen interactions available for model development and evaluation.
For more details, read the full article in bioRxiv.