A prospective, blinded competition involving 29 organizations and 511 AI-designed antibody sequences has producedone of the largest prospective, blinded benchmarks of computational antibody design to date — and found that performance varies so widely across groups and tasks that no single algorithmic approach dominates, according to a study published August 19 in Nature Biotechnology.
The benchmark, organized by Santa Fe-based Specifica (an IQVIA business) with experimental validation by Carterra and Sapidyne Instruments, tested three distinct tasks against the SARS-CoV-2 receptor-binding domain (RBD): affinity maturation of a known binder, rank-ordering of sequences within defined clusters, and de novo design of sequences outside the training library. All submissions were evaluated blindly against experimental affinity measurements from high-throughput surface plasmon resonance (HT-SPR) and orthogonal kinetic exclusion assay (KinExA) confirmation, plus a five-assay developability panel.
AI delivers in affinity maturation — but inconsistently
The clearest positive result emerged in Challenge 1, affinity maturation, where the best-performing submission — from California-based Aureka Biotechnologies using its pairformer-based AuraIDE platform — achieved a KinExA-confirmed equilibrium dissociation constant (KD) of 95 pM against the RBD, statistically indistinguishable from the best antibody produced by experimental selection (113 pM) and representing approximately a 2,000-fold improvement over the parental antibody. The winning method used Direct Preference Optimization to learn relative fitness rankings from yeast-display enrichment data rather than regressing absolute affinities, then sampled 10,000 CDR variants to identify top candidates.
The performance gap across participants was substantial, however. A simple consensus sequence approach — identifying the most frequent residue at each CDR position with no machine learning — ranked third in Challenge 1, establishing a non-AI baseline that most submitted methods failed to exceed. The paper states that "when presented with deep, targeted, high-quality data, the best AI algorithms (but not the others) can already provide real value in affinity maturation."
Cluster ranking exposed a persistent weakness
Challenge 2, which asked participants to rank sequences by predicted affinity within three HCDR3-defined clusters, produced the benchmark's most cautionary finding. All but one strategy — from Washington University in St. Louis — performed worse than randomly selecting clones from the cluster. Nearly every computational strategy performed worse than random selection: only 9.8%–13.8% of AI submissions improved on the cluster control, compared with approximately 39% of random picks. Structure-aware methods showed some advantage in certain clusters but not universally, and no method achieved consistent positive Spearman rank correlations across all three clusters.
De novo design remains the hardest frontier
In Challenge 3, out-of-library generative design, Pasadena-based Xencor (Nasdaq: XNCR) produced the highest-affinity hit at 2.9 pM KD — but the antibody failed hydrophobic interaction chromatography (HIC), raising a significant developability concern. Approximately 30% of all Challenge 3 submissions produced no detectable binding, and the performance gap between top and median organizations was larger here than in any other challenge. Winning designs tended to make conservative changes in flanking CDRs while retaining HCDR3 sequences already present in the experimental dataset.
Developability not reliably predicted by any method
Across all three challenges, the five-assay developability panel — covering hydrophobicity, polyreactivity, self-interaction, thermal stability, and aggregation temperature — revealed that high affinity and clean developability did not reliably co-occur in AI-designed sequences. Some high-affinity designs flagged in one or more assays, and computational scores did not consistently predict developability outcomes. The paper identifies this as an ongoing gap.