Benchmarking Apotek Landscape: Zero-Shot Developability Prediction on a Public Standard
- Jeff Ma

- Jul 21
- 6 min read
Updated: Jul 23
How an AI–Physics Hybrid Generalized to Unseen Antibodies—Without Benchmark-Specific Training

Cross-validation can show how well a developability model performs on familiar data. But in real drug discovery, the molecules being evaluated are usually new. The real challenge is whether a model can remain reliable when it encounters molecules it has never seen before.
We tested this by evaluating Apotek Landscape zero-shot on GDPa3, the blinded test panel from the 2025 Ginkgo AbDev Competition. This post presents the benchmark design, the results for thermal stability and aggregation, and what they reveal about generalization to unseen antibodies.
Key Takeaways
An Independent, Public Test: We evaluated Apotek Landscape on GDPa3 — the blinded, held-out panel from the 2025 Ginkgo AbDev Competition — rather than an in-house set. The competition itself found that strong cross-validation scores often dropped on the held-out data, which makes out-of-distribution generalization the property worth testing.
A Clean, Zero-Shot Evaluation: Apotek Landscape was run zero-shot (without benchmark-specific fine-tuning) — no exposure to the competition's training data of any kind. That is a rigorous evaluation of whether a platform works on molecules it has never seen.
Thermalstability - Comparable to Top Reported Performance: On thermal stability (experimental Tm2, the Fab/CH3 domain), Landscape's zero-shot ranking reached Spearman ρ = 0.407 (p = 3.16 × 10⁻⁴) — comparable to the top reported performance (0.392), achieved without touching the training set.
Aggregation - A Strong, Significant Signal: On aggregation/hydrophobicity (experimental HIC), Landscape reached ρ = 0.540 (p = 2.76 × 10⁻⁷) — a strong, statistically robust correlation that, on a zero-shot basis, falls around the top 10 of the competition's hydrophobicity leaderboard.
The Power of AI-Physics Hybrids: Both results come from the same principle behind Landscape: pairing the speed of AI-driven prediction with physics-based structural validation, so assessments are fast and biophysically grounded.
The Antibody Developability Challenge
The Generalization Problem
Therapeutic antibody discovery has advanced for decades, yet deciding which candidates are worth carrying forward remains slow and expensive. Much of that cost comes from developability — whether a molecule is stable, soluble, and resistant to aggregation. If those liabilities could be predicted reliably in silico, teams could test fewer molecules and reach decisions faster.
The catch is trust. A predictor is only useful if its performance holds on molecules no one has measured yet, and that trust has to be earned against experimental data that is independent and sufficiently diverse. A model tuned on a familiar in-house dataset may simply be memorizing its training distribution rather than learning biophysics.
The 2025 Ginkgo AbDev Competition made this concrete, using a public training set of 246 clinical antibodies (GDPa1) and a blinded, held-out test set of 80 sequences (GDPa3). The organizers found that strong cross-validation scores routinely dropped on the held-out set — evidence that generalizing beyond familiar data is genuinely hard. The lesson for anyone building developability models: performing well on data you have seen says little about the molecule in front of you.
Why Ranking, Not Just Scoring
In early triage, teams rarely need an exact melting temperature or aggregation value — they need to know which candidates to test first. A model that reliably orders a panel from low-risk to high-risk turns an intractable screen into a short, prioritized shortlist, even when absolute values stay uncertain. That is why rank correlation against experimental ground truth is the metric that maps most directly to a real decision.
What Makes the Apotek Landscape Approach Different
Integrated AI-Physics Architecture
Landscape does not rely on a single black-box sequence predictor. It layers complementary forms of intelligence:
Deep-learning models that rapidly evaluate sequence-to-property relationships
Deep-learning structure prediction that turns sequence into a three-dimensional model in minutes
Physics-based structural validation that screens and, where needed, repairs predicted structures so downstream scores rest on a biophysically valid foundation
This hybrid is the point: the AI delivers reach and speed, paired with physics-based structural validation.
Independent, Structure-Based Property Assessment
Each validated structure is scored in parallel by specialized engines — including thermostability and aggregation propensity. Rather than compressing every signal into a single opaque risk number, each engine targets the distinct biophysical mechanism behind its property and reports independently, so teams can weigh trade-offs against their own priorities.
Table 1. Comparing Computational Approaches for Protein Developability Assessment
Approach | Key Limitation | Practical Trade-off |
Sequence-based ML models | Limited structural information and weaker out-of-distribution generalization. | Highest throughput, with limited structural validation. |
Physics-based pipelines (e.g., Rosetta/PyRosetta lineage) | Computationally intensive for large-scale screening. | High structural rigor, lower throughput |
Apotek Landscape (AI-Physics Hybrid) | Slower than sequence-only inference due to structural modeling. | Balanced throughput and structural validation, with adjustable modeling depth. |
Benchmark Study: The GDPa3 Held-Out Panel
Approach: A Zero-Shot Test on Blinded Data
We ran Apotek Landscape against GDPa3 — the blinded, held-out test set of 80 IgG antibodies (paired heavy + light chains) from the 2025 Ginkgo AbDev Competition. The run was fully zero-shot: no fine-tuning, no use of the GDPa1 training set, no exposure to the benchmark during development. Because the platform never sees the test data in any form, there was no benchmark-specific training or fine-tuning — the result reflects how Landscape behaves on molecules unseen within this benchmark.
For each antibody, the variable domain was folded by deep-learning structure prediction, screened by physics-based validation (with structural repair applied where a modeling artifact was detected), and scored by Landscape's specialized engines. Each property's predicted ranking was then compared against the experimental values using Spearman rank correlation.

Results: Generalization That Holds Up
• Thermal stability (Tm2): ρ = 0.407 (p = 3.16 × 10⁻⁴), comparable to the competition's top reported performance (0.392) — reached entirely zero-shot, without ever training on the dataset.
• Aggregation (HIC): ρ = 0.540 (p = 2.76 × 10⁻⁷) — a strong, highly significant correlation on the blinded panel. On a zero-shot basis, this would place around the top 10 of the competition's hydrophobicity leaderboard.

The significance is the context. The competition observed that models often performed well in cross-validation but less so on the blinded set, underscoring how difficult out-of-distribution generalization is. A result comparable to the top reported thermal-stability performance, without any training on the benchmark, is exactly the kind of generalization that is hard to achieve.
Why Might the Hybrid Approach Help?
While this benchmark does not establish causality, one possible explanation is that the physics-based structural validation step filters low-quality predicted conformations before downstream property prediction. This may improve robustness when evaluating antibodies outside the benchmark’s training distribution. Additional validation across independent datasets will be needed to confirm this hypothesis.
Key Value: Evidence of Generalization
Because these numbers were produced zero-shot on a blinded, independent panel, they provide evidence of generalization beyond the benchmark — not just how well the platform was tuned to a familiar dataset. That is the difference between a leaderboard score and a platform for practical discovery workflows.
Application Areas
Apotek Landscape is designed for developability assessment of protein-based therapeutics including antibodies, cytokines, enzymes, and hormonal proteins, providing a scalable solution for the most complex biologics in today's R&D portfolios.
Conclusion: Evidence Behind the Platform
Evidence of Generalization on an Independent Benchmark: A result comparable to the top reported thermal-stability performance, achieved with zero fine-tuning, shows the AI-physics hybrid generalizes to unseen antibodies — the property the benchmark showed is hardest to achieve.
Supports Early-Stage Prioritization: Statistically significant ranking performance is exactly what early-stage prioritization needs to concentrate experimental budget on the candidates most likely to succeed.
Honest Scope: These results come from a single blinded benchmark and reflect relative rankings, not absolute values; broader cross-dataset validation is ongoing. We report it this way on purpose — a reproducible, public baseline is a better foundation for trust than an over-claimed headline.
In a field where benchmark claims are rarely reproducible, we're publishing the dataset license, the method, and the honest limitations — because a platform you can trust starts with results you can verify.
About Apotek
Apotek integrates Causal AI and GenAI technologies to build an integrated platform that transforms multi-omics data into actionable insights — enabling novel target discovery, biosimulation, and generative drug design. By uncovering cause-and-effect relationships rather than spurious correlations, our approach enhances biological interpretability, improves model robustness, and accelerates decision-making across the drug discovery pipeline.
Data Attribution
Data: GDPa1 / GDPa3 antibody developability datasets, © Ginkgo Datapoints (Ginkgo Bioworks), available at https://datapoints.ginkgo.bio/dataset-access (Hugging Face: ginkgo-datapoints/GDPa1). Licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). No modifications were made to the underlying dataset.
References
1. PROPHET-Ab: A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training. bioRxiv 2025.05.01.651684.
2. Ginkgo Datapoints Antibody Developability Competition outcomes: limited model performance and a call for data standardization. mAbs. 2026. doi:10.1080/19420862.2026.2634216.


