AstraZeneca’s Puja Sapra says AI is shortening biologics cycle times with fewer dead ends
The build-measure-learn loop, proprietary multimodal data, and a Kendall Square “lab of the future” are reshaping drug R&D.

Puja Sapra, AstraZeneca’s senior vice president and head of R&D biologics engineering and oncology targeted discovery, says AI now computationally enhances design, make, test, and analyze to speed biologics R&D. For decision-makers, it signals a shift toward closed-loop, data-driven discovery and more urgent bets on safety prediction and evaluation.
AstraZeneca is trying to do something that sounds simple and is brutally hard: compress the time between designing a biologic and learning whether it works. Puja Sapra, the company’s senior vice president and head of R&D biologics engineering and oncology targeted discovery, says the result is real-world momentum. “Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” Sapra says. Her punchline: “The cycle times are getting shorter while productivity and innovation increase.”
That statement matters because biologics development already has a high cost and a high failure rate. A new medicine can take many years to develop, requires significant investment, and even then most candidates never reach patients. For biologics, the complexity is even greater than synthetic-chemistry drugs: instead of traditional chemistry, biologic therapies are engineered proteins. Scientists must sift through vast quantities of possible molecules, looking for the rare few that will bind to the right target, stay stable in the human body, and be manufacturable at scale. AI is increasingly being used to narrow and refine the options for testing, because no human team can systematically explore the full combinatorial universe.
The core of AstraZeneca’s approach is what Sapra calls a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources on the top-ranked candidates instead of wasting effort on low-probability leads. In practical terms, that means fewer dead ends and faster iteration, because every experimental outcome feeds back into the next cycle. It also changes what targets teams consider realistic, since the tighter feedback loop can help pursue disease areas previously viewed as “undruggable” by conventional medicine.
But the next phase goes beyond speeding up today’s workflow. Sapra points to AI as a way to discover entirely new classes of medicines, specifically next-generation biologics that can hit multiple targets simultaneously or deliver therapeutic payloads precisely to specific cells. That kind of multi-specific optimization is not a single lever problem. It requires balancing many variables at once, such as potency, stability, manufacturability, and safety. Sapra explains that future AI-driven models could help identify which two or three targets to prioritize based on underlying biology, then optimize across multiple parameters to strike that balance.
All of this runs into an unavoidable constraint: AI models only perform as well as their training data. Sapra is explicit that in drug discovery, data quality and coverage are the differentiator. McKinsey estimates that generative AI combined with other computational tools could cut drug discovery timelines by as much as 50%, but Sapra’s point is that the work cannot outpace the data foundation. AstraZeneca’s datasets are proprietary and multimodal, including molecular structures, binding measurements, safety profiles, and manufacturing outcomes. The company says it has built an intentionally diverse portfolio across multiple disease areas and drug types so it can fine-tune frontier AI models with richer, more representative training sets. It has also invested in deep screening technologies to generate additional datasets “in volume” to constantly refine and validate models.
Then comes the infrastructure bet. AstraZeneca is building a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts, where AI and robotic automation can run a continuous, closed-loop discovery system. Sapra compares it to a self-driving car: a car uses sensors and models to navigate; this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data. The data then feeds directly back into the models to accelerate each subsequent cycle. Sapra says automated high-throughput systems could eventually make and evaluate thousands of molecular interactions on a weekly basis, generating AI-ready data at a scale that traditional workflows cannot match. Robotic sample handling, automated quality checks, and integrated data pipelines are meant to accelerate early drug development timelines.
The strategic endpoint is “de novo” design, where AI generates entirely new protein sequences that fit desired drug properties. Sapra describes designing structure, predicting safety, forecasting behavior in the body, and ensuring manufacturability. The field is making progress toward the idea of a completely AI-generated biologic, designed from scratch all the way to a clinical candidate, she says, adding that it is “a matter of time.” Getting there requires three ingredients: richer and more standardized training data across the industry, robust evaluation benchmarks for AI-generated candidates, and teams that can operate at the intersection of machine learning and biology. And if there is one constraint that could make or break de novo, Sapra says it is safety prediction, calling it one of the hardest problems in de novo design and perhaps the least discussed.
AstraZeneca’s response is to run what amounts to “virtual clinical trials.” These are advanced cell systems and micro-scale organ models that function as physical testbeds, paired with AI that learns from their outputs. Sapra argues these systems can generate enhanced biological signals without traditional testing bottlenecks, helping close the loop between AI-generated designs and clinical-ready candidates. She also notes a shift toward agentic AI systems that can simultaneously generate molecule candidates and predict likely efficacy and safety. The aim is to connect disease-level insights directly to molecule design, bridging data silos that used to be separate. Throughout, Sapra insists scientists stay central for oversight, ensuring outputs are explainable, ethical, and directed toward potential patient benefit.
For executives watching from the sidelines, this is not just a science story. It is an operating model story. When cycle times shrink and data becomes the moat, the winners are the organizations that can sustain high-quality multimodal datasets, build closed-loop infrastructure, and develop trustworthy safety evaluation. The board-level question is whether your organization is investing in the unglamorous parts that make AI useful: data standards, evaluation benchmarks, and safety signal generation. AstraZeneca is betting that biologics R&D is moving from artful iteration to continuously optimized computation, but it also makes clear the next frontier is safety and validation. That is the lever that determines whether faster learning turns into faster patients.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Science
CERN’s 2012 Higgs discovery may not be the finale, ML narrows the hunt
New machine-learning work at CERN sharpens searches for additional Higgs-family particles beyond the one found in 2012.

Testicular tissue grown from a transplant produced sperm in an infertile cancer survivor
A successful, more humane path for fertility restoration emerges after a long history of messy testicular transplant attempts.

Caliskan’s 55% silver “missing” is back in the sun, after July 2026 modeling
New simulations with non-equilibrium physics suggest the Sun actually contains 55% more silver than measured.

