
Prefer to listen instead? Here’s the podcast version of this article.
Artificial intelligence has written code, generated images, predicted protein structures and helped researchers sift through mountains of scientific data. Now, a team of AI “scientists” has taken on something considerably more ambitious: helping identify a promising therapeutic strategy for lung cancer.
The headline sounds almost science-fictional. A virtual biotechnology company staffed by tens of thousands of AI agents examined clinical evidence, investigated biological targets and proposed a treatment approach involving the protein B7-H3, also known as CD276.
But the real story is more interesting than “AI invented a cancer drug.”
According to [Nature] Stanford researchers developed a system called Virtual Biotech, capable of coordinating as many as 37,000 specialized AI agents under an AI “chief scientific officer.” The system was designed to imitate the organizational structure of a real biotechnology company, with different agents handling areas such as target discovery, genomics, clinical-trial analysis and therapeutic development.
What emerged is an important glimpse of where AI-powered drug discovery may be headed: not toward one giant chatbot replacing scientists, but toward networks of specialized AI agents working alongside human researchers.
Drug discovery is not one problem. It is hundreds of interconnected problems.
Researchers need to understand disease biology, identify promising targets, examine genetic evidence, study cellular behavior, assess toxicity, select an appropriate drug design, interpret previous clinical trials and eventually determine how a new candidate should be tested.
That fragmentation is exactly what the Stanford-led team wanted to address.
The researchers describe the architecture in their [Virtual Biotech research platform] as an organization of specialized AI scientists coordinated by a virtual Chief Scientific Officer, or CSO. Instead of asking a single language model one enormous question, the CSO breaks a scientific problem into smaller assignments and delegates them to specialist agents equipped with appropriate datasets and analytical tools.
Think of it less like ChatGPT answering a biology question and more like a digital research organization holding a very, very crowded staff meeting.
The peer-reviewed research, published in Science, involved Harrison G. Zhang, Peter Eckmann, Jiacheng Miao, Andrew B. Mahon and James Zou. The [PubMed record for the Science paper] describes Virtual Biotech as a multi-agent framework designed to support therapeutic discovery and development across multiple scientific disciplines.
Before tackling lung cancer, the researchers gave the system a monster-sized homework assignment.
Virtual Biotech analyzed outcomes from 55,984 clinical trials. More than 37,000 clinical-trialist agents participated in the analysis, with agents extracting structured information and connecting trial outcomes with genomic and cellular evidence.
The scale matters because pharmaceutical research generates a staggering volume of fragmented evidence. Clinical-trial registries, scientific papers, genomic databases and other research repositories often contain information that could affect drug-development decisions, but connecting those dots manually is slow.
According to [Stanford Medicine’s account of the project] the agents catalogued roughly 50,000 trials in less than a week—a job James Zou said could otherwise take human researchers years.
The agents then found an intriguing pattern.
Drugs aimed at genes whose activity was highly specific to particular cell types appeared to perform better than drugs with less specific targets. In the research dataset, drugs targeting these cell-type-specific genes were 40% more likely to progress from Phase I to Phase II, 48% more likely to reach market, and associated with 32% fewer adverse events.
Those percentages should be interpreted carefully. They represent associations found in historical data, not proof that targeting cell-specific genes automatically causes a drug to succeed.
But as a drug-development signal, the finding is valuable. If researchers can identify biological characteristics that correlate with better clinical outcomes, AI systems could help scientists prioritize promising targets before enormous amounts of time and money are committed.
The researchers next asked Virtual Biotech to investigate B7-H3, a protein also known as CD276, as a possible therapeutic target in lung cancer.
B7-H3 was not an unknown protein magically discovered by the AI. Researchers were already studying it as an oncology target.
That distinction is important.
What Virtual Biotech did was assemble and analyze evidence from multiple sources to assess why the target might be useful and what therapeutic approach could make sense.
According to Stanford Medicine, the agents found evidence that B7-H3 was highly expressed in fibroblasts—connective-tissue cells that can exist around tumors. Further analyses suggested that B7-H3-expressing fibroblasts could participate in signaling interactions that suppress nearby immune cells, potentially helping cancer hide from immune defenses.
The AI agents then recommended an antibody-drug conjugate, or ADC.
If you are not spending your evenings reading oncology journals—and we will forgive you—an ADC can be thought of as a precision delivery vehicle. An antibody recognizes a protein found on particular cells, while a potent cancer-killing payload is attached to it. The objective is to bring that payload more directly to cells expressing the target instead of distributing chemotherapy as broadly throughout the body.
The Virtual Biotech agents proposed a B7-H3-directed ADC strategy. [Phys. org’s coverage of the study] offers another useful overview of how the agents arrived at the treatment approach and how researchers plan to test additional AI-generated hypotheses in physical laboratories.
There is one rather stubborn problem with computational drug discovery: biology does not care how impressive your demo looks.
A hypothesis eventually needs to survive contact with the physical world.
Cells must be tested. Experiments need to be reproduced. Toxicity has to be evaluated. Animal studies may be required. Manufacturing processes need development. Human clinical trials must establish safety and efficacy. Regulators then have to evaluate the evidence.
Nature’s coverage emphasizes this limitation: outside researchers noted that Virtual Biotech itself has not yet been validated through the complete real-world drug-development process and that many of its predictions still require experimental testing.
The researchers make essentially the same point. The [bioRxiv version of the Virtual Biotech research] presents the system as a framework for generating and evaluating therapeutic hypotheses—not a substitute for experimental validation.
The more autonomy we give scientific AI systems, the more important accountability becomes.
What happens when an AI agent cites a paper incorrectly? How should a pharmaceutical company audit a conclusion generated by thousands of interacting agents? Who is responsible for catching subtle biological assumptions? How do regulators evaluate research where important decisions were partly generated by AI?
These are not reasons to abandon the technology. They are reasons to build validation, transparency and human review into the technology from the beginning.
Agentic scientific systems will need traceable evidence, reproducible analyses, strong data governance and clear boundaries around which decisions require human authorization.
This is particularly important in healthcare, where “the model sounded confident” is not a scientific validation strategy.
The most fascinating thing about Virtual Biotech is not the number 37,000.
After all, spinning up thousands of software agents is much easier than hiring thousands of scientists.
The important development is coordination.
For years, AI systems have excelled at narrowly defined scientific tasks: predicting structures, classifying images, searching chemical space or summarizing papers. Virtual Biotech points toward systems that can combine many of those capabilities into a coordinated research process.
That could change the economics of early-stage drug development.
Instead of researchers manually spending months navigating fragmented datasets and literature, AI agents could continuously evaluate evidence and surface promising hypotheses for experts to investigate. Human scientists could spend more time designing experiments, challenging assumptions and making high-level scientific judgments.
The result would not be “AI replacing scientists.”
It could be scientists suddenly gaining access to a research organization that can expand or shrink on demand.
And that may be the real breakthrough.
The idea of thousands of AI agents working together to uncover promising cancer treatments may sound futuristic, but it is quickly becoming part of real scientific research. Stanford’s Virtual Biotech shows what can happen when AI moves beyond answering questions and starts coordinating complex research tasks across clinical trials, biology, and drug development.
This does not mean AI has replaced scientists—or that it has independently cured lung cancer. What it does show is that AI can help researchers process enormous amounts of information faster, connect evidence that might otherwise be missed, and generate stronger hypotheses for human experts to test.
The biggest opportunity may be collaboration. As AI systems become more specialized, transparent, and connected to real laboratory workflows, scientists could gain access to powerful digital research teams that accelerate discovery without removing the need for human judgment, validation, and oversight.
For biotech, pharmaceutical companies, and healthcare innovators, this is a glimpse of what comes next: AI not just as a tool, but as a research partner helping turn massive datasets into meaningful scientific progress.
WEBINAR