AI Designs 700,000 Virus Genomes, 16 Survive Lab Tests
In a groundbreaking experiment, researchers from Stanford University and the Arc Institute have used an AI model to design nearly 700,000 viral genomes. Of those, they synthesized 285, and 16 turned out to be fully functional bacteriophages—viruses that infect and kill bacteria. The study, published in Science on August 6, marks a significant leap from designing individual proteins to creating entire viral genomes from scratch.
The AI model, called Evo, is a genomic language model trained on about 9 trillion nucleotides. For this project, the team fine-tuned it on roughly 15,000 small virus genomes similar to the bacteriophage Phi X-174, which they used as a template. Phi X-174 is a well-studied virus that only infects bacteria, making it a safe and convenient choice for proof-of-concept.
The process wasn't as simple as pressing a button. Evo generated the sequences, but they were just strings of DNA. The researchers then had to screen them, synthesize the most promising candidates, and test them in the lab. Most of the synthesized sequences did nothing—no reaction when exposed to E. coli. But on some plates, clear spots appeared, indicating that the bacteria had been lysed, or destroyed, by new viruses.
What's even more impressive is that the 16 surviving viruses weren't just barely functional. In competition experiments, some of them actually outperformed the natural Phi X-174, lysing bacteria faster and more efficiently. The team also mixed AI-designed bacteriophages through generations, and after just 1 to 5 rounds, the viruses broke through three resistance barriers of E. coli. This could have implications for phage therapy, a potential alternative to antibiotics for treating drug-resistant infections.
But with great power comes great responsibility. The idea of AI designing viruses raises obvious security questions. While these particular viruses are harmless to humans, the technology could potentially be misused. Researchers at the University of Oxford point out that these AI-generated viruses are still very similar to natural ones, relying on the same biological mechanisms. An independent analysis in June found that Evo tends to favor sequences close to natural viruses, essentially improving search efficiency within the range of natural variation.
The research team took precautions: they removed viruses that infect humans and their close relatives from the training data, and they used non-pathogenic E. coli in the lab. They claim this version of Evo cannot generate human virus sequences. But what if someone else trains a similar model on different data? The answer is unknown. Researchers at Johns Hopkins University have expressed concerns that future genome models might be used to enhance the transmissibility or lethality of viruses like influenza.
In late July, the U.S. government announced a policy on high-risk life science research, prohibiting federally funded experiments that make biological agents more dangerous. However, there's no consensus yet on how reviewers can assess the risks of an AI-generated genome that has never appeared in nature. The path from a digital sequence to a dangerous pathogen involves multiple steps—synthesis, assembly, host adaptation—but this research serves as a proactive alarm for biosecurity governance.
Key Points
- AI model Evo generated 700,000 viral genome sequences; 16 became functional bacteriophages.
- The study demonstrates a leap from designing proteins to entire viral genomes.
- Some AI-designed viruses outperformed natural ones in lab tests.
- Security concerns remain, as future models could be misused to create dangerous pathogens.
- The U.S. government has issued new policies on high-risk life science research, but gaps in oversight persist.