AlphaFold database opens predicted protein complexes for 2,800 viruses
EMBL-EBI says about 30% of the protein interactions added have never been documented, and that none of them replace experimental work.
The AlphaFold Protein Structure Database took in predicted structures for the protein complexes of more than 2,800 viruses on Thursday, free to download, EMBL-EBI said. Google DeepMind's AlphaFold2 produced the set, running on NVIDIA's BioNeMo Inference Runtime, and it prioritises viral families known to infect humans.
About 30% of the protein interactions in the release are completely new to science, showing interaction shapes that have never been documented, NVIDIA said in its announcement. The figure is the project's own. The database now holds more than 260 million protein predictions in total, according to EMBL-EBI.
The coverage runs from Picornaviridae, the family behind the common cold, to Mpox. EMBL-EBI said the structures sit behind a Pandemic Preparedness Portal on the database's homepage. Eight institutions contributed: EMBL-EBI, Google DeepMind, NVIDIA, the Coalition for Epidemic Preparedness Innovations, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.
EMBL-EBI was explicit about what the data is not. The structures show what viral proteins may look like and how they might interact, it said. They do not predict the impact of genetic variation. They cannot say how a change makes a virus more deadly or more transmissible, and confirming any of it may require experimental work.
NVIDIA said its optimised AlphaFold2 pipeline predicts a structure in minutes, against crystallographic methods that can take years and cost thousands of dollars per structure. That comparison is the company's own, and no outside group has published a check of it. NVIDIA publishes the BioNeMo structure prediction pipeline on GitHub.