NVIDIA and Partners Release Open Viral Protein Dataset to Aid Pandemic Readiness

NVIDIA has teamed with Google DeepMind and EMBL-EBI to publish predicted 3D structures for protein complexes from more than 2,800 viruses through the AlphaFold Database. The dataset was generated using AlphaFold2 with optimization from NVIDIA BioNeMo Inference Runtime. NVIDIA also released its BioNeMo Structure Prediction Pipeline so researchers can generate structures for their own targets.
NVIDIA worked with Google DeepMind and EMBL-EBI to add predicted structures for protein complexes from over 2,800 viruses to the AlphaFold Database. The predictions used AlphaFold2, accelerated by NVIDIA's BioNeMo Inference Runtime, enabling analysis across many viral proteomes. NVIDIA also shared the BioNeMo Structure Prediction Pipeline, letting researchers predict structures for their own targets.
Roughly 30% of newly added protein interactions were previously unknown, with shapes absent from the Protein Data Bank. The project aims to build structural knowledge before future outbreaks, since most proteins act in multi-molecule complexes that vaccines or drugs may need to target. COVID-19 benefited from decades of coronavirus research; this effort seeks to reduce reliance on such prior knowledge.
The open dataset could help researchers, public health agencies, and drug developers investigate viral targets faster, potentially shortening early response times during future outbreaks. It may also lower barriers for labs with fewer resources by offering ready-made structural predictions. However, predicted structures are hypotheses, not experimental proof, so their societal benefit may depend on validation and how widely the tools are adopted. Patients and communities could gain if vaccine or treatment discovery accelerates, but no outcome is guaranteed.