Can artificial intelligence build a virtual version of a living cell—and eventually predict how the human body will respond to disease, drugs, or other interventions?
For decades, scientists have dreamed of creating a “virtual cell”: a computational model capable of representing the complex biological processes that occur inside living cells and predicting how those processes change when the cell is exposed to a drug, genetic modification, environmental change, or disease.
That vision is now receiving renewed attention as advances in artificial intelligence (AI), single-cell biology, genomics, medical imaging, spatial biology, and multimodal data integration begin to converge.
In September 2026, researchers at the University of Oxford’s Institute of Biomedical Engineering published a roadmap exploring how multimodal AI could help move virtual cells from largely correlational models toward systems that can make experimentally testable predictions. The review, Multimodal AI for predictive and programmable virtual cells, highlights the importance of combining different biological data types rather than relying on a single source of information.
The implications could be significant. Instead of studying biological systems one dataset at a time, future AI systems could potentially connect information from genomics, single-cell measurements, imaging, spatial data, electronic health records (EHRs), and wearable devices to build increasingly comprehensive models of health and disease.
But creating such a system is far more difficult than simply training a larger AI model.
What Is a Virtual Cell?
A virtual cell is an in-silico computational representation of a biological cell designed to model its molecular functions, interactions, processes, and behavior.
The basic idea is straightforward:
If researchers can represent enough of the biological information inside a cell, AI and computational models could potentially predict what happens when that cell changes.
For example, a virtual cell might eventually be used to investigate questions such as:
- How will a cancer cell respond to a particular drug?
- What happens when a specific gene is switched off?
- How might a genetic mutation change cellular behavior?
- Which biological pathway is responsible for a disease phenotype?
- What could happen if a therapeutic molecule interacts with a particular target?
- How might cells change under different environmental conditions?
Traditional biological research generally answers these questions through experiments. A virtual cell aims to complement that process by allowing researchers to simulate and prioritize experiments computationally before testing them in the laboratory.
Importantly, today’s virtual-cell systems are not digital copies of living cells. Researchers are still working toward models that are sufficiently accurate, interpretable, generalizable, and experimentally validated.
A 2026 Nature Reviews Genetics perspective describes virtual cells as in-silico models designed to simulate molecular functions and ultimately cellular behavior, while noting the field’s shift from traditional mechanistic models toward data-driven AI approaches as large-scale single-cell datasets have expanded.
Why Multimodal AI Matters
Biology is inherently multimodal.
A cell does not operate according to its DNA sequence alone. Its behavior emerges from interactions between genes, RNA, proteins, metabolites, cellular structures, neighboring cells, tissues, environmental signals, and other factors.
This creates a fundamental problem for AI.
A model trained on only one type of data may learn useful patterns, but it can miss information contained in other biological modalities.
Multimodal AI attempts to solve this problem by learning from multiple forms of information simultaneously.
Instead of asking:
“What does the genome tell us?”
a multimodal system could ask:
“What does the genome tell us, how is that information being expressed, what does the cell look like, where is it located, what has happened to it previously, and how is the wider biological system responding?”
This shift from isolated datasets to integrated biological representations is one of the central ideas behind the emerging virtual-cell roadmap.
Oxford researchers argue that progress will require richer multimodal datasets, causal perturbation information, stronger biological understanding, and continuous experimental validation—not simply larger AI models.
The Biological Data That Could Power a Virtual Cell
1. Genomics
Genomics provides information about an organism’s DNA.
Genetic sequences contain instructions that influence cellular processes, but interpreting those sequences is extremely complex. Two cells can contain essentially the same genome while behaving very differently because of differences in gene expression, epigenetic regulation, cellular environment, and other factors.
AI models can analyze large genomic datasets and identify patterns associated with:
- Genetic variants
- Gene regulation
- Disease-associated mutations
- Regulatory elements
- Gene expression
- Cellular states
Recent advances in multimodal genomic foundation models are beginning to combine information from genomic sequences, single-cell transcriptomics, and spatial datasets.
For virtual cells, genomics could provide the underlying biological blueprint.
But a blueprint alone is not enough.
2. Transcriptomics and Single-Cell Data
Genes are not simply “on” or “off.”
Cells continuously regulate which genes are expressed and at what levels. Single-cell RNA sequencing allows researchers to examine gene-expression patterns at the level of individual cells.
This is particularly important because cells within the same tissue can exist in very different states.
For example, a tumor may contain:
- Cancer cells
- Immune cells
- Stromal cells
- Blood-vessel-associated cells
- Different cancer subpopulations
A virtual-cell system could use single-cell datasets to understand these differences and model how cellular states change.
The rapid expansion of single-cell atlases has become one of the major foundations for data-driven virtual-cell research.
3. Proteomics and Molecular Information
Genes provide instructions, but proteins perform many of the functions required for cellular life.
Proteomics can therefore provide another layer of biological information.
A future virtual-cell model could potentially connect:
DNA → RNA → proteins → cellular processes
This would help AI models move beyond simple associations and toward richer representations of biological mechanisms.
However, protein interactions are highly complex. Protein abundance, location, structure, chemical modifications, and interactions with other molecules can all influence cellular behavior.
That makes proteomic integration another major technical challenge.
4. Medical and Cellular Imaging
Biology is not only about molecular sequences.
Cells have shapes, structures, locations, and spatial relationships.
Microscopy and medical imaging can capture information that sequence-based datasets cannot.
For example, imaging can reveal:
- Cell morphology
- Tissue organization
- Subcellular structures
- Cellular interactions
- Disease-related structural changes
- Spatial relationships between different cell types
Oxford researchers are also working on multimodal imaging approaches that combine three-dimensional tissue imaging, gene-expression information, and AI to better understand disease.
For virtual cells, imaging could provide the physical context surrounding molecular information.
5. Spatial Biology
One of the most important developments in modern biology is the ability to study where biological activity occurs.
Two cells can have similar molecular characteristics but behave differently because they occupy different locations within a tissue.
Spatial transcriptomics and related technologies allow researchers to connect molecular information with physical location.
This could help future virtual-cell models answer questions such as:
- Where are particular cell types located?
- Which cells are interacting?
- Where are disease-associated signals concentrated?
- How does gene expression change across tissue regions?
Adding spatial information could therefore transform a virtual cell from a collection of molecular measurements into a model embedded within a biological environment.
6. Electronic Health Records
The virtual-cell concept becomes even more ambitious when cellular models are connected with patient-level information.
Electronic health records can contain longitudinal information such as:
- Diagnoses
- Laboratory results
- Medication history
- Clinical measurements
- Procedures
- Disease progression
- Treatment responses
EHRs operate at a very different scale from molecular biology.
This creates a major challenge—but also an opportunity.
A future multimodal AI architecture could potentially connect information across several biological scales:
Molecules → Cells → Tissues → Organs → Patients
This is similar to the broader vision described in recent work on AI-driven digital organisms, which proposes integrated models spanning biological scales from molecules and cells to individuals.
7. Wearable and Continuous Health Data
Wearable devices introduce another dimension: time.
Unlike a genomic test that may be performed once, wearable devices can continuously collect physiological signals.
Depending on the device, these may include information related to:
- Heart rate
- Activity
- Sleep
- Movement
- Temperature
- Other physiological measurements
The combination of molecular, clinical, imaging, and longitudinal physiological information could eventually help AI models understand not just what a biological system looks like, but how it changes over time.
However, wearable data are noisy and highly variable. Different devices, populations, measurement conditions, and user behaviors can affect data quality.
Therefore, integrating wearable information into biological foundation models will require careful validation.

From Separate Datasets to a Unified Biological Model
The real challenge is not simply collecting more data.
It is connecting the data correctly.
Imagine a simplified biological information chain:
Genome
↓
Gene expression
↓
Protein activity
↓
Cellular state
↓
Tissue organization
↓
Organ function
↓
Clinical state
↓
Real-world physiological signals
A sophisticated multimodal AI system could attempt to learn relationships across these levels.
This is one reason researchers increasingly describe future systems as multiscale foundation models rather than conventional predictive algorithms.
The objective is not merely to predict one outcome. It is to create reusable representations that can support many biological tasks.
How Could a Virtual Cell Actually Work?
A future virtual-cell platform could operate as a continuous computational-experimental loop.
Step 1: Collect Biological Data
Researchers gather information from genomics, transcriptomics, proteomics, imaging, spatial biology, perturbation experiments, and other sources.
Step 2: Harmonize the Data
Different datasets have different formats, scales, resolutions, and levels of noise.
AI systems must therefore learn how to align these sources into a common representation.
Step 3: Build a Biological Representation
A multimodal model learns representations of biological states.
For example, it could learn that a particular combination of gene expression, protein activity, morphology, and spatial context corresponds to a specific cellular state.
Step 4: Introduce Perturbations
This is a critical stage.
Instead of simply asking the model to recognize patterns, researchers can ask:
“What happens if we change something?”
Examples include:
- Altering gene activity
- Introducing a genetic perturbation
- Applying a drug
- Changing environmental conditions
Perturbation data are particularly valuable because they can help models learn relationships that are closer to causal biology rather than simple correlations.
Step 5: Generate a Prediction
The model predicts how the biological system might respond.
Step 6: Test the Prediction Experimentally
Researchers then perform laboratory experiments.
If the prediction is incorrect, the experimental result can be fed back into the model.
This creates a closed learning loop:
AI prediction → Laboratory experiment → New data → Model improvement
Reviews of AI-driven virtual cells emphasize this computational-to-experimental validation cycle as an important pathway toward practical applications.
Why Causality Is More Important Than Correlation
This may be one of the biggest challenges facing virtual-cell research.
AI can become extremely good at recognizing correlations.
But biology often requires understanding cause and effect.
Suppose an AI model discovers that two biological signals frequently appear together in cancer cells.
That does not necessarily mean one causes the other.
A useful virtual cell needs to answer a more difficult question:
“What will happen if I change this biological variable?”
This is why perturbation experiments are so important.
The Oxford roadmap specifically emphasizes the need to move beyond largely correlational models toward systems grounded in experimental evidence and capable of predicting biological responses to interventions.
Virtual Cells and Drug Discovery
One of the most promising applications is pharmaceutical research.
Today, drug development is expensive and time-consuming. Researchers must identify promising targets, design molecules, evaluate biological activity, test safety, and eventually conduct clinical studies.
Virtual-cell models could potentially help earlier in this process.
For example, researchers might use an AI model to simulate how different cellular states respond to candidate interventions.
Potential applications include:
- Drug-response prediction
- Target identification
- Combination therapy research
- Disease modeling
- Biomarker discovery
- Personalized treatment research
Current research suggests that AI-driven virtual-cell models could support drug-response and gene-perturbation prediction, although significant validation and translation challenges remain.
The goal is not to eliminate laboratory science.
It is to make laboratory science more targeted and informative.
Virtual Cells and Cancer Research
Cancer is particularly suitable for multimodal biological modeling because tumors are highly heterogeneous.
A single tumor can contain multiple cellular populations with different genetic and molecular characteristics.
A future virtual-cell framework could potentially combine:
- Tumor genomics
- Single-cell sequencing
- Spatial information
- Histopathology images
- Treatment history
- Clinical records
The resulting model could help researchers understand how tumor populations respond to different interventions.
In the long term, researchers hope such approaches could contribute to more personalized cancer research.
However, predictions would still need to be validated experimentally and clinically before being used to guide patient care.
Virtual Cells and Gene Editing
Gene-editing technologies create another important use case.
Suppose researchers want to understand what happens when a particular gene is modified.
A virtual-cell model could potentially predict changes in:
Gene activity → cellular pathways → cellular state → phenotype
The model could then identify the most informative experiments to perform.
This could reduce the number of unnecessary experiments and help researchers prioritize promising hypotheses.
But again, prediction does not replace experimental confirmation.
From Virtual Cells to Digital Organisms
The virtual-cell concept may eventually become part of something even larger.
Instead of modeling only a cell, researchers are beginning to explore multiscale AI systems capable of representing biological processes from molecules through cells, tissues, organs, and individuals.
A recent Nature Medicine perspective describes an AI-driven digital organism as a modular system of integrated foundation models designed to represent biological scales and their connections.
The long-term vision could therefore look something like:
Molecular AI
↓
Virtual Cell
↓
Virtual Tissue
↓
Virtual Organ
↓
Digital Organism
Such a system would be enormously ambitious.
It would require models to understand biological processes across different scales and timeframes while preserving meaningful connections between them.
The Biggest Challenges
Despite the excitement surrounding virtual cells, researchers face major obstacles.
1. Biology Is Extremely Complex
Living systems contain enormous numbers of interacting components.
Even a highly detailed dataset represents only part of the underlying biology.
A model can therefore be powerful without being complete.
2. Data Quality and Standardization
Multimodal datasets often come from different laboratories, instruments, populations, and experimental conditions.
This creates problems with:
- Data compatibility
- Missing information
- Measurement noise
- Batch effects
- Different experimental protocols
A model can only be as reliable as the data used to train and validate it.
3. Limited Perturbation Data
Observational data are abundant compared with carefully controlled perturbation experiments.
But perturbations are essential for understanding causal relationships.
Generating high-quality perturbation datasets at scale remains difficult.
4. Biological Interpretability
A model may produce an accurate prediction without clearly explaining the biological mechanism behind it.
For researchers, that can be a serious limitation.
A useful virtual cell should ideally provide scientifically meaningful explanations rather than functioning entirely as a black box.
5. Generalization
A model trained using one cell type, disease, laboratory, or population may not perform equally well elsewhere.
Researchers therefore need robust validation across different biological contexts.
6. Computational Requirements
Multimodal biological models can require substantial computational resources.
As models grow larger and datasets become more complex, researchers must find ways to make these systems efficient and scalable.
7. Privacy and Ethics
Connecting molecular data with EHRs and wearable information introduces significant privacy concerns.
Health information is highly sensitive.
Future systems will therefore require strong safeguards for:
- Data security
- Consent
- Privacy
- Responsible AI
- Clinical accountability
These concerns become increasingly important as biological AI moves closer to patient-level applications.
The Oxford Roadmap: What Comes Next?
The recent Oxford review provides an important perspective on where the field is heading.
Rather than suggesting that simply building larger AI models will solve the virtual-cell problem, the researchers emphasize several interconnected requirements:
Richer Multimodal Data
AI systems need access to complementary biological measurements.
Causal Perturbation Data
Models need information about what happens when biological systems are intentionally changed.
Better Biological Understanding
AI must be connected to established biological knowledge rather than relying solely on statistical patterns.
Continuous Experimental Validation
Predictions must repeatedly be tested against real experiments.
Programmability
The ultimate goal is not merely to describe biological states but eventually to predict and potentially help engineer cellular behavior.
Oxford researchers describe this transition as a move from models that primarily observe biology toward systems that can help researchers understand, predict, and rationally engineer biological behavior.

A Possible Roadmap for AI-Powered Virtual Cells
The development of virtual cells is unlikely to happen through one giant breakthrough.
Instead, it may progress through several stages.
Stage 1: Better Biological Representations
AI models become increasingly capable of representing individual biological modalities such as genomic sequences, single-cell data, proteins, and images.
Stage 2: Multimodal Integration
Different data types are connected into unified representations.
Stage 3: Perturbation Prediction
Models begin predicting how biological states change following interventions.
Stage 4: Experimental Validation
Predictions are systematically tested in laboratories.
Stage 5: Closed-Loop Biological AI
AI proposes experiments, experiments generate new data, and those results improve the model.
Stage 6: Multiscale Biological Models
Virtual cells become connected to tissue-, organ-, and organism-level models.
This represents a gradual transition from AI that analyzes biological data to AI that can simulate aspects of biological systems.
What Could This Mean for the Future of Medicine?
If virtual-cell technology becomes sufficiently reliable, it could change how biomedical research is conducted.
Instead of starting every investigation with a large number of physical experiments, researchers could first explore hypotheses computationally.
A future workflow might look like:
Patient or biological data
→ Multimodal AI
→ Virtual biological model
→ Predicted intervention
→ Targeted laboratory experiment
→ Experimental result
→ Updated AI model
This could help researchers prioritize promising experiments and identify unexpected biological relationships.
However, this should be viewed as a research vision rather than an established clinical reality.
Today’s models remain limited, and biological predictions still require rigorous experimental and clinical validation.
Will Virtual Cells Replace Laboratory Experiments?
Probably not.
In fact, the emerging research direction suggests almost the opposite.
Virtual cells are most powerful when they work together with laboratory science.
AI can search enormous spaces of possibilities much faster than humans can experimentally test them.
Laboratory experiments can then determine whether those computational predictions are actually correct.
The two approaches can form a feedback loop:
Computation identifies possibilities.
Experiments establish reality.
New experimental data improve computation.
This combination could become one of the defining approaches of AI-driven biological research.
The Bigger Picture: From Prediction to Programmable Biology
The most ambitious vision for virtual cells goes beyond prediction.
Researchers ultimately want AI systems that can help explain biological mechanisms and predict how biological systems respond to interventions.
That could move biomedical AI from:
Observe → Predict
toward:
Observe → Understand → Predict → Test → Design
This is why virtual cells are attracting attention from researchers across computational biology, AI, genomics, medicine, and biotechnology.
The field is still developing, but the direction is becoming clearer.
The future of biological AI may not be a single model trained on a single dataset.
It may instead be a network of connected models and data modalities, operating across biological scales and continuously tested against experimental evidence.
Conclusion
The idea of a virtual cell once sounded like science fiction.
Today, advances in multimodal AI, single-cell sequencing, spatial biology, imaging, genomics, perturbation experiments, and foundation models are making the concept increasingly realistic—although a truly comprehensive virtual cell remains a major scientific challenge.
The recent Oxford roadmap highlights an important principle: progress will depend on more than larger AI systems. Researchers need richer multimodal data, causal perturbations, biological knowledge, and continuous experimental validation.
If these pieces can eventually be connected, virtual cells could become powerful computational laboratories for exploring biology.
The long-term vision is even larger: models that connect molecules, cells, tissues, organs, and individuals into integrated representations of biological systems.
The ultimate goal is not to replace biology with AI. It is to give scientists a new way to understand, predict, test, and potentially engineer biology.
And that could make the virtual cell one of the most important frontiers in biomedical AI.
Scientific References
- University of Oxford — Institute of Biomedical Engineering. “Researchers publish roadmap for AI-powered virtual cells.” September 2026.
- Wu, A. R. “Revisiting the blueprint for an interpretable virtual cell.” Nature Reviews Genetics, 2026.
- Song, L., Segal, E., & Xing, E. “How to build an AI-driven digital organism.” Nature Medicine, 2026.
- Ma, C. et al. “AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential.” npj Digital Medicine, 2026.
- Nature Biotechnology. “Generalist biological artificial intelligence in modeling the language of life.” 2026.
- Nature. “‘Virtual cells’ aim to turn raw data into predictive models of biology.” 2026.
- Nature Methods. “Multimodal foundation transformer models for multiscale genomics.” 2026.