Project · 2025 — present
AI Co-Scientist
An agentic AI system that runs bioinformatics analyses for you, learns your project and research preferences as it goes, and proactively finds the new datasets, papers and analyses worth pursuing.
The problem
A large share of computational biology is skilled but routine work: moving data between tools, running standard analyses, keeping track of jobs, and turning outputs into something a scientist can act on. It is where research tends to stall, in the gap between a raw dataset and a finished result. Meanwhile the field keeps moving, with new datasets and papers appearing faster than anyone can track.
What it does
The AI Co-Scientist, built at Academia Inc., takes on that work. Multiple AI agents, each with reasoning tuned to a specific area of bioinformatics, plan and carry out analyses end to end, and return insights rather than just files.
It also suggests where to go next. Given a research question, it identifies new datasets and analyses a scientist could use to investigate it, and can then run them.
When a task calls for a specialist model, the system reaches for one. For structural biology, it runs protein structure prediction through an integration with Chai-1, a state-of-the-art biomolecular structure model.
It learns as it goes
The Co-Scientist learns about each project and about the scientist behind it: the question they are chasing and how they prefer to work. It uses that context to get sharper over time, so its analyses and suggestions fit the research more closely the longer it is used.
It works in the background
The system is also passive by design: it doesn’t wait to be asked. It keeps watch for newly published papers and datasets relevant to each project, and brings them to the scientist.
How it’s built
The platform is organised around a Model Context Protocol (MCP) server, which gives the agents one modular, standard way to reach tools and models. That architecture is what lets it scale:
- Deterministic job orchestration. Every job moves through a defined life cycle, so each run behaves predictably.
- Enterprise-grade memory safety across long-running, multi-step workflows.
- Modularity. New tools and models plug in without reworking the core.
My role
I direct the scientific strategy and technical roadmap: which research workflows the system should take on, how its agents should reason about biology, and which external models to bring in. I work with the engineering team to turn that roadmap into a platform scientists can rely on.