Ἐπιστήμη · Knowledge
Morbius
An AI Co-Scientist for Autonomous Discovery.
Morbius doesn't just summarize your literature or suggest an idea. It ingests evidence, builds lasting research memory, reasons over it, and autonomously generates and tests hypotheses — closing the loop from raw literature to verified, citable output.
› query
10
Research Agents
210M+
Papers Indexed
<2s
Ingestion
How Morbius Discovers
Discovery is not a feature set. It is a loop — and Morbius runs every stage of it as one continuous process, carrying the output of each stage forward as the input to the next.
- 01 · Ingest01Ingest
Evidence in, structure out.
Morbius takes in PDFs, arXiv preprints, journal articles, technical reports, lecture recordings, and internal lab documents, then decomposes them. Sections, tables, figures, equations, citations, and individual claims are extracted and typed, so every later stage reasons over structured scientific evidence rather than undifferentiated text.
PDFarXivJournalsReportsLecturesInternal docs - 02 · Remember02Remember
Research memory that compounds instead of resetting.
Each ingested source is written into a persistent graph linking authors, concepts, methods, datasets, and results. This is the Co-Scientist's long-term memory: evidence read months ago remains available and, more importantly, remains connected. Context accumulates across sessions rather than being rebuilt from scratch in every conversation.
AuthorsConceptsMethodsDatasetsResultsPersistent - 03 · Reason03Reason
Grounded inference — the substrate, not the destination.
Questions resolve against the graph, and every claim returns with its provenance: the source paragraph, the figure it came from, the citation, a confidence score. This stage exists to make the next one trustworthy. Hypotheses are only worth testing if the evidence they rest on can be inspected and challenged.
Source paragraphFigureCitationConfidence scoreProvenance - 04 · Discover04Discover
Where the loop closes — hypotheses generated, then tested.
This is the core engine. Morbius runs ten specialized agents in parallel rather than prompting a single model repeatedly. Together they read the graph for structural weaknesses — where the literature contradicts itself, where a method has never been applied to the dataset that would test it, where a conclusion rests on one unreplicated study — and turn each into a candidate hypothesis with a design capable of falsifying it.
The loop advances without a human prompting each step. Hypotheses that fail the audit stage are discarded by Morbius rather than surfaced for a researcher to catch, so what reaches you has already survived internal scrutiny: a testable prediction, the experiment or analysis that would discriminate it from the alternatives, and the evidence trail behind both.
- Reviewer
- Audits the evidence base for gaps, weak support, and contradiction.
- Hypothesis
- Proposes candidate explanations for what the graph leaves unresolved.
- Planner
- Designs the experiment or analysis that would discriminate between them.
- Statistician
- Specifies power, controls, and the result that would falsify the claim.
- Writer
- Renders reasoning and outcomes into reviewable scientific prose.
- Auditor
- Checks every surviving conclusion back against its source evidence.
Gap detectedContradictionNovel hypothesisTestable predictionFalsification check - 05 · Publish05Publish
Verified findings, rendered publication-ready.
Findings that survive verification are composed into output: a manuscript draft, conference slides, a poster, an executive summary, teaching material. Citations resolve back to the graph entries that produced them, so every claim and figure in the export remains traceable to the evidence it came from.
ManuscriptSlidesPosterSummaryTeaching material
Verified output re-enters the graph as evidence, and the loop runs again — each pass reasoning over a larger, better-supported body of work than the last.
Most AI research tools stop at one part of this loop — summarizing papers, or answering questions, or proposing an idea for a human to validate. Morbius is built to run the entire loop autonomously, closing the gap between a plausible answer and a verified one. That gap — not raw language ability — is what separates AI-assisted research from AI-driven discovery.
Discovery is a loop. Morbius runs all of it.
Autonomy, Measured
This is the same autonomous engine benchmarked below against other AI research platforms. Across 20 research prompts, Morbius was scored on two co-primary outcomes — scientific quality and research execution — alongside the Claude Science Platform and Biomni Lab.

Chart A: Mean co-primary performance across 20 prompts. Co-primary means include 95% bootstrap CIs; diamonds mark means.

Chart B: Secondary composite distribution. Morbius achieves median score of 80.6/100.

Chart C: First-place finishes by co-primary outcome. Morbius leads with 36 total wins.
Scientific Quality
89.8%
Mean performance on scientific quality evaluation across 20 research prompts
First-Place Wins
36/40
Total first-place finishes (16 scientific quality + 20 research execution)
Median Composite
80.6
Secondary composite score out of 100 across all evaluation dimensions