Status: Live

Ἐπιστήμη · Knowledge

Morbius

An AI Co-Scientist for Autonomous Discovery.

Morbius doesn't just summarize your literature or suggest an idea. It ingests evidence, builds lasting research memory, reasons over it, and autonomously generates and tests hypotheses — closing the loop from raw literature to verified, citable output.

morbius · agent sessionRUNNING
Autonomous Discovery

› query

Scanning PubMed & arXiv corpus
Extracting entity & citation graph
Clustering by methodology
Identifying knowledge gaps & contradictions
Generating hypothesis candidates
Composing synthesis report
1/6 STEPS
17%

10

Research Agents

210M+

Papers Indexed

<2s

Ingestion

How Morbius Discovers

Discovery is not a feature set. It is a loop — and Morbius runs every stage of it as one continuous process, carrying the output of each stage forward as the input to the next.

  1. 01 · Ingest

    Evidence in, structure out.

    Morbius takes in PDFs, arXiv preprints, journal articles, technical reports, lecture recordings, and internal lab documents, then decomposes them. Sections, tables, figures, equations, citations, and individual claims are extracted and typed, so every later stage reasons over structured scientific evidence rather than undifferentiated text.

    PDFarXivJournalsReportsLecturesInternal docs
  2. 02 · Remember

    Research memory that compounds instead of resetting.

    Each ingested source is written into a persistent graph linking authors, concepts, methods, datasets, and results. This is the Co-Scientist's long-term memory: evidence read months ago remains available and, more importantly, remains connected. Context accumulates across sessions rather than being rebuilt from scratch in every conversation.

    AuthorsConceptsMethodsDatasetsResultsPersistent
  3. 03 · Reason

    Grounded inference — the substrate, not the destination.

    Questions resolve against the graph, and every claim returns with its provenance: the source paragraph, the figure it came from, the citation, a confidence score. This stage exists to make the next one trustworthy. Hypotheses are only worth testing if the evidence they rest on can be inspected and challenged.

    Source paragraphFigureCitationConfidence scoreProvenance
  4. 04 · Discover

    Where the loop closes — hypotheses generated, then tested.

    This is the core engine. Morbius runs ten specialized agents in parallel rather than prompting a single model repeatedly. Together they read the graph for structural weaknesses — where the literature contradicts itself, where a method has never been applied to the dataset that would test it, where a conclusion rests on one unreplicated study — and turn each into a candidate hypothesis with a design capable of falsifying it.

    The loop advances without a human prompting each step. Hypotheses that fail the audit stage are discarded by Morbius rather than surfaced for a researcher to catch, so what reaches you has already survived internal scrutiny: a testable prediction, the experiment or analysis that would discriminate it from the alternatives, and the evidence trail behind both.

    Reviewer
    Audits the evidence base for gaps, weak support, and contradiction.
    Hypothesis
    Proposes candidate explanations for what the graph leaves unresolved.
    Planner
    Designs the experiment or analysis that would discriminate between them.
    Statistician
    Specifies power, controls, and the result that would falsify the claim.
    Writer
    Renders reasoning and outcomes into reviewable scientific prose.
    Auditor
    Checks every surviving conclusion back against its source evidence.
    Gap detectedContradictionNovel hypothesisTestable predictionFalsification check
  5. 05 · Publish

    Verified findings, rendered publication-ready.

    Findings that survive verification are composed into output: a manuscript draft, conference slides, a poster, an executive summary, teaching material. Citations resolve back to the graph entries that produced them, so every claim and figure in the export remains traceable to the evidence it came from.

    ManuscriptSlidesPosterSummaryTeaching material

Verified output re-enters the graph as evidence, and the loop runs again — each pass reasoning over a larger, better-supported body of work than the last.

Most AI research tools stop at one part of this loop — summarizing papers, or answering questions, or proposing an idea for a human to validate. Morbius is built to run the entire loop autonomously, closing the gap between a plausible answer and a verified one. That gap — not raw language ability — is what separates AI-assisted research from AI-driven discovery.

Discovery is a loop. Morbius runs all of it.

Autonomy, Measured

This is the same autonomous engine benchmarked below against other AI research platforms. Across 20 research prompts, Morbius was scored on two co-primary outcomes — scientific quality and research execution — alongside the Claude Science Platform and Biomni Lab.

Mean co-primary performance across platforms - Morbius leads with 89.8% in scientific quality

Chart A: Mean co-primary performance across 20 prompts. Co-primary means include 95% bootstrap CIs; diamonds mark means.

Secondary composite distribution - Morbius median 80.6 out of 100

Chart B: Secondary composite distribution. Morbius achieves median score of 80.6/100.

First-place finishes by outcome - Morbius: 16 scientific, 20 execution firsts

Chart C: First-place finishes by co-primary outcome. Morbius leads with 36 total wins.

Scientific Quality

89.8%

Mean performance on scientific quality evaluation across 20 research prompts

First-Place Wins

36/40

Total first-place finishes (16 scientific quality + 20 research execution)

Median Composite

80.6

Secondary composite score out of 100 across all evaluation dimensions