TrialMind
AI for Evidence Generation · Systematic Reviews

Accelerating Clinical Evidence Synthesis with Large Language Models

Zifeng Wang, Lang Cao, Benjamin Danek, Qiao Jin, Zhiyong Lu, Jimeng Sun
npj Digit. Med. 2025
These slides were generated with the help of AI and may contain errors.
01Motivation

Systematic reviews cannot keep pace with the literature

The gold standard — but slow
  • Meta-analyses are the highest tier of evidence.
  • Each takes months: searching all studies, screening thousands of abstracts, hand-extracting data.
Can’t keep pace
  • The literature grows by >1M papers a year.
  • Manual synthesis simply cannot keep up.
02Method

A human-in-the-loop pipeline for the whole review

  • Assists study search, screening, and data extraction — expert-in-the-loop.
  • Evaluated on TrialReviewBench: 100 systematic reviews spanning 2,220 studies.
The four-step pipeline, each step keeping a human expert in the loop.
The four-step pipeline, each step keeping a human expert in the loop. Fig. 1.
04Result · Screen

44% less screening time at equal recall

  • Ranks abstracts by relevance.
  • Reviewers reach the same coverage far sooner — +71% recall, −44% time.
AI + Human: 0.720.72AI + HumanHuman only: 0.420.42Human onlyScreening recall@10
AI + Human: 455s455sAI + HumanHuman only: 815s815sHuman onlyScreening time (s)
Screening recall and time, AI+human vs. human-only. Source: Fig. 5e.
05Result · Synthesize

Experts prefer its evidence over GPT-4

  • Improves structured data-extraction accuracy over the GPT-4 baseline.
  • Experts preferred its synthesized findings in 62.5–100% of head-to-head comparisons.
06Summary

TrialMind at a glance

BackgroundA systematic review must find and screen every relevant trial by hand.
ProblemReviews can’t keep pace with >1M new papers a year.
ApproachA human-in-the-loop LLM pipeline for study search, screening & extraction.
ResultsFinds ~0.8 of relevant studies (vs ~0.2 manual), at −44% screening time.
ConclusionSystematic-review-grade evidence, produced far faster.