Qiao Jin
← Research
AI for Evidence Generation · Systematic Reviews
Accelerating Clinical Evidence Synthesis with Large Language Models
Zifeng Wang, Lang Cao, Benjamin Danek,
Qiao Jin
, Zhiyong Lu, Jimeng Sun
npj Digit. Med.
2025
PDF
Journal
These slides were generated with the help of AI and may contain errors.
01
Motivation
Systematic reviews cannot keep pace with the literature
The gold standard — but slow
Meta-analyses are the highest tier of evidence.
Each takes months: searching all studies, screening thousands of abstracts, hand-extracting data.
Can’t keep pace
The literature grows by >1M papers a year.
Manual synthesis simply cannot keep up.
02
Method
A human-in-the-loop pipeline for the whole review
Assists study search, screening, and data extraction — expert-in-the-loop.
Evaluated on TrialReviewBench: 100 systematic reviews spanning 2,220 studies.
The four-step pipeline, each step keeping a human expert in the loop.
Fig. 1
.
03
Result · Search
Recovers ~0.8 of relevant studies vs. ~0.2 for humans
Expands queries and mines citation links.
Finds the studies manual search misses.
TrialMind: 0.78
0.78
TrialMind
Manual: 0.19
0.19
Manual
GPT-4: 0.07
0.07
GPT-4
Study-search recall
Mean recall across four therapy topics (TrialMind 0.71–0.83 per topic). Source:
Fig. 2c
.
04
Result · Screen
44% less screening time at equal recall
Ranks abstracts by relevance.
Reviewers reach the same coverage far sooner — +71% recall, −44% time.
AI + Human: 0.72
0.72
AI + Human
Human only: 0.42
0.42
Human only
Screening recall@10
AI + Human: 455s
455s
AI + Human
Human only: 815s
815s
Human only
Screening time (s)
Screening recall and time, AI+human vs. human-only. Source:
Fig. 5e
.
05
Result · Synthesize
Experts prefer its evidence over GPT-4
Improves structured data-extraction accuracy over the GPT-4 baseline.
Experts preferred its synthesized findings in
62.5–100%
of head-to-head comparisons.
06
Summary
TrialMind at a glance
Background
A systematic review must find and screen every relevant trial by hand.
Problem
Reviews can’t keep pace with >
1M
new papers a year.
Approach
A human-in-the-loop LLM pipeline for study search, screening & extraction.
Results
Finds
~0.8
of relevant studies (vs ~0.2 manual), at
−44%
screening time.
Conclusion
Systematic-review-grade evidence, produced far faster.