Skip to content

The Best Near-Term Job for a Research Agent Is Reading the Whole Literature — Carefully

No human can hold a field anymore; the corpus outgrew comprehension decades ago. The highest-value near-term use of research agents is careful synthesis across the whole literature — but only if they read the primary source and keep the caveats.

By Mehdi8 min read
Share
On this page

The scientific literature passed the limit of human comprehension a long time ago, and everyone in research quietly works around it. No living immunologist has read immunology. No one holds oncology, or even a subfield of it, in a single head. My bet — and it is a forecast, so I will argue for it rather than assert it — is that the most valuable thing research agents will do in the next few years is not design experiments or propose theories, but read the whole corpus carefully and synthesize it: find what is already known but sits unconnected across two subfields, surface contradictions no single reviewer ever saw together, and map the actual state of the evidence. That is a job no human can do at scale, and it is the one closest to being real.

The catch is a discipline, not a capability, and it decides whether the same tool serves truth or industrializes noise. I will get to it. First, the size of the gap.

The corpus outgrew the reader decades ago

Do the arithmetic, because it is the whole premise. PubMed indexes north of 35 million citations and takes in well over a million new ones a year — on the order of a few thousand biomedical papers a day. A specialized subfield — say, epigenetic aging clocks, which is part of what I work on — might hold 50,000 papers. A diligent researcher who read ten papers a day, every day, with no weekends and no experiments to run, would clear that subfield in about fourteen years, by which point it would have roughly tripled. Nobody does this. What people actually do is read a few hundred papers deeply, track a few dozen groups, and infer the rest from reviews, talks, and the ambient sense of the field.

That inference layer is where knowledge goes missing. A review article is one author's compression of a slice they happened to see, written on a two-year lag. The "ambient sense of the field" is a social artifact — it tracks who is loud, who is at the right institutions, and which results were fun to repeat, none of which is the same as what the evidence says. The literature is not a body of knowledge that humans hold. It is a body of knowledge that no one holds, sampled thinly and unrepresentatively by each of us.

So a great deal of scientific value is not locked inside any paper. It is locked in the gaps between papers that nobody read together.

The value is in the gaps, and there is a real precedent for mining them

The clearest historical proof that these gaps are real and exploitable comes from Don R. Swanson, an information scientist at the University of Chicago who in 1986 coined the term undiscovered public knowledge. His argument: if one literature establishes that A affects B, and a separate literature — one that does not cite the first and is read by different people — establishes that B affects C, then "A may affect C" is a testable hypothesis already implied by the published record but stated by no one, because no one read both literatures. He worked the example by hand. Dietary fish oil was known to reduce blood viscosity and platelet aggregation; Raynaud's syndrome was known to involve high blood viscosity and vascular reactivity; the two literatures did not talk to each other. Swanson proposed, from reading alone, that fish oil might help Raynaud's. It was later supported clinically. He did the same for magnesium and migraine.

Swanson found these by manually cross-reading disjoint literatures — an artisanal process he could run a handful of times in a career. The field of literature-based discovery grew out of it, and it has stayed niche for one boring reason: the reading does not scale on human hands. This is precisely the shape of task that changes character when an agent can hold the whole corpus. The forecast is not that agents will invent a new kind of insight. It is that they will make Swanson's move cheap and routine instead of heroic.

The gaps come in at least three flavors, and each is a distinct near-term product:

  • Cross-subfield connection. A result in one community answers an open question in another that never cites it. This is Swanson's case, and it is a retrieval-and-inference problem at heart — exactly what large-scale reading is good at.
  • Contradiction detection. Study X reports an effect; studies Y and Z fail to find it under slightly different conditions; no reviewer ever placed all three side by side because they sit in different journals under different keywords. Surfacing the contradiction — and the boundary condition that might reconcile it — is pure collation at scale.
  • Meta-patterns. A regularity visible only across thousands of methods sections: a reagent that correlates with a class of results, a covariate that flips a sign whenever it is included, a p-value distribution that betrays selective reporting across a whole subfield. No individual reader has the sample size to see these. An agent that has read all of it does.

Systematic review is the existing, respectable version of this work, and it shows both the demand and the ceiling. A Cochrane-style review screens thousands of abstracts to include a few dozen studies, takes a year or more, costs a great deal, and is partly stale by publication. What agents threaten to change is not the rigor of that process but its throughput — turning a year into a week, and letting the review be re-run the day a new trial posts. That alone would reshape what "knowing the literature" means.

The discipline: read the primary source, keep the caveats

Here is where the whole thing turns, and it is the point I keep returning to. A summary is lossy compression, and the loss is not random — it deletes exactly the caveats, effect sizes, and boundary conditions that decide whether a claim is true. I have argued elsewhere that reading the primary source is becoming a superpower precisely because the codec of summarization is tuned to keep the memorable headline and drop the machinery of doubt. That argument does not soften when you point it at an agent. It sharpens into a design constraint.

A synthesis agent that reads abstracts is worse than useless. Not weaker — actively harmful. The abstract is already the authors' own lossy compression, written to maximize citations, and it keeps "intervention reverses epigenetic age" while dropping that n was small, the work was in vitro on one cell type, the clock was never validated on that tissue, and the measured reversal sits inside the clock's own error band. Every one of those caveats is fatal to the headline, and none survives to the abstract. Now run synthesis across ten thousand such abstracts. The agent aggregates ten thousand headlines and discards ten thousand sets of caveats, then hands you a confident map of "consensus" that is really a map of overclaims with the qualifications stripped out. It has laundered lossiness into authority. A human skimming abstracts at least knows they are skimming; an agent that produces a polished cross-paper synthesis erases that signal, and the synthesis gets cited as though it read the papers.

So the requirement is concrete and non-negotiable: the agent must read the methods and the results table, not the abstract, and it must carry the caveats forward with every claim it propagates. When it tells you "fourteen studies support X," it has to attach, per study, the effect size, the sample, the population, the conditions, and the one discussion sentence conceding what could break the result. A synthesis that says "significantly associated with mortality" without carrying "hazard ratio 1.05" has told you nothing and made you feel informed, which is the worse of the two states. The value of reading the whole literature exists only if the reading goes to the layer where nobody was trying to make you feel anything.

The peril: the same tool industrializes the literature's biases

Comprehensiveness is not neutrality. The published record is a biased sample of the experiments that were actually run — positive results get published, null results sit in file drawers, and p-hacking tilts the surviving effects. An agent that faithfully synthesizes the literature will faithfully reproduce every one of those distortions, and it will do so with a fluency and completeness that makes the bias more convincing, not less. Point the same capability at production rather than verification and you have built a machine for turning publication bias into apparent consensus at scale. This is the same fork I have argued runs through AI and the reproducibility crisis: the technology helps a checking problem only when it is aimed at checking, and deepens the problem the moment it is aimed at fluent production.

Which direction you get is decided entirely by whether the agent verifies or merely aggregates. Verifying looks like weighting by study quality and sample size instead of counting papers; hunting the null results in registries and preprints that never made it to a journal; reporting effect-size heterogeneity rather than averaging a real contradiction into a comfortable mean; and — again — carrying caveats so a reader can see when the confidence is thin. Aggregating looks like summing what the corpus says and presenting the sum as the state of nature. The first fights the literature's biases. The second industrializes them. Same model, same corpus, opposite epistemics, and the difference is a posture you have to demand explicitly, because nothing in the training objective supplies it for free.

The honest limit, and what to demand

The strongest counterargument is that judging a methods section is real expertise the machines mostly do not have yet, and I think that is true. Deciding whether an adjustment set is adequate, whether a clock transfers across tissues, or whether an effect is an artifact of one lab's protocol is domain judgment, and current systems do it unreliably when they do it at all. Citation hallucination is a live failure mode; a synthesis is worthless if some of its supporting papers do not say what it claims or do not exist. None of this is a reason to dismiss the forecast, but all of it bounds it.

The bound points at the right division of labor. The near-term win is not the agent adjudicating the evidence. It is the agent doing the one thing humans genuinely cannot — comprehensive, careful reading and cross-paper collation at corpus scale — and staging the primary-source passages so a human expert makes the call. Keep the machine on the reading and the retrieval, keep the human on the verification, and insist the whole chain stay auditable back to the methods.

So the thing to demand of any literature-synthesis agent is narrow and testable: cite the primary source, quote the methods-level passage, and preserve the caveat — never the abstract-level claim, never a naked "studies show." If it cannot show you the results table it is standing on, it did not read the literature. It read the headlines, and it is selling you the field's overclaims with the doubt filed off.

Frequently asked questions

Isn't 'read the whole literature' just a bigger version of the summarization tools we already have?
No, and the difference is the whole argument. A summarization tool compresses one paper into a shorter version of that paper. A synthesis agent reads many papers at their primary-source level and looks for relationships across them: the result in field A that answers the open question in field B, the effect four studies report and one refutes, the meta-pattern that only appears once you line up two thousand methods sections. The unit of value is the relationship between papers, not the shrinking of any single one. Don Swanson identified that category in the 1980s under manual conditions; the forecast is that agents make it routine rather than heroic. And it only works if the reading is deep. A system that ranks and connects abstracts is doing keyword matching with better manners, not synthesis.
Why is an agent that reads abstracts 'worse than useless' rather than just weaker?
Because it does not merely fail to add value — it manufactures false consensus. An abstract keeps the headline claim and drops the effect size, the sample, the boundary conditions, and the one sentence conceding residual confounding. When an agent synthesizes across thousands of abstracts, it aggregates the overclaims and discards every caveat that would have qualified them, then presents the pile as a settled state of evidence. A human skimming abstracts at least knows they are skimming. An agent that produces a confident cross-paper synthesis from abstracts launders that lossiness into something that looks authoritative and gets cited as if it read the papers. The failure is invisible at exactly the moment it matters, which is worse than an obvious gap.
Can an agent actually judge a methods section, or does that still require a domain expert?
Today, only partially, and that is the honest limit on this forecast. Judging whether a clock was validated on the tissue it was applied to, whether an adjustment set is adequate, or whether an effect sits inside its own error band is real methodological judgment, and current systems do it unreliably. The near-term win does not depend on the agent replacing that judgment. It depends on the agent doing the part humans cannot do at all — comprehensive retrieval and cross-paper collation at corpus scale — while surfacing the primary-source passages a human expert then adjudicates. The right design keeps the human on verification and lets the machine carry the reading load. An agent that claims to have done the judging, rather than staged it for a human, is overreaching.
If the same tool can serve truth or industrialize bias, what actually decides which one you get?
Whether it verifies against ground truth or just aggregates what the corpus already says. The literature carries known distortions — publication bias, selective reporting, p-hacking — and an agent that faithfully summarizes the published record will faithfully reproduce those distortions at scale, presenting a biased sample as the state of nature. To point at truth instead of noise it has to do the things that fight those biases: weight by study quality and sample, hunt for the null results that never got published or sit in registries and preprints, flag effect-size heterogeneity instead of averaging it away, and carry each claim's caveats with it. That is a verification posture, not a summarization posture, and it is the same design principle that separates AI that helps the reproducibility crisis from AI that deepens it.

Filed under Applied AI. AI that ships, not AI that demos.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Applied AI

The Jagged Frontier: AI Is Superhuman and Subhuman at the Same Time

AI capability isn't one number climbing toward "human level." It's a jagged frontier — superhuman at some tasks, worse than a child at others, with no smooth link between them — and that jaggedness, not the average, is what makes deployment hard and "AGI" a category error.

8 min read