Skip to content

Trusting an AI Is the 250-Year-Old Problem of Believing a Stranger

"Should I trust what the AI told me" is the epistemology of testimony applied to a testifier with unknown reliability, no stable identity, and no accountability — exactly the configuration where philosophers say default trust is unwarranted.

By Mehdi9 min read
Share
On this page

"Should I trust what the AI told me?" is not a new question, and treating it as one is why most answers to it are bad. It is the epistemology of testimony — the centuries-old problem of when it is rational to believe what someone else tells you — applied to a new kind of testifier. That tradition, from Hume and Reid down through modern social epistemology, has a worked-out account of when accepting testimony is rational. Run a language model through it and the model fails on the exact three conditions the tradition says matter most: unknown and unstable reliability, no stable identity across contexts, and no accountability for being wrong. Those are precisely the conditions under which the philosophy of testimony says default trust is unwarranted and you must fall back to independent evidence. A 250-year-old theory turns out to be a diagnostic instrument pointed straight at the thing everyone is anxious about.

Almost everything you know, you were told

Start with a fact about your own mind that is easy to underrate: almost nothing you know, you verified. You know the Earth is roughly 4.5 billion years old, that your blood carries oxygen on hemoglobin, that a city called Lima exists, that antibiotics were discovered in the twentieth century. You checked none of it. You were told — by teachers, textbooks, journalists, institutions, strangers who drew the maps — and you believed, rationally, without independent confirmation. C.A.J. Coady, whose 1992 book Testimony largely revived this corner of epistemology, made the point sharply: strip out everything you believe on the word of others and what remains is a pitiful stub. John Hardwig put the epistemic version bluntly in his 1985 paper "Epistemic Dependence" — in any developed field, even the experts depend on the testimony of other experts they cannot check. Knowledge is not mostly a solo achievement. It is mostly a relay.

This is why the AI trust question is not exotic. You have been outsourcing cognition to testifiers your whole life. The real question was never whether to rely on sources you can't fully verify — you have no choice — but which ones, how much, and when to stop. That question already has a literature. The novelty is only the testifier.

The two classic answers, and when each is rational

The philosophy of testimony has been organized for two centuries around a single disagreement, and it is exactly the one you feel when you stare at a confident paragraph of model output.

On one side is Hume's reductionism. In "Of Miracles" (Section X of the Enquiry), Hume treats the credit we give testimony not as a basic entitlement but as an ordinary empirical inference. We have observed, over our lives, a regular conjunction between what people report and what turns out to be true — and we proportion our belief in any given report to that observed reliability, discounting for the reporter's competence, the report's inherent improbability, and any motive to deceive. Testimony is trustworthy only to the extent your own experience licenses an inductive bet that this kind of source, on this kind of claim, tends to be right. You reduce the testimony to independent evidence about the testifier.

On the other side is Thomas Reid's anti-reductionism. Reid held that we come equipped with two native dispositions: a principle of veracity (a natural pull to tell the truth) and a principle of credulity (a natural default to believe what we are told). For Reid, testimony is a basic source of justification the way perception is — you are entitled to believe it by default, no prior track record required, unless you have positive reason to doubt. Without something like this, he argued, a child could never learn a language or a first fact, because a child has no independent stockpile of confirmations to run Hume's induction over.

Modern social epistemology mostly lives between the poles. Elizabeth Fricker's "local reductionism" allows default trust but requires the hearer to actively monitor for signs of unreliability and withhold when they appear. Jennifer Lackey's hybrid view holds that testimony transmits knowledge only when the source is reliable and the hearer has no defeaters — no available reason to think this source, here, is wrong. Both moderate positions share a structure: default trust is licensed only in an environment where testifiers are generally reliable and where unreliability leaves detectable traces you can monitor for. Take away either condition and everyone in the literature, reductionist and anti-reductionist alike, converges on the same fallback — you are back to Hume, and you must reduce to independent evidence.

The condition the tradition keeps circling: accountability

A second thread runs underneath the reliability debate, and it turns out to be decisive for AI. When someone tells you something, they are not merely emitting a true-or-false string. They are performing an act — an assertion — and assertion comes loaded with a norm and a stake. Richard Moran calls this the assurance view: in telling you that p, a speaker openly assumes responsibility for its truth, gives you their word, and thereby invites you to believe it on their say-so. Part of why it can be rational to believe a stranger is that they have made themselves answerable. If they are wrong, they are on the hook — reputationally, socially, sometimes legally. They have skin in the transaction.

This is not a moral garnish on top of the reliability story. It is load-bearing. Accountability is what makes a source's reliability trustworthy going forward rather than merely true in the past. A track record tells you how a source behaved; accountability tells you the source has a standing reason to keep behaving that way, because errors land on them. It is the mechanism that converts "has been reliable" into "can be relied upon." A witness under oath, a physician who can be sued, a journalist whose byline absorbs the blame — their answerability is a large part of why deferring to them is rational rather than credulous.

Hold the two requirements together, because they are the whole test. Rational default trust in testimony requires (1) a source whose reliability is known and stable enough to bet on, and (2) a source that is answerable for being wrong. Now bring the model to the bar.

The model fails all three probes at once

Reliability is unknown and non-stationary. Hume's rule needs a stable base rate to proportion belief against. A model doesn't offer one. Its accuracy is savagely domain-dependent: strong on well-represented topics, quietly catastrophic just off-distribution, and — crucially — it reports both with identical fluency and confidence. Do the arithmetic on what that means. Suppose a model is right on 92% of questions in some domain, a genuinely strong number. If the 8% of errors were flagged, or clustered somewhere you could see, you could route around them. They aren't. They are interleaved with the correct answers and delivered in the same assured register, so the conditional reliability you actually need — the probability that this specific claim is true — is not 92% and is not recoverable from the aggregate. Worse, the base rate is non-stationary: it shifts with phrasing, with topic, with a silent version update that can move behavior overnight. You cannot proportion belief to a reliability that won't hold still. Hume's induction has nothing stable to bite on.

There is no stable identity across contexts. Testimonial trust attaches to a someone — a persistent testifier who accrues a reputation you carry from one encounter to the next. "Dr. Reyes was right about the last three referrals" is a usable fact because Dr. Reyes persists. A model is closer to a new stranger each session: no memory of what it told you yesterday, no continuity of commitment, a different effective system underneath after each update. The unit that reputation is supposed to bind to keeps dissolving. You cannot build the track record the whole apparatus runs on, because there is no stable entity for it to accumulate against.

There is no accountability. This is the deepest failure and the one the assurance view names. A model gives no assurance in Moran's sense. It cannot vouch, cannot stake anything, bears no consequence for misleading you. When it is confidently wrong, nothing lands on it. This is not a gap a bigger model closes — it is structural, and it is the same missing piece I traced from the other direction in The Coming Agent Trust Crisis: intelligence is commoditizing while the answerability that would make it safe to defer has no home in the system. A testifier with no skin in the game is, in the exact terms of the literature, a source you are not entitled to take on its say-so.

Three probes, three failures — and they are not probes invented to indict AI. They are the standing conditions the field settled on for human testimony, written down decades before the models existed. The configuration a model presents — unknown non-stationary reliability, no persistent identity, no accountability — is the textbook description of the case where default trust is irrational and reduction to independent evidence is mandatory.

Does it even know? — and why that's a separate question

One might think all of this collapses into a prior question: can the thing know anything at all, such that there is knowledge to transmit? I take that seriously enough to have given it its own treatment in Can a Machine Know Anything?. But the testimony literature holds a surprising result that keeps the two questions apart. Lackey's central cases show that testimonial knowledge can flow from a source that does not itself know or even believe what it conveys. Her example is a schoolteacher who reliably teaches evolution straight from the textbook while privately, as a creationist, disbelieving it — the students learn biology anyway. What did the work was the reliability of the channel, not the mental state of the source.

The payoff is direct. Even if a model understands nothing — no belief, no knowledge, no inner grasp behind the words — its output could still convey knowledge to you provided it is reliable and you have no defeaters. The trust case does not have to wait on settling machine cognition. It rests on the two things you can assess from the outside: demonstrated reliability, and checkable provenance. Which is fortunate, because those are also the two things you can engineer.

What would make an AI a testifier you could rationally trust

The tradition doesn't just diagnose; it prescribes. Convert its two requirements into demands on the system.

Reduce, by default, to independent evidence. This is Hume operationalized. Treat any load-bearing claim as a lead, not a verdict, and route it through a source that has the reliability and answerability the model lacks. In practice: citations you can open and confirm actually say what the model claims — provenance you can check without trusting the party that benefits from the answer. An assertion should not be permitted to be its own warrant.

Scope trust to domains of demonstrated reliability. Don't ask "is the model trustworthy," which has no answer, but "is it trustworthy here — on this kind of question, where I have personally watched it be reliable, and where a wrong answer is cheap to catch." A calibrated, narrow default beats a global one. This is Fricker's monitoring, mechanized: default trust confined to a region where defeaters are detectable.

Build the accountability the model can't personally hold. Since the testifier cannot be answerable, answerability has to live in the surrounding structure — logged and inspectable outputs, an operator who bears consequences, staked reputation that a fresh instance cannot fake. This is what the assurance view forces: where there is no natural stake, trust is rational only once an artificial one is installed.

Strip the abstraction and the practical rule is almost rude in its simplicity. Treat the model as testimony from a stranger of unknown reliability, and demand exactly what you would demand of one. You already know how to deal with a confident, articulate stranger who might be right and might be selling you something: you check the checkable parts, you trust them narrowly and only where you have seen them deliver, and you never let them be the only witness on anything that matters. That instinct is not naïveté about AI. It is 250 years of epistemology, and it was right about strangers long before the strangers learned to type.

Frequently asked questions

Doesn't a model with high benchmark accuracy earn trust the way a reliable expert does?
Not from a benchmark alone. Hume's rule is to proportion belief to a source's observed reliability, but that only works when reliability is stable enough to form a prior. A model's accuracy is domain-specific and non-stationary: it can score 90% on a public benchmark and fail systematically on inputs slightly off-distribution, on a new topic, or after a silent version update, while stating every answer with the same confidence. A benchmark tells you the average over a fixed distribution; trust requires knowing the reliability on this question, which is exactly what you don't have. High measured accuracy scopes where trust is defensible; it does not license default trust everywhere.
If a model doesn't really 'know' anything, can its output still give me knowledge?
Yes, and this is the useful subtlety. Jennifer Lackey's work on testimony shows that testimonial knowledge can flow from a source that lacks the corresponding belief or knowledge, provided the source is reliable and the hearer has no defeaters. Her case is a creationist teacher who reliably conveys evolution from a textbook while privately disbelieving it; the students still learn. So whether a model knows anything and whether its testimony can rationally be believed come apart. A model could transmit knowledge by being reliable even if it understands nothing, which is why the trust case rests on demonstrated reliability plus checkable provenance, not on the model's inner states.
What does 'reduction to independent evidence' actually require me to do?
Treat the model's claim as a lead, not a conclusion, and route it through a source that has the reliability and accountability the model lacks. Concretely: demand a citation you can open and verify says what the model claims; check load-bearing facts against a source that bears consequences for error; confine your default trust to domains where you have personally seen the model be reliable and where a wrong answer is cheap to catch. This is Hume's reductionism operationalized. You are not refusing to use the output; you are refusing to let the assertion be its own warrant.

Filed under Cross-Disciplinary Deep Essays. Where biology, computation, markets, and philosophy collide.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Cross-Disciplinary Deep Essays

"AGI by Year X" Is Unanswerable Until You Name the Definition

"Are we close to AGI?" is incoherent because AGI names at least four incompatible criteria that come apart in practice. Separate them and the timeline debate dissolves into concrete, checkable questions.

9 min read