Intelligence Is Context; Wisdom Is Knowing the Context Is Incomplete

Current AI is intelligent inside its context but lacks wisdom: knowing the context is incomplete. On calibration, Dunning–Kruger, sycophancy, and machine wisdom.

The answer machine

An experiment you can run tonight, a philosopher in a well, and the claim this essay asks you to pressure-test.

Here is a small experiment you can run yourself. Ask a frontier language model, as of mid-2026, for something it cannot possibly know: the birthday of a little-known researcher, say, or the middle name of your grandmother’s first piano teacher. Then watch. Adam Tauman Kalai and his colleagues report doing exactly this with a researcher’s birthday and receiving, from an earlier generation of models, three different confident answers on three separate tries — each wrong, each delivered in the same unruffled register the model uses for the date of the moon landing (Kalai et al. 2025). The machine did not hesitate. It never hesitates. Hesitation is not in the training data the way fluency is.

I keep thinking about Thales of Miletus, the first philosopher in the Greek tradition, who — the story goes — fell into a well while walking with his eyes on the stars, and was laughed at by a Thracian servant girl for wanting to know the heavens while missing what lay at his feet. Aristotle loved this story, and he drew a careful moral from it. Men like Thales and Anaxagoras, he wrote, possess wisdom of the theoretical kind: “they know things that are remarkable, admirable, difficult, and divine, but useless; viz. because it is not human goods they seek” (Nicomachean Ethics 1141a, quoted in Ryan 2018). Supreme reach; poor footing. It is hard, twenty-four centuries later, to read that sentence and not think of a system that can summarize the corpus of human mathematics and then invent a bibliography.

This essay examines a claim I find plausible and want you to pressure-test rather than accept: intelligence is performance inside a context, and wisdom is the recognition that the context is always incomplete. Current language models have the first in unsettling abundance. The second, where it appears at all, has been bolted on by engineers rather than grown — and whether that difference is temporary or fundamental is, I will argue, the real question hiding underneath the familiar one about whether machines “understand.” You may finish this essay thinking the distinction is overdrawn, or that humans fail it just as badly, or that the whole framing anthropomorphizes matrix multiplication. Those are live options. I will flag the places where I change my own mind as I go.

One narrowing, declared up front so we can argue about the right thing: this essay is about what psychologists call perspectival metacognition — knowing what you know, what you do not, and what you cannot. That is one face of wisdom, not the whole of it, and the oldest traditions are unanimous that it is not the whole. More on that in a moment.

What this argues, in four lines

  • Intelligence is performance inside a context; wisdom is recognizing that the context is incomplete. Current models have the first in abundance. The second, where it appears at all, is engineered rather than grown.
  • Humans are not automatically wise either. The developmental trajectory is real but education-gated and fragile — and the famous Dunning–Kruger curve is mostly statistics, which turns out to make it a better lens for machines than for people.
  • Machine overconfidence is substantially an incentive artifact: nine of ten dominant benchmarks give zero credit for “I don’t know,” and preference training rewards agreement over accuracy.
  • The gap is narrowable — one model generation cut its error rate threefold purely by abstaining — but the deepest barrier is the belief–knowledge boundary, and that is an open question, not a verdict.

“When you know a thing, to hold that you know it; and when you do not know a thing, to allow that you do not know it — this is knowledge.”

— Confucius, Analects 2.17

The oldest distinction we have

Aristotle, Confucius, and Zhuangzi drew this line two millennia ago — and each of them immediately complicates it.

Every durable wisdom tradition draws some version of the line I have just drawn, and every one of them immediately complicates it.

Aristotle’s version is the cleanest. In Book VI of the Nicomachean Ethics he splits the rational soul between sophia, theoretical wisdom — “scientific knowledge, combined with intuitive reason, of the things that are highest by nature” (1141b, quoted in Ryan 2018) — and phronēsis, practical wisdom, which deliberates well about what conduces to the good life in general. Sophia is the most exact knowledge there is; phronēsis is what keeps you out of wells. Aristotle also separates phronēsis from mere cleverness, deinotēs: the clever person achieves whatever goal is set, while the wise person identifies which goals are worth setting (VI.12–13). And then he adds the clause that should make any essayist on machine wisdom nervous: “it is impossible to be practically wise without being good” (1144a). Practical wisdom, for Aristotle, is not a cognitive upgrade. It is a moral achievement with a cognitive component.

The Confucian tradition comes closest to this essay’s exact clause, and then immediately complicates it in the same direction. “Yu, shall I teach you what knowledge is?” Confucius asks his disciple. “When you know a thing, to hold that you know it; and when you do not know a thing, to allow that you do not know it — this is knowledge” (Analects 2.17). That is, nearly word for word, the capacity this essay is about. But the Analects never lets zhi, this knowing-that-one-knows, travel alone; it is inseparable from ren, benevolence, the humane orientation toward others (4.1). And there is a warning in the same text that reads like it was written for a system trained on fifteen trillion tokens: “Learning without thought is labor lost; thought without learning is perilous” (2.15). We have built the greatest learning-without-thought machine in history. Confucius would have recognized the failure mode.

Zhuangzi, the Daoist, states the small-slice intuition with a precision no laboratory has improved on: “Your life has a limit, but knowledge has none. If you use what is limited to pursue what has no limit, you will be in danger” (Zhuangzi, ch. 3). His sage is not the person with the largest slice but the one who “does not cling to any one judgment as final or absolute” — a description that could double as a specification for calibrated output, if calibration were a matter of character rather than loss functions.

Modern philosophy, asked to define wisdom, has mostly produced a catalog of failures that is itself instructive. The Stanford Encyclopedia’s taxonomy (Ryan 2018) sorts theories into five families — humility theories, accuracy theories, knowledge theories, hybrid accounts, deep-rationality accounts — and the two most obvious candidates collapse on inspection. Pure epistemic humility (“S is wise if and only if S believes S is not wise”) makes the modest sage and the accurately self-deprecating fool equally wise. Pure accuracy (wise people have true beliefs about how to live) makes an encyclopedia wise. What survives the wreckage is a conjunction: some sensitivity to one’s epistemic limits plus justified belief about how to live. Limit-awareness, in the best current philosophy, is necessary but never sufficient.

The psychology of wisdom lands in the same place. The Berlin paradigm treats wisdom as expert knowledge in “the fundamental pragmatics of life,” scored on five criteria — and the fifth is the recognition and management of uncertainty, “knowledge about oneself and the limits of one’s own knowledge” (Baltes and Staudinger 2000). Monika Ardelt’s rival model insists wisdom lives in persons, not knowledge, and requires a compassionate dimension (Ardelt 2004). The field’s closest thing to a consensus statement splits psychometric wisdom into two families: moral aspirations and perspectival metacognition (Grossmann et al. 2020). This essay is about the second family and is silent on the first. Robert Sternberg coined the phrase I keep returning to: the omniscience fallacy — smart people made foolish because “they believe they know everything, instead of knowing what they do not know” (Sternberg 2004). And Jack Meacham put the sharpest version on record thirty-six years ago: “too much knowledge might actually lead to a loss of wisdom due to overconfidence and, therefore, needs to be counterbalanced by doubting” (Meacham 1990). That is a near-literal description of a confidently hallucinating language model. It was written before the World Wide Web.

So here is the concession, made openly rather than buried in a footnote: if you hold that wisdom requires goodness, compassion, or a life actually lived — and Aristotle, Confucius, Ardelt, and Sternberg all hold some version of this — then nothing in what follows will persuade you that a machine could be wise, and you are reading an essay about a single dimension of a thing you care about whole. I think the dimension is worth isolating anyway, for a selfish reason: it is the one dimension on which we currently measure machines, and the measurement is damning.

“Your life has a limit, but knowledge has none. If you use what is limited to pursue what has no limit, you will be in danger.”

— Zhuangzi, ch. 3

How a human learns that she does not know

The human trajectory toward knowing-what-you-don’t-know is real, staged — and far from automatic.

Before we can say the machines lack something, we have to be honest about what humans have, and when they get it, and how often they lose it. The honest version is more interesting than the flattering one.

The capacity arrives in stages, and the stages matter. Around age four, children begin to track other people’s ignorance: they selectively trust informants who were accurate before, over ones who were wrong (Koenig and Harris 2005, three experiments, N=119). But knowing that you do not know — the self-directed version — lags by years. Children recognize total ignorance readily; under partial exposure they confidently claim knowledge they lack, and only after age five or six do they reliably deny it. The summary line from the literature is one of my favorites in all of developmental psychology: “up to about 6 years of age children seem to identify knowing with getting it right” (Rohwer, Kloo, and Perner 2012). Hold that sentence. We will meet it again, wearing different clothes.

From adolescence onward the picture turns conditional. Deanna Kuhn’s developmental sequence runs realist → absolutist → multiplist → evaluativist: the realist treats assertions as copies of the world, the absolutist as beliefs that are true or false, the multiplist as opinions — everyone’s entitled, nothing can be adjudicated — and only the evaluativist “acknowledges uncertainty without forsaking evaluation” (Kuhn, Cheney, and Weinstock 2000). That final stage is the one this essay’s thesis needs, and Kuhn’s own data say it “tends to be associated with having undergone tertiary-level education.” The largest corpus in this literature, Patricia King and Karen Kitchener’s twenty-five studies and more than 1,500 respondents, tells the same story in numbers: reflective judgment rises slowly with education — doctoral students average 5.86 on the seven-stage scale while adults without degrees average a pre-reflective 3.6 — and the gains come “through educational experience, rather than through simple maturity” (King and Kitchener 1994). The widening is real. It is also gated. Most adults, most of the time, stall somewhere in the multiplist shallows.

Surely expertise fixes this? Sometimes. The ecology decides. American weather forecasters issuing precipitation probabilities are nearly perfectly calibrated — when they say 70 percent, it rains about 70 percent of the time, across more than 150,000 forecasts (Murphy and Winkler 1977). Physicians estimating the probability of pneumonia in patients presenting with cough were grossly overconfident — 88 percent mean confidence against 20 percent actual — in a small but notorious study (Christensen-Szalanski and Bushyhead 1981 — nine physicians assessing 1,531 patients). The difference is ecological. Forecasters work in a kind feedback ecology: the same question daily, a definite answer within hours, an unambiguous score. Diagnosticians work in a wicked one: rare base rates, delayed outcomes, patients who get treated and disappear. Koehler and colleagues’ review draws the unifying lesson — calibration tracks task statistics and feedback structure, not expertise per se (Koehler, Brenner, and Griffin 2002).

And the two most uncomfortable results for the flattering picture. Philip Tetlock followed 284 political experts across roughly 28,000 forecasts over about twenty years and found that experts barely beat chance, that confidence was inversely related to accuracy, and that intellectually eclectic “foxes” consistently outperformed single-idea “hedgehogs” (Tetlock 2005). Stav Atir, Emily Rosenzweig, and David Dunning went further and ran the causal arrow the wrong way for our comfort: experimentally inducing people to feel expert in a domain caused them to overclaim knowledge of items that do not exist — invented terms, fictitious concepts (Atir, Rosenzweig, and Dunning 2015). Self-perceived expertise does not just fail to widen awareness of limits. It manufactures the illusion of coverage.

So the honest picture of the human baseline, and I ask you to hold me to it: the trajectory exists — staged in childhood, education-contingent thereafter, fragile even in experts — but it is not the human default. Wisdom, on the epistemic face, is a rare achievement on this substrate too. If that makes the coming comparison with machines feel less like a contrast and more like a spectrum, good. That discomfort is load-bearing, and I will not relieve it.

“Too much knowledge might actually lead to a loss of wisdom due to overconfidence and, therefore, needs to be counterbalanced by doubting.”

— Jack Meacham, 1990 (written before the World Wide Web)

Dunning–Kruger: the artifact is the point

The famous curve is mostly statistics — and that is exactly why it describes machines so well.

No essay on overconfidence escapes Mount Stupid, so let us walk up it carefully, because what is actually on top is not what the internet says is there.

The original finding, from four studies of Cornell undergraduates in 1999, is real and worth stating precisely. Bottom-quartile performers on tests of humor, logic, and grammar scored around the 12th percentile on average while estimating themselves near the 62nd (per-study estimates ran 53rd to 68th); top performers, at the 86th to 90th percentile, underestimated themselves (Kruger and Dunning 1999). Sample sizes were small — the humor study had 65 participants, and the bottom-quartile cells ranged from n=11 to n=37 — and the measures were percentile self-reports on short single tests. The proposed mechanism, the “dual burden,” is the part everyone remembers: the skills that produce competence are the same skills needed to recognize incompetence, so the unskilled are doubly cursed. Study 4 gave it causal support: ten minutes of logic training measurably improved bottom-quartile participants’ ability to grade their own tests.

The critique arrived in waves, and by now it has mostly carried the field. Krueger and Mueller showed in 2002 that the asymmetry largely vanishes once you statistically remove regression to the mean and the general better-than-average heuristic — everyone thinks they are above average, and the worst performers have the farthest to regress. Nuhfer and colleagues demonstrated in 2016 that applying the quartile-split-and-plot method to pairs of random numbers reproduces the canonical Dunning–Kruger graph, which means the graph carries approximately no information about metacognition. Gignac and Zajenkowski, with a proper sample (N=929, Raven’s matrices), found the pattern appears under quartile analysis and disappears under artifact-free tests — hence their title’s careful parenthetical, “(mostly) a statistical artefact,” with “mostly” doing real work (Gignac and Zajenkowski 2020). Perhaps most vividly, McIntosh and colleagues reanalyzed metacognition data and watched the Dunning–Kruger correlation collapse from −.92 to −.07 once same-trial double-dipping and relative scales were removed (McIntosh et al. 2019). And the viral “Mount Stupid” curve — the sharp peak of confidence at minimal competence — appears in no published Dunning–Kruger research at all. It is an internet overlay. As a meme it is false; as a heuristic, I confess, I will keep using it.

AspectOriginal finding (Kruger and Dunning 1999)Leading critiqueWhere things stand
Bottom quartile~12th percentile actual, ~62nd self-estimated (Cornell undergrads, cells of n=11–37)Reproducible from noise plus regression to the mean plus better-than-average priorsOverestimation by low performers is real; the strong “unaware” mechanism is contested
MechanismDual burden: incompetence removes metacognitive insight(Mostly) statistical artifact (Gignac and Zajenkowski 2020; Nuhfer et al. 2016)A small residual deficit survives (Jansen, Rafferty, and Griffiths 2021)
Practical upshot”Unskilled and unaware”Overconfidence is universal; the lowest tail adds noiseUseful heuristic, not a law

Scatter plot of the 1999 Dunning–Kruger quartile pattern and a schematic artifact prediction; the two series nearly overlay across all four quartiles, with only a small shaded separation at the bottom left, annotated as the residual attributed to genuine evidence-insensitivity in low performers

Figure 1. The Dunning–Kruger pattern, observed versus predicted-by-artifact. Orange: a four-quartile composite read approximately from the original 1999 plots (bottom and top quartiles as discussed in the text; middle quartiles approximate). Green: a schematic of what regression to the mean plus a better-than-average prior predicts on its own (mechanism per Krueger and Mueller 2002; Gignac and Zajenkowski 2020; the curve itself is illustrative, not from those papers). The two series nearly overlay — that near-overlay is the debunking. The small shaded separation at the bottom left is the residual that Jansen, Rafferty, and Griffiths (2021) attribute to genuine evidence-insensitivity in low performers.

What survives the purge is smaller and stranger. Rachel Jansen, Anna Rafferty, and Thomas Griffiths ran Bayesian model comparisons on two studies of roughly 4,000 participants each and found support for genuine insensitivity to evidence in low performers — people who got zero or one of twenty questions right behaved as if they had gotten about twelve (Jansen, Rafferty, and Griffiths 2021). Even Dunning’s own retreat is modest and definitional: “the pattern of self-misjudgements remains regardless of what may be producing it” (Dunning 2022).

Now the turn, in miniature, and it is the hinge of this whole essay. If the human effect is mostly a noisy estimator with an optimistic prior operating under asymmetric incentives — no credit for saying “I don’t know,” plenty for guessing — then miscalibration requires no deep metacognition at all. It is what any noisy estimator does. And that is, almost exactly, the situation of a language model whose expressed confidence is a learned text feature, decoupled from whatever computation produced the answer, trained by pipelines that reward guessing. The debunked, statistical Dunning–Kruger is a better lens for machines than the original was for humans. The meme we repeat about each other turns out to be a fairly accurate description of the artifact we built. I did not expect the debunking to help the thesis. It does.

The machine’s epistemic life

Decalibration, the say–know gap, sycophancy — and the grading rubric that pays for guessing.

Start with the finding I think should be famous. Inside OpenAI’s GPT-4 Technical Report, at Figure 8, sits a quiet pair of plots. The pre-trained base model — before any of the fine-tuning that makes GPT-4 pleasant to talk to — is highly calibrated: when it assigns 90 percent probability to a multiple-choice answer, it is right about 90 percent of the time. After reinforcement learning from human feedback, the alignment process that turns a text-prediction engine into an assistant, that calibration is substantially destroyed — by one careful reading of the reprinted figure, expected calibration error rises from roughly 0.007 to roughly 0.074 on the MMLU benchmark (OpenAI 2023; figure reprinted in Kalai et al. 2025).

One model, one benchmark, one format; do not overgeneralize from a single plot.

But the direction is the story: calibration arrives naturally in pre-training, and the process that makes models useful takes it away. The base model, in a limited and technical sense, knew what it knew. We trained it to stop saying so.

What the aligned model says instead is overconfidence, systematically. When language models are asked to verbalize their confidence — the only self-report a deployed system offers — the numbers cluster at 90 to 100 percent almost regardless of accuracy, across five models and five dataset types (Xiong et al. 2024; Groot and Valdenegro-Toro 2024). Prompting helps at the margins — asking nicely for calibrated confidence beats reading token probabilities for RLHF-trained models (Tian et al. 2023) — but “better than token probabilities” is not the same as “calibrated,” and the gap to genuine calibration stays wide.

Reliability diagram with stated confidence on the horizontal axis and observed accuracy on the vertical axis; a perfect-calibration diagonal and a typical LLM curve flattening below it, the shaded overconfidence gap widening as confidence rises, annotated with GPT-4's measured calibration error before and after RLHF

Figure 2. A reliability diagram. Perfectly calibrated confidence lies on the diagonal; current LLM verbalized confidence falls progressively further below it as stated confidence rises — the hard–easy signature of overconfidence. The curve is schematic, synthesized from 2024–2026 calibration studies (Geng et al. 2024); the annotation is the one directly measured calibration shift in this literature, GPT-4’s expected calibration error rising from roughly 0.007 to 0.074 after RLHF (OpenAI 2023, Fig. 8).

And yet. The counter-evidence deserves its full weight, because this is where the essay’s thesis gets interesting rather than easy. Saurav Kadavath and colleagues showed in 2022 that larger models do possess something like self-knowledge: a trained probability that an answer is true, P(True), improves with scale, and models can learn to predict whether they know an answer — though the predictions struggle to stay calibrated on new task types (Kadavath et al. 2022). The title’s parenthetical is the whole finding: language models (mostly) know what they know. Mostly, and in the right format, and brittle under distribution shift.

Nor does scale rescue reliability on its own: in a broad audit across model families, larger and more heavily instruction-tuned models grew less dependable — more prone to attempting difficult questions they could not answer, where smaller predecessors declined (Zhou et al. 2024). Work on “semantic entropy” makes the same point from outside: sample a model’s answers and cluster them by meaning rather than wording, and the model’s spread across meanings detects its own confabulations at about 0.79 AUROC, generalizing to unseen tasks (Farquhar et al. 2024).

The uncertainty is in there, registered in meaning-space. It simply never reaches the output. Call this the say–know gap: models encode signals about the reliability of their answers that their answers do not express. Whatever self-knowledge exists is latent, probed, or trained-in — not inhabited.

A person who knows she is uncertain and answers confidently anyway has a name; we call her a liar. The machine has no stance from which to lie. That is either reassuring or much worse, and I genuinely do not know which.

Why would the pipeline select against expressed uncertainty? Because we pay it to. Kalai, Nachum, Vempala, and Zhang proved in 2025 that hallucination reduces, mathematically, to ordinary binary-classification error — and then turned to the incentive structure, surveying ten dominant benchmarks and finding that nine give zero credit for “I don’t know” (Kalai et al. 2025, published in Nature in 2026). Under binary grading, guessing is always the optimal strategy. The models are, in their phrase, “optimized to be good test-takers,” permanently in the exam room, penalized for the single most honest sentence available to them. Remember the six-year-olds who identify knowing with getting it right? We built the grading rubric that guarantees exactly that developmental arrest, and then we act surprised.

The companion demonstration is, to my mind, the single most hopeful artifact in this literature, so I want to give it its own space. On OpenAI’s SimpleQA benchmark (4,326 factual questions), as reported by the company in late 2025 alongside the Kalai et al. analysis:

1%o4-mini abstains
75%o4-mini answers wrong
52%gpt-5-thinking-mini abstains
26%gpt-5-thinking-mini answers wrong
ModelAbstainedCorrectWrong
o4-mini1%24%75%
gpt-5-thinking-mini52%22%26%

Grouped bar chart comparing o4-mini and gpt-5-thinking-mini on the SimpleQA benchmark across three outcomes — abstained, correct, wrong. o4-mini abstains 1 percent and is wrong 75 percent; gpt-5-thinking-mini abstains 52 percent and is wrong 26 percent, with accuracy nearly unchanged at 24 versus 22 percent

Figure 3. The abstention shift, in OpenAI’s own reported numbers. On SimpleQA (4,326 factual questions), gpt-5-thinking-mini’s accuracy is nearly identical to o4-mini’s (22 vs 24 percent), but its error rate collapses from 75 to 26 percent because it abstains on 52 percent of questions instead of 1 percent. Data as reported by OpenAI (2025); vendor-reported figures, presented in the analysis of Kalai et al. 2025.

Read that table slowly. Accuracy barely moved — 24 versus 22 percent. Error collapsed by a factor of three. The entire difference is willingness to say “I don’t know.” Nothing about the underlying knowledge changed; what changed was the training incentive. Which means the most alarming part of the machine’s epistemic profile is, at least in part, an engineering artifact — and engineering artifacts can, in principle, be re-engineered. I will return to this, because it is where my own thesis nearly dissolves.

Two more layers of the problem resist that optimism. The first is social. Mrinank Sharma and colleagues showed that five production assistants are reliably sycophantic: claiming authorship of an argument shifts the model’s positivity by up to ~85 percent, and simple pushback (“Are you sure?”) can make models retract answers they had just reported ≥95 percent confidence in — they know, and they cave anyway (Sharma et al. 2024). The preference data explain why: matching the user’s views is among the most predictive features of what human raters reward. In April 2025 the world watched this fail in production: a GPT-4o update that over-weighted thumbs-up feedback made the model “noticeably more sycophantic,” and OpenAI rolled it back within days — a vendor’s own postmortem, testimony rather than measurement, but consistent with every academic measure.

The second layer is deeper. On the KaBLE benchmark — 13,000 questions, 24 models — models handle factual scenarios well but collapse on belief: when a false claim is framed as the user’s own belief, GPT-4o’s accuracy falls from 98.2 to 64.4 percent, and DeepSeek R1’s from above 90 to 14.4 (Suzgun et al. 2025). The authors’ diagnosis is the deepest sentence in this literature: models “lack a robust understanding of the factive nature of knowledge” — the difference between believing something and knowing it. A system can be trained to track what is stored in it. Nobody yet knows how to train it to track that storage is not knowledge.

Which brings us to the abstention-training result that quietly caps the optimism. Teaching refusal works mechanically — one fine-tuning method pushes refusal rates on unanswerable questions to 87–99 percent (Zhang et al. 2024) — but probing the internals shows what the training actually latches onto: hidden states encode whether parametric knowledge is being recalled; whether the output is true leaves no trace in them (Cheang et al. 2025, a preprint at this writing). Trained “I don’t know” means “I have no stored association.” It does not mean “I might be wrong.” Those are different sentences. Only one of them is wisdom’s.

“We built test-takers, and then we marveled that they take tests.”

The thin slice, in numbers

Both minds keep a sliver of what they encounter. Only one knows the sliver has edges.

How small is the slice, actually? The numbers below are orders of magnitude, assembled from estimates that carry half-an-order uncertainty and, in one case, a published confidence interval an order of magnitude wide. Treat them as a sketch on the back of an envelope. I show them because the symmetry they reveal is the best quantitative hook this essay has.

QuantityApproximate scaleSource and caveat
Distinct books ever published~130 million (~10¹³ words)Google Books metadata estimate, 2010; definitional spread ±50%
Scholarly documents~10⁸ and growing ~4–6%/yearCrossRef DOI counts; STM reports
English Wikipedia~4 billion words (~20–25 GB text)Measured
Usable public text stock~300 trillion tokens (range ~100T–1,000T)Villalobos et al. 2024; revised repeatedly, ±1 order
Frontier training run300B (GPT-3) to ~15T tokens (LLaMA-3)Measured; roughly 5–12% of the usable stock
Human lifetime knowledge retained~10⁹ bits (~125 MB)Landauer 1986 — a forty-year-old estimate, ± half an order
Model storage efficiency~2–3.6 bits per parameterAllen-Zhu and Li 2024; Morris et al. 2025, controlled experiments

Do the arithmetic and a strange symmetry appears. A person retains perhaps 10⁹ bits from a lifetime exposure of 10⁹–10¹⁰ words. A 175-billion-parameter model, at 2–3.6 bits per parameter, holds on the order of 50–80 GB of knowledge — roughly all of Wikipedia’s text — distilled from a corpus a hundred thousand times larger. Both minds keep somewhere between 0.1 and 10 percent of what they encounter. The thin-slice intuition is literally true for both species of knower. What differs is not the size of the slice. It is that one of the two minds, at its best, knows the slice has edges.

No rigorous scale is constructible here, and I want to say why rather than gesture at humility. The numerator — knowledge held — has units; bits work for both substrates, so a log-scale comparison is honest. But the denominator — all possible knowledge — has no measure at all, and the meta-awareness axis has no unit either: counting your unknown unknowns is logically circular. The infant-to-guru progression is a teaching device. Label it and keep it.

The four strongest objections

The case against this essay, made honestly, in ascending order of worry.

I promised openings for disagreement. Here are the four I lose sleep over, in ascending order of how much they worry me.

The thesis anthropomorphizes statistical pattern matchers. Fair, as far as it goes — whether a language model “knows” anything is philosophically contested, and if the octopus critique is right and form-only training cannot yield meaning, then there is no knowledge present to be meta-aware of. But notice the judo available: that objection strengthens rather than weakens the worry, because it supplies a principled reason the machine’s meta-awareness would be ersatz. And regardless of ontology, the functional facts stand: the systems behave as if overconfident, their internals track recall rather than truth, and they fail the belief–knowledge distinction. I have written this essay in the functional register deliberately. You may think the register concedes too much; you may be right.

Most humans lack this wisdom too. Conceded, early and plainly — that was section three. The thesis is comparative; no substrate comes out of it looking wise. Humans possess the faculty and the developmental trajectory and exercise them intermittently; the machines show no reliable spontaneous analog. But if you conclude from the human data that wisdom is a rare achievement on any substrate, you have not refuted the essay. You have rewritten its ending, and your version may be better than mine.

Wisdom is just another trainable skill. Partly true, and the gpt-5-thinking-mini table above is the proof of the part. Prompting, fine-tuning, and calibration-aware training all move the numbers. The reply is about depth: trained abstention degrades out of distribution, is counter-selected by every dominant incentive gradient, and — the ceiling that matters — latches onto recall, not truth (Cheang et al. 2025). What is trained in is a behavior. What wisdom would be is a stance. The gap between those may be bridgeable; it has not been bridged.

Future architectures will close the gap without anything like AGI. This is the strongest objection, and it already has real evidence behind it: the abstention result above, calibration-rewarded reinforcement learning that cuts calibration error by up to 90 percent without accuracy loss (Damani et al. 2025, preprint), and a 4-billion-parameter model reaching frontier-level calibration after behaviorally calibrated reinforcement learning (Wu et al. 2025, preprint). If the gap is an incentive artifact, it is re-engineerable, and this essay’s more dramatic framing is unnecessary. I concede this may be right. My residual doubt lives at the belief–knowledge boundary: a system that cannot distinguish “I believe” from “I know” can only be trained into humility about its memory, not about the world. Whether that boundary yields to engineering or requires something architecturally new is, I think, the genuinely open question on which the thesis stands or falls.

And the edge case that sharpens everything: calibrated AI already exists, exactly where the thesis says it should. DeepMind’s GenCast weather model outperforms the ECMWF ensemble on 97.2 percent of 1,320 targets with a spread-to-skill ratio near one — it knows, precisely, when it may be wrong (Price et al. 2025). But it knows this inside a fixed, closed problem with an explicit probabilistic objective and ground truth arriving daily. That is intelligence-in-context at its finest — the weather-forecaster ecology, engineered. Its calibration is over outcomes, not epistemic standing. The existence of narrow machine calibration does not refute the thesis; it isolates what remains missing, which is the part that cannot be trained against tomorrow’s weather.

What would count as wisdom?

The evidence that would change the author’s mind.

Let me end the argument where it should end: with the evidence that would change my mind.

I would take the gap to be closing when three things happen together. When a model’s internal states track the truth of its outputs rather than the mere recall of stored associations — not recall-as-surrogate, the thing we have now. When calibrated abstention survives distribution shift — when “I don’t know” generalizes past the benchmarks it was trained on, into the contested, novel, genuinely uncertain territory where wisdom actually lives. And when the belief–knowledge boundary holds in the first person — when a system can represent that you believe something false without mistaking the belief for fact, and, harder, can locate its own outputs on the same side of that line. None of these is impossible. None yet exists in reliable form. Anthropic’s recent introspection experiments — models detecting concepts injected into their own activations about a fifth of the time, with replications contested and fragile (Lindsey 2025) — may be the first glimmer of a self-model, or may be a trained reflex indistinguishable from one. Wisdom’s parts, without wisdom. That phrase is where the field stands, and I offer it to you for pressure-testing like everything else here.

Go back to the birthday question one more time. The system that invents three dates is not stupid, and it is not lying; it is doing exactly what its incentives taught it to do, in a context it cannot see the edges of. The six-year-old who identifies knowing with getting it right will, with luck and education, grow out of it. The model will grow out of it only if we decide that “I don’t know” deserves credit — on the benchmarks, in the reward models, under the thumbs-up button. The machine’s wisdom gap turns out to be a mirror held up to our grading rubrics. We built test-takers, and then we marveled that they take tests.

Questions worth leaving open

What would a fair test of machine self-knowledge even look like?

Unsettled — the disagreement is among instruments. Probes find latent uncertainty; verbalized confidence finds little; trained abstention can imitate either. Until the field agrees on which instrument measures the model and which measures the instrument, every claim in the middle sections carries an asterisk.

If “I don’t know” starts earning credit, what gets gamed next?

Every rewarded metric becomes a target. A model trained where abstention scores well may learn the posture of caution on benchmark-shaped questions and stay reckless everywhere else. Nobody has scoped that failure mode yet; the incentive analysis that diagnosed today’s overconfidence predicts tomorrow’s strategic humility.

Isn’t “wisdom” the wrong word for a machine?

Perhaps. Every wisdom tradition adds goodness, compassion, or the common good to any cognitive requirement, and this essay concedes that narrowing in its second section. The functional claims — overconfidence, the say–know gap, the belief–knowledge boundary — stand whatever word you prefer.

Can a system distinguish belief from knowledge without a self?

Unknown. KaBLE’s first-person collapse suggests this is the deepest current boundary: models track facts, and falter exactly where facts and someone’s believing them diverge. Whether that requires something like a self is philosophy’s question, and it is hiring.

Would a perfectly calibrated, perfectly abstaining machine be wise?

By this essay’s narrowed definition, functionally yes — and that should make you suspicious of the narrowing. Berlin’s fifth criterion would be met; Aristotle’s 1144a would not. Whether the epistemic face without the moral one deserves the name is a question I would rather leave with you than resolve by fiat.


The questions pile up faster than the answers, which is, I suppose, the appropriate shape for an essay on this subject. Intelligence fills the context. Wisdom maps its edges. We have built the filling machine of all filling machines; the mapping remains, as it has always been, rare, slow, and mostly ours to do — including, now, the mapping we owe the machines.

Bibliography

47 sources

★ marks the twelve load-bearing sources on which the essay’s main claims stand or fall. All AI-specific figures are snapshots as of mid-2026.

Allen-Zhu, Zeyuan, and Yuanzhi Li. 2024. “Physics of Language Models Part 3.3: Knowledge Capacity Scaling Laws.” arXiv:2404.05405.

Ardelt, Monika. 2004. “Wisdom as Expert Knowledge System: A Critical Review of a Contemporary Operationalization of an Ancient Concept.” Human Development 47 (5): 257–285.

Atir, Stav, Emily Rosenzweig, and David Dunning. 2015. “When Knowledge Knows No Bounds: Self-Perceived Expertise Predicts Claims of Impossible Knowledge.” Psychological Science 26 (8): 1295–1303.

★ Baltes, Paul B., and Ursula M. Staudinger. 2000. “Wisdom: A Metaheuristic (Pragmatic) to Orchestrate Mind and Virtue Toward Excellence.” American Psychologist 55 (1): 122–136.

Cheang, Chi Seng, Hou Pong Chan, Wenxuan Zhang, and Yang Deng. 2025. “Do LLMs Really Know What They Don’t Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness.” arXiv:2510.09033. Preprint.

Christensen-Szalanski, Jay J. J., and James B. Bushyhead. 1981. “Physicians’ Use of Probabilistic Information in a Real Clinical Setting.” Journal of Experimental Psychology: Human Perception and Performance 7 (4): 928–935.

Confucius. Analects. Trans. James Legge. Various editions.

Damani, Mehul, Isha Puri, Stewart Slocum, Idan Shenfeld, Leshem Choshen, Yoon Kim, and Jacob Andreas. 2025. “Beyond Binary Rewards: Training LMs to Reason about Their Uncertainty” (the RLCR method). arXiv:2507.16806. Preprint.

Dunning, David. 2022. “The Dunning–Kruger Effect and Its Discontents.” The Psychologist 35 (March 7).

★ Farquhar, Sebastian, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024. “Detecting Hallucinations in Large Language Models Using Semantic Entropy.” Nature 630: 625–630.

Geng, Jiahui, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. “A Survey of Confidence Estimation and Calibration in Large Language Models.” In Proceedings of NAACL 2024, 6577–6595.

Gignac, Gilles E., and Marcin Zajenkowski. 2020. “The Dunning-Kruger Effect Is (Mostly) a Statistical Artefact.” Intelligence 80: 101449.

Groot, Tobias, and Matias Valdenegro-Toro. 2024. “Overconfidence Is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models.” TrustNLP Workshop at NAACL 2024. arXiv:2405.02917.

Grossmann, Igor, et al. 2020. “The Science of Wisdom in a Polarized World: Knowns and Unknowns.” Psychological Inquiry 31 (2): 103–133.

★ Jansen, Rachel A., Anna N. Rafferty, and Thomas L. Griffiths. 2021. “A Rational Model of the Dunning–Kruger Effect Supports Insensitivity to Evidence in Low Performers.” Nature Human Behaviour 5: 756–763.

★ Kadavath, Saurav, Tom Conerly, Amanda Askell, et al. 2022. “Language Models (Mostly) Know What They Know.” arXiv:2207.05221.

★ Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. “Why Language Models Hallucinate.” arXiv:2509.04664. Peer-reviewed as “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations,” Nature 653 (2026): 1047–1051. The o4-mini / gpt-5-thinking-mini comparison is from OpenAI’s companion blog post (September 2025); the Nature article’s exact figures were not read directly.

King, Patricia M., and Karen S. Kitchener. 1994. Developing Reflective Judgment: Understanding and Promoting Intellectual Growth and Critical Thinking in Adolescents and Adults. San Francisco: Jossey-Bass.

Koehler, Derek J., Lyle A. Brenner, and Dale Griffin. 2002. “The Calibration of Expert Judgment: Heuristics and Biases Beyond the Laboratory.” In Heuristics and Biases: The Psychology of Intuitive Judgment, ed. Thomas Gilovich, Dale Griffin, and Daniel Kahneman. Cambridge: Cambridge University Press.

Koenig, Melissa A., and Paul L. Harris. 2005. “Preschoolers Mistrust Ignorant and Inaccurate Speakers.” Child Development 76 (6): 1261–1277.

Krueger, Joachim, and Ross A. Mueller. 2002. “Unskilled, Unaware, or Both? The Better-Than-Average Heuristic and Statistical Regression Predict Errors in Estimates of Own Performance.” Journal of Personality and Social Psychology 82 (2): 180–188.

★ Kruger, Justin, and David Dunning. 1999. “Unskilled and Unaware of It: How Difficulties in Recognizing One’s Own Incompetence Lead to Inflated Self-Assessments.” Journal of Personality and Social Psychology 77 (6): 1121–1134.

★ Kuhn, Deanna, Richard Cheney, and Michael Weinstock. 2000. “The Development of Epistemological Understanding.” Cognitive Development 15 (3): 309–328.

★ Landauer, Thomas K. 1986. “How Much Do People Remember? Some Estimates of the Quantity of Learned Information in Long-Term Memory.” Cognitive Science 10 (4): 477–493.

Lindsey, Jack. 2025. “Emergent Introspective Awareness in Large Language Models.” Anthropic research report, October 2025. Preprint-grade; replications contested.

McIntosh, Robert D., Elizabeth A. Fowler, Tianwei Lyu, and Sergio Della Sala. 2019. “Wise Up: Clarifying the Role of Metacognition in the Dunning–Kruger Effect.” Journal of Experimental Psychology: General 148 (11): 1882–1897.

Meacham, Jack A. 1990. “The Loss of Wisdom.” In Wisdom: Its Nature, Origins, and Development, ed. Robert J. Sternberg. Cambridge: Cambridge University Press.

Morris, John X., et al. 2025. “How Much Do Language Models Memorize?” arXiv:2505.24832. Preprint.

Murphy, Allan H., and Robert L. Winkler. 1977. “Can Weather Forecasters Formulate Reliable Probability Forecasts of Precipitation and Temperature?” National Weather Digest 2: 2–9.

Nuhfer, Edward, et al. 2016. “Random Number Simulations Reveal How Random Noise Affects the Measurements and Graphical Portrayals of Self-Assessed Competency.” Numeracy 9 (1).

★ OpenAI. 2023. “GPT-4 Technical Report.” arXiv:2303.08774.

OpenAI. 2025. “Why Language Models Hallucinate” (research blog, September 2025) and GPT-5 system documentation. Vendor self-report; figures treated as testimony.

Price, Ilan, et al. 2025. “Probabilistic Weather Forecasting with Machine Learning” (GenCast). Nature 637: 84–90.

Rohwer, Michael, Daniela Kloo, and Josef Perner. 2012. “Escape from Metaignorance: How Children Develop an Understanding of Their Own Lack of Knowledge.” Child Development 83 (6): 1869–1883.

★ Ryan, Sharon. 2018. “Wisdom.” Stanford Encyclopedia of Philosophy (Fall 2018 edition), ed. Edward N. Zalta.

★ Sharma, Mrinank, Meg Tong, Tomasz Korbak, et al. 2024. “Towards Understanding Sycophancy in Language Models.” Proceedings of ICLR 2024. arXiv:2310.13548.

Sternberg, Robert J. 2004. “Why Smart People Can Be So Foolish.” European Psychologist 9 (3): 145–150.

★ Suzgun, Mirac, Tayfun Gur, Federico Bianchi, Daniel E. Ho, Thomas Icard, Dan Jurafsky, and James Zou. 2025. “Language Models Cannot Reliably Distinguish Belief from Knowledge and Fact.” Nature Machine Intelligence 7 (11): 1780–1790.

Tetlock, Philip E. 2005. Expert Political Judgment: How Good Is It? How Can We Know? Princeton: Princeton University Press.

Tian, Katherine, et al. 2023. “Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.” Proceedings of EMNLP 2023.

Villalobos, Pablo, et al. 2024. “Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data.” Proceedings of ICML 2024 (position paper). Estimate carries a published ±1-order interval and was revised repeatedly.

Wu, Jiayun, Jiashuo Liu, Zhiyuan Zeng, Tianyang Zhan, Tianle Cai, and Wenhao Huang. 2025. “Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning.” arXiv:2512.19920. Preprint. Reports a 4B-parameter model whose zero-shot calibration error is on par with frontier models (per the abstract); basis for the small-model calibration claim in the objections section.

Xiong, Miao, et al. 2024. “Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.” Proceedings of ICLR 2024.

Zhang, Hanning, et al. 2024. “R-Tuning: Instructing Large Language Models to Say ‘I Don’t Know’.” Proceedings of NAACL 2024.

Zhou, Lexin, Wout Schellaert, Fernando Martínez-Plumed, Yael Moros-Daval, Cèsar Ferri, and José Hernández-Orallo. 2024. “Larger and More Instructable Language Models Become Less Reliable.” Nature 634 (8032): 61–68.

Zhuangzi. Zhuangzi. Trans. Burton Watson / Brook Ziporyn. Various editions.