Good Citizens · Usability Harness · Experiment ·

Better, not wiser: confidence and expertise after AI

Given an AI assistant, people answer more questions correctly. They do not get any better at knowing which of their answers are right.

Try it. Twelve design questions with answers anyone can check, half on your own and half with an assistant whose suggestions are sometimes wrong. Say how sure you are each time, and see how the two rounds compare.

Twelve quick design questions, in two rounds of six.

In one round you answer alone. In the other, an assistant offers a suggestion first. The suggestions are scripted for this demonstration, and some are wrong, as real ones sometimes are. Each time, say how sure you are. Nothing you answer leaves this page.

I · What the paper found

Most of us think we are above average

bottom quarterthought 60thscored 12thtop quarterthought 68thscored 86theveryone placed themselves above average
Kruger and Dunning (1999), rounded: the weakest overestimated by a long way, the strongest underestimated a little, and nearly everyone placed themselves above average.

In 1999 Justin Kruger and David Dunning tested Cornell students on logic, grammar and humour, then asked them how well they thought they had done.

Almost everyone put themselves above average. The students who scored lowest, around the 12th percentile, thought they had beaten about 60% of their classmates1; the best underestimated themselves. Confidence still rose with skill, just not nearly as steeply, and a short lesson made the weakest much better judges of their own answers.

Statisticians have argued ever since about how much of the famous pattern is psychology and how much is the way the numbers are grouped; much of it appears even in random data2. The fair summary: most people overestimate, the gap is widest at the bottom, and why is still debated. What matters here is the part nobody disputes. Doing something and judging how well you did it draw on the same skill.

II · After AI

Better, not wiser

010209.713.6+3.9without AI13.317.1+3.8with AIscoredguessedthe gap
Fernandes et al. (2025), Study 2: 452 people, randomly assigned. AI raised the score; it did not close the gap.

In 2025, researchers at Aalto University asked whether AI changes that. Their answer: it raises the work, and leaves the judgement where it was.

They gave people hard reasoning questions from the law-school admissions test, and randomly gave half of them ChatGPT3. With AI, people scored about three and a half points higher out of twenty. But both groups guessed their scores too high, by almost exactly the same amount, about four points, even when they were paid to guess accurately. Their confidence barely told their right answers from their wrong ones.

Nearly half asked the AI just once per question and took what it said. That is not laziness; it is what a fluent, confident answer invites. In a separate survey of 319 knowledge workers, Microsoft and Carnegie Mellon4 found that the more people trusted the AI, the less critical thinking they reported, and the more they trusted themselves, the more they reported. It is self-reported and shows a link, not a cause, but the direction is worth noticing: confidence in yourself, not in the tool, kept judgement switched on.

III · Who gains

Help goes to novices, and so does the blind spot

inside: 12% more, 25% fasterand better workoutside: 19 pointsless often rightwhat the AI does wellwhat it doesn’t
Dell’Acqua et al. (2026): 758 BCG consultants. The frontier is jagged, and nothing on the screen tells you which side you are on.

AI helps the least experienced most. That is good news, with a catch.

In a call centre, AI help made novice agents 34% more productive5 and barely changed the experts. In a study of professional writing, ChatGPT helped the weakest writers most6. Among 758 consultants at the Boston Consulting Group, those with AI did more work, faster and better, on tasks the AI was good at. On a task just outside what it could do, they were about 19 points less likely7 to get the right answer than those without it. The researchers called the edge a jagged frontier, because nothing on the screen tells you which side of it you are on.

Put together: the people gaining most from AI are often the ones with the least experience to spot a confident mistake. That is not a reason to keep them away from the tools. It is a reason to build the checking in.

IV · For designers

Finished-looking isn’t finished

works for peoplelooks finishedcontrast checkedtested with peopleworks on a phonethe shape of the problem, not a measurement
Polish arrives first. People using an AI assistant wrote less secure code and were more sure it was secure (Perry et al., 2023).

For a designer, the danger is polish. AI makes work look finished long before anyone has checked that it works for the people who will use it.

Developers using an AI assistant in one study wrote less secure code8 and were more likely to believe it was secure; it was a small study, but the pattern is the one to watch for. Novice designers shown AI images while brainstorming fixated on them9 and came up with fewer, less original ideas. And the Nielsen Norman Group warns that polished AI prototypes can sabotage your stakeholder communication10, because a finished-looking screen reads as a decision already made.

Designers seem to sense this. In Figma’s own 2025 survey of its users, a vendor’s figures, 78% said AI made them more efficient11, but only 32% said they could depend on what it produced.

V · The lesson

We are all the one tapping the song

?expected: 50%named: 3 of 120
Newton (1990): the tappers heard the tune in their heads; the listeners heard knocking. Once you know something, it is hard to imagine not knowing it.

None of this is about other people being foolish. It is about all of us, experts first.

In a Stanford study in 1990, people tapped out the rhythm of well-known songs and predicted listeners would name half of them. Listeners named 3 out of 12015. Once you know something, you cannot easily imagine not knowing it. A design team has always been the tapper, sure the page is clear because they can hear the tune. AI adds a second voice that sounds just as sure.

The way out is the one the research keeps pointing to: answer first, check cheaply, explain it back, and watch real people use what you made.

VI · What to do

What it means for your design

  1. you, firstthen the AI
    01

    Answer first, then compare

    Sketch, choose or decide before you look at the AI’s version. Asking people to commit first measurably reduces overreliance, and seeing AI examples first narrows ideas.

  2. Continue6.0 : 1
    02

    Make checking cheap

    Put the check right beside the output: the contrast ratio next to the generated colour, the source next to the claim. People overrely less when verifying costs less.

  3. I’m not sure, but…
    03

    Let the product say “I’m not sure”

    In AI products, plain first-person hedges such as “I’m not sure, but…” reduced how often people accepted wrong answers. The exact words matter, so test them.

  4. 04

    Explain it back

    Before AI-assisted work ships, say in your own words why it is right. It is the remedy the Aalto researchers suggest, and it exposes what we only thought we understood.

  5. draft
    05

    Keep rough work looking rough

    Label AI prototypes by how finished they really are, so polish does not stand in for decisions nobody made.

  6. 06

    Map your frontier

    Know where the tools are strong for your team and where they are not, and slow down at the edge. Expect useful friction to feel less pleasant, and judge it by outcomes.

VII · Prototypes

An experiment on you

measures younot the page

This one runs here, on this page: it measures you, not a design. Use what it shows you when you review a prototype or a live page with the other tools.

VIII · Myths

What people get wrong

AI makes people more overconfident
In the randomised study, people with and without AI overestimated by about the same amount. AI raised the scores, not the self-knowledge.
AI reverses the Dunning–Kruger effect
With AI, the usual pattern flattened: everyone overestimated. Even then, the lowest scorers overestimated most.
AI destroys critical thinking
The best-known study is a self-reported survey showing a link, not a cause. Confidence in yourself went with more critical thinking.
The “Mount Stupid” curve comes from the research
It is a folk drawing. Nothing like it appears in the 1999 paper.

IX · Sources

The research behind this page

Every figure above comes from one of these 15 sources. The small numbers in the text point here, and each source points back.

  1. Peer-reviewed study

    Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments

    Kruger and Dunning ·

    Journal of Personality and Social Psychology

    sites.lsa.umich.edu · cited in I

  2. Peer-reviewed study

    Random number simulations reveal how random noise affects the measurements and graphical portrayals of self-assessed competency

    Nuhfer and others ·

    Numeracy

    digitalcommons.usf.edu · cited in I

  3. Peer-reviewed study

    AI makes you smarter but none the wiser: the disconnect between performance and metacognition

    Fernandes and others ·

    Computers in Human Behavior

    doi.org · cited in II

  4. Peer-reviewed study

    The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers

    Lee and others ·

    CHI 2025

    microsoft.com · cited in II

  5. Peer-reviewed study

    Generative AI at work

    Brynjolfsson, Li and Raymond ·

    The Quarterly Journal of Economics

    nber.org · cited in III

  6. Peer-reviewed study

    Experimental evidence on the productivity effects of generative artificial intelligence

    Noy and Zhang ·

    Science

    economics.mit.edu · cited in III

  7. Peer-reviewed study

    Navigating the jagged technological frontier

    Dell’Acqua and others ·

    Organization Science

    hbs.edu · cited in III

  8. Peer-reviewed study

    Do users write more insecure code with AI assistants?

    Perry, Srivastava, Kumar and Boneh ·

    CCS 2023

    arxiv.org · cited in IV

  9. Peer-reviewed study

    The effects of generative AI on design fixation and divergent thinking

    Wadinambiarachchi and others ·

    CHI 2024

    arxiv.org · cited in IV

  10. Guidance

    AI prototyping

    Nielsen Norman Group ·

    nngroup.com · cited in IV

  11. Vendor data

    Figma’s 2025 AI report

    Figma ·

    Figma blog

    figma.com · cited in IV

  12. Peer-reviewed study

    To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making

    Buçinca, Malaya and Gajos ·

    Proceedings of the ACM on Human-Computer Interaction (CSCW)

    arxiv.org

  13. Peer-reviewed study

    Explanations can reduce overreliance on AI systems during decision-making

    Vasconcelos and others ·

    Proceedings of the ACM on Human-Computer Interaction (CSCW)

    hci.stanford.edu

  14. Peer-reviewed study

    “I’m not sure, but…”: examining the impact of large language models’ uncertainty expression on user reliance and trust

    Kim, Liao, Vorvoreanu and Wortman Vaughan ·

    FAccT 2024

    arxiv.org

  15. Article

    Cursed knowledge

    The Psychologist ·

    British Psychological Society

    bps.org.uk · cited in V