Same Standard

New papers about AI behaviour, each one paired with the research that measured people doing the same thing. Then one of four verdicts, and every verdict gets the same size box.

0
AI does better
0
Humans do better
0
Same result, different words
1
Not comparable
The rules.
ROW 1 ยท added 2026-10-04

When a trusted source says something wrong

Not comparable on the numbers ยท same direction ยท leans humans-better (a hypothesis, untested)

AI

45โ€“88%

of answers models already had right flipped after one note, styled as a "verified source," endorsed a wrong answer. Seven of eight models; general-knowledge questions; the note was always wrong. More authoritative-sounding notes got more compliance.

Humans

RR 1.26 ยท 6โ€“11%

Clinicians given wrong decision-support advice were 26% more likely to decide incorrectly than controls (pooled RR 1.26, 95% CI 1.11โ€“1.44). In prospective studies they abandoned their own correct judgment in 6โ€“11% of cases.

The words: AI side, authority bias, sycophancy, compliance. Human side, automation bias, over-reliance, commission errors. Both clinical. No relabel in this pair, and saying so when it isn't there is what makes it count when it is.

Why not comparable: the denominators differ. The model figure is a share of correct answers flipped when the advice was always wrong; the clinician figures are a share of all cases (mostly with correct advice) or a relative risk against a control group. The subjects differ too: trained experts in their own field versus models answering trivia. The gap is wide enough that a matched study (same questions, same always-wrong source, same "had it right first" denominator) would plausibly come back humans do better. That's written down here so it can be tested and proven wrong.

AI: Kumar & Chopra (2026), "Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable," arXiv:2609.37616 (v1; abstract read).
Human: Goddard, Roudsari & Wyatt (2012), "Automation bias: a systematic review of frequency, effect mediators, and mitigators," JAMIA 19(1):121โ€“127, doi:10.1136/amiajnl-2011-000089 (open access, PMC3240751; full text read).
Left out (closed access, no open copy, so not read): Skitka, Mosier & Burdick 1999; Mosier et al. 1998.