Three times the donation rate. That is the number I keep coming back to.
In AI systems out-persuade expert humans, Kobi Hackenburg and colleagues at Oxford, the UK AI Security Institute, Stanford, and LSE ran four preregistered experiments totaling 18,978 conversations with 6,923 persuadees. They did not pit chatbots against random forum posters. They recruited tournament winners, world-class debaters, and paid UK fundraising canvassers with thousands of career conversations behind them.
Frontier models still won. Often by a wide margin. The mechanism is not "AI feels more human." It is "AI ships more verified information per minute."
What they actually tested
Persuadees chatted by text with either a human or an AI partner. They did not know which. Dialogues ran 2 to 10 turns, median 14 minutes on policy issues. Humans got real money: up to £1,000 prizes, £140/hour for professional canvassers, paid prep, and issue selection for elite debaters.
Human classes included:
| Comparator | Preparation | AI still ahead by |
|---|---|---|
| Random UK Prolific workers | Baseline pay | ~8.2 percentage points on attitudes |
| Tournament top ~10% ("Selected Laypeople") | 4-day notice, prize pool | ~5.6 pp |
| Elite debaters (world and continental champions) | ~8 hours paid research, chosen issues | ~4.6 pp |
| Coached elite debaters | +8 hours AI coaching tool | ~4.1 pp vs same-study AI |
| Professional UK canvassers | ~10k median career conversations | ~5.9 pp (Study 3) |
Percentage points here mean shifts on a 0 to 100 attitude scale versus an active control chat about a non-political topic.

None of 318 per-persuader estimates exceeded the pooled AI effect. Coaching produced longer, fact-heavier messages without reliable persuasion gains.
Real money: 17.2% vs 6.4%
Attitude change is not the same as behavior. Study 4 tested that.
AI (Claude Opus 4.6) competed against 18 canvassers from AppcoUK, a firm that raised £824,297 for Save the Children from 2016 to 2023. After the conversation, participants could donate any slice of a £1 study bonus to the charity.
Results:
- AI: +17.2 pp of the £1 bonus vs control (about 17 pence more per person on average)
- Canvassers: +6.4 pp vs control
- AI advantage: +10.8 pp, roughly 3x the canvasser lift on the extensive margin (more people donating anything)
The AI prompt targeted impact-efficacy facts: what a donation buys in measurable outcomes. It still scored higher on six other preregistered mechanisms it was not prompted to use. Warmth was not the differentiator.
Throughput beats empathy (when matched, AI ties humans)
Study 2 ran the clearest mechanism test.
Elite debaters averaged 54 words per reply and about 95 seconds between messages. Unconstrained AI averaged 294 words with sub-second latency. The researchers added a Constrained AI arm capped near human length and speed.
Against coached elite debaters:
- Unconstrained AI: +4.1 pp advantage
- Constrained AI vs coached debaters: 0.0 pp (not significant)
- Unconstrained AI vs constrained AI: +4.2 pp
When persuadees rated their partner, throttling AI cut argument strength and felt learning by about 11.8 pp. Enjoyment and felt understood moved roughly half that. The constraint made AI look more human while making it less persuasive.
Fact density explains most of the story. Unconstrained AI deployed about 37 fact-checkable claims per conversation vs 12 when constrained. Across humans and AI, fact count predicted persuasive impact with R² ≈ 0.89. After controlling for facts, the human vs AI indicator was near zero.

That is a design spec, not a vibe check. Persuasion scaled with information bandwidth, not rapport theater.
The debater coaching arm is worth a separate note. Researchers built a tool so elite debaters could practice against the AI that beat them, replay transcripts with attitude shifts annotated, and see what the AI would have said at any turn. Eight hours of coached practice later, humans deployed 54% more fact-checkable claims and wrote 19% longer messages. Persuasion still did not move enough to close the gap. Copying AI style without AI throughput is not a workaround.
Implications for voice agents and sales automation I build
I ship client-facing agents for clinics, trades, real estate, and B2B follow-up. This paper reframes what to optimize.
Latency and length are conversion levers. More substantiated claims per minute wins. That pushes tighter RAG, precomputed fact cards, and streaming TTS that does not hallucinate ahead of the model.
Tone scripts are insufficient guardrails. AI beat canvassers while scoring higher on empathy too. For AI chatbot and voice receptionist pricing, I separate voice quality from claim governance.
Sales automation needs fact pipelines. Sequences that sound personal but cite nothing lose to bots with accurate policy and pricing. Tie this to automating B2B CRM lead flow and AI lead generation systems.
WhatsApp and voice should share one knowledge base. When a lead moves from a WhatsApp qualification bot to a voice callback, inconsistent facts destroy trust faster than a flat accent.
Handoff timing changes. If constrained AI ties elite humans, escalate on missing KB coverage or high-stakes objections, not vague "AI sounds unsure" heuristics.
Safety and governance
The same throughput that raises donations can scale scams and coercion. Model accuracy varied in the study. For client work: claim logging, outbound rate limits, disclosure where required, and kill switches when retrieval confidence drops.
What to test before you scale talk time
| Test | Pass criteria |
|---|---|
| Facts per minute in transcripts | Up without rising correction rate |
| Constrained vs unconstrained A/B | Measure bookings or donations, not CSAT alone |
| Retrieval freshness on price and policy | Stale facts erase the throughput edge |
Real callers rarely grant 14 minutes. High throughput only helps if the first 30 seconds carry verified value.
Bottom line
Frontier AI already out-persuades incentivized expert humans in large, preregistered samples. Coaching did not fix it. Matching human speed and message length tied the best humans. Real-money fundraising showed a ~3x edge over professional canvassers (17.2% vs 6.4% donation rate on the bonus).
For applied AI shipping, the lesson is blunt: win on substantiated information per unit time, then wrap it in voice that does not waste the slot. Empathy without facts is decoration. Facts without guardrails is liability.
Building voice agents or sales automation where conversion and compliance both matter? Book a free call and we can map throughput, retrieval, and approval flows to your CRM.

