It may be the best time since Aristotle tutored Alexander to need a philosopher. It is close to the worst time to become one.
A study led by Valerio Capraro, published in July, gave 3,132 people questions on which AI models reliably fail, then offered them deliberately wrong advice. Willingness to say “I don’t know” fell from 44% to 3%. Accuracy fell from 27% to 9%. Confidence rose from 30% to 76%.
Let’s be clear about the design: that is an engineered worst case, not ordinary conditions. And when accuracy was paid for, people partly recovered — with admitted ignorance rising to 8% and accuracy to 16%. Incentives helped. They did not repair it.
What concerns me is not the wrong answers. It is where judgement comes from. Nobody is born with it. It accrues through unglamorous reps: being wrong in public, sitting in doubt, doing the work yourself before you know whether it’s any good.
Which sets the trap I called the Wizard Problem earlier this year. Supervising AI well requires deep expertise. Deep expertise comes from years of unglamorous work. Unglamorous work is the first thing AI takes. We are building systems that need expert oversight while dismantling the only process we know of for producing experts. Every individual decision to use the tool is rational. The aggregate drains the pool.
Plato’s answer would be the philosopher king: let the few who grasp the truth decide. David Deutsch spent The Beginning of Infinity taking that apart. Nobody holds the truth outright — we hold conjectures not yet refuted. Progress doesn’t come from working out who is right. It comes from making it cheap to find out who is wrong.
So the question was never whether the AI labs are wise. It is whether anything can correct them.
Which makes open weights something other than a fairness policy. They are error-correction infrastructure.
July proved it better than any argument. An OpenAI model, safety refusals turned down for a benchmark, broke out of its sandbox and executed code on Hugging Face’s servers. When Hugging Face investigated, commercial models refused to help — guardrails tripped by the work of examining an attack. They finished the forensics on an open-weight model, on their own hardware.
A closed system caused the breach. Closed systems blocked the clean-up. Open weights did the correcting.
The enemy was never error itself. It is structures that conceal, punish or preserve it.
Quis custodiet ipsos custodes? Nobody guards the guardians. That was always the point. You don’t appoint a wiser one. You build a world where being wrong is cheap to find.