
AI Did Not Suddenly Get Smarter — It Finally Got Reliable Enough
A popular reading of OpenAI post-training lead Yann Dubois’s remarks: the “overnight leap” many people felt was a reliability threshold being crossed, not a break in the intelligence curve.
Ordinary users, developers, and industry watchers have all described a similar feeling lately: AI suddenly became useful enough to trust with real work — agent coding, messy automation, longer multi-step jobs. The easy story is that models had another “intelligence explosion.” Dubois’s account points somewhere else.
Capability, he argues, has been rising on a continuous curve. What felt like a step change was reliability crossing a threshold — roughly dated to late 2025 in his telling — after which models became stable enough to take over chunks of researchers’ daily work. The outside world then experienced that as overnight strength.
Reliability is an error-rate problem, not a vibes problem
Dubois gives an engineering picture that is easy to remember. For agent-style models, there is some chance of a mistake about every two minutes of running. The longer the task chain, the more those small risks compound — until the whole job collapses. A major R&D goal is to keep driving that per-interval error probability down across architecture, training, and product layers.
That is why reliability is not “raw IQ” alone, and not only ops polish. Until the cumulative error is low enough, AI stays in demo land — impressive in short clips, fragile in real workflows. Once it clears the line, people stop babysitting every step and start handing over work.
Three forces stacked at once
In Dubois’s summary, the recent jump in felt usefulness was not one magic trick. Three things stacked:
- Reliability crossed the threshold — long, multi-step agent work became stable enough to trust.
- Stronger models sped up the researchers themselves — AI writing tools and training the next models, compounding progress.
- Reinforcement learning tools that once worked mainly on clean, verifiable rewards began to generalize into messy real-world tasks.
Together, those shift AI from “usable in demos” toward “usable as a tool.” The intelligence curve did not need to snap; the system finally became trustworthy enough that people noticed.
What to ask before you call it a leap
For companies and teams, the useful reframe is simple. Do not only ask whether the model got smarter on a benchmark. Ask whether error rates on your long workflows fell enough that humans can step back. If a job needs twenty quiet steps without a human rescue, reliability — not a flashier demo — is the gate.
That is the quiet claim underneath the hype cycle: 2026’s “sudden” AI felt less like a break in intelligence, and more like reliability finally catching up to ambition.