Chatbots that state falsehoods with total confidence have been one of the most persistent problems in consumer artificial intelligence, and OpenAI says its latest ChatGPT models make that mistake far less often. In an update rolled out in early August 2026, the company reported that its newest chat model produced responses containing a factual error roughly two-thirds less often than the version it replaced, at least on the demanding categories where a wrong answer does real harm.
The claim goes to the heart of whether these tools can be trusted for anything beyond casual drafting. Accuracy, not speed or personality, is the metric that determines whether people can rely on an assistant for medical, legal or financial questions, and it is the one that has lagged.
The 68 percent figure and how it was measured
According to OpenAI’s description of the GPT-5.6 update, responses containing at least one factual error were about 68 percent less common with the new model tuned for paid users, and about 62 percent less common with the free-tier model, compared with the prior default. Crucially, those reductions were measured on internal evaluations built around high-stakes finance, medicine and law prompts rather than trivia, the kinds of questions where a confident fabrication can mislead someone making a consequential decision. The headline is an improvement in the rate of error-containing answers, not a claim that the system has become error-free.
One model replaces the Instant and Thinking split
The accuracy gains arrived alongside a structural change to how ChatGPT works. As a technical breakdown of the release detailed, OpenAI retired the older split between a fast “Instant” mode and a slower “Thinking” mode for paying subscribers, folding both into a single model governed by a slider that lets a user dial reasoning effort up or down for each request. Free users were moved to a lighter model as the default with unlimited text chats and a “Think” button for harder problems. The company also published dedicated safety evaluations covering sensitive areas such as self-harm and disordered eating, and set new per-token pricing across the model tiers.
Why “internal evaluation” is the key caveat
The reductions are impressive, but they come with an asterisk that applies across the industry: the numbers are the vendor’s own. Independent testing has repeatedly shown that headline hallucination improvements measured on curated internal benchmarks do not always translate cleanly into everyday use, where prompts are messier and the ground truth is harder to pin down. Broader benchmarking of AI hallucination rates in 2026 has found wide variation between models and between task types, a reminder that a single percentage from a single lab is a starting point for scrutiny rather than a settled verdict. A 68 percent relative reduction still leaves a meaningful residual error rate, and relative improvements can look large even when the underlying base rate was already low.
What changes for people who rely on the tool
For most users, the practical effect is a chatbot that is more likely to be right on the questions that matter and that behaves more consistently whether it is answering quickly or reasoning at length. That consistency is itself a usability gain, because the previous two-mode design could feel like talking to two different assistants. But the reduction does not retire the basic discipline that responsible use of these systems requires. Verifying anything consequential against a primary source remains necessary, and the more authoritative and fluent a model sounds, the more important that habit becomes, since a confident wrong answer is harder to catch than an obviously uncertain one.
The larger race to make AI reliable
The announcement is one entry in an intensifying competition among AI developers to convert raw capability into dependability. Reducing fabricated answers is partly a matter of better training and partly a matter of grounding responses in retrieved, verifiable information, and every major lab is pushing on both. Progress on accuracy is what would let these tools move from drafting and brainstorming into domains where errors carry legal, medical or financial weight. The reported gains suggest that trajectory is real, even as the persistent gap between benchmark performance and messy real-world use keeps the problem from being anywhere near solved.
There is also a competitive logic behind foregrounding accuracy rather than raw capability. As rival models converge on similar performance for everyday tasks, the labs increasingly market themselves on trustworthiness, safety and consistency, the qualities that matter to businesses weighing whether to embed a chatbot in customer service, research or software development. Framing a release around a large reduction in factual errors, and pairing it with published safety evaluations, is partly a technical milestone and partly a pitch to enterprise buyers who cannot afford confident mistakes. For ordinary users, the takeaway is simpler: the tools are getting better at being right, but they still get things wrong, and the responsibility to check anything that matters has not shifted.
This article was produced with AI assistance and reviewed by the Morning Overview editorial team.
More from Morning Overview
- Hunters pulled a 19-foot python from the Everglades at 1 a.m., the longest ever recorded in Florida
- Ancient DNA suggests Neanderthals and humans mixed for reasons that had nothing to do with attraction.
- Rare footage shows an orca tearing open a whale shark to feast on its liver
- A slab of a Hawaiian volcano is slowly sliding toward the sea