by Denkstrom
All storiesChatGPT Makes 52% Fewer Errors in Crucial Domains

ChatGPT Makes 52% Fewer Errors in Crucial Domains

OpenAI switched ChatGPT to the new GPT-5.5 Instant model. In medical, legal, and financial queries, it produces 52.5% fewer fabricated facts than its predecessor. How this is possible and where limitations remain.

Anyone asking ChatGPT about proper medication dosage has received answers from a different model since May 5. The error rate for such critical queries was previously 18.7 percent. With GPT-5.5 Instant, OpenAI says it has dropped to 8.9 percent. This improvement of more than half stems from altered training procedures designed to make the model more cautious when uncertain. What this number means and where GPT-5.5 Instant still hits limitations becomes clear from the nature of language models.

Why language models sometimes invent facts

Language models like ChatGPT were trained to generate probable text. This makes them good at writing coherently, structuring arguments, and answering questions. But it makes them poor at distinguishing between things they actually learned and things they find statistically plausible.

The phenomenon is called hallucination. The model does not invent answers from malice, but because its architecture lacks a mechanism to distinguish between verified facts and plausible formulations. This becomes particularly problematic with questions where false information has consequences: medication dosages, legal questions, investment decisions. Here, a confidently stated error can have serious consequences.

What OpenAI changed in GPT-5.5 Instant

OpenAI released the full GPT-5.5 model on April 23, 2026. GPT-5.5 Instant is the lighter and faster version that the company rolled out on May 5 as the new standard for all ChatGPT users. According to OpenAI, thousands of hours of external evaluation by physicians, lawyers, and financial professionals were incorporated into training to improve model accuracy in these areas. On the internal benchmark for high-risk queries, the error rate fell from 18.7 to 8.9 percent, a reduction of 52.5 percent.

OpenAI describes the method as a combination of refined training data and a reward model specifically tuned for factual accuracy in high-risk domains. When the model lacks reliable grounding in its training knowledge, it should now more frequently communicate this explicitly instead of constructing a plausible-sounding statement. The model was trained to be more cautious in phrasing when uncertain. OpenAI reports ChatGPT is used by approximately 900 million people weekly, measured in February 2026.

What practically changes for users

For general queries, text drafting, or research assistance, the difference is barely noticeable. The improvements are specifically targeted at three domains where errors carry particular consequences.

In medical queries, the model more frequently issues explicit uncertainty warnings and rarely invents medication names or dosages. When asked whether two active ingredients can be combined, GPT-5.5 Instant should more often explicitly point out that such a question requires a physician or pharmacist, according to OpenAI. In legal questions, errors in citing court rulings and legal paragraphs have declined. In financial topics, the model less frequently invents stock prices or balance sheet figures.

Yet at 8.9 percent error rate in high-risk domains, statistically almost one in ten answers is still wrong. Independent tests after release showed GPT-5.5 Instant still tends to formulate incorrect answers with confidence rather than admit uncertainty. This is especially true for highly specific questions: exact dates, direct quotations, and concrete figures.

What AI language models fundamentally cannot do

Germany's Federal Office for Information Security (BSI) warns in its AI deployment guidelines against using language models as sole information sources for security-critical decisions. This reasoning also applies to GPT-5.5 Instant. A model knows only what was in its training data. It has no knowledge of events after its training cutoff and cannot reliably assess the quality of its own answers.

Medical and legal professional associations have repeatedly emphasized that professional responsibility for AI-assisted recommendations remains with humans. GPT-5.5 Instant improves the quality of support but does not change this fundamental principle.

Until fall: what independent tests will show

GPT-5.5 Instant is the first OpenAI model introduced to the mass market with an explicit promise of reduced hallucinations. The benchmark figures so far come exclusively from OpenAI. Whether 8.9 percent error rate holds up in practice or runs higher will be shown by external evaluations currently underway. Results are not expected before the third quarter of 2026.

For most of the 900 million ChatGPT users, daily life changes little. Texts sound similar, speed is comparable. Whether the model truly responds more cautiously to the next medical query, users will discover themselves in practice. One fundamental principle of handling AI-generated information does not change with GPT-5.5 Instant: critical facts should be verified in primary sources.