The Dangerous Delusion of 'Clinician-Level' AI

AI-generated image · US National Wire
OpenAI is betting on a medical revolution by integrating fragmented health data into ChatGPT, but the gap between corporate marketing and clinical safety is a chasm that could cost lives.
OPINION: In the race to dominate the artificial intelligence landscape, OpenAI has decided that the next frontier is the human body. With the rollout of ChatGPT Health to all U.S. users aged 18 and older, the company is attempting to transform a general-purpose chatbot into a personalized medical oracle. But as someone who has covered the intersection of tech and biology for years, I find the company's framing not just optimistic, but dangerously misleading.
Let's start with the rhetoric. During a briefing reported by The Verge, Ashley Alexander, OpenAI’s vice president of health product, claimed that the company’s models are now capable of reasoning at levels that are "better than clinician level." This is the kind of vaporware-adjacent branding that should make every patient and provider shudder. When the actual evidence was requested, Karan Singhal, OpenAI's health lead, walked the claim back, telling The Verge he would "temper" the assertion, though he noted that some individual studies from Stanford and Harvard had pointed in that direction.
This pivot—from "better than clinicians" to "some studies suggest it might be"—is a classic example of the gap between AI marketing and medical reality. In medicine, "reasoning" isn't just about predicting the next most likely token in a sentence; it's about diagnostic accuracy, physical examination, and the nuanced understanding of a patient's history. To suggest a generative LLM can out-reason a trained physician is a claim that lacks rigorous, universal validation.
OpenAI is backing this push with its new GPT-5.6 Sol, which the company describes as its "strongest model yet for health." TechCrunch reports that even the smallest model in the latest release, GPT 5.6-Luna, outperforms GPT 5.5 on "HealthBench," an open-source benchmark developed by OpenAI itself. Relying on an internal benchmark to prove clinical superiority is a circular logic that would never pass a peer-reviewed medical journal. It is a measure of how well the AI mimics the expected answer, not necessarily how safely it treats a patient.
Then there is the data integration. ChatGPT Health allows users to connect a fragmented mosaic of health information: post-visit notes, care team notes, lab results, and medications. It pulls from hospital systems like Oracle Health and Epic, and health platforms such as One Medical, Function Health, and Function. It also integrates consumer-grade data from Apple Health, MyFitnessPal, and Weight Watchers.
On paper, this looks like a convenient health dashboard. In practice, it is a recipe for hallucinated diagnoses. Medical records are often contradictory, incomplete, or coded in shorthand that requires human context to interpret. Feeding this fragmented data into a generative model—which is designed to be fluid and creative—creates a high risk of the AI "connecting dots" that shouldn't be connected, leading to medical advice that is not only wrong but potentially lethal.
We are already seeing the consequences. As reported by both The Verge and TechCrunch, a Florida pastor has sued OpenAI after the chatbot allegedly provided "extremely dangerous medical recommendations." The lawsuit claims these suggestions led the man to delay seeking care for a pulmonary embolism, resulting in a near-fatal experience.
OpenAI’s defense, as noted by TechCrunch, is to point to its terms of service, which state the service is "not intended for use in the diagnosis or treatment of any health condition." The company further maintains that ChatGPT Health "supports, not replaces, professional care."
There is a profound cognitive dissonance in claiming your AI reasons "better than clinician level" while simultaneously hiding behind a legal disclaimer that says the tool shouldn't be used for diagnosis. You cannot market a product as a revolutionary medical tool and then claim it is merely a suggestion engine the moment a patient suffers. If the tool is not for diagnosis, why frame its capabilities in comparison to clinicians?
Beyond the diagnostic risks, there is the privacy nightmare. OpenAI claims that conversations are "encrypted at rest and at transit" and that data connected via the Health feature receives "additional encryption protection." TechCrunch reports that OpenAI says it does not use user data to train its models. However, the sheer volume of sensitive data being aggregated—from lab results to medication lists—creates a honeypot of unprecedented proportions. The idea that a single interface can seamlessly pull from Epic, Oracle Health, and Apple Health suggests a level of data permeability that should terrify anyone concerned with medical privacy.
OpenAI is clearly feeling the pressure from competitors like Google and Anthropic, who are also launching health-related AI features. But the medical field is not a place for "move fast and break things." When you break things in a social media app, you lose a few hours of productivity; when you break things in healthcare, people die.
As TechCrunch reports, users are already leaning into this. OpenAI noted that health-related queries grew from 230 million per week in January to 300 million currently. Furthermore, 70% of these queries happen outside the dedicated health hub, meaning users are treating the general chat as a medical consultant. By integrating health data directly into the main chat experience, OpenAI is encouraging users to bypass the very guardrails the company claims to have in place.
Integrating fragmented health data into a generative LLM is not a medical revolution; it is a dangerous experiment in corporate liability. Until these models can demonstrate a level of reliability that doesn't require a legal disclaimer to protect the company from pulmonary embolism lawsuits, they have no business claiming to reason at a "clinician level."

