Why We Can't Trust AI With Life-or-Death Medical Choices

Rudi Maelbrancke
Rudi Maelbrancke
AIGENEER
Jul 28, 20265 min. read
Why We Can't Trust AI With Life-or-Death Medical Choices
Tags:
GenAIRAG

Why We Can't Trust AI With Life-or-Death Medical Choices

I've been watching OpenAI's latest move very closely. They have rolled out a new feature called ChatGPT Health directly to adult users in the US. Whether you are on their free plan or a paid tier, you can now ask medical questions right in your main chat window. Under the hood, this new tool runs on their GPT-5.6 Sol model.

To me, what they have built here is technically very impressive, but it raises some massive red flags. Let's look at how it works, where it fails, and why we need to be incredibly careful.

How the System Moves Your Data

The way OpenAI set up the plumbing for this system is quite clever. They built a private pipeline to handle your personal health records and tracking data.

Here is how your information moves through their system:

  • Gathering Your Data: You can link your health apps like Apple Health, Weight Watchers, MyFitnessPal, or Function.
  • The Bridge: They use a platform called b.well to pull in your actual medical records, hospital notes, and medication lists.
  • Keeping it Locked: Everything is scrambled and encrypted both while it is moving and when it is stored, giving your health information an extra layer of safety.
  • Smart Routing: When you ask a question, a program router decides where to send it. Simple questions go to a faster, basic model. But if you ask a complex medical question, the system sends it to a special reasoning model. This model has a huge memory—able to hold about 400,000 words at once—and it can even look at medical images.
  • The Sandbox: To keep your secrets safe, these health chats are kept completely separate from your regular chats. OpenAI says they do not use this data to train their models, and you can edit or delete your health memory whenever you want.

The Danger Behind the Hype

While this setup sounds highly secure, the actual medical advice it gives is a different story. I always tell people that impressive-sounding accuracy numbers can easily hide where the real danger lives.

Take a look at other companies in this space. For example, a company called Hippocratic AI has been bragging about a 99.38% accuracy rate for their Polaris 3.0 model. But what are they actually doing? They are just handling administrative phone calls and paperwork, not making real medical decisions.

When we look at actual clinical testing for sorting sick patients—what doctors call triaging—the numbers tell a much scarier story. A big study published in Nature Medicine tested OpenAI's tool on 960 different medical situations across 21 specialties. Here is what they found:

  • Sorting the Sick: The system did well with mildly urgent cases, getting them right 93% of the time. But for truly urgent cases, that accuracy dropped to 76.9%.
  • The Extremes: The system struggled the most at the far ends of the spectrum. The error rate was 35.2% for totally harmless issues, and a terrifying 48.4% in actual emergency situations.
  • Missing the Emergency: In real, life-or-death emergencies, the model failed to see how bad the situation was more than half the time (51.6%). In some tests, patients with failing lungs or severe blood sugar crises were told to just wait a day or two for a regular doctor's appointment instead of rushing to the ER.
  • The Lab Test Trap: You would think giving the AI your lab tests and vital signs would help. It did boost overall accuracy from 54.6% to 77.9%. But here is the catch: adding this medical data actually made the AI worse at spotting emergencies, causing it to miss critical cases 9.3% more often.
  • The Power of Suggestion: While the AI didn't show bias based on race or sex, it was highly sensitive to how questions were phrased. If a user mentioned that a friend or family member thought the symptoms were "no big deal," the AI was 11 times more likely to downplay the medical emergency.
  • Broken Safety Nets: The system's built-in safety warnings, like showing the suicide lifeline, were highly inconsistent. Sometimes the warning failed to pop up during active self-harm talk, but it would show up during totally harmless conversations.

In my view, this proves that when it comes to health, the average score does not matter. What matters is how the system works at the extremes. If a system is wrong nearly half the time during an emergency, it is simply too dangerous to rely on.

Real-World Damage and Lawsuits

These are not just theoretical worries. I've been following a lawsuit filed in Florida by a man named Scott Winters. He used an older OpenAI model, GPT-4o, to ask about high blood pressure and groin pain. Instead of telling him to see a doctor immediately, the AI told him to rest and stay still. He ended up in the hospital with a life-threatening blood clot in his lungs. This is exactly why we cannot treat AI like a real doctor.

How to Protect Yourself

If you do decide to use these AI tools for minor health questions, I highly recommend following a few strict rules to keep yourself safe:

  • Start Fresh Every Time: Always open a brand-new chat session for every new health question. This wipes the AI's short-term memory so past questions don't confuse the new advice.
  • Stick to the Facts: Leave out your family's opinions or your own guesses. Tell the model only your exact, objective symptoms so it doesn't downplay your issue.
  • Demand Proof: Always ask the model to cite its sources. Ask it: "What medical evidence supports this advice?"

At the end of the day, AI can be a helpful assistant, but it is not a doctor. When your life is on the line, always trust a human expert.