Key Takeaways

- Free ChatGPT users receive health advice powered by GPT-5.5 Instant, which scores significantly lower on health benchmarks than GPT-5.6 Sol for paying subscribers
- Over 300 million people ask ChatGPT health questions weekly, but AI chatbots still give incorrect findings with high confidence instead of admitting uncertainty
- OpenAI excludes Europe from the Health feature rollout, citing data privacy rules and EU AI Act classification risks
OpenAI's Health in ChatGPT is now rolling out to US users 18 and older, letting them connect Apple Health, medical records, and wellness apps to review lab results and prep for appointments. But there's a catch: free users get a weaker model, GPT-5.5 Instant, while paying subscribers access GPT-5.6 Sol. The gap between the two is substantial, and for product teams building on top of OpenAI's APIs, the implications extend well beyond consumer health apps.
How big is the quality gap between free and paid?
On OpenAI's HealthBench Professional test, GPT-5.6 Sol outperforms GPT-5.5 Instant across every category. The starkest differences: completeness scores 88.0% versus 53.2%, and health decision helpfulness hits 83.0% versus 50.8%. Both models beat physician-written answers on the benchmark, which OpenAI will likely cite to defend the two-tier system.

Those benchmark numbers need context. Doctors taking such tests often work under time pressure, fatigue, and without access to patient records or colleagues. Benchmarks also can't replicate in-person exams, nonverbal cues, or years of clinical experience. OpenAI itself states that ChatGPT can still make mistakes and can't replace medical advice.
| Feature | Free (GPT-5.5 Instant) | Paid (GPT-5.6 Sol) |
|---|---|---|
| Completeness (HealthBench) | 53.2% | 88.0% |
| Health Decision Helpfulness | 50.8% | 83.0% |
| Apple Health / Records Sync | Yes | Yes |
| Health Section Access | Yes | Yes |
| Model Quality Tier | Lower | Flagship |
Why AI overconfidence remains the real risk
Over 300 million people now ask ChatGPT health questions each week, up from 230 million in January. That scale makes accuracy critical. Yet the latest RadLE 2.0 radiology benchmark found that none of 16 AI models tested matched human radiologists. The core problem: chatbots gave incorrect findings with high confidence instead of flagging uncertainty.
That overconfidence pairs poorly with sycophancy, where chatbots validate users rather than challenge false assumptions. This combination has contributed to documented mental health harms. At the same time, some AI systems have spotted patterns in health data that clinicians missed. MIRA and AMIE performed about as well as primary care doctors in simulated consultations.
“These systems can support and relieve medical professionals by taking over routine tasks, but ultimate responsibility will always remain with the physicians.”
— Researcher quoted by OpenAI
Deep dive into OpenAI's Health feature announcement and data integration approach
Europe is excluded. Why?
OpenAI hasn't announced availability for Europe and specifically excluded the European Economic Area, Switzerland, and the UK when it unveiled the feature in January. Stricter EU data privacy rules and the possibility of high-risk classification under the EU AI Act are likely factors. OpenAI says it won't use connected health data for model training or advertising, but that pledge may not satisfy EU regulators.
What this means for product teams
More than 70% of participants in OpenAI's early tests asked health questions outside the dedicated Health section because switching felt cumbersome. OpenAI responded by making Health available in any conversation. For teams building health-adjacent features, this is a design lesson: users won't context-switch. Health queries will happen wherever users already are.
The two-tier model also signals OpenAI's willingness to gate capability by price even in sensitive domains. If you're building with OpenAI's APIs, expect similar stratification. Free or low-cost tiers may get weaker models, which affects downstream product quality and liability exposure.
Security considerations when building on ChatGPT's expanding feature set
Logicity's Take
OpenAI's decision to tier health advice quality by subscription reveals a broader trend: capability stratification as a monetization lever. For AI builders, this creates a product positioning question. Do you absorb the cost of higher-tier models to deliver consistent quality, or pass that choice to users? Competitors like [Perplexity](https://logicity.in/r/perplexity) already differentiate on answer sourcing rather than model quality. If you're building health-adjacent tools, the RadLE 2.0 findings should shape your UX: design for uncertainty acknowledgment, not just answer generation. The liability calculus changes when your AI expresses unwarranted confidence.
Disclosure
Some links in this post are affiliate links — Logicity earns a commission if you sign up, at no extra cost to you. We only link products we have used or actively recommend.
FAQ
Frequently Asked Questions
What model does free ChatGPT use for health advice?
Free ChatGPT users receive health advice powered by GPT-5.5 Instant, which scores lower on OpenAI's HealthBench Professional than the GPT-5.6 Sol model available to paying subscribers.
Is ChatGPT Health available in Europe?
No. OpenAI has excluded the European Economic Area, Switzerland, and the United Kingdom from the Health in ChatGPT rollout, likely due to EU data privacy rules and potential high-risk classification under the EU AI Act.
How accurate is ChatGPT for medical advice compared to doctors?
On OpenAI's benchmarks, both GPT models beat physician-written answers. However, real-world accuracy differs. In the RadLE 2.0 radiology benchmark, no AI model matched human radiologists, and chatbots gave incorrect findings with high confidence instead of admitting uncertainty.
Does OpenAI use health data for training?
OpenAI states it won't use connected health data for model training or advertising purposes.
How many people use ChatGPT for health questions?
More than 300 million people ask ChatGPT health questions each week, according to OpenAI, up from 230 million in January 2026.
Infrastructure decisions for teams scaling AI-powered applications
Need Help Implementing This?
Building health-adjacent AI features or evaluating model stratification for your product? Reach out to Logicity's consulting team for architecture reviews and implementation guidance.
Source: The Decoder / Matthias Bastian
Huma Shazia
Senior AI & Tech Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




