DR. INFOvsChatGPTUpdated 5 August 2026

ChatGPT for Clinical Use: Strengths, Limits and a Grounded Alternative

ChatGPT is a general-purpose model, and clinicians already use it widely for drafting, summarising and explaining. The question is not whether it is useful but where its output can be relied on. A general model writes from its own parameters rather than from a retrieved, citable source, so it does not by default show you the guideline or paper an answer rests on, it is not a regulated medical device, and its consumer version is hosted outside the EU. Those are the things that decide whether an answer can be checked, approved and used at the point of care. What to look for in any tool used for a clinical question: whether the answer is built from retrieved sources with citations you can open; whether it covers the guidelines in use where you practise; its regulatory status and where it hosts data; and how it behaves on the acuity-sensitive tasks where an under-call is unsafe. On a structured triage benchmark, the tested general-purpose configuration (GPT-5.1, the closest equivalent to ChatGPT) assigned a lower urgency than warranted far more often than a retrieval-grounded system on the most urgent cases [1]. That is one setup on one benchmark, not a verdict on general-purpose models as a class, and triage error varies widely between models [2]. This is not the only signal: an independent study in Nature Medicine found that when patients queried ChatGPT directly for emergency self-triage, it undertriaged 51.6% of cases [3]. DR. INFO is one grounded option: it writes a cited answer from guidelines and literature, is a CE marked medical device under EU MDR 2017/745 hosted in the EU, and has a free tier for clinicians [4]. The honest summary is that a general model is a good drafting and learning aid, and a grounded, cited tool is the safer choice when the answer has to be verified against a source.

DR. INFO and ChatGPT, side by side

Category

DR. INFO Clinical evidence answer engine with citations

ChatGPT General-purpose assistant

Best used for

DR. INFO Answering a specific clinical question with a source you can open

ChatGPT Drafting, summarising, explaining, learning

Answer grounding

DR. INFO Retrieved guidelines, literature and regulatory sources, inline citations with links

ChatGPT The model's own parameters; browsing or citations vary by version and setup

Guideline coverage

DR. INFO Guideline search across specialties, including German AWMF and DGVS

ChatGPT No dedicated guideline retrieval; depends on training data and any browsing

Regulatory status

DR. INFO CE marked medical device under EU MDR 2017/745

ChatGPT Not a medical device; not intended for clinical decision-making

Data residency

DR. INFO Hosted in the EU under GDPR

ChatGPT Consumer version hosted outside the EU

Behaviour on urgent triage

DR. INFO Undertriaged none of the most urgent cases on a structured benchmark

ChatGPT Undertriaged a majority of the most urgent cases on the same benchmark [mtsbench]

Languages

DR. INFO Answers in any language; interface in EN, DE, PT, NL

ChatGPT Answers in many languages

Cost

DR. INFO Free tier for physicians and students

ChatGPT Free tier; paid tiers for more capability

Frequently asked questions

Is ChatGPT safe for clinical use?
It depends on the task. For drafting, summarising and explaining, a general model like ChatGPT is useful. For a clinical question where the answer must be verified, its limits matter: it is not a regulated medical device, it does not by default cite the guideline or paper behind an answer, and on a structured triage benchmark the tested GPT-5.1 configuration assigned a lower urgency than warranted on most of the most urgent cases [1]. That result is specific to the tested setup, not general-purpose models as a class, since triage error varies widely between models [2]. For tasks where the answer must be checked against a source, a retrieval-grounded, cited tool is the more reliable category.
Can doctors use ChatGPT?
Many do, for drafting and learning. What ChatGPT is not is a medical device intended for clinical decision-making, so an answer used at the point of care should be checked against a primary source. Tools like DR. INFO are built for that: the answer arrives with citations you can open, and the tool is CE marked under EU MDR 2017/745 [4].
Does ChatGPT cite its sources?
Not reliably by default. Depending on the version and setup it may browse and link, but a general model primarily writes from its own parameters, so the source behind a statement is not guaranteed. A clinical answer engine like DR. INFO builds the answer from retrieved documents and shows the citation inline, so you can open and verify it.
Is patient data safe in ChatGPT?
The consumer version is hosted outside the EU, which matters for GDPR and for clinical data governance. Avoid entering identifiable patient data into any general consumer tool. DR. INFO is hosted in the EU under GDPR and is a CE marked medical device [4].
What is the difference between ChatGPT and a clinical AI?
ChatGPT is a general assistant; a clinical AI like DR. INFO is built to answer a clinical question from retrieved, citable sources and to cover the guidelines in use. In practice many clinicians use a general model to draft and a grounded tool to get an answer they can check against the source.

ChatGPT is a strong drafting and learning aid. When the answer has to be verified against a source, covered by a guideline, and handled under EU data rules, a retrieval-grounded, CE marked tool like DR. INFO is the safer fit. The triage benchmark behind this comparison is available as a preprint on medRxiv [1], so the method and per-priority results can be checked directly.

References

  1. 1.Ravichandran S, Romano M, Corga da Silva R, Mendes T, Absi N, Isidoro M, Kumar S, van der Heijden M, Gnanapragasam VE. MTS-Bench: A Manchester Triage System Benchmark for Language Model Triage Safety. medRxiv 2026.08.04.26359651; doi: 10.64898/2026.08.04.26359651.
  2. 2.Linzmayer R, Ramaswamy A, Hugo H, Nadkarni G, Elhadad N. Aggregate benchmark scores obscure patient safety implications of errors across frontier language models. medRxiv 2026.03.18.26348695; doi: 10.64898/2026.03.18.26348695. Preprint. Not yet peer reviewed.
  3. 3.Ramaswamy A, Tyagi A, Hugo H, et al., Nadkarni GN. ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine. 2026. doi: 10.1038/s41591-026-04297-7.
  4. 4.Regulation (EU) 2017/745 of the European Parliament and of the Council on medical devices (EU MDR). 2017.