AI chatbot for health questions: how far you can trust the answer
The essentials at a glance
An AI chatbot classifies; it does not recognize anything. It can explain what a symptom may indicate, which questions make sense to ask in a doctor’s office, and how a laboratory value is obtained. It cannot know what you have: it lacks your medical history, your examination, and every measured value. The boundary lies where an explanation becomes a diagnosis.
This is not a warning about the technology, but a description of how it is built. A language model generates text that fits your question. Whether that text is correct is a separate question, and it is not answered while the text is being generated. Anyone who knows this can use a chat window sensibly. Anyone who does not will mistake a fluent answer for a verified one.
Chapters 1 and 2 explain what actually happens when an answer is generated. Chapter 4 draws the line between guidance and self-diagnosis. Chapter 6 reveals where AI itself provides advice in this shop and what it cannot do. Chapter 7 identifies the cases in which a chat window is no longer the right place to turn.
What to expect in this article
1. What a chatbot does when you describe a symptom
2. Why a wrong answer sounds just as confident as a correct one
3. Three statements about AI answers that are simply not true
4. Guidance or self-diagnosis: the line in detail
5. What can be measured cannot be replaced by a language model
6. The in-house AI in the shop: what Cloé can and cannot do
7. When the question belongs in a doctor’s office rather than in the chat window
8. How to ask questions without going down the wrong path
9. Who benefits from the chat window and who does not
10. What you can do differently next time
Frequently asked questions
Sources
What a chatbot does when you describe a symptom
You type three lines about fatigue, pressure in your upper abdomen, and poor sleep. Two seconds later, there is a neatly structured text with possible causes, recommendations, and a friendly closing sentence. What happened in those two seconds had nothing to do with examining you.
A large language model calculates which word is most likely to follow the previous one. It has read a great deal of text, including medical literature, forum posts, advertising copy, and advice pages. From this, it forms an answer that fits your question. This process does not include a verification step that checks the statement against a source.
An image for this: The model works like an exceptionally well-read night porter. He has read every magazine in the building, knows the familiar sentences on every topic, and formulates them calmly and clearly. He has never examined you, and he will never say that he does not know you.
Key message
A language model assesses which words fit together. It does not check whether a statement is true.
How widespread use has become by now
You are not alone in having this question. According to a representative survey of 1,145 people aged 16 and over, 45 percent of people in Germany use AI chatbots for questions about symptoms and health; of those who do, 55 percent trust the answers (Bitkom e. V., 2025). This means that for almost one in two people, the chat window is the first place where a complaint is put into words.
That is understandable. A chatbot is available at three in the morning, it does not interrupt you, and it requires no preparation. Those three qualities make it a good starting point and a poor endpoint.
Why a false answer sounds as confident as a correct one
With a human, you recognize uncertainty from hesitation. A language model does not hesitate. It formulates an invented study citation in the same calm tone as a correct one. The technical term for this is hallucination, and it describes not an exceptional case but a property of how it is designed.
Documented source
“If the AI has no data that matches the user’s question, it may happen that it invents the answer. This is also called hallucination. The answer may sound plausible, but it can be entirely made up. This also applies to citations.”
Health Knowledge Foundation
Searching for health information with AI chatbots, as of 2025
The sentence about the citations is the most uncomfortable one. An invented claim can be checked if you are skeptical. An invented source cited after the claim removes precisely that skepticism because it looks like evidence. If you do not open the source, you consider the answer doubly verified.
The second error occurs during reading
In 2024, the World Health Organization published guidance on large multimodal models in healthcare. It warns about statements that are “false, inaccurate, biased, or incomplete” and can harm people, and it identifies a second mechanism: “automation bias,” meaning the tendency of healthcare professionals and patients to overlook errors in machine-generated answers (World Health Organization, 2024).
Translated, this means: The text does not seem convincing because it is good, but because it comes from a machine and expresses no discernible doubts. With a human, you would ask where they got it from. When the answer has no face behind it, you do not ask.
Three statements about AI answers that are not true as stated
The following three assumptions regularly come up in conversations about chatbots. Each is understandable, and each leads one astray in a different way.
Common assumption
If the chatbot cites sources, they have been checked.
What actually applies
Sources can be invented just like the sentence before them. The Stiftung Gesundheitswissen explicitly states this in 2025. A source becomes verifiable only when you open it and find the cited sentence there.
Common assumption
A detailed answer is a reliable answer.
What actually applies
Length comes from the amount of matching text patterns, not from certainty. A model often has more to say about a rare set of symptoms than a common one, because rare conditions are described in greater detail in the specialist literature. Length therefore says nothing about accuracy.
Common assumption
If I enter my values, the AI incorporates them correctly.
What actually applies
A model knows reference ranges as text, not the ones from your laboratory. Reference ranges depend on the laboratory, method, and age. A value classified as abnormal in an answer may fall within the stated range in your test result.
Guidance or self-diagnosis: the line in detail
The boundary is not a feeling; you can tell it from the question you ask and what you do with the answer. The same conversation can begin on one side of the line and end on the other without the tone changing.
| Criterion | Guidance — holds up | Self-diagnosis — does not hold up |
|---|---|---|
| Your question | What could this symptom indicate? What is usually examined in this case? | What do I have? Is it dangerous? Do I need to see a doctor or not? |
| What comes back | Terms, connections, possible examination methods, vocabulary for the next conversation | An assessment that looks like a finding but is not based on any examination |
| Data basis | General knowledge about medical conditions, independent of you | Your typed description, without an examination, medical history, or measurement |
| What you do afterward | You write down three questions for the appointment and go | You postpone the appointment, stop taking something, or start taking something on your own |
| How you can tell | You have more questions than before, but better ones | You feel relief or fear, but not a single new question |
| The honest test | You could show the answer to a doctor without having to justify yourself | You would rather not mention where you got it in the conversation |
The last line is the most useful. If you would not openly take an answer into the consultation room, you have long since treated it as a substitute rather than as preparation.
What is measured cannot be replaced by a language model
There is a difference between a statement about ferritin and your ferritin level. The first sentence appears in thousands of texts from which a model can draw. The second exists only once someone has drawn blood and a laboratory has analyzed it.
A language model knows sentences about your body. It does not know a single value from it.
That is why, if a deficiency is suspected, the path leads through a measurement rather than a conversation with a machine. We have compiled which symptoms are suitable for measurement at all in our article on symptoms of nutrient deficiency. What a blood count covers is explained in the article what is examined in a complete blood count.
The reverse is equally true, and it is the more uncomfortable part. A measured value also does not answer the question you came with. It is a data point, not a verdict. Which diseases a blood count can and cannot indicate is explained in the article which diseases can be detected in a complete blood count; the same logic applies to the gut, as the article testing the microbiome shows.
An example from our own product range to make the scale tangible: the VitalCheck | Nährstoff & Mineralstoff-Test Complete by mybody®x (MYBODY Lab GmbH) measures 18 biomarkers from capillary blood—that is, from a few drops taken from the fingertip—and costs €169.00 (as of September 3, 2026; subject to change). It provides numerical values for vitamins, minerals, and iron status. It does not tell you what is causing your fatigue, and it does not replace a medical examination. A numerical value is a starting point, not a result.
The difference between a chatbot and a laboratory is therefore not one of reliability, but of responsibility. One generates text, the other generates numbers. In both cases, interpretation remains the responsibility of a trained person.
The shop’s own AI: what Cloé can and cannot do
An article about the limits of AI that conceals its own AI would be worthless. So, in order: this shop uses an AI consultation tool. It is called Cloé, appears as a question bar between the products, and answers questions about the health tests.
Key message
Cloé provides advice on tests and products. It does not diagnose, and it does not read test results.
Since 2024, the European AI Act has required people to be informed when they are speaking with a machine. Article 50(1) puts it this way: Providers shall ensure that AI systems intended for direct interaction with natural persons “are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system” (Regulation (EU) 2024/1689, Article 50(1)).
Cloé is identified as an AI in the consultation; this identification cannot be disabled. It works with a stored knowledge base from its own product range, not with the open internet. What it recommends are tests in the areas of DNA, blood, and the gut.
What Cloé explicitly does not do
Four things the consultation cannot do
No diagnosis
It provides context and narrows things down. It does not tell you what you have, and it rules nothing out.
No reading of test results
It does not accept laboratory results or interpret values. A result belongs with the doctor who ordered it.
No view beyond its own fence
Its knowledge base ends with its own product range. It is a sales consultation with mandatory professional information, not an independent source of information about the market.
No children, no desire to have children
On these two topics, it refers you to a medical practice and deliberately does not recommend a substitute product.
This means that the same applies to Cloé as to every other language model discussed in the preceding chapters. It is faster than emailing customer service and worse than speaking with a professional. You cannot have both at once. Those who use it for the former save time; those who use it for the latter lose something more important.
When the question belongs in a medical practice, not in the chat window
There are complaints where every minute spent researching is one minute too many. The following list is not a substitute for an assessment; it is a stopping condition: if one line applies, the search ends in the chat window and the next step is a medical practice or an emergency department.
Where self-research ends
Sudden, severe, different from usual
Pain that comes out of nowhere and feels unlike anything you have experienced before. Especially in the chest, head, or abdomen.
Blood where it does not belong
In your stool, urine, vomit, or when coughing. Even if it happens only once and even if it stops afterward.
Shortness of breath or pressure in the chest
This is not the time to research or wait. Call 112.
Unintentional weight loss or persistent fever
If your weight is dropping without trying or a fever persists for days, it needs to be investigated rather than interpreted.
Children, pregnancy, ongoing treatment
These three situations have the greatest gap between general information and your situation. Never stop taking a medication based on a chat response.
The doctor is not presented anywhere in this article as the chatbot’s opponent. The doctor is the next point of contact, and a well-prepared conversation is worth more than a perfectly worded answer at midnight.
How to ask without getting lost
The following four steps do not change the technology, but your role in using it. They are intended as preparation for a medical consultation and explicitly not as self-diagnosis.
Ask about terms, not diagnoses
Instead of “What do I have?” ask “Which technical terms describe this symptom, and what is usually examined in this context?” This gives you vocabulary instead of a diagnosis. Source: Stiftung Gesundheitswissen, as of 2025.
Ask about the uncertainty too
Add to each question what is uncertain about the answer and what argues against it. Models rarely provide this part on their own because it does not look like an answer. Source: World Health Organization, 2024.
Check every source yourself
Open the cited source and look for the quoted sentence there. If you cannot find it, the claim does not hold. Fabricated citations are a documented phenomenon, not an isolated mistake. Source: Stiftung Gesundheitswissen, as of 2025.
Turning the answer into three questions
Write down three questions from the chat history for your appointment. That is the real benefit: a better first minute in the consultation room. Source: mybody®x Editorial Team, as of September 3, 2026.
At a glance
Ask about terms, request the uncertainty as well, check every source yourself, and take three questions to your appointment.
Who benefits from the chat window and who does not
The question is not whether a chatbot is good or bad. The question is what situation you are currently in.
Useful for you if …
… you have a medical report in front of you and do not know the technical terms in it. A model explains vocabulary well because it does not need any knowledge about you to do so.
… you have an appointment and want to prepare. Organizing questions is a language task, and language tasks are what this technology is good at.
… you want to know which examinations are usual for a set of symptoms, so that you are not at a loss during the conversation.
Probably not if …
… you are actually looking for reassurance. Then you read the answer in the way you need it to read, and the appointment is delayed by weeks.
… you are making a decision about a medication. Stopping it, changing the dose, or starting something new belongs in the hands of a doctor.
… one of the lines in the chapter on warning signs applies. In that case, the answer on the screen is the wrong thing to focus on.
What you can do differently next time
The common recommendation is to avoid chatbots for health questions. It misses the reality: 45 percent of people in Germany use them for this anyway (Bitkom e. V., 2025), and a ban would not change that. The more useful question is what task you assign to the window.
There is one observation from operating the advice service itself that helps here. An AI system does not stay silent when it is missing something. It keeps answering, only without a basis, and no one can tell from the answer. With Cloé, this once led to her giving detailed advice without recommending a single product, because her product list no longer matched the shop. The text was flawless; the foundation was gone.
Applied to your health question, this means: The risk is not the wrong answer, but the missing gap in the text. A person says, “I don’t know.” A model fills the gap.
So next time, take the two minutes and write down three questions from the chat history for your appointment. That is the whole trick. The advice in the shop costs you nothing, and you explicitly will not receive a diagnosis there.
Frequently asked questions
Can an AI chatbot make a diagnosis?
No. A chatbot generates text based on probabilities; it does not examine you and knows none of your measurements. In 2025, the Stiftung Gesundheitswissen stated that although you will always receive an answer to a question about a medical condition, it is not necessarily correct, and that a chatbot therefore cannot replace a visit to the doctor. What it can do is help you prepare: explain terms, identify possible examination methods, and help you ask the right questions in the consultation room.
How can I tell that an AI answer is fabricated?
You cannot tell from the text itself, and that is the core of the problem. An invented answer can be phrased just as fluently as a correct one. The only reliable way is to check the source: open the cited source and look for the quoted sentence there. If you cannot find it or the page does not exist at all, the statement is invalid. In 2025, the Stiftung Gesundheitswissen explicitly pointed out that source references can also be completely fabricated.
May I enter my laboratory results into a chatbot?
Technically, you can, but it is only advisable to a limited extent. A model knows reference ranges as general textual information, not those used by your laboratory. Reference ranges depend on the laboratory, method, and age; the information on your report is always authoritative. There is also the issue of data protection: health data are specially protected, and anything you enter into an external system leaves your sphere of control. The medical practice that requested the test is responsible for interpreting your results.
Does it have to be clear that I’m talking to an AI?
Yes. Article 50(1) of the European AI Regulation requires AI systems intended for direct interaction with natural persons to be designed so that the people concerned are informed that they are interacting with an AI system. The exception applies only when this is already obvious from the circumstances.
The Cloé consultation service in this shop is identified as AI; this identification cannot be disabled.
What can the AI consultation service Cloé do, and what can’t it do?
Cloé answers questions about health tests in the areas of DNA, blood, and gut health and suggests suitable products based on them. It works with a stored knowledge base from its own product range, not with the open internet. It does not make diagnoses, read or interpret laboratory results, or make statements about providers outside its own company. For topics concerning children and trying to conceive, it refers users to a medical practice and deliberately does not recommend a substitute product.
Next step
Turn a supposition into a value
If you want to know where you actually stand after the chat, a measurement will help more than the next question. Discuss what follows from it with a human afterward.
View VitalCheck Complete All articles about blood testsRead more
You might also be interested in this
What preventive care can achieve when it is not based on assumptions.
An example of how far individual values extend and where they stop.
Sources
- Stiftung Gesundheitswissen: Searching for Health Information with AI Chatbots (as of 2025) – stiftung-gesundheitswissen.de
- World Health Organization: WHO releases AI ethics and governance guidance for large multi-modal models (as of 2024) – who.int
- Bitkom e. V.: Dr. AI – When the Chatbot Becomes a Medical Adviser (as of 2025) – bitkom.org
- European Union: Regulation (EU) 2024/1689 laying down harmonized rules on artificial intelligence, Article 50 (as of 2024) – eur-lex.europa.eu
The statements about hallucinations, fabricated sources, and the fact that a chatbot does not replace a doctor's visit are based on [1]. The warning about false, inaccurate, distorted, or incomplete statements, as well as the term automation bias, comes from [2]. The usage and trust figures for Germany come from the representative survey of 1,145 people aged 16 and over in [3]. The wording concerning the labeling requirement is set out in [4]. The product name, price, and number of markers were checked on the mybody®x product pages on 03.09.2026.
mybody®x Editorial & Specialist Team
Digital Health Information Laboratory Diagnostics AI in the Shop
This article was created by the mybody®x editorial team, while the specialist team reviews the laboratory diagnostic statements. The contributors are listed on the authors page.
Published on 03.09.2026 · Last updated on 03.09.2026
The content is for general information and does not replace medical advice, diagnosis, or treatment. Reference ranges depend on the laboratory, method, and age—the information on your report is always authoritative. Our products are not intended to diagnose, treat, cure, or prevent diseases.





Share now:
Genotype and phenotype: the difference and what lies between them
Warum ist Abnehmen so schwer? Was der Körper dagegen tut