How to Fact-Check Supplement Claims With AI
- AI is a research assistant, not a source. Use it to find claims and studies, never as the final word on whether something works.
- Ask the AI for primary sources with identifiers, then confirm each one exists on PubMed or ClinicalTrials.gov yourself. It takes about a minute per citation, and both databases are free.
- A real citation is not the same as a supported claim. Check who was studied, at what dose, for how long, and who paid for the research.
- If you sell the product, this workflow is due diligence, not substantiation. Regulators expect actual human clinical evidence behind health claims, and an AI literature summary does not create it.
In 2023, two New York lawyers were sanctioned by a federal judge for filing a brief that cited six court cases. All six were invented by ChatGPT, complete with realistic case names, docket numbers, and quotes. The detail worth remembering: one of the lawyers asked ChatGPT whether the cases were real. It assured him they were.
That failure mode is not unique to law, and it is not gone. When researchers had ChatGPT-3.5 write 30 short medical papers and then checked all 115 references it produced, 47 percent were fabricated outright, 46 percent were real papers cited inaccurately, and only 7 percent were both real and accurate (Cureus, 2023, PMID 37337480). A separate analysis of 636 citations found that 55 percent of GPT-3.5's references were fabricated, falling to 18 percent for GPT-4 (Scientific Reports, 2023, PMID 37679503).
Two honest caveats before you swear off AI entirely. First, those studies tested 2023-era models, and newer tools, especially ones that search the live web and link their sources, fabricate less. Second, fabrication was only half the story. Another 46 percent of those references were real papers cited wrong: the study exists, but it does not say what the AI claims it says. That failure survives every model upgrade, because it is not a citation problem. It is a reading problem.
So the answer is not to avoid AI when you are evaluating a supplement. It is genuinely good at surfacing research you would never find on page one of Google. The answer is to change its job description.
Here is the workflow. It works whether you are deciding what to buy or deciding what to put on your label.
Step 1: Pin down the exact claim
"Does magnesium work?" is not a checkable question. "Does magnesium glycinate at this dose improve sleep quality in healthy adults?" is. Pull the specific claim off the label or the ad: the ingredient, the promised effect, and ideally the dose. Vague questions get vague answers, and vague answers are where AI fills gaps with confident guesses.
Step 2: Ask the AI for its receipts
Do not ask "is this true." Ask for the evidence behind it, with identifiers you can verify. A prompt that works:
List the specific human clinical studies behind the claim that [ingredient] [claimed effect]. For each study, give the title, the journal, the year, and the PMID or ClinicalTrials.gov NCT number. If you are not certain a study exists, say so instead of guessing.
That last sentence matters. It gives the model a permitted answer other than a list of studies.
Step 3: Look up every source yourself
This is the step people skip, and it is the whole game.
PubMed
The National Library of Medicine's free index of biomedical literature, at pubmed.ncbi.nlm.nih.gov. Paste the study title in quotes. A real paper resolves in seconds. If the title returns nothing, or the PMID points to a different paper, treat the citation as unverified and discard it.
ClinicalTrials.gov
The registry for clinical studies. Any NCT number the AI gives you should resolve to a real trial record, and the record shows you the design, the enrollment, and whether results were ever posted.
Google Scholar
A useful backup for material outside PubMed's scope. We wrote a full comparison of when to use each one here: Google Scholar vs PubMed.
Do not take our word for the numbers in this article either. Both PMIDs above are real, and looking them up yourself takes about a minute. That is the point.
Step 4: Read past the abstract
A citation that exists is the start, not the finish. Five questions separate "there is a study" from "the study supports the claim":
- Who was studied? Mice are not humans. Cell cultures are not humans. Twelve elite cyclists are not the general population.
- How many people, for how long? A two-week study of 15 people is a pilot signal, not proof.
- What was the dose? Ingredient studies often use doses far higher than what is in the product making the claim. Compare the study dose to the label dose.
- What was actually measured? A change in a blood marker is not the same as feeling or performing better. Check whether the study's endpoint matches the promise in the ad.
- Who paid? Industry funding does not automatically invalidate a study, but it is context you want. Funding and conflict disclosures are usually at the end of the paper.
Step 5: Weigh it honestly
One small positive trial means "early evidence," not "clinically proven." Several independent randomized controlled trials pointing the same direction is a different, stronger situation. Calibrate your confidence to the evidence, and be comfortable with the honest middle answers: "promising but thin" and "studied, but not at this dose" describe most of the supplement aisle.
- A citation that does not resolve on PubMed or ClinicalTrials.gov
- "Clinically proven" with no study identified anywhere on the site
- Human claims resting entirely on animal or cell data
- A study dose that does not match the label dose
- Endpoints that do not match the promise (a biomarker shifted, but the claim is about how you will feel)
- An AI answer with no sources, or sources it cannot give identifiers for
- The AI vouching for its own citations (remember the lawyers)
A 60-second check
Say a recovery drink claims it is "clinically shown to reduce muscle soreness." The workflow: ask for the specific human studies with identifiers (Step 2). Suppose the AI returns two titles. One resolves on PubMed, one does not exist (Step 3). The real one turns out to be a 10-day study of 18 trained athletes using three times the dose in the drink (Step 4).
You now know exactly what "clinically shown" means here, and it is a much smaller statement than the label implies (Step 5). No verdict needed on the ingredient itself. The claim, as written, is what failed the check.
If you sell the product, the bar is much higher
Everything above is consumer due diligence. If you are a founder or brand operator, the same workflow is useful for a different reason: it shows you what a motivated customer, competitor, or regulator will find when they run it on your claims.
And the regulatory bar sits well above "an AI summary looked supportive." The FTC's operative standard for health-related claims is "competent and reliable scientific evidence," and its Health Products Compliance Guidance is explicit that this generally means randomized, controlled human clinical testing, and that animal and in vitro studies alone do not substantiate health claims. An AI literature review can tell you whether evidence exists. It cannot create evidence that does not.
That gap is where a lot of brands quietly live: borrowing an ingredient study run at a different dose, in a different population, with a different endpoint than their product and their claim. Running the five steps above on your own label is the cheapest audit you will ever do.
Full disclosure of our bias here: this workflow is our editorial policy. Reputable runs IRB-governed, decentralized studies for wellness brands, and every claim we publish gets verified against PubMed and ClinicalTrials.gov before it ships. If it cannot be verified, it does not go out. When brands run this check on themselves and find the evidence behind their claim does not exist yet, generating it, with real participants, wearable data, and validated outcome measures, is the honest fix. That is the work we do.
The bottom line
Use AI the way a good researcher uses a sharp junior assistant: send it to gather, never let it conclude. The claim is only as good as the primary source, the primary source is only as good as its design, and both are checkable for free in minutes.
If you want one health claim broken down like this every week, that is what our Study Digest does.
Sources
- Mata v. Avianca, Inc., Opinion and Order on Sanctions, U.S. District Court, Southern District of New York, June 22, 2023.
- High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus, 2023. PMID 37337480.
- Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 2023. PMID 37679503.
- Health Products Compliance Guidance. Federal Trade Commission, December 2022. ftc.gov/business-guidance/resources/health-products-compliance-guidance
Stay in the loop. No hype.
Evidence-based health insights and study updates, delivered weekly.
Your info is safe. We never share or sell your email.