Join a Study Heartbeats Join a Study Run a Study Why Reputable Our Process Verified Evidence Program Clinical Research Research and Articles Blog Posts Case Studies News and Media White Papers Contact Schedule a Call
For Brands
Should I Run a Pilot First?
By Reputable Health Team 8 min read August 2026
The short answer
  • A pilot answers can we run this study, and what should the real one look like. It does not answer does the product work.
  • The costly version of this mistake is designing a pilot as though it will double as substantiation. The choices that make a pilot cheap can be the same ones that make it unusable for a claim.
  • The crux is the control group. FTC guidance states plainly that improvement over time in a treatment group alone can come from placebo effect, spontaneous changes in health, or the practice effect.
  • A pilot is not a regulatory shortcut. Subject count has no bearing on whether the IND regulations apply, and FDA's definition of a clinical investigation reaches studies with as few as one subject.
  • And the risk that rarely gets priced in: a pilot becomes part of your evidence base whether it helps you or not.

There are two versions of this question, and they get very different answers.

The first is: we want evidence, and a small study is cheaper, so can we start there? The second is: we are going to run a real study eventually, and we want to de-risk it first.

The second question has a good answer. The first one runs into a structural problem, and it is worth understanding before you spend anything.

What a pilot is actually for

A pilot is an operational and statistical instrument. It answers questions about your study rather than about your product.

Can we recruit these people?

Not in theory. Actually, at the rate and cost you assumed, from the channels you have.

Will they stay?

Retention, compliance with the protocol, and whether the burden you designed is one real people will carry for the full duration.

Does the measurement work?

Whether your instruments capture what you thought, whether your data pipeline is clean, and whether the endpoint you chose can even move in the window you chose.

What does the real study need to look like?

A pilot gives you information about how much your outcome varies between people, which is one input to sizing a larger study. Treat any effect size a small pilot produces as a rough signal rather than a number to plan around, because a small sample estimates it imprecisely.

Each of those is worth real money, because each is a way the expensive study fails if you get it wrong. That is the case for a pilot, and it is a strong one.

Notice what is not on the list.

What a pilot cannot buy you

It cannot buy you a claim, and the reason is structural rather than about size.

The FTC's guidance sets out the basic principles it expects human clinical research to satisfy, and the first is a control group. The efficacy of a product should be demonstrated by comparing the treatment group to the control group. Improvement over time in the treatment group alone, the guidance says, could result from a placebo effect, from spontaneous changes in subjects' health, from improvements on a test measure that come purely from practice or repetition, or from other variables unrelated to the product.

A single-arm design can be the right choice for the questions a pilot is built to answer. It is also the design that cannot separate your product from those alternative explanations.

A single-arm pilot showing that a metric improved is doing exactly what it was built to do. It is also not evidence that the product caused the improvement.
Why this is not about sample size

The FTC's own example is a 200-person study

The guidance walks through a marketer who commissioned a study of a shake for osteoarthritis symptoms. Two hundred subjects. Randomized. Placebo-controlled. Double-blind. A validated symptom measure, assessed at intervals over ninety days. By any standard, a serious study.

The author reported a statistically significant improvement in the treatment group from baseline to day ninety. The guidance says that result does not substantiate the claim, because it does not compare the treatment group's improvement against the control group's. Both groups improved over time, and the treatment group's improvement was not statistically greater.

The marketer then searched the data for significant outcomes outside the original protocol and found one, in a subgroup with the mildest disease. The guidance treats that as unreliable evidence of benefit, and says further research in that population would be needed to verify anything.

Two hundred subjects, a rigorous design, and a within-group comparison still failed. A twenty-person single-arm pilot reporting the same kind of before-and-after number is making a weaker version of an argument that already did not work.

A pilot is not a regulatory shortcut either

Founders sometimes reach for a pilot on the theory that a small study attracts less scrutiny. It does not, at least not in the way they mean.

FDA's guidance addresses this directly. The number of subjects enrolled has no bearing on whether a study falls under the IND regulations, and the definition of a clinical investigation reaches studies with as few as one subject.

The same logic runs through the IRB question. Whether review is required turns on the nature of the investigation and where its results are headed, not on how many people are in it. We wrote about both separately: the IND question here and the IRB question here.

Small does not mean exempt. It just means small.

Three ways a pilot becomes a liability

  1. Analyzing your way to a result. When the pre-specified endpoint does not move, the temptation is to go looking. The FTC guidance names this: analysis that departs from the original protocol can indicate data mining, and the more comparisons you examine, the more likely a significant difference is just chance. It says such analysis may identify areas for future exploration but does not generally provide reliable evidence to substantiate a claim.
  2. Measuring everything. A pilot may carry a wide panel of outcomes, which is reasonable when you are learning what moves. It becomes a problem when one of eight measures hits significance and gets reported as the finding. The guidance works through exactly that scenario and points out that with eight outcomes and no statistical correction, a single positive result could easily be chance.
  3. Reporting selectively. The guidance is explicit that studies using multiple outcome measures should report all of them rather than only the positive ones, and that public trial registration is a generally accepted practice partly because it reduces the chance of that happening.

Each of these turns a pilot from a useful internal instrument into something that damages you if anyone reads it carefully. And on a claim, someone will.

The risk nobody prices in

Here is the part that gets left out of the case for running a pilot, and it is the one worth sitting with.

The FTC guidance works through an advertiser with two controlled, double-blind studies showing a modest but statistically significant loss of body fat at six weeks. There is also an equally well-controlled twelve-week study showing no statistically significant difference between treatment and control. Taken together, the guidance says, the studies suggest that if the product has any effect it would be very small and may not persist. Given the totality of the evidence, the claim is unsubstantiated.

The same section states that advertisers should consider all relevant well-conducted research relating to the claimed benefit, rather than focusing only on research that supports an effect while discounting research that does not.

A pilot becomes part of your evidence base whether it helps you or not.

Which means a promising pilot followed by a null confirmatory study does not return you to where you started. It leaves you with a documented inconsistency you now have to account for, and you cannot quietly set it aside because it came out the wrong way.

That is not an argument against pilots. It is an argument for being clear-eyed that you are generating evidence, not generating an option on evidence.

What makes a pilot worth running
  • Write the protocol first, with primary and secondary outcomes defined in advance. This is what separates a pilot from a fishing expedition.
  • Register it. Public registration helps ensure the study is conducted and analyzed the way the protocol said, and that all the data gets reported.
  • Plan the analysis to include everyone, including people who dropped out or did not fully comply. An intent-to-treat analysis is the honest denominator.
  • Choose a duration that can show what you want to claim, including any follow-up needed to show an effect persists.
  • Say what it is. Label the pilot a pilot, in the protocol, in the writeup, and in anything that quotes it later.

That last one costs nothing. A pilot described accurately is an asset. The same pilot described as proof is the thing a competitor screenshots.

So: should you?

Run a pilot when you are genuinely uncertain about feasibility, when the protocol has a piece nobody has tested operationally, when you need to know whether your endpoint can move at all in the window you have, or when you need internal evidence to justify a bigger spend.

Skip to the controlled study when you already know the study runs, you have a defensible basis for sizing it, and the actual goal is a claim. In that case the pilot is a delay wearing the costume of prudence.

Do neither yet when the honest answer is that you do not know what you would claim. That is a positioning problem, and no study design fixes it. Work out the claim first, then work backward to the evidence it needs. We wrote about that half of the problem here.

Where we sit in this

We run pilots. So read the above knowing we have a commercial interest in you deciding to run one, and notice that the article argues a pilot cannot deliver a claim.

That is the point. The reason to run a pilot is that it makes the real study cheaper, faster, and more likely to work. The reason not to run one is to get a number for the label. If that is what you need, a pilot is the wrong instrument, and we would rather say so before you spend the money than after.

Sources

  1. Health Products Compliance Guidance. Federal Trade Commission, December 2022. ftc.gov
  2. Investigational New Drug Applications (INDs): Determining Whether Human Research Studies Can Be Conducted Without an IND. U.S. Food and Drug Administration, September 2013, revised October 2015 to reflect a stay of parts of subsection VI.D.
Weekly Newsletter

Stay in the loop. No hype.

Evidence-based health insights and study updates, delivered weekly.

Your info is safe. We never share or sell your email.