P-Value Statistical Significance: What Lab Results Actually Show
Few numbers in a lab report attract as much attention, or as much misunderstanding, as the p-value. A result below 0.05 is often treated as proof and anything above it as failure, yet p-value statistical significance was never meant to carry that much weight on its own. For bench scientists and QC staff reviewing assay data, a clear picture of what a p-value does and does not say makes it much easier to judge whether a finding is solid. This article explains the definition in plain language, lists the most common misreadings, and suggests what to report alongside it.
The definition in plain language
A statistical test starts from a null hypothesis, usually that there is no difference between groups or no effect of a treatment. The p-value is the probability of seeing data at least as extreme as what you observed, assuming that null hypothesis and the test’s other assumptions are true.
A small p-value means the data would be surprising if there were truly no effect. It does not measure how big the effect is, how important it is, or how likely the hypothesis is to be correct. Those are separate questions that need separate evidence.
Where the 0.05 threshold came from
The 0.05 cut-off is a convention, not a law of nature. It sets the rate of false positives you are willing to accept when there is no real effect: about one in twenty tests. Nothing special happens at 0.049 that does not happen at 0.051. Treating results either side of the line as fundamentally different leads to dichotomous thinking that statisticians have warned against for decades. Reporting the exact p-value, rather than just “significant” or “not significant”, lets readers weigh the evidence themselves.
Common misreadings of p-value statistical significance
| Misreading | Why it is wrong |
|---|---|
| “p = 0.03 means a 3% chance the result is a fluke.” | The p-value assumes the null is true; it is not the probability that the null is true. |
| “p > 0.05 proves there is no effect.” | It means the data were not strong enough to rule out no effect. Small experiments often miss real effects. |
| “A smaller p-value means a bigger effect.” | With enough replicates, trivial effects produce tiny p-values. Size must be read from the effect estimate. |
| “Significant means biologically important.” | A statistically detectable change can be too small to matter for the question at hand. |
| “The result will replicate with 95% certainty.” | Replication probability depends on true effect size and power; a result at p just under 0.05 often fails to replicate. |
Effect sizes and confidence intervals
The single most useful companion to a p-value is the effect size with its confidence interval. Reporting that a treatment raised a reporter signal by, say, 1.8-fold, with a 95% interval of 1.2 to 2.6-fold, tells the reader both the likely magnitude and the uncertainty. Once those numbers are shown, the p-value adds little; without them, it says very little on its own.
For potency data, the same idea applies: an EC50 with its confidence interval says more than a statement that two curves differ significantly.
How lab practice inflates false positives
Several everyday habits raise the real false-positive rate well above the nominal 5%, often without anyone intending it:
- Multiple comparisons: testing twenty conditions against control produces about one “significant” result by chance alone. Corrections such as Bonferroni, Holm or false discovery rate methods account for this.
- Flexible analysis: trying several tests, transformations or exclusion rules and reporting the one that works.
- Optional stopping: adding replicates until p drops below 0.05.
- Pseudoreplication: counting technical wells as independent samples, which shrinks the error and the p-value.
- Selective reporting: showing only the experiments that reached significance.
Choosing the right test
A p-value is only as reliable as the test that produced it. A few checks before running any analysis:
- Identify the experimental unit, usually the independent biological replicate.
- Decide whether the design is paired (each experiment has its own control) or unpaired. Paired tests are often more powerful and more appropriate for plate work.
- Check whether the data are roughly normal on the scale analysed; ratios and potency values are often better analysed on a log scale.
- For more than two groups, use an ANOVA-type model with appropriate post-hoc comparisons rather than many separate t-tests.
- Specify the test, the comparisons and the significance level before seeing the results.
Power: the other half of the story
Statistical power is the probability of detecting an effect of a given size if it is really there. Small cell-based experiments with three replicates often have low power for moderate effects. Low power has two consequences: real effects are missed, and the effects that do reach significance tend to be overestimates. A quick power calculation using pilot variability helps decide how many independent experiments are worth running.
Controlling the inputs you can control
Much of the noise that weakens statistical tests comes from inputs rather than biology: different reagent lots, different cell passages, and different lots of the test material. Reducing that noise raises power without adding experiments. For labs buying research peptides in volume, reserving vials from one lot for a whole study, and logging each vial against its lot and the cap and crimp colour that matches its certificate, removes one avoidable source of variation. Bulk Peptides products are third-party tested for purity by HPLC, with certificates published for some products, and ship from within Canada.
Bulk Peptides supplies peptides for in-vitro research only. They are not for human or veterinary use, and this article discusses the statistics of laboratory experiments.

