Outlier Exclusion in Lab Data: When Removing a Point Is Justified
Every lab eventually faces the well that reads ten times higher than its neighbours, or the replicate that sits far from the rest for no obvious reason. Outlier exclusion is one of the quietest ways a dataset can be bent toward the answer someone hoped for, and also one of the most necessary steps in cleaning real data. The difference between the two lies almost entirely in when and how the decision is made. This article sets out the legitimate reasons to remove a data point, the statistical tests people reach for, and a documentation habit that keeps the process defensible when someone else audits your numbers.
An unusual value is not automatically an error
Biological systems are noisy, and some spread is expected. A value that looks extreme in a set of three replicates may be entirely ordinary once you have thirty. Before treating a point as a problem, it is worth asking whether it could simply be the tail of normal variation. If it could, removing it narrows the error bars artificially and makes the effect look more certain than it is.
Sometimes the odd value is the finding. A subset of wells responding strongly may point to a subpopulation of cells, a threshold effect, or a contamination issue that matters for the whole project. Deleting it hides information.
Legitimate grounds for outlier exclusion
The strongest case for removal is a documented, identifiable cause that is independent of the result itself. Typical examples:
- Recorded technical failure: a pipetting error noted at the time, a bubble in a well, a tip that failed to dispense, a cracked or contaminated plate.
- Instrument fault: a reader error flag, a temperature excursion logged by the incubator, a detector saturation warning.
- Failed control: positive or negative controls on that plate outside predefined acceptance limits, which invalidates the whole run rather than a single point.
- Sample integrity issue: a stock solution that precipitated, a vial that was mislabelled, or material stored outside its conditions.
- Protocol deviation: a well treated for the wrong time or with the wrong concentration.
What these have in common is that the reason could be written down before anyone looked at whether the point helped or hurt the hypothesis.
Grounds that do not hold up
Equally clear are the reasons that should not be used. Removing a point because it weakens the effect, because the p-value crosses 0.05 without it, or because it “looks wrong” with no identified cause all introduce bias. So does applying an exclusion rule to one group and not another, or trying several rules until one produces the desired outcome. Reviewers and auditors increasingly ask for the full, unfiltered dataset, and a pattern of convenient removals is easy to spot.
Statistical outlier tests and their limits
Several formal tests exist to flag values that are unlikely under an assumed distribution. They are useful screens but not verdicts.
| Method | Basic idea | Limitation |
|---|---|---|
| Grubbs’ test | Tests whether the most extreme value is too far from the mean | Assumes normality; handles one outlier at a time |
| Dixon’s Q test | Compares the gap to the nearest value with the total range | Designed for very small samples; low power |
| Interquartile range rule | Flags points beyond 1.5 × IQR from the quartiles | Arbitrary cut-off; common in plots, less so as a removal rule |
| Median absolute deviation | Uses a robust measure of spread to flag extreme values | Threshold still needs to be chosen in advance |
| ROUT method | Identifies outliers from a nonlinear regression fit | Sensitivity depends on a chosen false discovery rate |
With three or four replicates, which is typical in plate work, no test has much power. A single point can look extreme purely by chance. That is why technical cause matters more than a test statistic in small datasets.
Write the rules before you see the data
The most effective protection against biased exclusion is a pre-specified rule. Before the experiment, record in the protocol or lab notebook:
- The acceptance criteria for controls on each plate, and what happens to the plate if they fail.
- Which technical events justify excluding a well, and who records them.
- Any statistical screen to be used, with its threshold.
- That the same rule will be applied across all groups.
- That excluded data are retained in the raw file and reported.
Robust alternatives to deletion
Sometimes the better answer is to keep every point and use methods that are less sensitive to extremes. Medians instead of means, nonparametric tests, robust regression and weighted curve fitting all limit the influence of a single aberrant value without pretending it did not happen. Reporting results with and without the questionable point, when the decision is borderline, lets readers judge the effect of the choice themselves.
Traceability makes the call easier
Many outlier decisions come down to one question: did something about this well differ from the others? Good records turn that from a guess into a lookup. For labs working through large, multi-vial orders, that means logging the lot number of every vial, the cap and crimp colour that ties it to its certificate, the date each stock solution was made and how it was stored. If a cluster of odd values all trace to one vial or one freshly thawed stock, you have an identifiable cause; if they are scattered across vials from the same lot, the material is an unlikely explanation.
Bulk Peptides products are third-party tested for purity by HPLC, with certificates published for some products. Keeping that paperwork with your inventory log is a small habit that pays off the first time a result has to be defended.
Products from Bulk Peptides are sold strictly for in-vitro laboratory research. They are not for human or veterinary use, and this article is about laboratory data handling only.

