An outlier is a value that sits unusually far from the rest of the data. It might be a typo, a faulty sensor, a member of a different group, or a genuine rare event. This calculator flags outliers three ways: Tukey’s IQR fences, z-scores, and modified z-scores based on the median. It shows every value’s scores in a table, draws a box plot, and reports how the mean and standard deviation change when the flagged values are set aside.
How to use the outlier calculator
- Paste your data. At least three values are needed.
- Choose the method for the headline result: IQR fences, z-score, or modified z-score.
- Adjust the threshold if needed: the fence multiplier k (1.5 or 3), the z cutoff (often 2, 2.5 or 3), or the modified z cutoff (3.5 by default). For the IQR method you can also choose how quartiles are computed.
- Compare the methods on the tape, then scan the table to see each value’s z and modified z.
Outlier formulas
IQR fences (Tukey). With quartiles Q1 and Q3 and IQR = Q3 − Q1:
Z-score. Using the mean x̄ and sample standard deviation s, flag values with |z| above the cutoff:
Modified z-score (Iglewicz and Hoaglin). With the median x̃ and MAD = median(|x − x̃|), flag |M| > 3.5:
Worked example
Thirteen website response times, in seconds: 12, 14, 14, 15, 16, 16, 17, 18, 19, 20, 21, 22, 45.
- IQR fences: Q1 = 14.5 and Q3 = 20.5, so IQR = 6. The fences are 14.5 − 9 = 5.5 and 20.5 + 9 = 29.5. The value 45 is outside.
- Z-score: the mean is 19.15 and s = 8.31, so 45 has z = (45 − 19.15) ÷ 8.31 = 3.11, just over 3.
- Modified z: the median is 17 and MAD = 3, so 45 scores 0.6745 × 28 ÷ 3 = 6.30, far beyond 3.5.
All three agree. Without 45, the mean drops from 19.15 to 17.00 and the SD from 8.31 to 3.07: one value nearly tripled the apparent spread.
Now add a second slow response of 47. The IQR and modified z methods flag both 45 and 47, but the z-score method flags neither: the two values pull the mean up to 21.14 and the SD up to 10.91, so 47 only reaches z = 2.37. This masking effect is why robust methods are preferred for screening.
Choosing and using a method
Robust versus classical
The IQR rule and the modified z-score are built from quartiles and medians, which ignore how extreme the tails are. They work for skewed data and are hard to fool. The z-score method is best when the data are roughly normal and you expect at most one outlier.
Skewed data
In right-skewed data such as incomes or reaction times, many legitimate large values will be flagged by any symmetric rule. Consider transforming the data (for example with logarithms) before screening, or judge outliers against the expected shape.
Report what you did
State the method, the cutoff, which values were flagged, and whether they were removed. The box plot follows the standard 1.5 × IQR convention regardless of the cutoff you choose. For quartiles under different conventions, see the quartile calculator, and for individual standardized scores, the z-score calculator.
Frequently asked questions
Which outlier method should I use?
For a quick, assumption-free screen, use the IQR rule; it relies on quartiles, so the outliers themselves cannot distort it. The modified z-score is similarly robust and gives each value a score. The ordinary z-score assumes roughly normal data and can miss outliers because extreme values inflate the standard deviation used to judge them.
Why does the z-score method sometimes find no outliers?
Two reasons. Outliers inflate the mean and standard deviation, which shrinks their own z-scores (masking). And in small samples no z-score can exceed (n − 1) ÷ √n; with 10 values that is 2.85, so a cutoff of 3 can never be reached.
What is the difference between 1.5 × IQR and 3 × IQR?
Tukey called values beyond 1.5 × IQR from the quartiles outside values and those beyond 3 × IQR far out. The 1.5 rule flags about 0.7% of values from normally distributed data; the 3 rule flags essentially none, so a hit is a strong signal.
Should I delete outliers?
Only with a reason. Check first for data-entry errors, measurement faults or values from a different population; correct or remove those and document it. Genuine extreme values belong in the data and may be the most interesting part. If unsure, report results with and without them.
What is the median absolute deviation (MAD)?
MAD is the median of the absolute distances from the median. Because it uses medians twice, a few wild values barely change it. Multiplying by 1.4826 turns it into an estimate of the standard deviation for normal data, which is where the 0.6745 constant in the modified z-score comes from.