Descriptive Statistics for Sales
~11 min read Β· Use mean, median, mode, range and standard deviation on sale-price data sets.
Descriptive statistics turn a pile of sales into usable numbers: mean, median, mode for the center; range and standard deviation for the spread. The exam computes small datasets and tests WHICH center to trust when outliers lurk.
Measures of center
Mean: sum Γ· count β uses every value, and therefore follows outliers. Median: the middle value when sorted (average the middle two for even counts) β robust to outliers, which is why housing markets report median prices. Mode: the most frequent value β useful for typical lot sizes, bedroom counts, unit rents. Skewed data (a mansion among tract homes) pulls the mean away from the median; the gap DIAGNOSES the skew.
- Mean follows outliers; median resists them
- Median = housing's standard center
- Mean > median β right (high-end) skew
Measures of spread
Range: max β min β quick, fragile. Variance: average squared deviation from the mean; standard deviation: its square root, in the data's own units β the workhorse. In roughly normal data, ~68% of values sit within 1 SD of the mean, ~95% within 2 SD β the basis for judging whether a comp is 'in the market' or an outlier worth investigating. Coefficient of variation (SD Γ· mean) compares spread across datasets of different scales.
- SD: typical distance from the mean, same units as the data
- 68/95 rule under normality
- COV compares variability across markets
Appraisal use
Sale-price arrays, price-per-square-foot distributions, DOM and ratio studies all reduce to center + spread. Uses: defining the competitive range, spotting data-entry errors and non-arm's-length sales (3-SD outliers), supporting adjustment ranges, and describing the neighborhood's predominant value. Statistics support judgment β they never replace verification of the individual sale.
Worked example
Seven tract sales (thousands): 452, 458, 461, 466, 470, 475, 720. Compute mean, median, and range; diagnose; then reconsider with the outlier handled.
Mean: (452+458+461+466+470+475+720)/7 = 3,502/7 = $500.3k. Median: sorted middle (4th) value = $466k. Range: 720 β 452 = 268. Diagnosis: the mean sits $34k above the median β heavy right skew from the 720 sale, which sits far outside the cluster (the other six span just 23k); verification finds it a renovated model on a double lot β a different product, not a data error, but not this market segment either. Excluding it: mean = 2,782/6 = $463.7k, median = 463.5k β center measures now agree, describing the tract honestly. The lesson the exam wants: when mean and median diverge, find out why before quoting either.
Common exam pitfalls
Quoting the mean of skewed data.
Outlier-influenced means mislead β housing convention reports medians, and the mean-median gap flags the skew.
Deleting outliers without investigation.
An outlier is a question, not garbage β verify first; it may be a different segment or a data error.
Reading SD as a maximum deviation.
SD is the TYPICAL deviation β about a third of normal data sits beyond 1 SD.
Mean chases the mansion, median holds the middle, and the standard deviation says how wide the market really is.
Recap
- Mean uses all values; median resists outliers; mode = most frequent
- Mean-median gaps diagnose skew
- Range quick and fragile; SD = typical deviation
- 68/95 rule frames normal markets
- Outliers get investigated, then included or excluded with reason
- Statistics support β never replace β verification
Prove it: 10 questions on this topic
Every lesson ends with a ten-question check in the free course β your progress syncs between the web and the EstatePass app.
