Averages Are Liars

Business Analytics Simplified · Beginner ·📊 Data Analytics & Business Intelligence ·2mo ago

About this lesson

In ordinary speech, the word average usually refers to what statisticians call the mean the sum of all the values divided by how many there are. If you have ten customers and they spent a combined ten thousand dollars last year, the mean spending is one thousand dollars per customer. That number is precise, easy to compute, and, in many cases, deeply misleading. It is misleading because the mean treats all customers as if they were interchangeable contributors to a single pool. If nine of those customers spent five hundred dollars each and one spent five thousand five hundred, the mean is still one thousand. But the typical customer in that group spent five hundred. The "average customer" meaning the customer you are most likely to encounter, the customer whose experience drives your retention rates and support volumes is not spending one thousand dollars. The mean is, in this case, a description of the entire revenue pool divided across heads, not a description of any individual customer's behavior. The median fixes this by reporting the middle value: line up all the customers from least to most spending, and the median is the one in the center. In our example, the median is five hundred dollars, which is a more honest summary of what a typical customer does. The mode is the most common value the spending level you would encounter most often if you sampled customers at random. For continuous business metrics like revenue, the mode is rarely useful directly, but for categorical data most common product purchased, most frequent support request, most likely day of the week for a transaction it can be the single most informative summary.

Full Transcript

Welcome to this explainer. Look, as senior management, your daily decisions literally dictate the trajectory of your entire enterprise. You rely on dashboards, right? You trust the reports. But what if the very metrics you trust most are systematically lying to you? Today, we're stepping in as your strategic consultants to diagnose a massive blind spot in your executive decision-making. We're looking at a core concept from What the Numbers Are Really Saying, and the premise is uncompromising. Averages are liars. Let's uncover the ground truth actually hiding in your data. Picture this. It's a very familiar 20-second boardroom scenario. You're at the annual leadership offsite. Your VP of sales stands up, chest puffed out, and proudly delivers the headline, "Average customer revenue increased 15% this year." The room claps. Bonuses are safe. Expensive golf trips are mentally booked. But then, a junior analyst quietly raises her hand. She's been looking at the exact same data set, but she reports a rather grim finding. The typical account, the median, is actually spending slightly less than last year. The paradox here? Neither person is lying. Both of these figures are mathematically flawless. So, how exactly can both statements be true at the same time? And more importantly for you, as the executive steering the ship, which number reveals the actual health of your business? I assure you, depending on which summary statistic you believe, your entire strategic trajectory could be compromised. If you assume the whole base is growing when it's really just a tiny fraction, your resource allocation is going to be disastrously wrong. This isn't just academic theory, folks. It is a fundamental vulnerability in corporate governance. Section one, strategic risks of skewed data, diagnosing the lopsided reality. To solve our little boardroom paradox, we have to pivot sharply to the actual shape of your data. The fundamental issue is that business data is almost never perfectly symmetrical. Look at the distribution on the left. Here, the mean and the median are perfectly aligned right in the center. Beautiful, but entirely fictional for business. On the right, that's your reality. Almost no metric you care about, whether it's revenue, deal size, or support response times, looks like the left side. It's skewed. A few massive outlier values violently pull the mean away from the bulk of the data, creating this incredibly dangerous illusion of what's normal. Let's look at the exact same 100 customers from your offsite. See that red dash line? That's the mean, sitting up at $973. Now, look at the green dash line. That is the median, way down at $354. So, what actually happened here? Well, a few massive enterprise deals came up for renewal at much higher tiers. Their enormous lopsided weight simply dragged the mean up by 15%. But that green line, that shows your mid-market reality. The bulk of your clients are actually shrinking their budgets. The mean completely masked a structural decay in your business. We must be absolutely precise with our definitions here. These aren't just high school math terms. They are critical strategic instruments. The mean is simply the total revenue pool divided by the number of heads. It tells you about the total. The median, on the other hand, requires you to line up all your customers from smallest to largest and pick the literal middle one. It isolates the typical experience. Listen to me on this. Confusing the mean for the median is the single most common way management decisions go disastrously wrong. Section two, executive traps, weights, and rates. Just knowing the definitions is only your first line of defense. We have to identify the specific rhetorical traps that get woven into everyday corporate reporting. The most prevalent one? I call it the silent substitution. It happens all the time. The presenter displays the mean on the projector, but verbally they use phrases like "the typical customer" or "the average employee." You're nodding along, forming a mental picture of the typical experience, completely unaware that the figure on the screen doesn't represent that experience at all. It is a silent, lethal substitution of reality. Now, a much subtler manipulation of influence happens with weighted averages. Think about your regional customer satisfaction scores. Do you weight the average by region, treating a tiny 12-customer office in São Paulo the exact same as a 12,000-customer office in New York? Or do you weight it by individual customer? Both are perfectly legitimate mathematically. Both will be presented to you simply as the average score. But a small, rapidly improving region can entirely dominate a regional average, completely obscuring stagnation in your primary market. You always, always have to ask, "Who chose these weights, and why?" All right, let's confront another severe asymmetry. Prepare for a strict mathematical reality check. Look at this. 25% loss. You know, most people possess a highly flawed intuition about percentages. They honestly believe that if a metric increases by 50% and then subsequently decreases by 50%, they've broken even. They think they're right back where they started. I assure you, this is a mathematical fallacy that routinely destroys corporate capital. Here is the absolute truth of percentage asymmetry, the base rate trap. Let's say your bookings started at 100. They rise 50% to 150. Excellent work. But then, the market shifts and they drop 50%. Well, that 50% drop is calculated against the new, larger base of 150. You lose 75. You are now left at 75, which is a net 25% loss from where you began. If a sales team loses 30% of its pipeline, they don't need 30% to recover. They need 43%. Never, ever let a dashboard convince you that equal percentages magically cancel each other out. And then, of course, we have index numbers, consumer price indexes, customer experience indexes, performance indexes. They're wonderfully convenient, right? Everything is normalized to a baseline of 100. But the supreme danger lies entirely in the choice of that baseline. If 2020 was an artificially massive year due to some temporary market tailwinds, setting it to 100 makes today's 95 look like a severe deterioration. But set the baseline to a weaker year like 2018, and suddenly today's metric looks like a triumph of management. The underlying data literally never changed. Only the narrative did. Section three, AI analytics, gaining your competitive edge. We've diagnosed the disease. Now, we shift to the cure. A modern executive doesn't just passively accept a dashboard. They interrogate it. And today, artificial intelligence provides an absolutely unparalleled strategic advantage for this exact task. Here is your strict, step-by-step executive workflow. Step one, paste the data set or summary tables directly into your secure AI platform. Step two, explicitly request it to calculate both the mean and the median. Do not let it default to just the mean. Step three, demand a plain language description of the distribution's shape. Are outliers driving the narrative? In mere seconds, AI performs the rigorous pushback that used to take a junior analyst an entire week. You must institutionalize these precise questions into your management culture. Treat this as your mandatory dashboard interrogation checklist. First, what is the median? Second, how exactly does the median compare to the mean? Third, what does the distribution actually look like? Is it roughly symmetric or is it skewed by a long tail? And fourth, the ultimate strategic query. If the mean and median diverge, which one actually answers the business question this dashboard was built to support? Force your teams and force your AI to answer these four specific questions. Let's cement this with the ultimate key insight from our source material. The mean tells you about the total. The median tells you about the typical. In skewed business data, these are two entirely different stories. Confusing them is the single most common statistical error in management. From this day forward, whenever an average is presented to you in a meeting, your immediate reflex must be to ask if the median tells a contradictory story. That distinction is absolutely your primary defense mechanism. Which brings us to the close of our consultation today. I leave you with a direct, uncompromising question. What vital strategic truth is the average currently hiding in your own organization? Do not wait for the next quarterly review. Tomorrow morning, I want you to open your primary operational dashboard, look at the highest-level average, and deploy the AI interrogation techniques we just covered. Seek out the true variation in your data, because the average is lying to you, and the truth is waiting in the median. Thank you for joining me for this explainer. Keep questioning the data.

Original Description

In ordinary speech, the word average usually refers to what statisticians call the mean the sum of all the values divided by how many there are. If you have ten customers and they spent a combined ten thousand dollars last year, the mean spending is one thousand dollars per customer. That number is precise, easy to compute, and, in many cases, deeply misleading. It is misleading because the mean treats all customers as if they were interchangeable contributors to a single pool. If nine of those customers spent five hundred dollars each and one spent five thousand five hundred, the mean is still one thousand. But the typical customer in that group spent five hundred. The "average customer" meaning the customer you are most likely to encounter, the customer whose experience drives your retention rates and support volumes is not spending one thousand dollars. The mean is, in this case, a description of the entire revenue pool divided across heads, not a description of any individual customer's behavior. The median fixes this by reporting the middle value: line up all the customers from least to most spending, and the median is the one in the center. In our example, the median is five hundred dollars, which is a more honest summary of what a typical customer does. The mode is the most common value the spending level you would encounter most often if you sampled customers at random. For continuous business metrics like revenue, the mode is rarely useful directly, but for categorical data most common product purchased, most frequent support request, most likely day of the week for a transaction it can be the single most informative summary.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

Up next
How to Prompt Your LLM Directly from SQL
Ian Wootten
Watch →