Reading Economic Data
Samples and Surveys
How statisticians learn about millions of people by asking a few thousand, and what makes a sample trustworthy or biased.
Much of what we know about the economy comes from asking people questions. Nobody can interview every single person in a country every month, so statisticians use a sample - a smaller group chosen to represent the whole population, which in statistics means the entire group you want to learn about. If the sample is chosen well, a few thousand answers can tell you a surprising amount about millions of people. If it is chosen badly, even a huge sample can be misleading.
Surveys behind the headlines
Many of the numbers you hear in economic news come from surveys. Unemployment rates in many countries are estimated from regular household surveys that ask people whether they worked or looked for work. India’s Periodic Labour Force Survey plays this role for jobs data, and similar surveys exist in most countries. Surveys also measure household spending, business confidence, and how prices are changing in shops.
A census, by contrast, tries to count everyone. Census data is extremely valuable but expensive and slow, which is why most countries run a full population census only about once every ten years, relying on samples in between.
What makes a sample representative
The key idea is random sampling: every person or household in the population has a known chance of being selected, and the choice is made by chance rather than convenience. Think of stirring a large pot of soup before tasting a spoonful. If the soup is well stirred, one spoonful tells you how the whole pot tastes. If you only skim from the top, you might miss everything that has sunk to the bottom.
Statisticians often go further, dividing the population into groups - such as rural and urban areas, or different states - and sampling from each group in proportion, so no important group is accidentally left out.
Imagine a researcher wants to know how many households in a region have savings accounts, and surveys 5,000 people by calling numbers from an online directory. The answers suggest that 90 percent of households have accounts. But people listed in an online directory tend to be better off and more connected than average, and households without phones never get asked at all. A well-designed random sample of just 1,500 households, visited in person across villages and towns, might find that only about 70 percent have accounts. The larger survey was less accurate because of who it could reach.
Sampling bias and non-response
Sampling bias happens when the way a sample is gathered makes some kinds of people more likely to be included than others. Online polls, call-in surveys, and street interviews in one neighbourhood are common sources of bias.
A related problem is non-response: even in a well-designed survey, some people decline to answer. If the people who refuse differ systematically from those who agree - for example, if very busy workers or very wealthy households answer less often - the results can be skewed. Good statistical agencies track response rates and adjust for them where they can.
A common mistake is trusting a survey just because it asked a very large number of people. How the sample was chosen matters far more than its size. A small random sample usually beats a large, self-selected one, such as an online poll where anyone can click to take part.
- A sample is a smaller group used to learn about a whole population.
- Many headline statistics, like unemployment rates, come from household surveys.
- Random sampling gives every member of the population a known chance of selection.
- Sampling bias and non-response can skew results even in large surveys.
- How a sample is chosen matters more than how big it is.
No recording for this one yet - EconReader can read it aloud for you.