Introduction
Econometrics begins with a simple problem: economic life produces numbers, but numbers do not explain themselves.
A worker earns a wage. A firm changes a price. A government raises a tax. A central bank changes an interest rate. A family decides whether to send a child to college. A researcher observes data on these events and asks: what can be learned from them? Econometrics is the discipline that develops quantitative methods for answering economic questions using data, probability, statistical inference, and economic reasoning. In the standard modern view, econometrics is not merely a collection of formulas; it is a way of connecting economic models, observed data, and uncertainty in order to estimate relationships and test claims about the world (Wooldridge, 2010; Hayashi, 2000).
This book is written for a graduate student who wants to learn econometrics from first principles, with a path inspired by the clarity of Wooldridge-style foundations: start with the economic question, define the population object of interest, state the assumptions needed to learn it from data, choose an estimator, and then assess uncertainty carefully. The goal is not only to run regressions. The goal is to understand what a regression, likelihood, moment condition, panel estimator, instrumental variable, or treatment-effect estimator is actually doing.
Econometrics is powerful because it lets us reason quantitatively under uncertainty. It is dangerous when used mechanically. The same command in statistical software can produce a useful estimate, a misleading association, or a completely invalid causal claim. The difference is not the command. The difference is the question, the assumptions, the data-generating setting, and the interpretation.
What econometrics tries to answer
An economic question is a question about behavior, allocation, institutions, markets, policy, or welfare. Econometric questions are economic questions made precise enough to confront data.
Consider four examples:
- Prediction: Which loan applicants are most likely to default?
- Description: How unequal are household incomes in a country this year?
- Causal inference: What is the effect of an additional year of schooling on wages?
- Structural estimation: What are the parameters of a model of consumer demand that can be used to simulate a new tax policy?
These questions are related, but they are not the same.
A prediction problem asks for accurate forecasts of an outcome. If a bank predicts default, it may care mainly whether the predicted probability is accurate for future applicants. A prediction model can be useful even if its variables do not have causal interpretations.
A descriptive problem summarizes patterns in data. For example, estimating the average wage gap between two groups is descriptive unless we interpret the gap as caused by group membership, discrimination, education, occupation, or some other mechanism.
A causal problem asks what would happen to an outcome if one factor were changed while other relevant conditions were held fixed in an appropriate sense. The phrase “what would happen” already points beyond the directly observed data, because for a given person we cannot observe both the wage after completing college and the wage that same person would have earned without completing college. This missing alternative is called a counterfactual. Modern causal inference is largely built around carefully defining such counterfactual comparisons and stating assumptions under which data reveal them (Imbens and Rubin, 2015; Angrist and Pischke, 2009).
A structural problem tries to estimate parameters of an economic model that describes deeper behavioral mechanisms. For example, a demand model may include preferences, substitution patterns, and price sensitivity. The researcher may then ask what would happen under a policy not yet observed in the data. Structural work usually requires stronger modeling assumptions, but in return it can answer policy questions that simpler reduced-form comparisons cannot answer directly.
A first principle of this book is therefore:
The econometric method must match the question.
A method that is excellent for prediction may be insufficient for causal inference. A credible reduced-form causal estimate may not answer a structural policy simulation question. A beautiful mathematical estimator may be irrelevant if the parameter it estimates is not the parameter the researcher needs.
Data, populations, and samples
Econometrics uses data, but data are only a partial view of the world.
A population is the full collection of units or outcomes about which we want to learn. The population might be all working-age adults in a country, all firms in an industry, all counties observed over many years, or all possible outcomes generated by a stable economic process.
A sample is the part of the population that we actually observe. If we observe wages and education for 5,000 workers, those 5,000 workers are the sample. The larger population might be all workers in the labor market.
A random variable is a quantity whose value is uncertain before observation. In econometrics, wage, schooling, employment status, price, output, treatment assignment, and measurement error can all be treated as random variables. This does not necessarily mean the world is physically random in some deep philosophical sense. It means that, for purposes of statistical analysis, we represent uncertainty using probability.
For example, let \(Y\) denote a worker’s hourly wage and \(X\) denote years of schooling. Before we draw a worker from the population, we do not know that worker’s wage or schooling. We therefore treat \(Y\) and \(X\) as random variables. Once we observe a sample of workers, we have data:
\[ (Y_1, X_1), (Y_2, X_2), \ldots, (Y_n, X_n). \]
Here \(n\) is the sample size. The subscript \(i\) indexes the unit, such as worker \(i\).
A parameter is a fixed feature of the population that we want to learn. For example, the average wage in the population is a parameter. The slope in a population regression of wages on schooling is also a parameter. A statistic is a number computed from the sample. An estimator is a rule for computing a statistic from data, and an estimate is the numerical value obtained in one particular sample.
For instance:
- Parameter: the true population average wage, \(\mu = E(Y)\).
- Estimator: the sample mean, \(\bar{Y} = n^{-1}\sum_{i=1}^n Y_i\).
- Estimate: if the sample mean in our data is 24.70, then 24.70 is the estimate.
This distinction will appear throughout the book. Many econometric mistakes come from confusing the sample number we computed with the population object we hoped to learn.
Models are maps, not the territory
An econometric model is a simplified mathematical representation of a relationship among variables. For example, a simple wage equation might be written as
\[ \log(wage_i) = \beta_0 + \beta_1 education_i + u_i. \]
This equation says that the logarithm of wage is represented as a linear function of education plus an unobserved term \(u_i\). The unobserved term is often called the error term or disturbance. It contains factors affecting wages that are not explicitly included in the model: ability, local labor market conditions, family background, occupation, luck, measurement error, and many other influences.
The coefficient \(\beta_1\) is often the object of interest. In a purely predictive or descriptive regression, \(\beta_1\) summarizes how wages differ with education in the population linear approximation. In a causal interpretation, \(\beta_1\) would be read as the effect of education on wages, but that interpretation requires assumptions. In particular, it requires that the variation in education used to estimate \(\beta_1\) is not confounded by unobserved determinants of wages. Wooldridge emphasizes this distinction between the algebra of regression and the assumptions needed for causal interpretation throughout his treatment of cross-sectional and panel-data methods (Wooldridge, 2010).
A model is not automatically true because it is written with Greek letters. The equation above does not mean education is the only determinant of wages. It does not mean the relationship is exactly linear for every person. It does not mean causality has been established. It is a disciplined approximation whose usefulness depends on the question and assumptions.
A helpful habit is to ask three questions whenever you see an econometric model:
- What population relationship or parameter is being defined?
- What assumptions connect that parameter to the available data?
- What interpretation is justified if those assumptions hold?
These questions are more important than memorizing formulas.
Association, causation, and the counterfactual problem
Econometrics often begins with association. If workers with more schooling earn more, schooling and wages are associated. But association alone does not prove causation.
Suppose workers with more education earn higher wages. There are several possible explanations:
- Education may raise productivity, causing higher wages.
- People with higher ability may both obtain more education and earn higher wages.
- Family background may influence both education and labor market opportunities.
- Local labor markets may differ in both school access and wage levels.
- Measurement error may distort the observed relationship.
The causal question asks: what would happen to a given worker’s wage if that worker received more education than they otherwise would have received? The difficulty is that we observe each worker under one realized education level, not under all possible education levels. This is the fundamental counterfactual problem.
The potential outcomes framework makes this issue explicit. For a simple treatment such as attending college, let \(Y_i(1)\) be person \(i\)’s wage if they attend college and \(Y_i(0)\) be their wage if they do not attend college. For each person, we observe only one of these two outcomes. If the person attends college, we observe \(Y_i(1)\), not \(Y_i(0)\). If the person does not attend college, we observe \(Y_i(0)\), not \(Y_i(1)\). This notation is central in modern causal inference because it forces the researcher to define the missing comparison carefully (Imbens and Rubin, 2015).
Econometrics offers several strategies for dealing with this problem. Randomized experiments solve it by assigning treatment independently of potential outcomes. Instrumental variables use external variation that shifts treatment but is otherwise unrelated to unobserved determinants of the outcome. Difference-in-differences compares changes over time between treated and comparison groups under a parallel trends assumption. Regression discontinuity uses threshold rules. Matching and weighting compare units with similar observed characteristics. Panel-data methods use repeated observations to control for certain forms of unobserved heterogeneity.
Each strategy will appear later in this book. For now, the key lesson is simple:
Causal interpretation comes from a research design and assumptions, not from a regression coefficient alone.
This is one reason Angrist and Pischke stress the importance of research design in applied econometrics, especially when the goal is credible causal inference from nonexperimental data (Angrist and Pischke, 2009).
Identification before estimation
One of the most important words in graduate econometrics is identification.
A parameter is identified if it is uniquely determined by the population distribution of the observed data under the maintained assumptions. This definition is abstract, so let us unpack it.
Imagine that we had unlimited data from the same population. Sampling error would disappear. We would know the joint distribution of all observed variables perfectly. Could we then recover the parameter we care about? If yes, the parameter is identified. If no, more observations from the same kind of data will not solve the problem.
For example, suppose we want the causal effect of schooling on wages. If schooling is strongly related to unobserved ability, and ability is not observed, then simply collecting more wage-schooling data may estimate the association very precisely while still failing to identify the causal effect. The problem is not small sample size. The problem is that the observed data do not separate the effect of schooling from the effect of unobserved ability without additional assumptions or variation.
This distinction between identification and estimation is essential.
- Identification asks whether the target parameter can be learned from the population distribution and assumptions.
- Estimation asks how to use a finite sample to approximate that parameter.
- Inference asks how uncertain the estimate is because we observe only a sample.
Manski’s work on identification helped clarify that empirical conclusions depend on what can be learned from data under stated assumptions, and that weaker assumptions may imply bounds rather than a single point estimate (Manski, 1995). This book will mostly develop point-identification methods, but the habit of asking “what is identified?” will be present throughout.
Why ordinary least squares appears so early
Ordinary least squares, usually abbreviated OLS, is the most familiar method in econometrics. OLS chooses regression coefficients that minimize the sum of squared residuals, where a residual is the difference between an observed outcome and the fitted value predicted by the regression.
For a simple regression,
\[ Y_i = \beta_0 + \beta_1 X_i + u_i, \]
OLS chooses \(\hat{\beta}_0\) and \(\hat{\beta}_1\) to make
\[ \sum_{i=1}^n (Y_i - \hat{\beta}_0 - \hat{\beta}_1 X_i)^2 \]
as small as possible.
OLS appears early in the book not because all econometrics is OLS, but because OLS is the best entry point into many deeper ideas. Through OLS we learn:
- what an estimator is;
- how sample moments estimate population moments;
- how assumptions imply unbiasedness or consistency;
- how standard errors measure sampling uncertainty;
- how omitted variables create bias;
- how controlling for variables changes interpretation;
- how matrix algebra clarifies projection and rank conditions;
- how robust inference responds to heteroskedasticity;
- how endogeneity motivates instrumental variables.
OLS is therefore both a method and a training ground. It teaches the grammar of econometrics.
But OLS is not magic. If the regressor is correlated with the error term, OLS generally does not estimate a causal effect. If the functional form is badly chosen, the coefficient may be hard to interpret. If the data are selected in a nonrandom way, the target population may be unclear. If standard errors ignore clustering or serial correlation, inference may be misleading. Learning OLS properly means learning both its power and its limits.
The role of probability and large-sample thinking
Econometrics uses probability because data are incomplete. We usually observe one sample, but we want to learn about a population or process. Probability gives us a language for describing how estimates vary across possible samples.
A sampling distribution is the distribution an estimator would have over repeated samples generated by the same process. We rarely observe repeated samples in practice, but the concept helps us understand uncertainty. If an estimator would usually be close to the true parameter in large samples, we call it consistent. If its scaled estimation error approaches a normal distribution as the sample size grows, we can often build approximate confidence intervals and hypothesis tests using asymptotic normality. These large-sample ideas are central in graduate econometrics because many important estimators are justified by asymptotic theory rather than exact finite-sample formulas (Hayashi, 2000; Wooldridge, 2010).
For example, suppose the sample mean \(\bar{Y}\) estimates the population mean \(E(Y)\). Under standard regularity conditions, as \(n\) grows, \(\bar{Y}\) tends to get closer to \(E(Y)\). Moreover, after suitable scaling, its distribution is often approximately normal. This is why large samples allow us to report standard errors and confidence intervals even when the underlying outcome is not normally distributed.
The book therefore spends early chapters on probability, conditional expectation, convergence, laws of large numbers, central limit theorems, and standard errors. These topics are not mathematical decoration. They explain why econometric procedures work when they work.
The path of this book
The first part of the book builds foundations. We begin with econometric questions and causal thinking, then develop probability and statistical inference. We then study simple and multiple regression, first in scalar notation and later in matrix form. This sequence is deliberate: before using compact matrix expressions, we first learn what the objects mean.
The second part develops inference and complications in the linear model. We study hypothesis tests, confidence intervals, heteroskedasticity, robust standard errors, generalized least squares, endogeneity, and identification. At this stage, the central question becomes: when does a regression coefficient have a credible interpretation?
The third part moves to major estimation frameworks: instrumental variables, two-stage least squares, generalized method of moments, maximum likelihood, and quasi-maximum likelihood. These methods may look different, but they share a common logic: define population conditions implied by a model or assumption, then choose parameter estimates that make the sample version of those conditions fit as well as possible.
The fourth part studies models for outcomes that are not naturally continuous and unrestricted: binary outcomes, multinomial choices, ordered responses, counts, censored outcomes, truncated samples, and selection models. Here the book emphasizes interpretation. In nonlinear models, coefficients are often not marginal effects themselves, so we must learn how to translate estimates into economically meaningful quantities.
The fifth part develops panel data, time series, and modern research designs. Panel methods use repeated observations to address unobserved heterogeneity. Difference-in-differences and event studies exploit policy timing and comparison groups. Regression discontinuity uses threshold rules. Matching and weighting use observed covariates to construct comparisons. Time series econometrics handles dependence over time, persistence, unit roots, cointegration, and forecasting.
The final part connects classical econometrics with modern predictive tools and empirical workflow. Machine learning can improve prediction and help manage high-dimensional controls, but causal estimation still requires identification assumptions. A credible empirical project also requires careful data cleaning, transparent code, robustness checks, honest reporting, and attention to threats to validity.
How to think while reading
As you read, do not try to memorize every formula on first contact. Instead, build a stable sequence of questions.
When you see a parameter, ask: What population object does this represent?
When you see an estimator, ask: What sample rule is being used to estimate the parameter?
When you see an assumption, ask: Why is it needed, and what could violate it?
When you see a standard error, ask: What uncertainty is it measuring, and does it account for the actual sampling or dependence structure?
When you see a causal claim, ask: What is the counterfactual comparison, and what identifies it?
When you see a table of regression results, ask: What would have to be true for these numbers to answer the stated economic question?
These questions turn econometrics from a set of techniques into a disciplined way of reasoning.
A small example to carry forward
Suppose we want to study the effect of job training on earnings. We observe data on individuals, including annual earnings \(Y_i\), a training indicator \(D_i\), education \(X_i\), and age \(A_i\). A simple regression might be
\[ Y_i = \beta_0 + \beta_1 D_i + \beta_2 X_i + \beta_3 A_i + u_i. \]
The coefficient \(\beta_1\) compares trained and untrained workers after linearly controlling for education and age. But whether \(\beta_1\) is a causal effect depends on how workers entered training.
If training was randomly assigned, then treated and untreated workers should be comparable before training, apart from random variation. In that case, \(\beta_1\) may have a credible causal interpretation.
If workers selected into training because they were more motivated, then motivation may be part of \(u_i\), and \(D_i\) may be correlated with \(u_i\). In that case, OLS may mix the effect of training with the effect of motivation.
If only unemployed workers were eligible for training, then the relevant population may not be all workers. The target parameter must be clarified.
If training participation is measured with error, the coefficient may be distorted.
If earnings are observed only for people who find jobs, selection into employment may matter.
This one example already contains many themes of the book: regression, controls, omitted variables, endogeneity, selection, measurement error, identification, and interpretation. Econometrics is the art and science of making such issues explicit rather than hiding them behind software output.
The promise and discipline of econometrics
Econometrics can help answer some of the most important questions in economics and public policy: Do schools raise earnings? Do minimum wages reduce employment? Do taxes change labor supply? Does health insurance improve health? Do firms respond to regulation by changing investment? Do central bank announcements affect inflation expectations?
But econometrics does not replace judgment. It disciplines judgment. It forces the researcher to define the question, state assumptions, examine data, quantify uncertainty, and explain what is and is not learned.
The spirit of this book is therefore practical and foundational at the same time. We will derive estimators, study assumptions, interpret coefficients, and discuss empirical design. The aim is not to make econometrics look easy. The aim is to make it understandable, usable, and honest.
If you learn the material well, you will not merely ask, “What command should I run?” You will ask:
What is the economic question, what is the parameter, what identifies it, what estimator is appropriate, and how credible is the resulting evidence?
That is the beginning of econometric thinking.
References
-
Angrist, J. D., and Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press.
-
Hayashi, F. (2000). Econometrics. Princeton University Press.
-
Imbens, G. W., and Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press.
-
Manski, C. F. (1995). Identification Problems in the Social Sciences. Harvard University Press.
-
Wooldridge, J. M. (2010). Econometric Analysis of Cross Section and Panel Data (2nd ed.). MIT Press.