Propensity Score Matching (PSM)

Propensity Score Matching (PSM) is a statistical matching method used to improve the comparability of treatment and control groups in observational or non-randomized research. It is particularly useful when researchers want to estimate a causal effect but treatment assignment was not randomized and the groups differ systematically in characteristics that may influence both treatment participation and the outcome.

The propensity score is the estimated probability that a unit receives a particular treatment, conditional on observed pre-treatment characteristics. Instead of attempting to match individuals directly across a potentially large number of covariates, PSM uses this probability as a balancing score to help construct treatment and comparison groups with more similar distributions of observed characteristics.

PSM does not recreate a randomized experiment. Matching can address imbalance in characteristics that researchers have observed and appropriately incorporated into the propensity-score analysis, but it cannot automatically eliminate bias caused by important unmeasured confounders. Credible use of PSM therefore requires much more than calculating propensity scores and finding matches: researchers must justify covariate selection, establish adequate overlap, assess balance after matching and consider the possibility of remaining unmeasured confounding.

On this page:

  • Propensity Score Matching Explained Simply
  • What Is Propensity Score Matching?
  • How Propensity Scores Work
  • PSM and the Counterfactual Problem
  • The Main Assumptions of PSM
  • How to Conduct Propensity Score Matching
  • Matching Methods
  • Common Support and Overlap
  • Covariate Balance After Matching
  • Dudovskiy PSM Causal Credibility Framework
  • Application of PSM: an Example
  • Advantages and Limitations of PSM
  • Common Mistakes When Using PSM
  • PSM in Business Research
  • PSM in the Age of AI and Digital Research
  • When to Use PSM
  • Dissertation Example
  • Exam Tip
Question What PSM can do What PSM cannot establish by itself
Pre-treatment differences Reduce imbalance in observed covariates Eliminate differences in unobserved characteristics
Group comparability Construct more comparable treated and untreated samples Reproduce random assignment
Selection bias Address selection related to appropriately modelled observed characteristics Automatically eliminate all selection bias
Common support Identify observations for which meaningful treated-control comparisons are available Create comparable controls where adequate overlap does not exist
Causal inference Strengthen a causal design under appropriate assumptions Make an observational association causal merely because matching was performed
Balance Improve balance on observed covariates Prove absence of unmeasured confounding

Propensity Score Matching Explained Simply

Imagine that a company introduces an advanced leadership programme for 500 managers and wants to determine whether the programme improves subsequent employee productivity. Comparing participating managers with managers who did not participate may initially appear straightforward, but programme participation was voluntary. Participants tend to be younger, have higher previous performance ratings, work in larger departments and have more senior-management support.

A direct comparison would therefore mix the possible effect of the leadership programme with differences that already existed between participants and non-participants. The researcher could attempt to find non-participating managers who resemble participants across these relevant pre-treatment characteristics, but matching simultaneously across many characteristics becomes increasingly difficult.

PSM approaches the problem by estimating each manager’s probability of participating in the programme based on observed pre-programme characteristics. A participating manager with an estimated propensity score of 0.72 could then be matched with a non-participating manager whose observed characteristics produce a similar estimated probability of participation.

After matching, the researcher does not simply proceed to compare productivity. The matched groups need to be checked to determine whether important pre-treatment characteristics have actually become sufficiently balanced. Only then does the researcher estimate the treatment effect in the matched sample and consider whether unmeasured factors could still provide an alternative explanation.

This final qualification is fundamental: PSM can make groups more comparable on observed characteristics, but it cannot demonstrate that they are equivalent on characteristics that were never observed.

Dudovskiy Research Assistant

Not sure if Propensity Score Matching is the correct choice for your dissertation?

Enter your topic to receive a free, tailored methodology preview.

Free · No account required
Learn more about Dudovskiy Research Assistant

What Is Propensity Score Matching?

The propensity score was introduced by Rosenbaum and Rubin as the probability of treatment assignment conditional on observed baseline covariates. One of its important properties is that, under appropriate conditions, conditioning on the true propensity score can balance the distribution of those measured covariates between treatment groups.

This provides a practical response to a dimensionality problem in observational research. Suppose researchers want to compare treated and untreated individuals according to age, income, education, previous performance, organisational tenure, department size and several other characteristics. Finding exact matches across every combination of these variables may be impossible, particularly as the number of covariates increases. The propensity score summarises information from those observed characteristics into a scalar measure that can be used for matching or other forms of adjustment.

Propensity scores can be used through several approaches, including matching, stratification, weighting and covariate adjustment. Propensity Score Matching specifically refers to using propensity scores to construct a matched sample in which treated observations are paired or grouped with sufficiently similar untreated observations. The broader family of propensity-score methods should therefore not be treated as synonymous with PSM.

The methodological purpose of matching is also important. Researchers are not trying to predict treatment participation as accurately as possible for its own sake. They are attempting to create a comparison sample in which the distributions of relevant observed pre-treatment covariates are sufficiently similar to permit a more credible treatment-effect comparison. This is why the predictive performance of the propensity-score model is not the principal criterion for judging whether PSM has worked.

How Propensity Scores Work

Suppose treatment status is represented by TT, where T=1T=1 indicates treatment and T=0T=0 indicates no treatment, and XX represents a set of observed pre-treatment covariates. Conceptually, the propensity score can be expressed as:

e(X) = P(T = 1 | X)

The score therefore represents the probability of receiving treatment given the observed characteristics included in XX. A propensity score of 0.70 does not mean that the individual received 70% of the treatment, nor does it indicate the probability that the treatment will succeed. It means that, given the variables incorporated into the model, the individual’s estimated probability of receiving the treatment is 70%.

Researchers commonly estimate propensity scores using logistic regression, although more flexible statistical and machine-learning approaches can also be used. The choice of estimation technique is secondary to the causal purpose of the model: it should help produce adequate balance between treatment groups on relevant observed pre-treatment characteristics.

Two individuals can have similar propensity scores despite differing on particular individual characteristics because different combinations of covariates may imply similar probabilities of treatment. This is precisely why balance must subsequently be evaluated on the covariates themselves. Similar propensity-score distributions do not, by themselves, demonstrate that all important observed characteristics are adequately balanced.

PSM and the Counterfactual Problem

PSM is best understood within the broader counterfactual logic of causal inference. For a treated individual, researchers observe the outcome after treatment but cannot simultaneously observe what would have happened to that same individual at the same time without treatment. The missing untreated outcome is the counterfactual.

In randomized experiments, random assignment provides a mechanism through which untreated units can, on average, represent what would have happened to treated units without treatment. Observational studies lack this protection. Individuals receiving treatment may differ systematically from those who do not receive it, and those same differences may influence the outcome.

PSM attempts to improve the counterfactual comparison by identifying untreated observations that resemble treated observations in terms of relevant observed pre-treatment characteristics. If two managers had similar pre-programme characteristics and therefore similar estimated probabilities of participating in a training programme, the untreated manager may provide a more informative comparison for the treated manager than an arbitrarily selected non-participant.

The quality of that counterfactual remains conditional on what has been observed and appropriately modelled. If managerial ambition strongly affects both participation and later performance but was never measured, matching on age, salary, tenure and previous performance cannot directly balance ambition. This limitation separates balance on observed covariates from the stronger and usually untestable claim that all relevant confounding has been removed.

The Main Assumptions of PSM

A causal interpretation of PSM depends on assumptions that should be understood before matching is performed. One of the most important is commonly described as conditional independence, conditional exchangeability, or no unmeasured confounding. In broad terms, after conditioning on the relevant observed pre-treatment characteristics, treatment assignment must not remain related to potential outcomes through unmeasured confounding.

This assumption cannot simply be demonstrated by showing excellent balance after matching. Balance diagnostics concern variables that have been measured. An omitted variable that strongly influences both treatment and outcome can remain imbalanced even when every variable appearing in the researcher’s balance table looks satisfactory.

A second requirement is positivity or overlap. For the relevant combinations of covariates, there must be some realistic possibility of observing both treated and untreated units. If every large multinational company in a dataset adopts a particular technology while every small company rejects it, there may be no meaningful untreated counterpart for the largest adopters. PSM cannot manufacture the missing counterfactual evidence.

Researchers should also consider the broader assumptions of the causal framework being used, including the definition and consistency of treatment and whether one unit’s treatment affects another unit’s outcome. The precise assumptions depend on the research setting, but they reinforce a general principle: matching is one component of a causal design, not a mechanical procedure that independently establishes causality.

How to Conduct Propensity Score Matching

The process should begin with a causal research question and a clear definition of treatment rather than with the matching software. Researchers need to establish which effect they intend to estimate and which pre-treatment characteristics are plausible confounders of the treatment-outcome relationship. Covariate selection should therefore be informed by substantive knowledge, theory and causal reasoning rather than simply by selecting variables that happen to be available or statistically significant. Methodological guidance particularly stresses understanding the relationships among covariates, treatment assignment and outcomes.

Propensity scores are then estimated using an appropriate model of treatment assignment. Logistic regression remains common for binary treatments, although alternative approaches may be appropriate where relationships are nonlinear or complex. Researchers should avoid interpreting this stage as an ordinary prediction exercise because the ultimate objective is covariate balance rather than maximum classification accuracy.

The next step is to examine whether sufficient overlap exists between treated and untreated observations. Units whose propensity scores lie in regions where no reasonable counterpart exists may need to be excluded from the matched analysis. Such restrictions can improve comparability but can also change the population to which the estimated treatment effect applies, so discarded observations should not be treated as methodologically inconsequential.

Researchers then select a matching algorithm and construct the matched sample. Once matching has been completed, covariate balance must be examined. If substantial imbalance remains, the propensity-score specification or matching procedure may need to be revised. Outcome analysis should come after an acceptable matched design has been established, preserving the conceptual separation between designing the comparison and estimating the treatment effect.

Matching Methods

Nearest-neighbour matching pairs each treated observation with one or more untreated observations having the closest propensity score. One-to-one matching is easy to understand and commonly used, while one-to-many matching can retain additional comparison observations. Decisions about whether matching occurs with or without replacement affect which controls can be reused and may influence both bias and precision.

Caliper matching places a maximum permitted distance between the propensity scores of matched observations. A treated observation is not matched if an untreated observation cannot be found within the specified caliper. This can prevent poor matches, although it may reduce the matched sample and alter the population represented by the analysis.

Radius matching allows a treated observation to be matched with multiple untreated observations falling within a specified propensity-score distance. Other approaches, such as optimal and full matching, attempt to construct matches using broader optimisation criteria.

There is no universally superior matching algorithm independent of the data and research objective. The important question is whether the selected procedure produces adequate overlap and covariate balance while retaining a population relevant to the causal effect being investigated. Researchers should therefore justify the matching strategy through its consequences for the causal design rather than merely reporting the default option offered by statistical software.

Common Support and Overlap

Overlap describes whether treated and untreated observations exist across sufficiently similar regions of the covariate or propensity-score distribution. It is a prerequisite for meaningful comparison because causal effects cannot be estimated through matching for treated units that have no plausible untreated counterparts in the available data.

Consider a study of the effect of adopting sophisticated AI-based inventory systems on firm productivity. If all firms with extremely high digital capability adopt the system and no comparable non-adopters exist, matching those firms to technologically unsophisticated companies would require comparisons across fundamentally different regions of the data. Restricting the analysis to a region of common support may provide a more defensible estimate for firms where adoption and non-adoption are both realistically observed.

The consequence is substantive as well as statistical. Removing observations outside common support may change the population represented by the analysis. An estimated effect among firms for which both adoption states were plausible should not automatically be interpreted as the effect for every firm in the original population.

Poor overlap can therefore reveal something important about the research question itself. Sometimes the data simply do not contain the counterfactual information required to answer the desired causal question for part of the population. Recognising that boundary is methodologically stronger than forcing every treated observation into an inappropriate match.

Covariate Balance After Matching

Matching is intended to create treatment groups with more similar distributions of relevant observed baseline covariates. Researchers therefore need to evaluate whether this objective was actually achieved. A propensity-score model should not be considered successful merely because it converged, generated plausible-looking probabilities or produced a large number of matched pairs.

Standardized mean differences (SMDs) are widely used to assess differences in covariate means or prevalences between treatment groups. Unlike conventional significance tests, standardized differences are not driven directly by sample size, making them useful for comparing imbalance before and after matching. Variance ratios, graphical comparisons and other distributional diagnostics can provide additional information because similar means do not necessarily imply similar overall distributions.

Balance should be evaluated for the substantive covariates used to construct the causal comparison, not only for the propensity score itself. Austin notes that comparing only the distributions of estimated propensity scores between treated and untreated groups can be uninformative as a balance diagnostic; researchers need to examine the underlying covariates.

If important imbalance remains, researchers should reconsider the propensity-score model or matching procedure rather than immediately estimate the treatment effect. Balance assessment is therefore not a reporting formality performed after PSM has succeeded. It is part of determining whether matching has succeeded at all.

Dudovskiy PSM Causal Credibility Framework

The Dudovskiy PSM Causal Credibility Framework provides a structured way to evaluate whether a propensity-score matched analysis supports the causal interpretation being proposed. It synthesises established principles of propensity-score methodology into a practical decision sequence rather than introducing a new statistical theory or matching method.

Causal Question → Confounder Selection → Propensity Score Estimation → Overlap Assessment → Matching → Covariate Balance → Unmeasured Confounding → Treatment-Effect Estimation → Defensible Causal Claim

Dudovskiy PSM Causal Credibility Framework showing the Propensity Score Matching process from causal question to defensible causal claim.

1. Causal Question

The process begins by specifying the treatment effect of interest. Researchers should determine what treatment is being compared with what alternative, for which population, and over what outcome period. This provides the substantive target for subsequent matching decisions.

2. Confounder Selection

Relevant pre-treatment covariates should be selected through knowledge of the treatment-assignment process and determinants of the outcome. The objective is not to maximise the number of variables in the propensity-score model, but to represent the confounding structure sufficiently well to support the intended comparison. Variables measured after treatment require particular caution because conditioning on post-treatment variables can distort the causal analysis.

3. Propensity Score Estimation

The probability of treatment is estimated conditional on the selected observed covariates. The model should be judged primarily according to whether it contributes to the construction of adequately balanced groups rather than according to conventional predictive performance alone.

4. Overlap Assessment

Researchers determine whether treated and untreated observations occupy sufficiently comparable regions of the data. Where appropriate matches do not exist, the causal question may need to be restricted to a population for which meaningful comparisons are possible.

5. Matching

An appropriate matching algorithm is used to construct the comparison sample. Choices concerning matching ratios, replacement, calipers and other specifications should be guided by their implications for comparability, sample retention and the treatment effect being estimated.

6. Covariate Balance

Observed baseline characteristics are compared after matching. Adequate balance provides evidence that the matching procedure has achieved its intended purpose for measured covariates. Inadequate balance signals that the design should be reconsidered before outcome effects are interpreted.

7. Unmeasured Confounding

Even excellent observed balance does not demonstrate that relevant unobserved characteristics are balanced. Researchers should identify plausible unmeasured confounders and, where feasible, use sensitivity analysis or other supporting evidence to investigate how strongly such factors could affect the conclusion.

8. Treatment-Effect Estimation

Once an acceptable matched design has been established, outcomes are compared using analytical procedures appropriate to the matched data and the causal estimand. The matched structure of the data should be respected rather than treated as if the observations had been sampled independently without matching.

9. Defensible Causal Claim

The final interpretation should reflect both what matching accomplished and what it could not establish. Strong observed balance and adequate overlap can strengthen causal credibility, but conclusions remain conditional on assumptions concerning unmeasured confounding and other features of the research design.

Four distinctions capture the logic of the framework:

Similar propensity scores ≠ automatically comparable groups

Successful matching ≠ randomization

Covariate balance ≠ absence of unmeasured confounding

Statistical significance ≠ causal credibility

The central principle is:

PSM strengthens causal inference when it constructs a credible comparison on observed pre-treatment characteristics and the remaining assumptions are substantively defensible—not simply when a matching algorithm successfully produces pairs.

Application of PSM: an Example

Suppose a commercial bank introduces an optional advanced analytics training programme for relationship managers. Management wants to determine whether participation increases the subsequent value of client portfolios, but employees were not randomly assigned to training. Participants tend to have more experience, stronger previous performance, larger portfolios and greater exposure to corporate clients than non-participants.

The researcher first defines the causal question as the effect of participating in the programme on portfolio growth during the following twelve months. Pre-training characteristics that plausibly influence both programme participation and subsequent performance are identified using organisational knowledge and existing research. These include tenure, previous portfolio growth, portfolio size, client composition, managerial grade and region. Only information measured before training is used for the matching design.

Propensity scores are estimated for participants and non-participants, after which the researcher examines their distributions for evidence of common support. Several highly experienced programme participants have propensity scores substantially above those of every non-participant. Rather than forcing inappropriate matches, these observations are excluded from the matched comparison, and the population represented by the eventual treatment-effect estimate is explicitly narrowed.

Participants are then matched with comparable non-participants. Balance diagnostics show that the substantial pre-matching differences in experience, previous performance and portfolio characteristics have been considerably reduced. However, one important covariate remains noticeably imbalanced, so the researcher revises the propensity-score specification and repeats the matching procedure before examining post-training outcomes.

Only after satisfactory observed balance has been established does the researcher compare subsequent portfolio growth. The final interpretation acknowledges that PSM has improved comparability on measured pre-training characteristics but cannot rule out unmeasured differences such as personal motivation. The resulting causal argument is therefore based on the entire design and its assumptions rather than on the existence of matched pairs alone.

Advantages and Limitations of PSM

PSM provides an intuitive way to improve comparisons when treatment groups differ substantially before an intervention. By focusing analysis on treated and untreated observations that are sufficiently similar in observed pre-treatment characteristics, researchers can avoid some of the direct comparisons that make naïve observational estimates misleading. Matching can also make the resulting research design more transparent because researchers can inspect which observations were retained, whether adequate overlap exists and how covariate distributions changed after matching.

The method is particularly useful when many observed characteristics influence treatment assignment. Matching directly on numerous variables can become impractical as dimensionality increases, whereas the propensity score provides a balancing score through which these characteristics can contribute to the matching process. This has made propensity-score methods useful across fields including management, information systems and programme evaluation, where randomized interventions are frequently unavailable.

Its principal limitation follows directly from the information used to construct the score. PSM can balance observed characteristics but cannot automatically balance variables that were not measured or appropriately represented. If entrepreneurial ability affects both participation in a business-support programme and subsequent company performance but is absent from the data, excellent balance on firm size, age, industry and previous revenue does not eliminate the possibility of confounding by entrepreneurial ability.

Loss of observations can create another trade-off. Enforcing common support and rejecting poor matches can improve internal comparability while reducing sample size and changing the population represented by the analysis. This is not necessarily a methodological failure; excluding incomparable observations may be preferable to retaining misleading comparisons. However, the resulting treatment effect must be interpreted for the population actually represented by the matched sample.

Results can also be sensitive to modelling and matching decisions. Covariate selection, propensity-score specification, caliper width, matching ratio and replacement rules can affect which observations remain and how well they balance. Researchers should therefore avoid treating PSM as an automatic transformation that turns observational data into experimental evidence.

Common Mistakes When Using PSM

Selecting covariates because they significantly predict treatment can misrepresent the purpose of the propensity-score model. Covariate selection should be informed by causal and substantive knowledge about treatment assignment and the outcome. A variable’s statistical significance in the treatment model is not the criterion that determines whether it is important for confounding control.

Examining the treatment outcome before establishing an acceptable matched design can encourage specification choices that favour a desired result. Conceptually, matching belongs to the design stage. Researchers should determine whether groups are comparable using pre-treatment information and balance diagnostics before allowing outcome differences to drive decisions about the matching procedure.

A large matched sample is not necessarily a good matched sample. Retaining every treated observation may appear desirable, but forcing matches in regions of poor overlap can compare units that are fundamentally dissimilar. Discarding observations without credible counterparts can sometimes strengthen the causal design even though the sample becomes smaller.

A successful software command does not demonstrate covariate balance. Matching algorithms will often return matched observations regardless of whether the resulting groups are sufficiently comparable. Researchers need to examine standardized differences and, where appropriate, other distributional diagnostics after matching.

Balance on the propensity score should not be confused with balance on the underlying covariates. Two groups can have similar score distributions while important covariate differences remain. The variables that matter for the causal comparison therefore require direct diagnostic attention.

Describing PSM as equivalent to randomization overstates what the method achieves. Matching can improve observed covariate balance, whereas randomization addresses both observed and unobserved characteristics probabilistically through the treatment-assignment mechanism. PSM remains dependent on assumptions about unmeasured confounding.

PSM in Business Research

PSM is particularly relevant to business research because organisational treatments are rarely allocated randomly. Firms choose whether to adopt technologies, employees decide whether to participate in training, customers select subscription plans and organisations choose whether to enter international markets. The same factors influencing those choices frequently affect subsequent outcomes, creating selection bias in simple treated-versus-untreated comparisons.

Management scholarship has therefore used propensity-score approaches to strengthen causal analysis with observational data, while information-systems research has used matching within a potential-outcomes framework to investigate questions such as returns to education among technology professionals. More recently, methodological work specifically addressing business marketing has emphasized both the potential usefulness of propensity-score modelling and the risk of treating it as a universal solution to endogeneity.

Consider research examining whether exporting improves SME productivity. Exporting firms may already differ from non-exporters in size, managerial capability, innovation, capital availability and previous productivity. PSM could be used to construct a comparison group of non-exporters resembling exporters on relevant observed pre-export characteristics. The resulting comparison may be more credible than comparing all exporters with all domestic firms, but it would still depend on whether important determinants of both exporting and productivity have been adequately observed.

This illustrates the method’s broader business-research value. PSM is most useful not as a way of making two datasets look statistically similar, but as part of an explicit argument about selection into a business treatment and the counterfactual outcome that would otherwise have occurred.

PSM in the Age of AI and Digital Research

AI and machine learning expand the range of techniques available for estimating propensity scores. Flexible algorithms can model nonlinear relationships and complex interactions that may be difficult to specify manually through conventional logistic regression. In high-dimensional datasets, this can help researchers model treatment assignment more flexibly.

Greater modelling power does not change the methodological objective. A machine-learning model that predicts treatment assignment extremely accurately may actually reveal severe separation between treatment groups rather than solve the causal problem. If treated and untreated observations occupy fundamentally different regions of the covariate space, no algorithm can create genuine counterfactual observations that the dataset does not contain.

AI can nevertheless contribute meaningfully to the PSM workflow. It can assist researchers in exploring covariate distributions, generating balance diagnostics, identifying regions of poor overlap, comparing matching specifications and conducting sensitivity analyses. Generative AI can also help explain why particular variables might represent confounders, although such suggestions require substantive verification rather than automatic inclusion.

The major AI-era risk is methodological automation without causal understanding. Researchers can increasingly ask software to “perform propensity score matching” and receive technically polished results without understanding why variables were selected, what population was discarded, whether balance improved or which unmeasured confounders remain plausible. Automating the matching procedure does not automate the justification for causal inference.

For this reason, AI makes the distinction between prediction and causal design more important rather than less important. The question remains whether the observed data, assignment process and assumptions provide a defensible counterfactual—not whether the propensity-score model is technologically sophisticated.

When to Use PSM

Propensity Score Matching may be appropriate when:

  • the research question concerns the effect of a treatment, intervention, programme or exposure;
  • random assignment is unavailable;
  • treated and untreated groups differ on observed pre-treatment characteristics;
  • researchers possess sufficiently rich pre-treatment information about plausible confounders;
  • both treatment conditions occur among sufficiently comparable observations;
  • adequate common support or overlap can be established;
  • matching can meaningfully improve observed covariate balance;
  • the researcher can justify the assumption that important confounding has been sufficiently measured;
  • the population and treatment effect represented by the matched sample correspond to the research question.

PSM becomes less attractive when treatment groups have little meaningful overlap, important confounders are known to be unmeasured or treatment assignment is driven predominantly by processes that the available data cannot represent. In these circumstances, a different causal design may be more defensible than attempting to force comparability through matching.

Dissertation Example

A dissertation titled “The Impact of Digital Transformation Grants on the Performance of Small and Medium-Sized Enterprises” could use PSM when grant recipients were selected or self-selected rather than randomly assigned. A simple comparison of recipient and non-recipient firms would be vulnerable to selection bias because firms applying for grants may already differ in size, previous growth, digital capability, management quality or access to finance.

The methodology chapter would define the treatment as receipt of the digital-transformation grant and specify the post-treatment performance outcome. Drawing on theory and knowledge of the grant-selection process, the researcher would identify relevant pre-treatment covariates and explain why they may influence both grant receipt and subsequent firm performance. Propensity scores would then be estimated using information measured before the grant was received.

The researcher would describe the matching procedure and examine common support before presenting balance diagnostics for the matched sample. Standardized differences could be reported before and after matching to demonstrate whether important observed characteristics became more comparable. Firms lacking credible counterparts might be excluded, with the implications for the population represented by the estimate explained explicitly.

Only after establishing the quality of the matched design would the dissertation estimate differences in post-grant performance. The methodology chapter would acknowledge that PSM addresses observed pre-treatment differences rather than guaranteeing equivalence on unmeasured factors such as entrepreneurial motivation. The causal conclusion would consequently be presented as conditional on the assumptions supporting the matching design rather than as a consequence of PSM alone.

Exam Tip

When defending PSM, do not say simply that it was used “to make the treatment and control groups similar.” Explain what kind of similarity the method can establish and why that matters for the causal question.

A stronger defence would state that PSM was used to improve comparability between treated and untreated observations with respect to relevant observed pre-treatment covariates, thereby constructing a more credible counterfactual comparison than the unmatched sample provided. You should then be prepared to explain how covariates were selected, whether adequate overlap existed, how matching was performed and how post-matching balance was assessed.

If an examiner asks whether successful matching removes selection bias, the answer requires qualification. PSM can reduce bias associated with appropriately measured and modelled observed characteristics, but it does not automatically eliminate bias caused by unmeasured confounding.

The principle worth remembering is:

Matching creates the comparison; balance diagnostics evaluate the comparison; causal assumptions determine how far that comparison can be interpreted causally.

Need to determine whether Propensity Score Matching is appropriate for your dissertation—and how to justify it within your causal research design?

Use Dudovskiy Research Assistant to develop a methodology tailored to your research topic, including the research design, sampling strategy, data analysis and academic justification for your methodological choices.

References

Rosenbaum, P.R. and Rubin, D.B. (1983) ‘The central role of the propensity score in observational studies for causal effects’, Biometrika, 70(1), pp. 41–55.

Dehejia, R.H. and Wahba, S. (2002) ‘Propensity score-matching methods for nonexperimental causal studies’, Review of Economics and Statistics, 84(1), pp. 151–161.

Austin, P.C. (2009) ‘Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples’, Statistics in Medicine, 28(25), pp. 3083–3107.

Austin, P.C. (2011) ‘An introduction to propensity score methods for reducing the effects of confounding in observational studies’, Multivariate Behavioral Research, 46(3), pp. 399–424.

Li, M. (2013) ‘Using the propensity score method to estimate causal effects: A review and practical guide’, Organizational Research Methods, 16(2), pp. 188–226.

Mithas, S. and Krishnan, M.S. (2009) ‘From association to causation via a potential outcomes approach’, Information Systems Research, 20(2), pp. 295–313.

Kainz, K., Greifer, N., Givens, A., Swietek, K.M., Lombardi, B.M., Zietz, S. and Kohn, J.L. (2017) ‘Improving causal inference: Recommendations for covariate selection and balance in propensity score methods’, Journal of the Society for Social Work and Research, 8(2), pp. 279–303.

[]