Regression Discontinuity Design
Regression Discontinuity Design (RDD) is a quasi-experimental research design used to estimate causal effects when treatment or intervention assignment changes at a known threshold or cutoff on a measurable variable. Rather than attempting to make all treated and untreated observations comparable, RDD focuses on observations located close to the cutoff and asks whether crossing that threshold produces a discontinuous change in the outcome.
For example, suppose businesses become eligible for a financial support programme when their eligibility score reaches 70. Businesses scoring 69.8 and 70.2 may be very similar in underlying characteristics, yet the programme’s eligibility rule places them on opposite sides of the treatment threshold. Comparing outcomes immediately around that cutoff can therefore provide a credible basis for causal inference when the assumptions of the design are satisfied.
RDD is regarded as a particularly important quasi-experimental design because treatment assignment is determined at least partly by whether an observed assignment variable lies above or below a fixed threshold. Its credibility depends on the assignment process and on what happens around that threshold, rather than simply on fitting a regression containing a cutoff variable.
On this page:
- Regression Discontinuity Design explained simply
- What is Regression Discontinuity Design?
- How RDD works
- Running variables and cutoffs
- Sharp vs fuzzy RDD
- Key assumptions of RDD
- Manipulation around the cutoff
- Bandwidth selection
- Estimating and interpreting the discontinuity
- Dudovskiy RDD Causal Credibility Framework
- Application of RDD
- Advantages and limitations
- Common mistakes
- RDD in business research
- RDD in the age of AI and digital research
- When to use RDD
- Dissertation example
- Exam tip
| Element | Meaning in RDD |
|---|---|
| Running variable | Variable that determines treatment eligibility relative to a cutoff |
| Cutoff | Threshold at which treatment assignment or probability changes |
| Treatment | Intervention, programme or exposure being evaluated |
| Outcome | Variable potentially affected by treatment |
| Discontinuity | Change at the cutoff used to identify the treatment effect |
| Bandwidth | Range of observations around the cutoff included in estimation |
| Main causal focus | Effect at or close to the cutoff |
| Central concern | Whether observations around the threshold provide a credible counterfactual comparison |
Regression Discontinuity Design Explained Simply
Imagine that a retailer offers intensive management support only to stores whose annual performance score falls below 60. A store scoring 59.8 receives the intervention, while another scoring 60.2 does not.
The two stores are on opposite sides of the rule, but their scores are almost identical. If managers cannot precisely manipulate their stores’ scores and no other important policy suddenly changes at 60, stores immediately around the threshold may provide a useful comparison. Researchers can examine whether subsequent performance changes discontinuously at exactly the point where eligibility changes.
The important comparison is therefore not simply:
stores receiving support vs stores not receiving support
but:
stores immediately on one side of the cutoff vs stores immediately on the other side.
RDD uses this local comparison to construct the counterfactual needed for causal inference.
What Is Regression Discontinuity Design?
Regression Discontinuity Design is a quasi-experimental approach in which treatment assignment is determined wholly or partly by whether a measured assignment variable—often called the running variable, forcing variable or score—crosses a predefined cutoff. The method was introduced by Thistlethwaite and Campbell and subsequently developed extensively in economics and other social sciences.
Suppose XX represents the running variable and cc represents the cutoff. Units with values on one side of cc are assigned, or become more likely to be assigned, to treatment. Researchers then examine whether the expected value of outcome YY changes discontinuously at cc.
The intuition is that observations extremely close to the cutoff may differ very little in relevant characteristics apart from their treatment status. Under the continuity-based interpretation of RDD, if potential outcomes would otherwise evolve smoothly through the cutoff, a discontinuity in observed outcomes at that point can be attributed to treatment. This identification logic underlies the canonical sharp RDD.
This makes RDD fundamentally a research design, rather than merely a particular regression equation. The statistical model is used to estimate the discontinuity created by an assignment mechanism whose causal credibility must first be justified.
How Regression Discontinuity Design Works
Consider a government programme providing subsidized consulting to small businesses with an eligibility score of 50 or below. The running variable is the eligibility score, 50 is the cutoff, programme participation is the treatment, and subsequent business performance is the outcome.
Researchers examine observations near 50 and estimate the relationship between the eligibility score and subsequent performance separately on either side of the threshold. If the expected outcome would have changed smoothly through 50 in the absence of the programme, but instead exhibits a distinct jump at 50, that discontinuity provides evidence about the programme’s causal effect near the cutoff.
Conceptually:
Running Variable → Cutoff Rule → Change in Treatment → Discontinuity in Outcome
The strength of the causal argument does not come merely from observing that firms above and below 50 have different outcomes. It comes from demonstrating that the difference appears specifically at the treatment threshold and that competing explanations for the discontinuity are implausible.
The Running Variable and Cutoff
The running variable is the variable that determines where each observation lies relative to the treatment threshold. Examples might include an income measure determining subsidy eligibility, a credit score determining access to a programme, an employee performance score determining training eligibility, or a firm-size measure determining whether a regulation applies.
The cutoff is the value at which treatment assignment changes. A credible RDD therefore requires a genuine assignment rule or institutional mechanism connecting the running variable to treatment. Simply choosing an interesting value from a continuous variable after examining the data does not create an RDD.
This distinction is crucial. RDD exploits a discontinuity in an assignment mechanism. A researcher cannot ordinarily convert a conventional observational dataset into a credible RDD merely by selecting an arbitrary threshold and fitting separate regression lines around it.
Sharp and Fuzzy Regression Discontinuity Designs
Two forms of RDD should be distinguished.
| Sharp RDD | Fuzzy RDD | |
|---|---|---|
| Treatment rule | Treatment changes deterministically at cutoff | Probability of treatment changes at cutoff |
| Compliance | Perfect | Imperfect |
| At threshold | Assignment completely determines treatment | Assignment strongly influences treatment |
| Main complication | Estimating discontinuity appropriately | Distinguishing assignment from actual treatment |
| Connection to IV | Usually not required for basic identification | Closely related to instrumental variables estimation |
In a sharp RDD, crossing the threshold completely determines treatment status. In the canonical sharp design, compliance with treatment assignment is perfect.
Suppose firms with fewer than 50 employees are legally exempt from a regulation and all firms with 50 or more employees are subject to it. If the rule is perfectly implemented, treatment changes sharply at 50 employees.
In a fuzzy RDD, crossing the cutoff changes the probability of receiving treatment but does not perfectly determine treatment. Some observations assigned to treatment may not actually receive it, while others on the nominally untreated side may nevertheless receive treatment. Modern RDD methodology explicitly treats fuzzy designs as an important extension of the canonical sharp design.
Fuzzy RDD also creates an important connection with Instrumental Variables (IV). The threshold-based assignment can provide instrumental variation in actual treatment received, allowing researchers to estimate a local causal effect under the required assumptions. This connection should not be interpreted as meaning that RDD and IV are identical methods; rather, IV reasoning becomes important for causal estimation in fuzzy RDD.
Key Assumptions of Regression Discontinuity Design
RDD does not become credible simply because treatment changes at a numerical cutoff. The causal interpretation rests on assumptions about what would have happened to observations around the threshold in the absence of treatment.
A central continuity-based assumption is that the relevant potential outcomes would change smoothly through the cutoff without the intervention. If so, there should be no sudden outcome jump at exactly that value except for a change generated by treatment. The identification result in canonical RDD follows from continuity of the conditional expectations of potential outcomes at the cutoff.
Researchers must therefore consider whether other factors also change at the same threshold. If firms reaching 50 employees simultaneously become subject to several different regulations, for example, an observed outcome discontinuity at 50 cannot automatically be attributed to one particular regulation.
Another important issue concerns control over the running variable. If participants can precisely manipulate the value determining treatment, observations just above and below the cutoff may no longer provide the comparison that RDD requires.
Manipulation Around the Cutoff
Suppose businesses qualify for a valuable grant when their reported annual revenue is below $1 million. If firms can strategically adjust or report revenue so that they fall just below the threshold, firms immediately on either side of $1 million may systematically differ.
This threatens the RDD logic because treatment status is no longer generated by an assignment process that leaves observations close to the cutoff suitably comparable. Researchers should therefore understand the institutional process generating the running variable and examine whether sorting or manipulation around the cutoff is plausible.
Statistical diagnostics can help detect suspicious changes in the density of observations around a threshold, and modern RDD methodology includes manipulation testing as part of design validation. However, such tests should complement rather than replace substantive knowledge of how the score was generated. The absence of statistically detectable sorting does not by itself prove that the entire causal design is valid.
Bandwidth Selection in RDD
RDD is fundamentally concerned with observations near the cutoff. The bandwidth determines how far from that cutoff observations are included in the primary estimation.
A narrow bandwidth concentrates the analysis on observations that are more local to the threshold, potentially making the comparison more credible, but it also reduces the number of observations available and can decrease statistical precision. A wider bandwidth provides more data but introduces observations further from the threshold, where the relationship between the running variable and outcome may be more difficult to model and where treated and untreated units may be less comparable.
Bandwidth selection is therefore not merely a cosmetic statistical choice. It reflects a bias–precision trade-off and should be handled transparently rather than selected after repeatedly trying alternatives until a statistically significant effect appears. Practical RDD work consequently gives substantial attention to bandwidth selection and to checking whether substantive conclusions remain credible under reasonable alternative specifications. (World Bank Blogs)
Estimating and Interpreting the Discontinuity
The graphical representation of RDD is especially informative. Researchers plot the outcome against the running variable, identify the cutoff, and examine whether the fitted relationship exhibits a jump at that point.
Conceptually, the treatment effect at the cutoff can be represented as:
Treatment effect at cutoff = Expected outcome just above cutoff − Expected outcome just below cutoff
with the precise direction depending on which side receives treatment.
The apparent simplicity of this comparison should not obscure the methodological work required to support it. Researchers need to justify the functional form or local estimation strategy, select an appropriate bandwidth, examine observations around the threshold and conduct suitable validation and falsification checks. Modern practical treatments of RDD devote explicit attention to graphical presentation, estimation, inference and falsification rather than treating the method as a single regression specification. (Cambridge Assets)
The resulting causal estimate is usually local. It describes the treatment effect for observations at or close to the cutoff under the relevant RDD assumptions. Researchers should therefore be cautious about extrapolating the estimate to observations located far away from the threshold.
Dudovskiy RDD Causal Credibility Framework
The Dudovskiy RDD Causal Credibility Framework provides a structured way to evaluate whether a threshold-based research setting can support a defensible causal claim. It synthesizes established principles of regression discontinuity methodology into a practical sequence for designing, evaluating and defending an RDD study; it does not propose a new form of causal identification.
Causal Question → Assignment Variable → Cutoff Rule → Treatment Discontinuity → Manipulation Assessment → Continuity and Comparability → Bandwidth Selection → Outcome Discontinuity → Robustness Evidence → Local Causal Claim
The central principle is:
RDD becomes causally credible when crossing a meaningful threshold changes treatment while observations immediately around that threshold remain otherwise sufficiently comparable.

1. Causal Question
Begin by specifying the treatment, outcome and population. The research question should ask about the causal effect of a treatment for which a genuine threshold-based assignment mechanism exists.
2. Assignment Variable
Identify the variable determining position relative to treatment eligibility. Researchers need to understand how this variable is measured, generated and recorded because the assignment mechanism forms part of the causal argument.
3. Cutoff Rule
Establish the threshold and explain why it exists independently of the researcher’s analysis. A policy rule, administrative criterion, institutional regulation or predetermined eligibility requirement provides a stronger foundation than a cutoff chosen retrospectively because it produces an interesting result.
4. Treatment Discontinuity
Determine what actually happens to treatment at the threshold. If treatment switches completely, the setting may support a sharp RDD. If only the probability of treatment changes, fuzzy RDD may be more appropriate.
5. Manipulation Assessment
Ask whether individuals, firms or administrators could precisely influence their position relative to the threshold. Institutional reasoning and empirical diagnostics should be used together when evaluating this possibility.
6. Continuity and Comparability
Assess whether other determinants of the outcome are likely to change smoothly through the cutoff. Researchers should investigate whether other policies, eligibility conditions, behavioural responses or institutional changes occur at the same point.
7. Bandwidth Selection
Choose and justify the neighborhood around the cutoff used for estimation. The aim is to preserve the local comparison while retaining enough information for meaningful estimation.
8. Outcome Discontinuity
Estimate whether the outcome changes at the cutoff and quantify the magnitude and uncertainty of that discontinuity. The statistical estimate should be interpreted in the context of the assignment mechanism rather than treated as standalone proof of causality.
9. Robustness Evidence
Investigate whether the causal interpretation survives appropriate alternative specifications and falsification exercises. Depending on the study, these may include alternative bandwidths, examination of predetermined covariates, placebo cutoffs and tests for manipulation or sorting.
10. Local Causal Claim
The final claim should match what the design actually identifies. A credible discontinuity at the threshold supports a causal conclusion for the population represented around that cutoff; it does not automatically establish the same treatment effect for the entire population.
The framework can therefore be summarized as:
Threshold-based assignment + credible local comparability + appropriate estimation + supporting robustness evidence = defensible local causal inference
Importantly:
An outcome jump at a cutoff ≠ automatically a causal treatment effect.
The assignment mechanism and identifying assumptions must make the discontinuity causally interpretable.
Application of Regression Discontinuity Design: an Example
Suppose a government introduces a digitalization grant for manufacturing firms with an innovation-readiness score below 60. Eligible firms receive funding for new software, automation equipment and employee digital-skills training. A researcher wants to determine whether receiving the grant subsequently increases labour productivity.
Simply comparing grant recipients with all non-recipients would be problematic. Firms scoring 30 may differ substantially from firms scoring 85 in managerial capabilities, technological sophistication and previous investment. Those differences could influence productivity independently of the grant.
An RDD would instead concentrate on manufacturers immediately around the eligibility threshold. Firms scoring 59.7 and 60.3 may have very similar underlying innovation readiness, yet the eligibility rule places them on opposite sides of the grant threshold. The researcher would first verify the exact assignment mechanism and investigate whether firms or programme administrators could manipulate scores around 60. It would also be necessary to determine whether any other important programme or regulation changes at the same cutoff.
The analysis would then examine whether actual grant receipt changes at 60 and whether subsequent productivity exhibits a corresponding discontinuity. Appropriate bandwidth selection and robustness analyses would test whether the finding depends excessively on a particular specification or group of observations.
If these conditions are supported, the researcher could make a causal claim about the effect of grant eligibility or receipt for firms around the score of 60. It would be much harder to justify a claim that the estimated effect applies equally to manufacturers with scores of 20 or 90.
Advantages and Limitations of Regression Discontinuity Design
RDD can produce highly credible causal evidence without randomized treatment assignment when a genuine threshold rule exists. Its strength comes from the institutional assignment mechanism: observations immediately around a cutoff can sometimes provide a counterfactual comparison substantially more convincing than broad comparisons between treated and untreated populations. For this reason, RDD is often discussed as a quasi-experimental design with a particularly strong internal-validity logic when its assumptions are satisfied. (National Bureau of Economic Research)
The design is also relatively transparent. A well-constructed RDD graph can make the assignment threshold and estimated discontinuity visible, helping readers understand where the causal estimate originates. Researchers can supplement the primary estimate with manipulation diagnostics, covariate checks, alternative bandwidths and falsification exercises, creating a cumulative argument about design credibility.
Its applicability, however, is restricted by the need for a meaningful threshold-based assignment process. Many research questions simply do not have such a mechanism. Even when a cutoff exists, manipulation, simultaneous policies or discontinuities in other determinants of the outcome can undermine identification.
External validity presents another important limitation. The strongest causal interpretation ordinarily concerns observations near the cutoff. This is often exactly what makes RDD internally persuasive, but it also means that an effect estimated for borderline firms, employees or consumers should not automatically be generalized to substantially different members of the population.
RDD may also require substantial data around the cutoff. Restricting analysis to an appropriate bandwidth can discard observations located farther away, reducing the effective sample available for estimation. This illustrates why a large overall dataset does not necessarily imply high precision for an RDD estimate. (World Bank Blogs)
Common Mistakes When Using Regression Discontinuity Design
Treating any numerical threshold as an RDD is a fundamental conceptual error. The cutoff should form part of a genuine treatment-assignment mechanism. Creating a threshold retrospectively because an outcome appears to change at that value does not provide the quasi-experimental logic on which RDD depends.
A visually impressive jump in the outcome is also insufficient. Researchers need to explain why other determinants of the outcome should evolve smoothly through the threshold and investigate whether another policy, behavioural change or institutional rule occurs at the same value.
Using the entire dataset without considering locality can weaken the design. Observations far from the cutoff may be fundamentally different from one another, and complicated global polynomial models can make results sensitive to modelling choices. The central RDD comparison concerns behaviour around the threshold, not the best-fitting curve across every available observation.
Bandwidth searching creates a related problem. Trying many bandwidths and reporting only the one producing the desired statistical result converts a methodological decision into a form of specification searching. Researchers should justify their main approach and use alternative reasonable bandwidths to assess robustness.
Manipulation of the running variable should not be treated as a purely statistical issue. A density test can provide useful evidence, but researchers also need to understand whether participants know the rule, have incentives to manipulate their score and possess the ability to do so.
Finally, causal language should not exceed the population identified by the design. An RDD demonstrating an effect among firms around an eligibility threshold does not automatically demonstrate that the intervention would have the same effect among firms located throughout the entire score distribution.
Regression Discontinuity Design in Business Research
RDD can be particularly useful in business research because firms, employees and consumers frequently encounter threshold-based institutional rules. Credit scores can determine financing eligibility, employee performance scores can determine bonuses or training, firm size can trigger regulatory requirements, customer status thresholds can change benefits, and financial indicators can determine access to support programmes.
Such settings create opportunities to move beyond correlations. A researcher interested in whether financing improves investment, for example, may find that loan eligibility changes at a predetermined credit threshold. A human-resource researcher may encounter a performance threshold that determines access to additional training. A strategy researcher may investigate whether crossing a firm-size threshold activates a regulation that subsequently changes investment or employment decisions.
The presence of a threshold should nevertheless be treated as the beginning of the methodological assessment rather than proof that RDD is appropriate. The assignment process, opportunities for manipulation, competing changes at the cutoff, availability of observations near the threshold and local nature of the resulting causal estimate all need to be considered.
Regression Discontinuity Design in the Age of AI and Digital Research
Digital administrative systems may expand opportunities for RDD because organizations increasingly use numerical scores and algorithmic thresholds to allocate loans, promotions, advertising opportunities, platform benefits, risk interventions and other treatments. Such systems can create unusually precise records of running variables, cutoff rules and subsequent outcomes, making threshold-based causal designs feasible in settings where the underlying assignment process can be documented.
AI can also assist researchers with coding, visualization, specification checks and exploration of alternative bandwidths. These capabilities make technically demanding analyses more accessible, but they do not resolve the core identification problem. An AI system can fit hundreds of RDD specifications; it cannot establish from the regression output alone that a cutoff was institutionally meaningful, that decision-makers could not manipulate assignment, or that no other intervention changed at the same threshold.
Algorithmic assignment introduces additional complications. A threshold may appear simple in observed data while the underlying decision system incorporates hidden variables, manual overrides or periodically changing models. Researchers using digitally generated assignment scores therefore need to understand the operational process behind them rather than assuming that an apparent discontinuity represents a stable quasi-experiment.
AI makes statistical implementation easier. It does not make causal identification automatic. In RDD, understanding why the threshold exists and what actually changes when it is crossed remains more important than the sophistication of the software used to estimate the discontinuity.
When to Use Regression Discontinuity Design
RDD may be appropriate when:
- treatment eligibility or treatment probability changes at a clearly defined cutoff;
- the cutoff existed independently of the researcher’s analysis;
- a meaningful running variable determines position relative to that cutoff;
- sufficient observations are available around the threshold;
- units cannot precisely manipulate their treatment status around the cutoff, or manipulation can be convincingly addressed;
- no important competing intervention changes at exactly the same threshold;
- potential outcomes can reasonably be expected to evolve smoothly around the cutoff in the absence of treatment;
- the research objective concerns causal inference rather than simple association;
- a local treatment effect around the cutoff is substantively meaningful;
- the researcher can conduct appropriate validation, robustness and falsification analyses.
RDD should not be selected merely because the dataset contains a continuous variable that can conveniently be divided at some value.
Dissertation Example
A Master’s dissertation investigates whether access to subsidized management consulting improves the productivity of small businesses. Under the government programme studied, businesses receiving an administrative capability score below 50 become eligible for subsidized consulting, while businesses scoring 50 or above are not normally eligible.
In the methodology chapter, the student explains that conventional comparison of participating and non-participating firms could produce selection bias because businesses entering the programme may differ systematically from those that do not. The dissertation therefore adopts a regression discontinuity design centred on the predetermined eligibility threshold of 50. The administrative capability score is defined as the running variable, programme eligibility as the treatment-assignment mechanism, and labour productivity twelve months later as the principal outcome.
The student justifies the causal strategy by arguing that firms located immediately around the threshold should be substantially more comparable than firms located far apart in the score distribution. The methodology then explains how possible manipulation of scores will be investigated, how the bandwidth around the threshold will be selected, and how predetermined firm characteristics and alternative bandwidths will be examined as part of the robustness analysis.
The dissertation explicitly limits interpretation of the estimated effect to firms around the eligibility cutoff. Rather than concluding that subsidized consulting increases productivity for all small businesses, the student states that the RDD is designed to estimate the causal effect for businesses close to the programme’s eligibility threshold. This distinction demonstrates an understanding of both the internal strength and the external-validity limitation of the research design.
Exam Tip
When explaining Regression Discontinuity Design in an exam, do not define it simply as “comparing observations above and below a cutoff.” That description misses the causal logic.
A stronger answer explains that treatment assignment or treatment probability changes discontinuously at a predetermined threshold on a running variable. If potential outcomes would otherwise evolve continuously through the cutoff and units cannot precisely manipulate their position around it, observations close to the threshold can provide a credible counterfactual comparison. The resulting treatment effect should normally be interpreted as local to the cutoff.
A concise formulation to remember is:
RDD exploits a discontinuity in treatment assignment to identify a local causal effect at a threshold.
Build and defend your causal research design
If your dissertation involves an eligibility threshold, policy rule, score-based intervention or another quasi-experimental setting, Dudovskiy Research Assistant can help determine whether Regression Discontinuity Design is appropriate for your research topic and explain the methodological assumptions and limitations you would need to defend.
References
Cattaneo, M.D., Idrobo, N. and Titiunik, R. (2020). A Practical Introduction to Regression Discontinuity Designs: Foundations. Cambridge University Press. The canonical sharp RDD discussed by the authors involves a continuously distributed score, a single cutoff and perfect treatment compliance. (Cambridge University Press)
Cattaneo, M.D., Idrobo, N. and Titiunik, R. (2024). A Practical Introduction to Regression Discontinuity Designs: Extensions. Cambridge University Press. The extensions include local-randomization approaches, fuzzy RDD, discrete scores and multidimensional RDD. (Cambridge University Press)
Imbens, G.W. and Lemieux, T. (2008). “Regression discontinuity designs: A guide to practice.” Journal of Econometrics, 142(2), 615–635. (National Bureau of Economic Research)
Lee, D.S. and Lemieux, T. (2010). “Regression Discontinuity Designs in Economics.” Journal of Economic Literature, 48(2), 281–355. (National Bureau of Economic Research)
Thistlethwaite, D.L. and Campbell, D.T. (1960). “Regression-discontinuity analysis: An alternative to the ex post facto experiment.” Journal of Educational Psychology, 51(6), 309–317. The design’s origins are also documented in later methodological reviews. (National Bureau of Economic Research)
