Chi-Square Test
The chi-square (χ²) test is a family of statistical procedures used to compare observed frequencies with frequencies expected under a specified null hypothesis. In dissertation and business research, its two most common applications are the chi-square goodness-of-fit test, which examines the distribution of one categorical variable, and the chi-square test of independence, which examines whether two categorical variables are associated.
Unlike a t-test or ANOVA, the chi-square test does not primarily compare means. It works with counts of observations in categories. For a test of independence, expected cell counts represent what would be anticipated if the two categorical variables were independent; the observed and expected counts are then compared to determine whether the departures are greater than would reasonably be attributed to sampling variation. (Online Statistics at Penn State)
Selecting a chi-square test therefore begins with the categorical research question rather than with the presence of numbers in a dataset. Researchers must establish what the categories represent, whether observations are independent, whether expected frequencies support the chi-square approximation, and what a statistically significant result actually reveals about the pattern of association.
On this page:
- Chi-square test explained simply
- What is a chi-square test?
- How the chi-square test works
- Observed and expected frequencies
- Chi-square goodness-of-fit test
- Chi-square test of independence
- Goodness-of-fit vs independence
- Assumptions
- Expected frequencies and sparse cells
- A significant chi-square does not identify where the association occurs
- Standardized residuals
- Effect size and Cramer’s V
- Chi-square vs Fisher’s exact test
- Dudovskiy Chi-Square Test Selection Framework
- Application example
- Advantages and limitations
- Common mistakes
- Chi-square tests in business research
- Chi-square tests in the age of AI
- When to use a chi-square test
- Dissertation example
- Exam tip
| Research question | Appropriate direction |
|---|---|
| Does one categorical variable follow a specified distribution? | Chi-square goodness-of-fit test |
| Are two categorical variables associated? | Chi-square test of independence |
| Which cells contribute to a significant association? | Examine cell residuals after the omnibus test |
| How strong is the association? | Consider an appropriate effect-size measure such as Cramer’s V |
| Are expected frequencies too sparse for the usual approximation? | Consider whether an exact or alternative procedure is more appropriate |
Chi-Square Test Explained Simply
Suppose an online retailer asks 300 customers which of three delivery options they prefer: standard delivery, collection point or express delivery. If the company has a justified reason to expect preferences to be distributed equally across the three options, it would expect approximately 100 customers in each category.
Imagine instead that 150 choose standard delivery, 90 choose collection and 60 choose express delivery. The observed frequencies are clearly not identical to the expected frequencies. The chi-square goodness-of-fit test evaluates whether the overall discrepancy between what was observed and what was expected under the null hypothesis is sufficiently large to provide evidence against that expected distribution.
The same fundamental logic can be extended to two categorical variables. If the company wants to know whether preferred delivery option is associated with customer age category, the expected frequencies are calculated under a different null hypothesis: that delivery preference and age category are independent.
The central idea is therefore straightforward:
Observed frequencies → Expected frequencies under H₀ → Size of discrepancies → Statistical evidence
What Is a Chi-Square Test?
The chi-square test evaluates discrepancies between observed and expected frequencies. The general χ² statistic gives greater weight to discrepancies that are large relative to their expected frequencies, combining information across categories or contingency-table cells into an overall test statistic.
For categorical-data applications, the important methodological issue is what determines the expected frequencies. In a goodness-of-fit test, expected frequencies follow from a specified null distribution. In a test of independence, they are derived from the marginal frequencies of the contingency table under the assumption that the two variables are independent. Penn State describes the expected count for a cell in a test of independence as the row total multiplied by the column total and divided by the total sample size. (Online Statistics at Penn State)
The same mathematical family can therefore answer different research questions. Saying merely that “chi-square compares observed and expected frequencies” is correct but incomplete; a defensible analysis must also explain why those particular expected frequencies represent the null hypothesis relevant to the study.
How Does the Chi-Square Test Work?
Consider a study examining whether preferred payment method is associated with customer age group. Participants are classified into age categories and according to whether they prefer cash, card or mobile payment. The resulting observations can be arranged in a contingency table.
Under the null hypothesis of independence, the proportion preferring each payment method should not systematically vary according to age category beyond sampling variation. Expected cell frequencies can therefore be calculated from the row and column totals. The chi-square statistic summarizes how far the observed table departs from this independence model. Penn State describes the procedure as calculating expected counts under independence and then comparing those counts with what was actually observed. (Online Statistics at Penn State)
A larger χ² statistic indicates a greater overall discrepancy between observed and expected frequencies. Whether the discrepancy constitutes statistically significant evidence depends on the relevant chi-square distribution and degrees of freedom. For an r × c independence table, the conventional degrees of freedom are (r − 1)(c − 1). (Online Statistics at Penn State)
This omnibus statistic tells researchers whether the observed pattern is inconsistent with the null model to the specified statistical threshold. It does not, by itself, explain which categories created that discrepancy or whether the association is substantively important.
Observed and Expected Frequencies
An observed frequency is the actual number of cases appearing in a category or contingency-table cell. An expected frequency is the number that would be expected in that category or cell if the relevant null hypothesis were true.
This distinction is fundamental because researchers sometimes look only at the raw counts and infer an association from visible differences. Yet unequal observed counts are not necessarily surprising. Different row and column totals can produce very different cell frequencies even when the variables are statistically independent.
For a test of independence, expected frequencies account for those marginal distributions. If 80% of a sample belongs to one customer category, for example, larger counts involving that category will often be expected even under complete independence. The chi-square analysis asks whether the observed table differs from the appropriate independence table, not whether all cells contain approximately the same number of observations.
The same principle applies to goodness-of-fit testing. Expected categories do not need to be equally probable. A null hypothesis could specify proportions of 50%, 30% and 20%, in which case expected frequencies should reflect those probabilities rather than impose an artificial equal distribution.
Chi-Square Goodness-of-Fit Test
The chi-square goodness-of-fit test examines whether observed frequencies across categories are consistent with a specified expected distribution. Conceptually, it concerns one categorical classification and asks whether its observed pattern fits the pattern specified by the null hypothesis.
Suppose a retailer expects customer complaints to be distributed across four categories according to historically established proportions. New data can be compared with those expected proportions to investigate whether the current distribution differs from the benchmark. Equal expected frequencies are only one possible case; the null distribution should come from the research question, theory, established benchmark or another defensible source.
NIST describes the general goodness-of-fit logic as comparing the number of observations falling into categories or intervals with the numbers expected under the hypothesized distribution. It also notes that the chi-square approximation requires sufficient expected frequencies and can be sensitive to how categories or bins are defined. (NIST)
In dissertation research, the expected distribution should therefore be specified and justified rather than constructed after inspecting the observed frequencies. Otherwise, the null hypothesis risks becoming a post-hoc description of the data rather than a meaningful proposition against which the data are evaluated.
Chi-Square Test of Independence
The chi-square test of independence examines whether two categorical variables are statistically associated. The observations are organized into a contingency table, and expected frequencies are calculated according to what would be expected if the two variables were independent.
Suppose a researcher investigates whether employees’ preferred working arrangement—office, hybrid or remote—is associated with job level—junior, middle management or senior management. The null hypothesis states that the variables are independent in the relevant population. Under this model, the distribution of working-arrangement preferences should not systematically depend on job level.
The chi-square statistic aggregates discrepancies between observed and expected frequencies across the contingency table. If the resulting p-value is sufficiently small under the prespecified decision criterion, the researcher has evidence against the independence model. Penn State similarly defines the test as assessing the relationship between two categorical variables through expected cell frequencies calculated under the null hypothesis of independence. (Online Statistics at Penn State)
A significant result supports the conclusion that the variables are associated, subject to the design and assumptions of the analysis. It does not establish causation, specify the strength of the association or automatically reveal which combinations of categories are responsible for the overall result.
Goodness-of-Fit vs Test of Independence
The easiest way to distinguish the two procedures is to begin with the research question.
A goodness-of-fit question asks:
Does the observed distribution of one categorical variable correspond to a specified expected distribution?
An independence question asks:
Are two categorical variables associated?
Both procedures compare observed and expected frequencies, but the source of those expected frequencies differs. For goodness-of-fit, they arise from the hypothesized distribution. For independence, they arise from the marginal distributions under a model in which the two variables are unrelated.
This distinction is more important than memorizing two formulas because it determines what the resulting test can legitimately tell the researcher. A dataset containing two categorical variables should not automatically be analysed through an independence test if the actual research question concerns the distribution of only one of them.
Assumptions of the Chi-Square Test
Chi-square procedures rely on the sampling and data structure being compatible with the inferential model. Observations should contribute to the analysis in the manner assumed by the design, and independence between observational units is particularly important. Repeated responses from the same participant, clustered observations or other dependencies cannot simply be treated as unrelated cases because they appear as separate spreadsheet rows.
The variables must also correspond to appropriate categorical classifications for the question being tested. Numerical category codes do not convert categorical data into continuous measurements. Conversely, arbitrarily converting genuinely continuous data into categories merely to use a chi-square test can discard information and may create conclusions that depend on the selected cut-points.
The adequacy of expected frequencies is another important consideration because the familiar chi-square reference distribution is an approximation. Penn State’s introductory treatment requires expected values of at least five for its standard independence procedure, while broader methodological treatments use more nuanced rules depending on table structure and context. (Online Statistics at Penn State) The underlying principle is more important than memorizing a universal threshold: sparse expected cells can make the asymptotic chi-square approximation unreliable.
Expected Frequencies Matter More Than Observed Cell Counts Alone
A common student rule is that “every cell must contain at least five cases.” This wording confuses observed and expected frequencies. The validity concern relates to whether expected counts under the null model are sufficiently large for the chi-square approximation, not simply whether the numbers printed in the observed contingency table exceed five.
This distinction matters because an observed cell can contain relatively few cases while its expected count remains adequate, or the reverse can occur. Researchers should therefore inspect the expected frequencies produced by the analysis rather than applying a rule to the raw table.
NIST notes that the chi-square approximation can become inappropriate with small expected frequencies and discusses combining categories where substantively defensible. (NIST) Combining categories merely to obtain significance or satisfy a numerical rule, however, can distort the construct being measured. If categories have distinct substantive meanings, statistical convenience alone is not sufficient justification for merging them.
For sparse contingency tables, an exact procedure may sometimes be preferable. In particular, Fisher’s exact test is an established alternative for certain contingency-table settings when expected counts are small. (Online Statistics at Penn State)
A Significant Chi-Square Does Not Tell You Which Categories Drive the Association
Suppose a 4 × 3 contingency table produces a statistically significant chi-square test of independence. The result indicates that the observed table departs from what would be expected under independence, but the χ² statistic aggregates discrepancies across all twelve cells. The omnibus result therefore does not identify which cells are primarily responsible for the association.
This matters particularly when a table contains several categories. A significant result might be driven predominantly by one unusually high observed frequency and one corresponding deficit elsewhere, while most of the table remains relatively close to expectation. Reporting only χ² and p therefore leaves an important interpretive question unanswered.
Researchers can examine cell-level residuals to investigate the structure of the discrepancy. NIST specifically notes the use of standardized residuals and association plots in chi-square independence analysis. (NIST) These diagnostics help researchers identify where observed counts are higher or lower than expected, although post-hoc interpretation should remain disciplined, particularly when many cells are being inspected.
The analytical sequence should consequently move from “Is there evidence of an association?” to “Where does the observed table depart from independence?” and finally to “What does that pattern mean substantively?”
Standardized Residuals
A raw residual is the difference between an observed frequency and its expected frequency. Large raw residuals suggest cells in which the observed table departs noticeably from the null model, but raw differences are difficult to compare when expected frequencies vary considerably across cells.
Standardized forms of residuals scale these discrepancies, making them more useful for identifying cells that contribute unusually strongly to the overall pattern. Positive residuals indicate that more observations occurred than expected under the null model, while negative residuals indicate fewer.
Residual analysis should not be treated as permission to construct an attractive narrative around whichever cells look most interesting. The researcher should consider the size and pattern of deviations, the number of cells examined, the substantive meaning of the categories and any multiplicity issues arising from follow-up analysis. Its purpose is to understand the omnibus result rather than to manufacture additional hypotheses after significance has already been established.
Effect Size and Cramer’s V
Statistical significance answers whether the data provide evidence against the specified null model; it does not indicate whether the association is strong or important. With sufficiently large samples, relatively modest departures from independence can produce statistically significant chi-square statistics.
For contingency-table analyses, Cramer’s V is commonly used as a standardized measure of association derived from the chi-square statistic and sample size while accounting for table dimensions. It ranges from zero upward to one in standard applications, with values closer to zero indicating weaker association and larger values indicating stronger association.
Conventional benchmarks are sometimes used to describe effect sizes, but they should not substitute for substantive interpretation. The importance of an association depends on the research context, table dimensions, consequences of the pattern and the decisions informed by the study. An apparently modest association between customer characteristics and product adoption, for example, may still have considerable commercial implications in a large market.
A stronger dissertation analysis therefore reports not merely whether the association is statistically detectable but also how substantial it appears to be and what the pattern means for the underlying research question.
Chi-Square Test vs Fisher’s Exact Test
The chi-square test of independence uses an asymptotic approximation whose performance depends partly on expected cell frequencies. When a contingency table is sparse, particularly in small-sample settings, the approximation may be questionable.
Fisher’s exact test offers an exact approach for contingency tables and is especially familiar in the 2 × 2 setting. Penn State notes that exact p-values can be calculated when expected counts are small and that chi-square and Fisher procedures become close approximations when counts are sufficiently large. (Online Statistics at Penn State)
This should not be converted into the mechanical rule “small sample = Fisher, large sample = chi-square.” Table dimensions, sampling structure, expected frequencies and the exact inferential question all matter. The broader methodological lesson is that researchers should inspect whether the assumptions supporting their chosen reference distribution are defensible rather than selecting a procedure solely from the total sample size.
Dudovskiy Chi-Square Test Selection Framework
The Dudovskiy Chi-Square Test Selection Framework synthesizes established categorical-data principles into a practical sequence for deciding whether and how a chi-square procedure addresses a research question. It does not introduce a new statistical test. Its purpose is to prevent researchers from choosing chi-square simply because their software recognizes categorical variables.
Research Question → Identify Categorical Structure → One Variable or Two? → Goodness-of-Fit or Independence → Expected Frequency Assessment → Chi-Square Analysis → Examine Pattern of Departures → Effect Size → Substantive Interpretation
The central principle is:
Choose the chi-square procedure according to the categorical question being asked, then interpret where observed frequencies differ from expectation—not merely whether χ² is significant.

The first decision concerns the structure of the research question. If the researcher has one categorical variable and wants to determine whether its observed frequencies correspond to a justified expected distribution, the analysis leads toward a chi-square goodness-of-fit test. If the question concerns whether two categorical variables are associated, the relevant branch leads toward a chi-square test of independence.
The next stage assesses whether the observations and expected frequencies support the intended procedure. This step occurs before interpretation because a precisely calculated χ² statistic cannot compensate for a design that violates independence or for a reference approximation that is unsuitable for a sparse table.
Once the omnibus analysis has been conducted, interpretation should examine the pattern underlying the result. In an independence analysis, this can involve comparing observed and expected counts and investigating appropriate cell residuals. The researcher then considers effect magnitude and returns to the substantive question that motivated the analysis.
The framework therefore moves the researcher away from “My variables are categorical, so I will run chi-square” toward “What categorical hypothesis am I testing, what would I expect under that hypothesis, and what pattern of departures does the evidence reveal?”
Application of the Chi-Square Test: an Example
A subscription-based digital service investigates whether subscription plan is associated with cancellation status. Customers are classified into three plans—Basic, Standard and Premium—and into two outcome categories according to whether they cancelled during a specified period. The research question therefore concerns the association between two categorical variables.
The researcher constructs a 3 × 2 contingency table and calculates expected frequencies under the null hypothesis that subscription plan and cancellation status are independent. Before relying on the chi-square approximation, the researcher verifies the independence structure of the observations and examines the expected cell frequencies. The analysis then compares the observed table with the table expected under independence.
Suppose the chi-square test is statistically significant. The researcher does not conclude merely that “subscription plan affects cancellation.” First, association does not by itself establish a causal effect. Second, the omnibus result does not reveal how the plans differ. The researcher therefore examines the observed-versus-expected pattern and appropriate residuals, finding that cancellations are more frequent than expected among Basic subscribers and less frequent than expected among Premium subscribers.
Cramer’s V is reported to provide information about the magnitude of the overall association. The final interpretation combines statistical and substantive evidence: subscription plan and cancellation status are associated in the sample, the departures from independence are concentrated in particular plan categories, and the magnitude and commercial importance of that pattern must be considered separately from its statistical significance.
Advantages and Limitations of the Chi-Square Test
Chi-square procedures provide an accessible way to analyse frequency data when research questions concern categorical distributions or associations. They do not require researchers to pretend that nominal categories possess meaningful numerical distances, and contingency tables make the underlying observed patterns relatively transparent. The basic comparison between observed and expected frequencies is also conceptually straightforward enough to explain and defend in many dissertation settings.
The framework accommodates tables larger than 2 × 2, making the independence test useful for many business-research questions involving customer segments, behavioural categories, organizational characteristics and survey responses. Follow-up examination of observed and expected counts can also connect the omnibus statistical result to a recognizable substantive pattern.
That apparent simplicity has limits. Sparse tables can undermine the chi-square approximation, while large samples can make relatively weak associations statistically significant. The omnibus statistic also compresses a potentially complex contingency table into one number and therefore cannot explain the structure of an association without additional examination.
Categorizing continuous variables solely to enable chi-square analysis can sacrifice information, and association between categorical variables does not establish causation. More complex questions involving adjustment for other variables, prediction or multiple explanatory factors may require models such as logistic or multinomial regression rather than a standalone chi-square test.
Common Mistakes When Using Chi-Square Tests
Treating category codes as quantitative measurements can lead researchers toward the wrong statistical family. A survey variable coded 1 = dissatisfied, 2 = neutral, 3 = satisfied contains numbers, but those numbers primarily identify ordered categories. Software and AI systems may calculate means from them, yet numerical coding alone does not determine the variable’s measurement properties or the appropriate analysis.
Observed frequencies are also frequently confused with expected frequencies when assumptions are assessed. The relevant concern for the usual chi-square approximation is the expected table under the null model. Researchers should inspect those expected counts rather than merely checking whether each observed cell happens to contain five or more cases. (Online Statistics at Penn State)
A statistically significant independence test is sometimes interpreted as evidence that one categorical variable caused the other. Chi-square establishes an association under the assumptions of the analysis; causal interpretation requires an appropriate research design and identification strategy beyond the test itself.
Stopping interpretation at p < .05 creates a different problem. A significant table can contain a complex pattern of over- and under-representation, and the overall p-value does not identify that pattern. Examining appropriate residuals and an effect-size measure can therefore add information that the omnibus significance test cannot provide.
Finally, categories should not be merged indiscriminately merely to increase expected frequencies. Combining categories can sometimes be defensible, but only when the merged categories remain conceptually meaningful. Changing the classification solely to obtain a convenient statistical result can change the research question itself.
Chi-Square Tests in Business Research
Categorical data are pervasive in business research. Researchers may investigate whether purchase status is associated with marketing channel, whether employee turnover differs according to employment category, whether payment preference is associated with customer segment, or whether adoption of a new technology is related to organizational size category.
Goodness-of-fit questions also arise naturally. A company may ask whether current customer preferences continue to follow historically established proportions, whether complaints are distributed across categories according to an expected operational pattern, or whether observed selections differ from theoretically predicted market shares.
The managerial interpretation should extend beyond significance. If customer segment and subscription choice are associated, decision-makers usually need to know which segments are over- or under-represented in particular plans and whether the magnitude of the association is commercially meaningful. This makes contingency-table inspection, residual analysis and effect-size interpretation important complements to the omnibus chi-square statistic.
Chi-Square Tests in the Age of AI and Digital Research
AI can calculate a chi-square statistic almost instantly, but categorical data create a particularly important risk: numbers do not necessarily represent quantities. A spreadsheet may encode customer type as 1, 2 and 3 or education level as 1 through 5. An AI system that interprets these codes as continuous measurements could recommend mean comparisons or correlations that do not correspond to the constructs represented by the codes.
The methodological sequence should therefore begin before statistical calculation. AI needs to know what each variable means, its measurement level, how categories were defined, whether observations are independent and what hypothesis the researcher intends to test. Only then can it help determine whether a goodness-of-fit test, independence test, exact procedure or another analytical approach is appropriate.
AI can be particularly useful after this reasoning has been established. It can help construct contingency tables, calculate expected frequencies, identify sparse cells, produce residual diagnostics, calculate effect sizes and explain output. It can also assist researchers in exploring whether a significant omnibus association is concentrated in particular cells.
The principal risk is consequently not computational error but semantic error. A statistically flawless chi-square calculation answers little if the categories were misunderstood, the wrong null model was specified or the data structure violates the assumed independence. As statistical calculation becomes increasingly automated, researchers need to become more explicit about what their variables and observations actually represent.
When to Use a Chi-Square Test
A chi-square procedure may be appropriate when:
- the research question concerns frequencies or categorical classifications;
- a goodness-of-fit question compares observed frequencies with a theoretically or substantively justified expected distribution;
- an independence question concerns association between two categorical variables;
- observational units satisfy the relevant independence structure;
- categories are defined meaningfully and cases are assigned appropriately;
- expected frequencies are sufficiently adequate for the intended chi-square approximation;
- sparse cells and potential exact alternatives have been considered where relevant;
- the researcher intends to examine the pattern underlying a significant omnibus result rather than reporting only the p-value; and
- effect magnitude and substantive importance will be considered alongside statistical significance.
The presence of categorical variables alone does not automatically make a chi-square test the appropriate analysis.
Dissertation Example
A dissertation titled “The Relationship Between Flexible Working Arrangements and Employee Retention Intentions in Professional Service Firms” classifies employees according to their primary working arrangement—office-based, hybrid or remote—and whether they report an intention to remain with their employer during the next year—yes or no. Because both variables are categorical and the research question concerns their association, the methodology chapter proposes a chi-square test of independence.
The researcher explains that each employee contributes one observation to the contingency table and examines the expected cell frequencies before relying on the chi-square approximation. The methodology does not justify the test merely by saying that “the data are non-parametric”; instead, it connects the procedure directly to the measurement structure and the hypothesis of independence between the two categorical variables.
If the omnibus test is statistically significant, the analysis proceeds to examination of observed and expected counts and appropriate cell residuals to determine where the table departs most clearly from independence. Cramer’s V is reported alongside χ², degrees of freedom and the p-value to provide information about the magnitude of the association.
The dissertation then distinguishes association from causation. A significant relationship between working arrangement and retention intention does not establish that changing an employee’s working arrangement would cause their intention to remain to change. The statistical conclusion is limited to evidence concerning the association represented by the observed contingency table, while causal interpretation depends on the broader research design.
Exam Tip
Do not define the chi-square test simply as “a test for categorical data.” A stronger answer explains that chi-square procedures compare observed frequencies with frequencies expected under a specified null hypothesis.
Then distinguish the two common applications. A goodness-of-fit test examines whether the distribution of one categorical variable corresponds to a specified expected distribution. A test of independence examines whether two categorical variables are associated by comparing the observed contingency table with the table expected if the variables were independent.
Remember that a significant χ² result does not identify which cells drive the result, does not indicate the strength of association and does not establish causation. Expected frequencies should also be examined because the conventional chi-square reference distribution relies on an approximation that can become problematic with sparse expected cells.
Build a methodology you can explain and defend
Selecting a statistical test requires understanding what your variables represent and what hypothesis your research design can support. Dudovskiy Research Assistant can help connect your dissertation topic, variables and research design to an appropriate analytical approach and explain how that choice can be justified in your methodology.
References
Agresti, A. (2019). An Introduction to Categorical Data Analysis. 3rd ed. Wiley.
Cochran, W.G. (1952). The χ² test of goodness of fit. Annals of Mathematical Statistics, 23(3), 315–345.
Conover, W.J. (1999). Practical Nonparametric Statistics. 3rd ed. Wiley.
Pearson, K. (1900). On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Philosophical Magazine, 50(302), 157–175.
Penn State Eberly College of Science. Chi-Square Test for Independence. STAT 500. (Online Statistics at Penn State)
NIST/SEMATECH. Chi-Square Goodness-of-Fit Test. e-Handbook of Statistical Methods. (NIST)
