t-test
A t-test is a statistical procedure used to evaluate hypotheses about population means when the relevant population variance is unknown and must be estimated from sample data. Depending on the research design, a t-test can compare a sample mean with a specified reference value, compare the means of two independent groups, or evaluate the mean difference between paired observations.
The choice among these procedures depends primarily on what is being compared and how the observations are related. A researcher comparing customer spending between two unrelated customer groups faces a different statistical problem from a researcher comparing the same customers’ spending before and after a loyalty programme. Both studies involve two sets of values, but the first involves independent observations while the second involves paired observations. UCLA’s statistical guidance similarly distinguishes one-sample, independent-samples and paired t-tests according to these underlying comparison structures. (OARC Stats)
A defensible t-test analysis therefore begins with the research design rather than with a software menu. Researchers also need to consider the assumptions of the selected procedure, the distinction between Student’s and Welch’s independent-samples t-tests, confidence intervals, effect magnitude and the substantive meaning of the estimated difference.
On this page:
- t-test explained simply
- What a t-test is
- How a t-test works
- The t-statistic
- Types of t-test
- One-sample t-test
- Independent-samples t-test
- Paired-samples t-test
- Student’s t-test and Welch’s t-test
- t-test vs ANOVA
- Assumptions of t-tests
- Normality and robustness
- One-tailed vs two-tailed tests
- Confidence intervals
- Effect size
- Dudovskiy t-Test Selection Framework
- Application example
- Advantages and limitations
- Common mistakes
- t-tests in business research
- t-tests in the age of AI
- When to use a t-test
- Dissertation example
- Exam tip
| Research question | Potential procedure |
|---|---|
| Does a sample mean differ from a specified value? | One-sample t-test |
| Do two independent groups have different means? | Independent-samples t-test |
| Do two measurements from the same participants differ on average? | Paired-samples t-test |
| Are independent-group variances not assumed equal? | Welch’s t-test |
| Are there three or more independent group means? | ANOVA may provide a more appropriate omnibus framework |
| Is the objective to estimate how large the difference is? | Mean difference, confidence interval and appropriate effect size |
t-Test Explained Simply
Suppose a company introduces a new customer-service training programme and wants to determine whether it improves employees’ assessment scores. One group of employees receives the programme while a separate group receives the existing training. The new-programme group achieves an average score of 84, compared with 79 for the existing-training group. The observed difference is therefore five points, but this difference alone does not establish that the underlying population means differ because different samples would produce somewhat different averages even if the training programmes had identical population effects.
An independent-samples t-test evaluates the observed mean difference relative to the sampling uncertainty associated with that difference. Broadly, a difference that is large relative to its standard error produces a larger absolute t-statistic and stronger evidence against the relevant null hypothesis. A small difference accompanied by substantial sampling uncertainty produces weaker evidence.
The important question is therefore not merely whether 84 is different from 79. Numerically, it obviously is. The statistical question concerns how compatible an observed difference of this magnitude is with the null hypothesis, given the variability and amount of information in the samples.
What Is a t-Test?
The t-test originates from the problem of making inferences about means when population variability is unknown. Instead of knowing the population standard deviation, researchers generally estimate variability from the sample itself. This additional uncertainty is reflected in the t-distribution, whose shape depends on the degrees of freedom associated with the analysis.
The term t-test actually refers to a family of related procedures rather than one universal test. The one-sample t-test evaluates a sample mean against a specified reference value. The independent-samples t-test evaluates a difference between means from separate groups. The paired-samples t-test instead analyses differences within matched pairs or repeated observations. These procedures share mathematical foundations, but they answer different research questions and represent different data structures. (OARC Stats)
This distinction is methodologically important because statistical software can calculate all three procedures easily. The difficulty is not obtaining a t-statistic; it is establishing which comparison the research design actually implies.
How Does a t-Test Work?
At its core, a t-test compares an estimated effect with the uncertainty surrounding that estimate. For a mean-difference problem, the effect may be the difference between two group means, while the uncertainty is represented through the standard error of that difference. The resulting t-statistic indicates how far the estimated effect lies from the null-hypothesized value in standard-error units.
Suppose two independent customer groups have mean satisfaction scores of 72 and 76. A four-point difference might represent strong statistical evidence if both means have been estimated very precisely from large, relatively stable samples. The same four-point difference could provide much weaker evidence if the samples are small and satisfaction scores vary widely within each group. The magnitude of the raw difference therefore cannot be interpreted independently of its uncertainty.
The t-test formalizes this relationship. Once the t-statistic and appropriate degrees of freedom have been determined, the t-distribution provides the basis for calculating the p-value and confidence interval. These outputs address related aspects of the same inferential problem rather than representing unrelated statistical calculations.
The t-Statistic
The general logic of a t-statistic can be expressed as:
t = estimated difference from the null value / standard error of the estimated difference
The exact calculation of the numerator and standard error depends on the type of t-test. In a one-sample test, the numerator represents the difference between the sample mean and the hypothesized population mean. In an independent-samples test, it represents the difference between two group means. In a paired test, it represents the mean of the within-pair differences relative to the hypothesized mean difference, commonly zero.
A larger absolute t-value indicates that the estimated difference lies farther from the null value relative to its standard error. Whether that constitutes strong evidence against the null hypothesis also depends on the relevant degrees of freedom and whether the test is one-sided or two-sided. The sign of t indicates the direction of the estimated difference according to how the comparison has been defined; it should not be interpreted as a measure of effect importance.
Types of t-Test
The three principal forms encountered in dissertation research are the one-sample t-test, independent-samples t-test and paired-samples t-test. Their differences arise from the structure of the comparison rather than from arbitrary statistical terminology. UCLA’s documentation, for example, describes the one-sample procedure as comparing a sample mean with a specified value, the independent procedure as comparing means from two groups, and the paired procedure as comparing related measurements while accounting for their dependence. (OARC Stats)
Choosing among them therefore requires researchers to understand where each observation comes from. Two columns in SPSS or Excel do not establish that an independent-samples t-test is appropriate. Those columns could contain measurements from separate groups or two observations from the same people, and that distinction changes the statistical model.
One-Sample t-Test
A one-sample t-test examines whether a population mean is consistent with a specified reference value. Suppose a hotel chain has historically treated a customer-satisfaction score of 80 as an important performance benchmark. A researcher collects a sample of satisfaction scores from a newly opened hotel and investigates whether its population mean differs from 80. The null hypothesis can be expressed as H₀: μ = 80.
The reference value should have a substantive justification. It might originate from a theoretical expectation, established industry benchmark, historical value or clearly specified research hypothesis. Selecting a reference value after examining the sample mean simply because it produces an interesting test would undermine the logic of the analysis.
A statistically significant result indicates evidence against the hypothesized population mean under the assumptions of the procedure. It does not establish that the observed difference is commercially or theoretically important. The estimated mean difference and its confidence interval remain important for substantive interpretation.
Independent-Samples t-Test
An independent-samples t-test is used when the research question concerns the difference between means from two separate groups whose observations are independent across groups. A researcher might compare average purchase intention between consumers shown Advertisement A and different consumers shown Advertisement B. Each participant contributes an observation to only one of the two groups.
The null hypothesis commonly concerns whether the difference between the two population means is zero. The analysis estimates the observed mean difference and compares it with the uncertainty associated with that estimate. The precise form of the standard error and degrees of freedom depends on whether the analysis uses the classical equal-variance Student procedure or Welch’s unequal-variance procedure.
Independence is fundamentally a design issue. If observations from one group are naturally linked to observations in the other—for example, spouses, deliberately matched participants or repeated measurements from the same individuals—the ordinary independent-samples model does not represent the structure of the data correctly.
Paired Data Are Not Two Independent Samples
A paired-samples t-test, also called a dependent-samples t-test, is appropriate when observations occur in meaningful pairs. The most familiar example is a before-and-after study in which the same participants are measured twice, although pairing can also arise through matched subjects or other research designs in which each observation in one condition corresponds to a particular observation in another.
The paired t-test works by calculating a difference for each pair and then evaluating whether the mean of those differences differs from the hypothesized value, usually zero. UCLA’s explanation of the paired t-test emphasizes precisely this structure: the procedure accounts for the fact that scores from the same subjects are not independent and tests the mean of their within-subject differences. (OARC Stats)
This means that a paired t-test is conceptually close to a one-sample t-test conducted on the difference scores. The pairing is not an inconvenience to be ignored; it contains information. When measurements within pairs are correlated, analysing the pairing appropriately can provide a more precise estimate of change than pretending that the observations came from unrelated groups.
Student’s t-Test and Welch’s t-Test
For two independent groups, an important distinction exists between the classical pooled-variance form of Student’s t-test and Welch’s t-test. The pooled procedure relies on an equal-variance model, whereas Welch’s procedure does not require the two population variances to be assumed equal and adjusts the standard error and degrees of freedom accordingly.
This distinction is more consequential than the traditional software routine of first performing a variance-equality test and then choosing a t-test according to whether that preliminary test is significant. Methodological research has challenged this conditional strategy. Zimmerman (2004), for example, found that choosing between pooled and Welch procedures according to a preliminary variance test can worsen Type I error behaviour, while Hayes and Cai (2007) similarly concluded that several unconditional approaches, including the separate-variance test, performed as well as or better than the conditional decision rule across the distributions they examined. (PubMed)
Welch’s procedure has consequently been recommended more broadly in methodological literature for independent-group comparisons, particularly because it remains applicable when population variances differ. West (2021), for example, explicitly advocates Welch’s t-test as a preferable default for comparing two groups in the context discussed. (PubMed) This should not be turned into another mechanical rule, however. Test performance can depend on distribution shape, sample size, variance heterogeneity and the estimand of interest; simulation evidence shows that no simple procedure is universally optimal across every problematic distribution. (PubMed)
t-Test vs ANOVA
A t-test is often the most straightforward procedure when the research question concerns one mean or the difference between two means. ANOVA becomes particularly useful when three or more group means or more complicated factorial structures need to be analysed. With two independent groups under corresponding assumptions, the conventional one-factor ANOVA and two-sided t-test are mathematically connected, with the ANOVA F-statistic equal to the square of the corresponding t-statistic.
The distinction becomes practically important as research designs grow more complex. If a researcher needs to compare four independent customer segments, conducting six unadjusted pairwise t-tests would create a multiple-testing problem and fragment an overall research question. An omnibus ANOVA provides a more coherent framework for examining equality of the group means before relevant follow-up comparisons are considered.
The choice should therefore not be reduced to “two groups = t-test, three groups = ANOVA.” Repeated observations, multiple factors, covariates and other design characteristics may require different models even when only two means appear in a particular comparison.
Assumptions of t-Tests
The assumptions of a t-test depend on which t-test is being used. Independence of observations is central to independent-samples procedures and arises principally from the study design. A paired t-test deliberately represents dependence within pairs, while requiring the pairs themselves to satisfy the relevant independence structure. Confusing these two situations changes the statistical problem rather than merely creating a minor assumption violation.
Distributional assumptions also need to be understood in relation to the quantity being analysed. In a one-sample t-test, attention concerns the distribution relevant to inference about the sample mean; in a paired t-test, the important distribution is that of the within-pair differences, not whether each measurement occasion separately looks perfectly normal. Extreme observations and substantial skewness can be especially consequential in small samples because means and standard deviations are sensitive to them.
For the independent-samples procedure, the classical pooled Student test additionally incorporates an equal-variance assumption. Welch’s t-test relaxes this requirement and is specifically designed for inference when the two variances are not assumed equal. (PubMed) Assumption assessment should therefore correspond to the actual procedure rather than treating “the assumptions of the t-test” as one universal checklist.
Normality and Robustness
A common mistake is to interpret any departure from a perfectly normal distribution as proof that a t-test cannot be used. The issue is more nuanced. t procedures can be reasonably robust to some departures from normality, particularly as sample information increases and when the distributions do not contain severe skewness or influential outliers. Skovlund and Fenstad (2001), for example, found the two-sample t-test robust to departures from normality when sample sizes were not very small under the conditions they investigated. (PubMed)
Robustness is not absolute. Small samples combined with strong skewness, extreme observations or unequal variances can create more serious problems, and the appropriate response depends on the research question and the nature of the data. Fagerland and Sandvik’s simulation study demonstrates that performance under combined non-normality and variance heterogeneity depends on a complex interaction among skewness, variance differences and sample-size patterns. (PubMed)
Researchers should therefore avoid the simplistic sequence normality test significant → abandon t-test → use a non-parametric test. The alternative procedure may test a different hypothesis, and formal normality tests themselves do not determine whether the t-test’s inference is substantively trustworthy. The appropriate decision requires examination of the data, design, estimand and severity of the relevant departures.
One-Tailed vs Two-Tailed t-Test
A two-tailed t-test evaluates evidence for differences in either direction. If a researcher compares two training programmes, the alternative hypothesis may be that their population means are different without specifying which programme produces the higher score. A one-tailed test instead specifies a directional alternative, such as the mean under Programme A being greater than the mean under Programme B.
The direction should be justified by the research hypothesis before examining the results. Choosing a one-tailed test after observing that the sample difference happens to point in the expected direction artificially changes the evidential threshold in response to the data. A directional theoretical prediction alone also does not automatically make a one-tailed analysis appropriate if an effect in the opposite direction would still be scientifically or managerially meaningful.
In many dissertation settings, a two-tailed test is therefore the more defensible default unless the directional hypothesis and treatment of an opposite-direction result have been established clearly in advance.
Confidence Intervals and the t-Test
A confidence interval complements the hypothesis test by showing a range of parameter values compatible with the data under the statistical procedure. In a two-group comparison, a confidence interval for the mean difference indicates both the estimated direction and magnitude of the difference and the uncertainty surrounding that estimate.
Suppose a study estimates that customers exposed to a redesigned website spend, on average, $8 more than customers using the existing design. Reporting only p = .03 says relatively little about whether an $8 difference is important or how precisely it has been estimated. A confidence interval might show that the plausible values under the model range from a very small increase to a considerably larger one, providing information that the binary significant/non-significant classification conceals.
Confidence intervals and p-values are mathematically related for corresponding tests, but they encourage different interpretive questions. The p-value asks about compatibility with a particular null value, whereas the confidence interval encourages researchers to consider the range and precision of plausible effect estimates.
Effect Size and Practical Importance
A statistically significant t-test does not establish that the difference is large enough to matter. With a sufficiently large sample, even a small difference may be estimated precisely enough to produce a small p-value. Conversely, a potentially important difference in a small sample may be accompanied by substantial uncertainty and fail to reach a conventional significance threshold.
Researchers can therefore complement the mean difference and confidence interval with an appropriate standardized effect-size measure, such as a version of Cohen’s d or Hedges’ g, when standardization is substantively useful. The choice and interpretation of the effect size should correspond to the research design; paired and independent-group studies, for example, can involve different standardization decisions.
Generic labels such as “small,” “medium” and “large” should not replace contextual interpretation. A small standardized difference in customer retention may have substantial financial implications when applied across millions of transactions, whereas a statistically large standardized effect on a minor outcome may have limited managerial relevance.
A Non-Significant t-Test Does Not Demonstrate Equality
Suppose a researcher compares two service models and obtains p = .18. The correct conclusion is not that the two service models are “the same.” Failure to reject a null hypothesis of zero mean difference does not establish that the population means are equivalent. The data may be consistent with a range of effects because the sample provides limited precision.
If the substantive research question is whether two means are sufficiently similar to be considered equivalent within a predefined margin, an equivalence-testing framework is more appropriate than attempting to prove equality through a non-significant conventional t-test. Research on equivalence procedures explicitly distinguishes this objective from traditional difference testing. (PubMed)
This distinction is especially important in dissertations because “there was no significant difference” is frequently transformed incorrectly into “there was no difference.” The former is a statement about the statistical evidence produced by a particular study; the latter is a much stronger claim about the underlying populations.
Dudovskiy t-Test Selection Framework
The Dudovskiy t-Test Selection Framework synthesizes established statistical principles into a practical sequence for selecting and interpreting a t-test. It does not introduce a new statistical procedure. Its purpose is to ensure that the test follows the research design rather than being selected merely because two columns of numerical data happen to appear in the dataset.
Research Question → What Mean Is Being Compared? → Relationship Between Observations → Appropriate t-Test → Assumption Assessment → Estimate Difference → t-Test and Confidence Interval → Effect Size → Substantive Interpretation
The central principle is:
Choose the t-test according to what is being compared and whether the observations are independent or paired—not simply because the study contains two means.

The first decision concerns the comparison itself. When one sample mean is compared with a justified reference value, the relevant branch leads to a one-sample t-test. When the research question concerns two groups, the researcher must establish whether those groups contain genuinely independent observations or whether the measurements are paired.
Separate participants in two groups lead toward an independent-samples t-test, with the choice of variance model requiring appropriate consideration. Repeated measurements from the same participants or genuinely matched observations lead toward a paired-samples t-test, in which inference concerns the mean of the within-pair differences.
Once the appropriate structure has been established, the researcher assesses assumptions relevant to that particular procedure and estimates the difference of interest. The analysis should then combine the hypothesis test with a confidence interval and, where useful, an appropriate effect-size measure. Interpretation returns to the original research question by asking not only whether the data provide evidence against the null hypothesis, but also how large the estimated difference is, how uncertain it is and whether it matters in the research context.
The framework therefore replaces the shortcut “I have two means, so I need a t-test” with the more defensible question “What exactly is being compared, and how are the observations related?”
Application of a t-Test: an Example
Consider a retail company investigating whether a redesigned product page increases customers’ average order value. During an experiment, one set of customers is shown the existing product page and a separate set is shown the redesigned version. Each customer appears in only one condition, making the two groups independent. The dependent variable is order value, measured in monetary units.
The research question concerns the difference between two independent population means, so an independent-samples t-test is considered. Before conducting the inferential analysis, the researcher examines the distributions, potential influential observations, sample sizes and variance structure rather than automatically applying the pooled-variance procedure. Based on the planned analytical approach, the researcher uses Welch’s t-test, which does not require the two population variances to be assumed equal.
The redesigned-page group has a higher average order value, and the resulting analysis provides evidence against a zero mean difference. The researcher reports the estimated difference and confidence interval alongside the t-statistic, degrees of freedom and p-value. An appropriate effect-size measure is also considered to determine whether the difference is meaningful from a commercial perspective.
The conclusion therefore does not stop at “the t-test was statistically significant.” It explains what was compared, why the observations were treated as independent, why the selected form of t-test represented the design, how large the estimated difference was and what that difference means for the business question.
Advantages and Limitations of t-Tests
t-tests provide a direct and interpretable framework for common research questions involving means. The estimated mean difference has an intuitive substantive interpretation, while the associated t-test and confidence interval connect that estimate to sampling uncertainty. The family also accommodates several important research structures through one-sample, independent-samples and paired procedures.
The methods are widely implemented and mathematically transparent, making them particularly useful for dissertation research in which researchers need to explain and defend their analytical decisions. Paired designs can also make efficient use of repeated observations by incorporating the relationship within each pair rather than discarding that information.
Their apparent simplicity can nevertheless encourage inappropriate use. A t-test designed for independent groups does not become valid merely because the dataset contains two columns, and the classical pooled procedure may be inappropriate when its variance model does not represent the populations being compared. Severe skewness, influential observations and small samples can also make inference more sensitive to distributional assumptions.
The framework becomes increasingly limited as the research design grows more complicated. Multiple groups, several predictors, interactions, clustering, covariate adjustment and longitudinal structures may be represented more effectively through ANOVA, regression, mixed-effects models or other statistical approaches. The t-test is powerful partly because it answers a focused question; it should not be stretched to answer a research question that requires a richer model.
Common Mistakes When Using t-Tests
Treating all two-group comparisons as independent is particularly problematic. Before-and-after measurements from the same participants are paired, as are many deliberately matched designs. Ignoring that relationship misrepresents the data structure and may alter the standard error and resulting inference.
Another recurring problem is selecting between Student’s and Welch’s procedures solely through a preliminary significance test of equal variances. Simulation research has shown weaknesses in this conditional strategy, and methodological literature provides reasons for considering Welch’s procedure without first requiring a significant variance test. (PubMed) The analytical choice should therefore reflect the statistical model and established methodological guidance rather than a mechanical two-stage software routine.
Researchers also sometimes interpret p > .05 as proof that two means are identical, or p < .05 as proof that the difference is important. Neither conclusion follows. The first confuses absence of sufficient evidence against equality with evidence of equivalence, while the second confuses statistical detectability with effect magnitude.
Assumption checking can create further errors when formal tests replace statistical judgment. Rejecting a t-test automatically because a normality test is significant, or using a non-parametric procedure without considering whether it addresses the same research hypothesis, may solve the wrong problem. A stronger analysis considers the design, distribution, sample size, influential observations, estimand and robustness of the proposed procedure together.
t-Tests in Business Research
Business research frequently involves questions that map naturally onto t-tests. Marketing researchers may compare average purchase intention between consumers exposed to two advertisements, human-resource researchers may compare employee engagement between two working arrangements, and operations researchers may evaluate whether a process change alters average completion time. One-sample procedures can also compare performance against a benchmark, while paired tests are useful for evaluating change within the same employees, customers, stores or organizational units.
The simplicity of these questions should not obscure their design differences. Comparing two independent branches after different interventions is not statistically equivalent to measuring the same branch before and after an intervention. Similarly, comparing observed performance with a corporate target requires a one-sample structure rather than inventing a second group.
Business interpretation should also extend beyond statistical significance. A small increase in conversion rate, spending or processing efficiency can be economically meaningful at scale, while a statistically detectable difference in an employee questionnaire may have little managerial importance. The value of the t-test therefore depends on connecting statistical evidence to the substantive decision that motivated the comparison.
t-Tests in the Age of AI and Digital Research
AI makes t-tests extraordinarily easy to perform and surprisingly easy to misuse. A researcher can upload two spreadsheet columns and ask an AI system to “compare these groups,” receiving a polished t-test, p-value and interpretation within seconds. The critical information may nevertheless be absent: whether the columns represent separate participants, repeated measurements from the same participants, deliberately matched cases or even variables that should not be compared through means at all.
This creates a distinctive methodological risk. The numerical calculation can be perfectly correct for the test that AI has selected while the test itself is inappropriate for the research design. A before-and-after dataset analysed as two independent samples is a clear example: nothing in the arithmetic necessarily alerts the researcher that the dependency structure has been discarded.
AI is more useful when the sequence is reversed. Researchers can first require the system to identify the unit of analysis, outcome variable, comparison being made, relationship between observations, relevant variance assumptions and intended estimand before generating statistical code. It can then assist with calculations, diagnostics, sensitivity analyses and interpretation while the researcher retains responsibility for the methodological justification.
The increasing ability of AI to perform statistics therefore makes research-design literacy more important, not less. When calculation is automated, the defensibility of the analysis depends increasingly on whether the researcher understands why that calculation answers the research question.
When to Use a t-Test
A t-test may be appropriate when:
- the research question concerns a population mean or a difference between two population means;
- a sample mean is being compared with a substantively justified reference value;
- two independent groups are being compared on a suitable numerical outcome;
- two measurements are paired because they come from the same participants or genuinely matched units;
- the independence structure has been established from the research design rather than inferred from spreadsheet layout;
- assumptions relevant to the particular t procedure are sufficiently defensible;
- the selected form of the independent-samples test appropriately represents the variance structure;
- the researcher intends to report the estimated difference and uncertainty rather than only a p-value; and
- effect magnitude and substantive importance will be considered alongside statistical significance.
A t-test should not be chosen simply because two means appear in the dataset.
Dissertation Example
A dissertation titled “The Effect of Digital Sales Training on the Performance of Retail Sales Employees” measures each participating employee’s monthly sales performance immediately before training and again three months after completing the programme. The methodology chapter identifies the two measurements as repeated observations from the same employees rather than independent samples. A paired-samples t-test is therefore selected to evaluate whether the population mean of the within-employee change scores differs from zero.
The researcher examines the distribution of the paired differences and investigates potentially influential observations because the inferential procedure concerns these differences rather than treating the pre-training and post-training distributions as unrelated samples. The methodology chapter explicitly explains this point instead of stating generically that “the data were normally distributed.”
The analysis estimates the average within-employee change and reports its confidence interval together with the t-statistic, degrees of freedom and p-value. An appropriate effect-size measure is considered to evaluate the magnitude of the observed change. If the result is statistically significant, the dissertation describes evidence of a mean change following the training period but avoids claiming automatically that the training caused the improvement if the research design lacks a credible counterfactual for other changes occurring over the same period.
The methodological justification consequently connects the research design directly to the statistical procedure: the same employees were observed twice → observations were paired → within-person differences were analysed → the mean difference and its uncertainty were interpreted in relation to the research question.
Exam Tip
A t-test should not be defined merely as “a test used to compare two groups.” That definition excludes the one-sample procedure and obscures the crucial distinction between independent and paired observations. A stronger answer is that a t-test evaluates a hypothesis about a population mean or mean difference by comparing the estimated departure from a null value with its estimated standard error using the t-distribution.
When selecting a test, first ask what is being compared and then ask whether the observations are independent or paired. One sample against a reference value suggests a one-sample t-test; two independent groups suggest an independent-samples procedure; and repeated or matched observations suggest a paired-samples procedure.
Interpretation should then move beyond the question of whether p is below .05. A strong answer considers the estimated mean difference, confidence interval, relevant effect size and substantive importance. Remember also that a non-significant difference test does not prove equivalence.
Build a methodology you can explain and defend
Selecting a t-test requires more than identifying two sets of numbers. Dudovskiy Research Assistant can help determine which statistical approach fits your dissertation design and explain how the choice should be justified in your methodology.
References
Gosset, W.S. [Student] (1908). The probable error of a mean. Biometrika, 6(1), 1–25.
Hayes, A.F. and Cai, L. (2007). Further evaluating the conditional decision rule for comparing two independent means. British Journal of Mathematical and Statistical Psychology, 60(2), 217–244. (PubMed)
Skovlund, E. and Fenstad, G.U. (2001). Should we always choose a nonparametric test when comparing two apparently nonnormal distributions? Journal of Clinical Epidemiology, 54(1), 86–92. (PubMed)
Welch, B.L. (1947). The generalization of “Student’s” problem when several different population variances are involved. Biometrika, 34(1/2), 28–35.
West, R.M. (2021). Best practice in statistics: Use the Welch t-test when testing the difference between two groups. Annals of Clinical Biochemistry, 58(4), 267–269. (PubMed)
Zimmerman, D.W. (2004). A note on preliminary tests of equality of variances. British Journal of Mathematical and Statistical Psychology, 57(1), 173–181. (PubMed)
