Content Analysis

Content analysis is a systematic research method used to classify, analyse and interpret recorded communication such as documents, interview transcripts, websites, advertisements, social-media posts, annual reports and other textual or visual materials. It can be quantitative, where predefined categories are coded and their occurrence examined numerically, or qualitative, where systematic categorisation is used to interpret meanings and patterns within the material. The appropriate form of content analysis therefore depends on the research question, the nature of the data and what the researcher ultimately wants the analysis to explain.

On this page

  • Content analysis explained simply
  • What content analysis is
  • Qualitative vs quantitative content analysis
  • Conventional, directed and summative content analysis
  • Manifest vs latent content
  • Dudovskiy Content Analysis Design Framework
  • Units of analysis
  • Developing a coding frame
  • The content analysis process
  • Intercoder reliability and quality
  • Content analysis vs thematic analysis
  • Reporting content analysis in a dissertation
  • Application example
  • Advantages and limitations
  • Common mistakes
  • Content analysis in business research
  • Content analysis in the age of AI
  • When to use content analysis
  • Dissertation example
  • Exam tip

Content Analysis at a Glance

Methodological decision Key question Main possibilities
Analytical purpose What should the analysis reveal? Description, comparison, quantification, interpretation
Analytical orientation What form of content analysis fits the purpose? Quantitative, qualitative
Category development Where will categories come from? Data-driven, theory-driven, mixed
Level of content What kind of meaning will be coded? Manifest, latent
Coding system How will material be classified consistently? Units, categories, coding rules
Quality How will analytical quality be demonstrated? Reliability, credibility, transparency, methodological coherence


Content Analysis Explained Simply

Imagine a company receives 2,000 online customer complaints. Management wants to understand what customers complain about most frequently.

A researcher develops categories such as delivery, product quality, pricing, customer service and refunds, and systematically assigns relevant complaints to these categories. The analysis might show that 38% concern delivery and 24% concern product quality.

The researcher could go further and examine the meanings within complaints—for example, whether delivery complaints primarily concern speed, inaccurate tracking or damaged products.

This illustrates an important characteristic of content analysis: it can range from systematic classification and counting to more interpretive examination of meaning, depending on how the study is designed.

Dudovskiy Research Assistant

Not sure if content analysis is the correct choice for your dissertation?

Enter your topic to receive a free, tailored methodology preview.

Free · No account required
Learn more about Dudovskiy Research Assistant

What Is Content Analysis?

Content analysis is a research method for making systematic and defensible interpretations from recorded material by classifying relevant elements of that material according to an analytical framework. Krippendorff conceptualises content analysis as a research technique for making replicable and valid inferences from texts and other meaningful matter to the contexts of their use.

The term covers a broader range of approaches than is sometimes assumed. Content analysis can examine the frequency of clearly defined categories, but it can also investigate meaning qualitatively. Accordingly, saying merely that “content analysis was used” leaves several methodological questions unanswered.

A researcher should be able to explain what material was analysed, how it was selected, what constituted the unit of analysis, how categories were developed, how coding was conducted, what level of meaning was examined and how the quality of the analysis was established.

Qualitative vs Quantitative Content Analysis

Quantitative content analysis systematically classifies content into categories so that patterns can be examined numerically, whereas qualitative content analysis uses systematic categorisation primarily to understand and interpret meaning.

Suppose a researcher analyses the sustainability reports of 100 companies.

A quantitative study might record whether each report contains carbon-reduction targets, diversity targets, independent assurance and supply-chain commitments. The researcher could then compare the prevalence of these categories across industries.

A qualitative study could examine how companies construct the meaning of corporate responsibility, including how responsibility is attributed, how environmental commitments are justified and how uncertainty is communicated.

The difference is therefore not simply whether the researcher uses numbers. It concerns the purpose and logic of the analysis.

Qualitative content analysis still requires systematic procedures. Schreier, for example, places the coding frame at the centre of qualitative content analysis and emphasizes issues including segmentation, category construction, trial coding and revision.

Quantitative and qualitative approaches can also coexist within a study. Category frequencies may reveal patterns worth investigating interpretively, while qualitative analysis can generate categories that are subsequently compared numerically. The researcher should nevertheless explain the purpose of combining these analytical operations rather than treating them as interchangeable.

Conventional, Directed and Summative Content Analysis

A particularly useful distinction within qualitative content analysis is between conventional, directed and summative approaches, associated especially with Hsieh and Shannon.

Conventional content analysis is appropriate when existing theory or research provides limited guidance about the phenomenon. Categories are developed primarily through engagement with the data rather than being imposed in advance.

For example, a researcher studying employees’ experiences of a newly introduced four-day working week might develop categories from recurring meanings in employee interviews.

Directed content analysis begins with existing theory or previous research that informs the initial coding structure. The researcher uses established concepts to guide attention while remaining alert to material that does not fit the initial framework.

For example, a study of online customer loyalty might use dimensions derived from an established trust model as initial categories.

Summative content analysis begins by identifying and comparing particular words or content and then moves beyond frequency towards interpreting their contextual use. Counting is therefore a starting point rather than necessarily the final analytical objective.

The choice between these approaches should reflect the relationship between the study and existing knowledge. A researcher should not describe analysis as inductive while simultaneously constructing the entire coding framework from a pre-existing theory.

Manifest vs Latent Content

Manifest content concerns what is explicitly present in the material, whereas latent content concerns underlying meanings, assumptions or implications.

Consider the following statement in a company’s annual report:

“We remain committed to protecting long-term shareholder value while responding proportionately to emerging environmental expectations.”

At the manifest level, the researcher could code explicit references to shareholder value and environmental expectations.

At a more latent level, the researcher might interpret the statement as constructing environmental responsibility as something that needs to remain subordinate to—or balanced against—shareholder interests.

Latent analysis therefore requires greater interpretive involvement. This does not automatically make it better than manifest analysis. If a research question asks how frequently companies disclose measurable emissions targets, manifest coding may be entirely appropriate.

The methodological requirement is alignment: the level of interpretation should follow from the research question and analytical purpose.

Dudovskiy Content Analysis Design Framework

The Dudovskiy Content Analysis Design Framework is a practical tool for converting the broad statement “I will use content analysis” into a more explicit and defensible analytical design.

Decision Ask yourself Methodological implication
1. Analytical purpose Do I need to describe, compare, quantify or interpret content? Establishes what the analysis needs to produce
2. Analytical orientation Is the primary purpose quantitative or qualitative? Influences coding, analysis and quality criteria
3. Category development Should categories originate mainly from data, existing theory or both? Positions category development as more data-driven, theory-driven or mixed
4. Level of meaning Am I examining explicit content or underlying meaning? Determines emphasis on manifest or latent content
5. Coding design What exactly will be coded and according to what rules? Requires explicit units, categories and coding instructions
6. Quality strategy What evidence would demonstrate trustworthy analysis for this design? Determines appropriate emphasis on reliability, credibility, transparency and methodological coherence

These decisions should not be treated as six independent switches. They interact.

For example, a quantitative study examining the frequency of risk disclosures across 500 annual reports may require clearly operationalised categories and strong evidence of coding reliability. An interpretive qualitative analysis of how a smaller sample of companies constructs the concept of responsibility may require a different account of analytical quality.

Dudovskiy Content Analysis Design Framework

The central principle is:

Start with the inference you need to make, then design the coding system capable of supporting that inference.

This prevents researchers from constructing elaborate coding schemes before deciding what analytical problem those schemes are supposed to solve.

Units of Analysis in Content Analysis

A content-analysis study needs to specify what exactly is being analysed and coded.

The unit may be a word, sentence, paragraph, image, social-media post, advertisement, article, speech, document, organisation or another meaningful element defined by the research design.

Suppose a researcher investigates how newspapers portray entrepreneurship.

If the unit is the article, each article may be classified according to its dominant portrayal of entrepreneurship.

If the unit is the paragraph, several different portrayals could be coded within the same article.

If individual words are counted, the resulting analysis answers yet another kind of question.

Unit selection therefore affects the findings that can be produced. Krippendorff treats unitizing as one of the fundamental components of content-analysis design rather than a minor technical detail.

Researchers should also distinguish where necessary between the sampling unit, the material selected for the study, and the coding unit, the element to which a coding decision is applied.

How to Develop a Coding Frame

A coding frame translates the research question into a systematic way of classifying material.

Strong categories should be clearly defined and sufficiently distinct for the intended analytical purpose. Each category needs an explanation of what qualifies for inclusion and, where ambiguity is likely, what should be excluded.

Consider a study of customer reviews. A category labelled service quality is probably too broad if it encompasses employee politeness, response speed, expertise, complaint handling and availability. Breaking the concept into analytically meaningful subcategories may produce a more useful coding frame.

Examples can also be included in the coding instructions. They help clarify category boundaries and are particularly useful when several coders are involved.

The initial coding frame should not automatically be treated as final. Trial coding can expose categories that overlap, fail to capture important material or are interpreted inconsistently. Qualitative content-analysis guidance specifically treats piloting and revising the coding frame as important stages of the process.

The Content Analysis Process

Content analysis is better understood as a sequence of connected methodological decisions than as a mechanical coding exercise.

Stage Researcher decision Main output Typical failure
Define the question What inference should the analysis support? Analytical objective Coding without a clear purpose
Construct the corpus/sample What material should be included? Dataset Convenience selection without justification
Define units What exactly will be coded? Unit specification Changing units during analysis
Develop coding frame How should content be classified? Categories and coding rules Overlapping or vague categories
Pilot coding Does the coding system work in practice? Revised coding frame Moving directly to full coding
Conduct coding How will rules be applied consistently? Coded dataset Coding drift
Analyse results What patterns, differences or meanings matter? Findings Reporting frequencies without answering the research question
Evaluate and report How will analytical quality be demonstrated? Defensible methodology Unsupported claims of reliability or credibility

The process can be iterative. Problems discovered during pilot coding may require the researcher to redefine categories or coding rules. Qualitative approaches may involve further refinement as understanding of the material develops.

The goal is not to eliminate analytical judgement. It is to make the relationship between research question → material → coding → inference sufficiently explicit that the resulting conclusions can be evaluated.

Intercoder Reliability and Quality in Content Analysis

Intercoder reliability concerns the extent to which independent coders apply a coding system consistently. It is particularly important when the research design depends on multiple coders assigning predefined categories to material.

Agreement should not simply be assumed because coders received the same codebook. Categories may contain ambiguous boundaries, coders may interpret instructions differently and coding practices may drift over time. Pilot coding, coder training and refinement of coding rules can therefore be important.

Krippendorff’s alpha is one established reliability coefficient and can accommodate different numbers of coders, levels of measurement and missing data. Krippendorff’s methodological treatment places reliability and validity among the central issues in content analysis.

However, intercoder reliability should not become a universal ritual applied identically to every form of content analysis.

In more interpretive qualitative approaches, analytical quality may additionally or alternatively depend on transparent category development, systematic engagement with the material, reflexivity, credible interpretation and clear documentation of analytical decisions.

The appropriate question is therefore not merely:

“Did you calculate intercoder reliability?”

It is:

“What evidence of quality is appropriate for the particular content-analysis design you claim to have used?”

Content Analysis vs Thematic Analysis

Content analysis and thematic analysis both involve coding qualitative material and identifying patterns, which explains why they are frequently confused. However, their typical analytical objectives differ.

Dimension Content Analysis Thematic Analysis
Core analytical operation Systematic classification of content Development of patterns of shared meaning
Orientation Quantitative, qualitative or combined Primarily qualitative
Typical output Categories, frequencies, comparisons and/or interpretations Themes organised around central meanings
Quantification Can be central Usually not the primary purpose
Coding structure Often uses an explicit category/coding framework Codes contribute to theme development
Quality emphasis May place strong emphasis on coding reliability Depends substantially on the form of thematic analysis
Particularly useful when Content needs systematic classification or comparison Patterned meaning across qualitative accounts needs interpretation

The distinction should not be reduced to “content analysis counts while thematic analysis interprets.” Qualitative content analysis can also interpret latent meaning, and thematic analysis necessarily involves systematic engagement with data.

A more useful decision rule is to consider the desired analytical output.

If the researcher primarily wants to classify a large body of corporate reports according to systematically defined disclosure categories and compare those categories, content analysis is likely to be attractive.

If the researcher wants to develop patterns of shared meaning concerning how managers experience organisational change, thematic analysis may be more appropriate.

How to Report Content Analysis in a Dissertation

A dissertation methodology chapter should explain enough of the analytical design for the reader to understand how raw material became research findings.

Simply writing:

“The documents were analysed using content analysis.”

does not accomplish this.

A stronger methodology identifies the type of content analysis used and explains why it fits the research question. It specifies the corpus or sample, unit of analysis, development of categories or coding frame, coding procedure, treatment of manifest or latent content where relevant, analytical procedure and the strategy used to establish quality.

If categories were derived from theory, the theoretical basis should be identified. If categories were developed from the data, the process should be described. If several coders were involved, the dissertation should explain how coding consistency was managed and, where appropriate, assessed.

The objective is not procedural detail for its own sake. The reader needs enough information to judge whether the inferences presented in the findings are actually supported by the analytical process.

Application of Content Analysis: an Example

Consider a study examining how major listed companies communicate environmental responsibility in their sustainability reports.

The researcher selects sustainability reports published by 50 companies during the same reporting year. The study uses the company report as a sampling unit and relevant statements or paragraphs as coding units.

An initial coding frame contains categories including carbon-reduction targets, renewable-energy commitments, supply-chain environmental standards, environmental investment, independent assurance, stakeholder engagement and measurable performance indicators.

Some categories are coded quantitatively. For example, the researcher records whether each company publishes a numerical carbon-reduction target and whether that target has a specified deadline.

Other parts of the analysis are more interpretive. Statements concerning environmental responsibility are examined to understand whether companies frame sustainability primarily as an ethical responsibility, regulatory obligation, risk-management issue or source of competitive advantage.

The resulting analysis might reveal that environmental commitments are widespread but vary substantially in specificity. Companies may frequently use language of commitment and leadership while providing fewer measurable targets against which future performance can be evaluated.

The value of the analysis lies not simply in counting sustainability-related words. The coding system allows the researcher to make systematic comparisons and connect those observations to the study’s research questions.

Advantages and Limitations of Content Analysis

One of the major advantages of content analysis is its ability to examine communication systematically. Materials that would otherwise remain an unstructured collection of reports, posts, advertisements or interview transcripts can be transformed into an analytical dataset using explicit categories and coding rules.

The method is also unusually versatile. Researchers can analyse historical documents without interacting with participants, compare communication across organisations or periods, combine qualitative interpretation with quantitative patterns and work with data that were originally created for purposes unrelated to academic research.

Its systematic character can improve transparency, particularly when the researcher clearly specifies sampling decisions, units, categories and coding procedures.

However, the apparent precision of a coding frame can create false objectivity. A category does not become theoretically valid simply because coders can apply it consistently. Researchers themselves decide what is relevant, how concepts are operationalised and what distinctions the coding system makes visible.

Context can also be lost when complex communication is fragmented into small coding units. A sentence removed from the document surrounding it may acquire a different meaning, while frequency can be mistaken for significance. A concept mentioned rarely may be strategically or theoretically important.

Content analysis can therefore be rigorous without being mechanically objective. Its quality depends on whether the analytical design supports the inferences the researcher ultimately makes.

Common Mistakes When Using Content Analysis

A particularly serious mistake occurs when researchers start coding before deciding what their units and categories mean. Categories then evolve inconsistently as new material is encountered, making comparisons across the dataset difficult to defend.

Another problem is confusing frequency with importance. If innovation appears 500 times in corporate reports and employee wellbeing appears 100 times, the numerical difference does not by itself demonstrate that innovation is five times more important to those organisations.

Category overlap can also undermine the analysis. If the same statement could equally be coded as customer satisfaction, service quality and customer experience without clear rules distinguishing them, frequency comparisons become difficult to interpret.

Conversely, an excessively rigid coding frame can make researchers blind to important content that was not anticipated when the framework was constructed.

A further mistake is reporting a reliability statistic as if it establishes the validity of the entire study. High coder agreement indicates that coding rules can be applied consistently; it does not prove that the categories appropriately represent the underlying concept or that the researcher’s substantive conclusions are correct.

Finally, researchers sometimes call any examination of documents content analysis. Reading company reports and selecting interesting quotations is not automatically content analysis. The method requires a systematic relationship between the research question, material, units, categories, analytical procedures and resulting inference.

Content Analysis in Business Research

Content analysis is particularly valuable in business research because organisations continuously produce recorded communication that reflects strategy, positioning, stakeholder relationships and organisational priorities.

Potential datasets include annual reports, sustainability reports, earnings-call transcripts, corporate websites, advertisements, job advertisements, press releases, customer reviews, social-media posts and internal organisational documents.

For example, researchers investigating corporate sustainability can compare whether firms publish measurable environmental targets. Marketing researchers can analyse how competing brands position products in advertisements. Human-resource researchers can examine job advertisements to identify changing skill requirements. Strategy researchers can analyse annual reports to investigate how companies communicate risk.

Content analysis therefore allows researchers to study what organisations communicate, how communication differs and, depending on the analytical approach, what those patterns may mean.

Its unobtrusive character can be especially useful because many business documents already exist independently of the research project. Researchers can therefore examine organisational communication without relying entirely on participants’ retrospective accounts.

Content Analysis in the Age of AI and Digital Research

AI materially changes the possibilities for content analysis because researchers can now process quantities of textual material that previously made manual coding prohibitively expensive. Large language models can propose categories, apply coding instructions, summarise material and identify patterns across thousands of documents.

But increased scale does not remove the methodological problem at the centre of content analysis:

Who determines what the categories mean and whether the resulting classification supports a valid inference?

If an AI system labels 100,000 customer reviews, the size of the dataset does not compensate for an ambiguous coding scheme. Researchers still need to establish what is being measured, define categories, test coding rules and validate whether automated classifications correspond sufficiently to the intended concepts.

The challenge becomes even greater with latent content. AI can generate plausible interpretations of implied meaning, but plausibility is not the same as methodological validity. Model outputs can be affected by prompts, training data, contextual limitations and model updates. A classification process may therefore appear systematic while remaining difficult to reproduce or theoretically justify.

Recent methodological research on AI-assisted qualitative analysis consequently emphasizes human oversight and the need to integrate AI into a coherent research methodology rather than treating automated coding as a substitute for methodological reasoning.

A defensible AI-assisted content analysis should therefore document what the AI was asked to do, which model or system was used where relevant, how categories and prompts were developed, how outputs were checked, what decisions remained with researchers and how coding quality was evaluated.

AI can dramatically increase the scale of content analysis. It does not eliminate the researcher’s responsibility for the validity of the inference.

When to Use Content Analysis

Content analysis is particularly appropriate when:

  • the research question requires systematic examination of recorded communication;
  • documents, reports, media, advertisements, websites, transcripts or other recorded materials constitute important research data;
  • the researcher needs to classify content according to explicit categories;
  • comparisons need to be made across organisations, groups, sources or time periods;
  • frequencies or distributions of particular content are analytically meaningful;
  • qualitative interpretation needs to be conducted within a systematic categorisation framework;
  • existing theory can usefully guide a directed coding framework, or categories can defensibly be developed from the data;
  • the research requires analysis of naturally occurring materials without direct interaction with participants.

Content analysis is less appropriate when the research question primarily concerns individual lived experience, detailed conversational interaction, linguistic construction or the development of patterned shared meanings better addressed through another qualitative analytical tradition.

The decision should therefore begin with what the researcher needs to infer from the material, rather than with the simple fact that textual data are available.

Dissertation Example

Consider a dissertation titled “How UK Universities Communicate Generative AI and Academic Integrity: A Qualitative Content Analysis of University Policies.”

The researcher collects publicly available academic-integrity and generative-AI guidance from a purposive sample of UK universities. A directed qualitative content analysis is selected because the study seeks to examine how institutions frame recurring issues identified in existing literature, while remaining open to additional categories emerging from the documents.

The methodology chapter explains that initial categories include permitted AI use, prohibited use, disclosure requirements, authorship, assessment design, verification and consequences of misuse. Relevant policy statements constitute the principal coding units. The coding frame is piloted on a subset of documents and refined where category boundaries prove ambiguous or important content is not adequately captured.

The subsequent analysis compares how institutions define acceptable AI assistance and identifies differences in the allocation of responsibility between students, educators and institutions. The dissertation therefore demonstrates not merely that content analysis was used, but which form was chosen, how categories were developed, what was coded and how the analytical design addressed the research question.

Exam Tip

If an exam asks you to define content analysis, avoid describing it simply as “counting words in documents.” A stronger answer explains that content analysis systematically classifies recorded material to support research inferences and can be quantitative or qualitative. Distinguishing qualitative from quantitative content analysis—and, where relevant, manifest from latent content—demonstrates that you understand the methodological range covered by the term. For a higher-level answer, explain that the quality of content analysis depends on alignment between the research question, units, coding framework and the inference ultimately made.

Not sure if content analysis is the right method for your dissertation—or how your coding strategy should be designed?

Use Dudovskiy Research Assistant to get a methodology tailored to your research topic, with clear academic justification for the choices you need to explain and defend.

References

Bengtsson, M. (2016) ‘How to plan and perform a qualitative study using content analysis’, NursingPlus Open, 2, pp. 8–14.

Hsieh, H.-F. and Shannon, S.E. (2005) ‘Three approaches to qualitative content analysis’, Qualitative Health Research, 15(9), pp. 1277–1288.

Krippendorff, K. (2019) Content Analysis: An Introduction to Its Methodology. 4th edn. Thousand Oaks, CA: SAGE.

Neuendorf, K.A. (2017) The Content Analysis Guidebook. 2nd edn. Thousand Oaks, CA: SAGE.

Schreier, M. (2012) Qualitative Content Analysis in Practice. London: SAGE.

Vaismoradi, M., Turunen, H. and Bondas, T. (2013) ‘Content analysis and thematic analysis: Implications for conducting a qualitative descriptive study’, Nursing & Health Sciences, 15(3), pp. 398–405.

[]