A secondary-data dissertation uses evidence that already exists, but the study still has to prove that the source can answer its research question. Availability alone is not enough.
Start with the evidence need rather than the available dataset
Define what the research question requires before deciding that a dataset, archive or document collection is suitable.
Verify who collected the evidence and why
Understand the source organisation, original purpose, population or cases covered, time period, definitions and collection process. These facts affect interpretation.
Source provenance helps determine what the evidence can support. Identify the organisation or researchers responsible for collection, the original purpose, the population or cases included, the collection period, and the definitions used. Administrative data created to operate a service may have very different strengths and weaknesses from data collected specifically for research.
Where documentation is available, examine methodology notes, codebooks, questionnaires or data dictionaries rather than relying only on the dataset title or download page.
Check whether variables or concepts match the question
A dataset may contain a familiar variable name but define or measure it differently from the dissertation. This matters particularly in quantitative research, while documentary and textual sources can create comparable interpretation problems in qualitative research.
Conceptual fit matters as much as data availability. A variable called “income,” for example, could refer to individual income, household income, gross income or a grouped income band. A dissertation that needs one definition cannot silently substitute another.
For qualitative secondary material, the same issue applies to context. Documents written for a different purpose may not contain the detail needed to answer the dissertation question even when they mention the same topic.
Examine coverage, missingness and measurement limits
Check which cases or periods are absent, how missing values are represented and whether key measures changed over time.
Coverage should be checked across people, organisations, geographic areas and time. A dataset may omit hard-to-reach groups, begin after an important policy change, end before the relevant period, or contain different levels of completeness for different variables.
Missingness also needs interpretation. A blank value can mean that information was unavailable, not applicable, refused, not collected or lost. Those situations should not automatically be treated as equivalent during cleaning.
Confirm access, licensing and ethics
Public availability does not automatically remove usage conditions or ethical responsibilities. Review source terms and university guidance, and use dissertation ethics help where identifiable, sensitive, restricted, or permission-controlled data are involved.
Plan data preparation without changing meaning
Cleaning may involve formats, duplicates, missing values or inconsistent labels, but changes should preserve meaning and be documented.
Keep a record of recoding, exclusions, merging, transformations and treatment of missing values so the analytical dataset can be explained. Where categories are combined, confirm that the new grouping remains substantively meaningful rather than being created only to make analysis easier.
Retain an unchanged copy of the original source where permitted. This makes it possible to distinguish original evidence from decisions introduced during preparation.
Keep analysis inside the source limits
Do not make claims the original coverage or measurement cannot support. For analysis support, use dissertation data analysis help.
Report limitations created by secondary evidence
Explain the consequences of unavailable variables, inherited measurement choices, selection, time coverage, missing data or restricted context.
For whole-methodology integration, return to Methodology Chapter Help.