# Research Ethics And Data Protection

> Research Ethics and Data Protection

- Skill: `ingridleiria/research-ethics-and-data-protection` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ingridleiria/research-ethics-and-data-protection`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ingridleiria/research-ethics-and-data-protection/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Research & Search
- Author: ingridleiria (https://skillmd.com/u/ingridleiria)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ingridleiria/research-ethics-and-data-protection

---


# Research Ethics and Data Protection

Two documents get treated as paperwork and neither is. An ethics application is a design document: a study that cannot describe its risks specifically has usually not thought about them, and committees are good at detecting the difference between a considered risk section and a reassuring one. A data management plan is an operational commitment: it says what will happen to real data on real systems for years, including after the researcher has left the institution, and it is enforceable.

The failures are predictable and they are all expensive in the same way, which is that they arrive after the work has been scheduled. An application submitted in the wrong month sits for eleven weeks and the fieldwork season is gone. An application returned for revision because recruitment ran through the participants' own manager costs another cycle. Consent that did not ask about reuse means the funder's data sharing requirement cannot be met, and there is no retrospective fix because the participants have dispersed. A dataset described as anonymised turns out to be pseudonymised, so it is still personal data, so the archive will not take it and the paper's data availability statement is wrong. And in the worst case a free-text field nobody read contains a participant's employer, role and town, and someone is identified.

None of this is bad luck. All of it is decided in the two weeks before submission, by a researcher who thinks of these documents as an obstacle rather than as part of the design.

Rules differ by country, institution, discipline and funder, and they change. What follows is the shape of the work and the judgement inside it. Verify every specific requirement against your own committee, your own institution's data protection office, and your own funder, rather than trusting any general description including this one.

## When to use this, and when not to

Use it before any study involving people, their data, their tissue, or their institutions: interviews, surveys, experiments, observation, and also secondary analysis of existing data, which needs review far more often than researchers expect because the data feels public. Use it when writing a data management plan for a funding application, which is usually the earliest deadline and the one most often written last. Use it when linking two datasets, which changes the risk profile even when neither dataset alone is sensitive. Use it when a committee has returned an application, because the stated reason is frequently a symptom of a structural problem in the recruitment or the consent rather than a wording issue.

Use it when the study involves people over whom the researcher, or the recruiting partner, holds any power. That case is the single most common reason for a returned application and it is invisible to the person who holds the power.

Do not use it to decide the research question or the design; those are `research-question-ideation` and `research-design`, and the ethics work should follow a design rather than substitute for one. Do not use it to write the survey or interview instrument itself, which is `survey-and-instrument-design`, although the instrument has to be attached to the application and its content drives half the risk assessment. Do not use it for the reproducibility side of sharing, which is `replication-package`; that skill handles code, documentation and archiving, and this one handles whether sharing is permitted at all and in what form.

Do not use it as legal advice. Where a question turns on the lawful basis for processing, on a cross-border transfer, or on a contract with a data provider, the answer comes from the institution's data protection officer or research office, and the correct action here is to identify the question precisely and route it.

## What you need before starting

**The design.** Population, recruitment route, what participants will do and for how long, what data is collected, and where it goes. Missing: the application cannot be written, and attempting it produces a document full of "may" and "as appropriate" which committees read as a lack of planning. Run `research-design` first.

**The list of committees with jurisdiction, and their real timelines.** Not the published turnaround, the actual one, which the department administrator or a recent applicant will know. Missing: ask, this week, before anything is scheduled. Ethics timing is the most predictable cause of a delayed project and the least often planned for.

**The data categories, named precisely.** Which fields are personal, which are sensitive under the rules that apply to you, whether any are about children or health or protected characteristics, and whether identifiers are held. Missing: list every variable and classify each. This exercise takes an hour and repeatedly surprises people, because free-text fields and open-ended survey questions carry identifiers that no variable list predicts.

**The recruitment route, and who holds power over whom.** Missing: map it explicitly. Who first contacts the participant, who knows who agreed, and what a refusal costs the participant. If the answer to the last is anything other than nothing, the route needs changing before the application is written.

**The storage and transfer reality.** Which systems are approved by the institution for this category of data, what the institution's position is on personal devices, on recording apps, on transcription services, and on cloud storage whose servers sit in another jurisdiction. Missing: ask the research computing or information governance team, and do it before choosing tools, not after buying a subscription.

**The funder's data sharing and retention requirements.** Missing: read the grant conditions now, because they determine what consent must ask for, and consent is the one thing that cannot be repaired later.

**The instrument, or a good draft.** Committees review what participants will actually see. Missing: submit the draft and say it is a draft, but expect a condition requiring the final version, and plan the extra cycle.

## The method

1. **Establish jurisdiction and calendar before anything else.** Which committee or committees, what the submission deadlines are, how long review really takes, and whether a second approval is needed for another site, another country, a data provider, or a health system. Put the dates in the project plan. The judgement call: where two committees might apply, apply to both rather than choosing, because a study reviewed by the wrong body has not been reviewed.

2. **Classify the data before writing anything.** Every variable, marked as identifying, indirectly identifying, sensitive, or none of these. The rule: treat a variable as indirectly identifying if it could be combined with two other variables in the dataset to narrow a person to a small group. This classification drives the storage decision, the anonymisation plan, the sharing route, and half the risk section, so doing it first saves rewriting all four.

3. **Write the risk assessment specifically, one row per risk.** Physical, psychological, social, legal, economic, and breach. For each: what could happen, to whom, how likely, how it is minimised, what happens if it occurs anyway, and what support exists. The rule that separates a real assessment from a reassuring one: name at least one risk that cannot be fully mitigated, because every study involving people has one, and a risk section with no residual risk reads as unexamined. Distress from being asked about a difficult subject is a real risk and naming it plainly is better received than minimising it.

4. **State the benefit honestly.** To participants where there is one, to knowledge where there is not. Most social science offers no direct participant benefit and that is acceptable and normal. Claiming a benefit that does not exist is not, and committees notice, because the claimed benefit is usually vague where a real one would be concrete.

5. **Fix the recruitment route so that declining is invisible and costless.** Where the researcher, an employer, a clinician, a teacher or a manager holds power over potential participants, route the invitation through someone who does not, let participants self-select by contacting the researcher rather than being nominated, and ensure the powerful party never learns who took part. Say in the application how the two roles are separated. The judgement call: if separation is genuinely impossible, say so, describe the compensating protections, and let the committee decide rather than presenting the arrangement as unproblematic.

6. **Write the information sheet for the participant, not the committee.** Plain language, in their language, at a length someone will read, typically one to two pages. It must cover who is doing this and who funds it, what participation involves and how long it takes, what happens to the data and who sees it and for how long, the risks, that participation is voluntary, how to withdraw and until when, whether the data may be reused or shared and in what form, and who to contact including for a complaint. The rule on withdrawal: state the limit explicitly, because after anonymisation or publication withdrawal is usually impossible, and a participant discovering that afterwards is a complaint waiting to happen.

7. **Ask for reuse and sharing consent at the outset, specifically.** Not as a general permission but as named uses: deposit in a named repository, availability to other researchers under controlled access, use in teaching, quotation in publications. Retrospective consent for sharing is rarely obtainable, so a consent form that omits this forecloses the funder's requirement permanently.

8. **Plan anonymisation as a risk judgement, not a deletion step.** See the section below. The decision to record: what will be removed, what will be generalised, what will be aggregated, who holds any key, where, and when it is destroyed.

9. **Write the data management plan as operations.** Named systems, named people, named dates. See below. The test: could a colleague execute it in your absence, and could an auditor check whether it happened?

10. **Write the four procedures before fielding.** Disclosure of harm, incidental findings, breach, and withdrawal. Each is a situation that arrives without warning and is handled badly when improvised. Each takes twenty minutes to write and is the difference between a bad afternoon and a serious incident.

11. **Read the whole application as the committee will.** They are looking for three things: does the researcher understand the risks, can the participant understand the sheet, and will the data actually be protected as claimed. Every inconsistency between sections is read as evidence against all three. The most common inconsistency is a sheet promising anonymity while the plan describes a key.

12. **Submit early enough to absorb one revision cycle.** Assume it will be returned, because most are, and a first submission timed to leave no slack turns a routine condition into a lost term.

## Anonymisation, done properly

Removing names is not anonymisation. Re-identification happens through combinations: a role, an organisation, a location, a date, and a rare characteristic together identify a person when none of them alone does. The relevant question is never whether direct identifiers were deleted; it is whether a motivated person with plausible auxiliary information could narrow a record to one individual.

Distinguish two states carefully, because in many jurisdictions the legal consequences differ sharply and researchers routinely conflate them. Data is anonymised when re-identification is not reasonably possible by anyone, including you, which usually means the key has been destroyed. Data is pseudonymised when identifiers have been replaced but a key exists somewhere, in which case it remains personal data and the full protections continue to apply. Most research data described as anonymised is pseudonymised. Getting this wrong in an information sheet is a promise that cannot be kept.

Four specific hazards recur.

Small cells identify people. A table with three respondents in a category, or a subgroup analysis on a rare characteristic, can expose an individual even in a large dataset. Set a minimum cell size in advance, as a number, and apply it to outputs as well as to deposited data.

The conventions are settled enough to quote. Five is the common default for published frequency tables, and it is also the usual k in a k-anonymised microdata file; most health and education statistics producers suppress counts below it. Ten is what most secure research environments require for anything exported, and it is the right choice for any subgroup defined by a sensitive attribute: health, ethnicity, sexuality, immigration status, union membership, a criminal record. Three appears in some producers' rules for plainly non-sensitive counts and is rarely defensible on its own. Where the cell holds a total rather than a headcount, for instance total pay, the governing rule is a dominance rule instead: suppress where one or two contributors account for most of the total, commonly written as no single contributor above 85 percent.

Three consequences follow and are routinely missed. Suppressing one cell without suppressing a second in the same row or column leaves the first recoverable by subtraction, so secondary suppression is part of the rule rather than an extra step. A percentage computed on a denominator below the threshold discloses as much as the count, so the threshold applies to denominators too. And two tables drawn from the same data can be differenced to recover a suppressed cell, so the rule applies across the whole set of outputs rather than table by table. If your data provider, agency or committee sets a number, use theirs and cite it; if nobody does, write five for published counts and ten for sensitive subgroups into the design and hold to it.

Quotations carry unique details. A sentence naming a role, a length of service and a decision is identifying inside the organisation even when the organisation is pseudonymised, and it is colleagues rather than strangers who will do the identifying. Judge each quotation against the internal reader, not the external one.

Free text hides identifiers. Open-ended survey responses, interview transcripts and notes fields contain names, places and employers that no variable-level plan predicts. They have to be read, by a person, before any sharing, and the time for that has to be in the schedule.

Linkage defeats generalisation. A dataset safe alone may be identifying once joined to another, and the risk assessment must consider what else a recipient plausibly holds.

Where a key exists, record who holds it, where it is stored separately from the data, who else can access it, and the date and method of destruction. Where the plan is to destroy it, say what verifies the destruction.

## The data management plan

Write it as an operational document that names things. Every clause below should contain a proper noun or a date.

What data is collected, in what format, and at what volume. Where it is stored, on which specific system, administered by whom, and whether that system is approved by the institution for this category of data. Who has access, how access is granted, and, the clause most often missing, how access is removed when someone leaves the project. How data is transferred, since transfer is where most breaches happen, and with the explicit statement that email is not a transfer method for personal data. What the backup arrangement is and who verifies it works. The retention period, what triggers destruction, how destruction is carried out and how it is verified. And who is responsible for the data after the project ends or the researcher leaves, named as a role at minimum and preferably as a person.

Three areas need explicit attention rather than assumption: cloud storage whose servers sit in another jurisdiction, cross-border transfer of any kind, and personal devices including phones used for recording. A transcription service is a data processor and needs the same scrutiny as any other, including where its servers are and what it retains.

Where the funder requires a specific plan format, write the operational version first and map it onto the form, because forms invite general answers and the general answer is the failure.

## The four procedures to write before fielding

**Disclosure.** What happens if a participant discloses harm to themselves or others, or something you are obliged to report. The limits of confidentiality must be stated in the information sheet before anything is disclosed, not explained afterwards. Decide the threshold, the person to consult, and the timeline.

**Incidental findings.** What happens when the research reveals something about a participant that they did not know and might want to, or might not. Decide in advance whether findings are returned, which ones, and through whom.

**Breach.** Who is notified, in what order, within what deadline, and who decides whether participants are told. Deadlines for notifying an institution or a regulator are often short, measured in days or hours, and are not the moment to start reading policy.

**Withdrawal.** What withdrawal means at each stage: before fielding, during, after transcription, after anonymisation, after analysis, after publication. The point at which it becomes impossible is the number that goes in the information sheet.

## Worked example

**Situation.** A doctoral student, Priya Ramaswamy, at Kelso University, planned 34 interviews with care assistants across six residential homes run by a single provider, on how staff interpret and apply a new medication recording procedure. She also wanted to link interview participants to two years of shift records held by the provider's human resources system. Her funder required a data management plan and ten year retention. The provider's operations director was enthusiastic and had offered to circulate the invitation to staff and collect the consent forms himself. Fieldwork was planned for a twelve week window starting in nine weeks.

**Task.** An approvable application, a consent process that permitted the linkage and the funder's archiving requirement, and a data management plan that the provider's own information governance lead would sign.

**Action.** The variable classification in step 2 produced the first surprise. Shift records contained not only hours but sickness absence codes, which in that system encoded whether an absence was health related. That moved a chunk of the linked dataset into a sensitive category, changed the storage requirement, and meant the linkage had to be justified rather than assumed. The alternative, requesting shift records with the absence reason field suppressed, was adopted, and the student's own view later was that the field would not have been used anyway. Requesting less data was the cheapest risk reduction available and it was found only because the classification was done variable by variable.

The first wrong turn was the recruitment route. Having the operations director circulate invitations and collect consent forms was convenient, and it meant the employer would know exactly which staff had agreed to be interviewed about compliance with a procedure their employer had introduced. Care assistants on shift patterns with limited job security are not in a position to decline a request from that quarter and be sure it costs nothing. This was not a wording problem and no amount of reassurance in the information sheet would have repaired it.

The route was rebuilt. The provider circulated a notice describing the study with the researcher's contact details and no requirement to respond through the employer. Staff contacted the researcher directly. Interviews were held off site or in a private room booked by the researcher, and the provider was told only the total number of participants per home and only after fieldwork ended, with a floor of three so a single participant home could not be identified. The operations director was initially unhappy and accepted it once the reason was explained as protection for his staff rather than distrust of him.

The second wrong turn was in the draft information sheet, which promised that all data would be fully anonymised. The design required a key, because interview transcripts had to be linked to shift records, and the key would exist for at least the duration of the analysis. The sheet was promising something the study could not deliver. It was rewritten to say that the recordings and transcripts would be pseudonymised, that a key linking the pseudonym to the participant would be held by the researcher alone on an institutional encrypted drive, that it would be destroyed once linkage was complete and no later than a stated date, and that after that point withdrawal would no longer be possible because the data could no longer be located. That last sentence is the one that matters and it was in the sheet from that version onwards.

The third repair was storage. The original plan had recordings going to a personal cloud account and a commercial transcription service chosen for price. Research computing confirmed the institution had an approved encrypted store and an approved transcription arrangement, and that the cheaper service retained audio for thirty days on servers in another jurisdiction, which would have been a cross-border transfer of personal data with no assessment behind it. The switch cost about four hundred pounds against the original budget line and removed a clause the committee would have queried.

Consent was rebuilt around named uses, with separate boxes: participation, audio recording, linkage to shift records, use of anonymised quotations in publications, and deposit of the anonymised transcript set in a named repository under controlled access. Separating these mattered: of the twenty-nine people eventually interviewed, eleven agreed to everything and eighteen declined the repository deposit, and two of those eighteen also declined recording and were interviewed with notes.

The four procedures were written in an afternoon. The disclosure procedure earned its place: the study asked about medication recording in a care setting, so a participant describing a practice that endangered a resident was foreseeable, and the threshold, the person to consult and the obligation to act were fixed in advance and stated in the information sheet.

**Result.** The committee returned the application once, with two conditions: a plain-language pass on the information sheet, which had drifted back towards the register of a legal document, and a request for the provider's written confirmation of the linkage arrangement. Both were cleared in nineteen days. Total elapsed time from submission to approval was eight weeks against a published turnaround of six.

Fieldwork started five weeks later than originally planned, which was absorbed because the calendar in step 1 had been built with a revision cycle in it. Twenty-nine of the intended thirty-four interviews were completed.

The deposit was smaller than hoped, at eleven transcripts, because most participants declined that box. That is the correct outcome of asking properly, and the data availability statement in the resulting paper said so in one sentence rather than claiming more.

The unglamorous detail that mattered most: reading the transcripts for identifiers before deposit took eleven hours, because care assistants describe incidents by naming residents, colleagues and shift dates. No variable-level plan would have caught that, and the time had been scheduled only because the free-text hazard was in the plan.

### A second scenario, where it goes differently

A postdoctoral researcher at the same university proposed a study using linked national administrative records on education and earnings, accessed inside a secure research environment run by the statistical agency. No participants are recruited, no consent is obtainable, and no information sheet exists.

Almost every step changes content while the structure holds.

Step 1 still runs and produces a different list: the institution's committee, which may treat this as low risk or exempt, plus the agency's own data access panel, which is the real gate and which reviews the research question rather than the researcher. Its timeline was fourteen weeks and its rejection rate was material, so a fallback question was prepared in advance.

Consent is replaced by lawful basis. The question is not what the participant agreed to but under what legal authority the data is processed for research, and that question goes to the institution's data protection officer rather than being answered by the researcher. The application must justify why the research cannot be done with less data, which is a real constraint: a request for the full population where a cohort would do is refused, and refusals cost a cycle.

Anonymisation is replaced by output disclosure control. The data never leaves the environment, so the risk sits entirely in what is exported: tables, coefficients, figures. Rules are set by the agency, typically including a minimum cell count, suppression of dominant contributors, and restrictions on maxima and minima. The practical discipline this imposes on the analysis is easy to underestimate. A heterogeneity table with a subgroup of four cannot be exported, so subgroups have to be designed against the disclosure rule in advance, which means the disclosure rule belongs in the research design rather than being discovered at export.

The data management plan shrinks to almost nothing on storage, because the agency holds the data, and expands on outputs, code and derived files. Retention becomes the agency's project end date rather than the funder's ten years.

The four procedures reduce to one. There is no disclosure risk, no incidental finding and no withdrawal, but there is a breach procedure, and in a secure environment the realistic breach is an accidental export or a screenshot, which is a discipline problem rather than a technical one and is worth writing down for that reason.

What did not change: the variable-by-variable classification, the requirement to state a residual risk honestly, the calendar built with a revision cycle in it, and the rule that a promise in a document must be one the study can keep.

## Output

Four deliverables, and they must agree with each other.

```
ETHICS APPLICATION
Study, question, design, and the instrument attached.
Committees:      [each, with submission date, expected decision date, revision slack]
Participants:    [population, number, how identified, inclusion and exclusion]
Recruitment:     [who invites, who never learns who agreed, what declining costs]
Consent:         [process, timing, who takes it, language, capacity considerations]

| Risk | To whom | Likelihood | Mitigation | If it happens | Residual? |

Benefit:         [to participants, or none; to knowledge, specifically]
Data:            [see classification table]
Procedures:      [disclosure, incidental findings, breach, withdrawal: attached]
```

```
PARTICIPANT INFORMATION SHEET
Plain language, participant's language, one to two pages.
Who we are and who funds this. What you will be asked to do, and for how long.
What happens to what you tell us: who sees it, where it is kept, for how long.
Risks. Voluntary participation. How to withdraw, and the date after which you cannot,
and why. Whether your data will be shared or reused, and in what form.
Limits of confidentiality, stated before anything is disclosed.
Who to contact, including for a complaint.
```

```
CONSENT FORM
Separate boxes, each initialled:
[ ] I agree to take part
[ ] I agree to the interview being audio recorded
[ ] I agree to my data being linked to [named records]
[ ] I agree to anonymised quotations being used in publications
[ ] I agree to my anonymised data being deposited in [named repository] under [access route]
```

```
DATA MANAGEMENT PLAN
| Dataset | Format | Category | System | Administered by | Who has access | Access removed by |
Transfer method: [ ]        Backup: [ ]        Verified by: [ ]
Key: held by [ ] at [ ], destroyed on [date] by [method], verified by [ ]
Anonymisation: [what is removed, generalised, aggregated; minimum cell size]
Retention: [period] triggered by [event]. Destruction by [method], verified by [ ]
Responsible after project end: [named role or person]
Sharing route: [open / on request / controlled access via named committee]
```

## Failure modes

**Recruitment through the person with power.** Recognise it when the employer, clinician, teacher or supervisor distributes the invitation or learns who agreed. The commonest cause of a returned application. Fix by rebuilding the route so participants self-select and the powerful party never learns who took part.

**Promising anonymity while holding a key.** Recognise it by comparing the information sheet against the data management plan. Fix by using the accurate term, describing the key, and stating the point after which withdrawal is impossible.

**Consent that does not cover sharing.** Recognise it when the form has one box. Fix before fielding, because there is no fix afterwards, and separate the uses so that a participant can agree to some and not others.

**The generic risk section.** Recognise it because the risks could belong to any study and none is residual. Fix by naming risks specific to this population and this topic, and by admitting the one that cannot be eliminated.

**Timeline optimism.** Recognise it when the fieldwork date assumes first-time approval. Fix by building one revision cycle into the plan as a default, not as a contingency.

**A tool chosen before governance was consulted.** Recording apps, transcription services, survey platforms and cloud drives all process personal data. Recognise it when a subscription already exists. Fix by asking research computing what is approved before choosing, and expect the approved option to cost more.

**Free text assumed clean.** Recognise it when the anonymisation plan lists variables and says nothing about open responses or transcripts. Fix by scheduling a human read of every free-text field before any sharing.

**Small cells in published outputs.** Recognise it in a heterogeneity table with single-digit counts. Fix by setting a minimum cell size at design stage, as a number (five for published counts, ten for sensitive subgroups, unless your provider sets its own), and applying it to the paper as well as the archive.

**Nobody responsible after the researcher leaves.** Recognise it when the plan's retention period outlasts the contract. Fix by naming a role that persists, and by telling that person.

**Treating the committee as an adversary.** Recognise it in an application written to be hard to refuse rather than easy to assess. It reads as evasion and it produces more conditions, not fewer. Fix by stating the difficult things plainly and proposing the mitigation.

## Edge cases

**Secondary analysis of existing data.** Frequently still requires review, and the licence or data use agreement usually imposes obligations beyond the ethics approval, including on outputs and on retention. Read the agreement, not a summary of it, and note that some agreements forbid linkage or require approval of publications.

**Covert or partially deceptive designs.** Some fields permit deception with a debriefing; many do not permit covert observation at all. Where deception is proposed, the application must justify why the question cannot be answered otherwise, describe the debriefing, and offer withdrawal after debriefing. Do not assume your committee's position; ask before designing around it.

**Minors, people without capacity, and other formally protected groups.** Specific rules apply, they vary considerably by jurisdiction, and they are not a place to improvise. Establish the assent and consent requirements, the safeguarding checks required of the researcher, and the reporting obligations, before designing recruitment.

**Research in another country.** Expect local approval as well as home approval, expect it to take longer, and expect the local requirements to differ in substance rather than only in form. Budget the time and identify a local collaborator who knows the committee.

**Online and social media data.** Public availability is not consent, terms of service are not ethics, and a quotation from a public post is often traceable back to its author by a search engine. Decide whether posts will be quoted verbatim, paraphrased, or aggregated, and treat identifiability rather than accessibility as the test.

**The study changes after approval.** An added site, a new question, a new linkage, an extended period. Amend before doing it. Committees process amendments faster than applications, and proceeding without one is the kind of finding that jeopardises a whole project rather than one element.

**A participant withdraws after anonymisation.** If the sheet stated the limit, explain it kindly and honestly, and remove them from anything still removable such as future contact. If the sheet did not state the limit, the promise governs, which is why step 6 insists on it.

**Data the participants would not want shared, and a funder that requires sharing.** Do not resolve this by over-promising to either. Use a controlled access route with a data access committee, describe the restriction and its reason in the paper in one honest sentence, and record the reason in the plan so a successor does not have to reconstruct it.

**A breach happens.** Follow the written procedure, notify inside the institution first and immediately, do not investigate alone, and document the timeline as it occurs rather than afterwards. The severity of the outcome depends far more on the speed of notification than on the size of the breach.

## Quality bar

- Every committee with jurisdiction is identified, with real timelines and one revision cycle in the project plan.
- Every variable is classified, and free-text fields are treated as identifying until read.
- Each risk is specific to this study, with a mitigation and a response, and at least one residual risk is named.
- The recruitment route makes declining invisible and costless to the participant, and the application says how.
- The information sheet is in the participant's language, readable in a few minutes, and states the date after which withdrawal is impossible.
- Anonymised and pseudonymised are used accurately, and every promise in the sheet is one the design can keep.
- Consent asks separately for each intended use, including linkage, quotation and deposit.
- The data management plan names systems, people and dates, including who is responsible after the project ends.
- Disclosure, incidental finding, breach and withdrawal procedures exist in writing before fielding.

## Adapting this to your context

The vocabulary here is European: a research ethics committee, a data protection officer, lawful basis, pseudonymised. The structure holds elsewhere; the names and some of the obligations do not.

- **The regulatory frame.** In the United States this is an IRB under the Common Rule, with exempt, expedited and full board routes, plus HIPAA and its eighteen identifiers for health data. Brazil has the CEP and CONEP system and the LGPD, Canada TCPS 2 and an REB, Australia an HREC. Map each term to your own before writing, because the exempt route in particular has no European equivalent.
- **Trials and clinical work.** Prospective registration on ClinicalTrials.gov, ISRCTN or a WHO primary registry is a condition of publication in most medical and many psychology journals, and it happens before the first participant is enrolled. Add adverse event reporting and, for anything with risk, a data monitoring arrangement.
- **The cell size thresholds.** Five and ten above are general disclosure conventions. If your data provider, secure environment or statistical agency publishes its own, that number governs and yours is irrelevant.
- **What not to change.** Every promise in the information sheet must be one the design can keep, and consent asks separately for each intended use, including deposit, before fielding, because none of it can be obtained afterwards.

## Related skills

`research-design` produces the design this application describes and should run first. `survey-and-instrument-design` supplies the instrument that is attached to the application and that drives much of the risk assessment. `preregistration-and-analysis-plan` runs in parallel on a shorter timetable and its exclusion and subgroup rules interact with disclosure control where outputs are restricted. `qualitative-coding-and-analysis` handles the transcripts this skill protects, including the identifier read before sharing. `replication-package` takes over once sharing is permitted, covering code, documentation and deposit. `research-proposal-and-grant` carries the data management plan into the funding application, usually before the ethics application is written.

