# Research Question Ideation

> Turns a topic, a dataset, or a vague interest into a shortlist of three research questions that could survive a committee, each written as one testable sentence, each checked against the nearest existing paper by verified citation, each with a named source of variation and a rough minimum detectable effect before anyone commits a semester to it. Enforces the rule that a question is killed on identification or feasibility before it is ranked on interest. Use this skill when someone says "I want to work on X", "is this a good paper idea", "what can I do with this dataset", "help me find a research question", "is this novel", "I need a third chapter", "my supervisor says this is too broad", or brings enthusiasm and no question. Trigger at the start of a project, before research-design, and whenever an existing idea has stalled and needs to be stress-tested or replaced.

- Skill: `ingridleiria/research-question-ideation` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ingridleiria/research-question-ideation`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ingridleiria/research-question-ideation/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: ingridleiria (https://skillmd.com/u/ingridleiria)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ingridleiria/research-question-ideation

---


# Research Question Ideation

Doctoral projects fail slowly. A student spends eleven months cleaning data, running specifications, and writing a chapter around a question that was never identifiable, or was answered in 2019 by a paper they had not found, or was so narrow that no examiner could say who should care. The execution is often competent. The question was wrong at the start and nobody stopped it, because at the start it sounded interesting and everyone was being encouraging.

The cost is measured in years, not weeks. A killed question in week two costs an afternoon. The same question killed in month eleven costs the chapter, the conference submission built on it, the data access application, and usually the student's confidence in their own judgement, which is the expensive part. Supervisors who let bad questions through are almost never being careless; they are being kind at the wrong moment.

This skill is the demanding conversation that should happen in week two. It generates candidates fast, kills most of them faster, and produces three that have already survived the objections a committee will make.

## When to use this, and when not to

Use it at the point where there is enthusiasm and no question: a new project, a new chapter, a dataset that has just become available, a supervisor meeting where the feedback was "this is too broad", a proposal that has to be written and has nothing to build on yet. Use it also when an existing question has stalled, because a stalled question usually has a defect that ideation would have caught and the honest move is to test it again rather than push harder.

Use it when someone arrives with a method rather than a question, which is common and is a specific failure worth catching. "I want to do a regression discontinuity" is a tool looking for a job.

Do not use it once the question is settled and what is needed is hypotheses, an estimating equation, a sample definition and an exhibit list; that is `research-design`, and this skill hands work to it. Do not use it to fix a paper that already has results, where the real task is reframing rather than reinvention; `introduction-writer` and `journal-targeting` cover that. Do not use it as a substitute for reading the literature. The searching here is enough to see the shape of a conversation and to locate the nearest paper; it is not enough to make anyone expert, and a question that survives this process still needs its literature read properly by `literature-verification`.

Do not use it where the question is set by a funder, a contract, or a supervisor's grant. There the task is to find the identifiable version of an imposed question, which is a narrower exercise: run the kill round only, on the one candidate and two near variants.

## What you need before starting

**The seed.** A topic, a dataset, an observation, a policy someone finds interesting, or a paper that annoyed them. Rank the seed types honestly: a dataset in hand is the strongest seed because feasibility comes free, a policy reform with known timing is next because identification often comes free, and a topic is the weakest because it constrains nothing. Missing: there is always something. Ask what they read last that made them want to argue with it.

**The constraints.** Time to submission, the venue being targeted, the methods the person can actually execute today rather than aspire to, and whether they have or can get the data. Missing: assume the hardest realistic case, which is eighteen months to a defensible chapter using methods they have already used once, and say you have assumed it. Designing for skills someone has not yet acquired is how a two year project appears.

**Data access reality.** Whether the data is downloadable now, requires an application with a stated turnaround, requires a secure room, or does not exist yet and would have to be collected. Missing: treat it as the first thing to check, and make the recommendation conditional on it. A question that depends on restricted microdata with a nine month approval queue is a different project from the same question on public files.

**The field's current conversation.** What is being asked now, what is contested, what the venues the person is targeting have published on the topic in the last five years. Missing: search before generating, not after. Ideation without a search produces questions that were answered while the student was reading for their qualifying exam, and the embarrassment of that discovery in a committee meeting is avoidable in ninety minutes.

**Whether anything has already been tried.** Prior dead ends, questions the supervisor has already rejected, chapters that exist. Missing: ask directly. Regenerating a candidate the supervisor killed last term wastes the meeting and damages credibility.

## The method

1. **Fix the seed and name its type.** Write one line: this is a dataset seed, a policy seed, a topic seed, or a contradiction seed. The type determines which generators will be productive. Dataset seeds produce good feasibility and weak novelty; topic seeds produce the reverse. Knowing which you have tells you where the candidates will be weak before you write them.

2. **Search the current conversation, for forty to ninety minutes, and stop.** Two or three searches per subtopic, varying vocabulary because fields name the same construct differently. Read abstracts, not papers. What you need out of this is a map: the three or four questions currently being asked, the two findings that are contested, and the names that recur. Record every reference with an identifier, verified, following the standard in `literature-verification`. Do not extend this stage; the point is to generate against a real conversation, not to become expert, and open-ended reading at this stage is the most common way ideation quietly becomes a literature review that never produces a question.

3. **Generate eight to twelve candidates using the generators deliberately, in turn.** Working through the generators in order rather than free-associating is what prevents a list of eight variations on one idea, which is the usual output of unstructured brainstorming and which feels productive while foreclosing the alternatives.

4. **Write each candidate as one sentence with population, variation, and outcome named.** If the sentence cannot be written, the candidate is not yet a question and either gets fixed on the spot or dropped. This step alone eliminates about a quarter of what generation produces.

5. **Run the kill round, in order, and stop scoring a candidate as soon as it fails.** Identification first, feasibility second, then novelty, interest and fit. The order matters and is not the order people naturally use. Scoring interest first makes everyone attached to a candidate before they discover it cannot be identified, and attachment is what keeps dead questions alive for months. Kill on identification or feasibility without negotiation; there is no partial credit.

6. **Estimate the sample and the minimum detectable effect for each survivor, roughly, in writing.** Ten minutes with a back of the envelope calculation. The rule: if the effect the literature reports is smaller than the effect the design could detect, the question is dead regardless of how good it sounds. This is where most surviving candidates actually die, and it is the step people skip because it feels premature.

7. **Rank the survivors on interest and fit and cut to three.** Interest has a test: name the reader, or the policy decision, that changes on either answer. If the only honest answer is "the literature", the candidate ranks last.

8. **For the top candidate, name the single cheapest thing that would falsify it,** and say it should be done before anything else. Usually a pre-trends plot, a count of treated units, a check that the key variable is populated in the years needed, or an email asking whether the data access even exists. This one check, done in the first week, is what separates this method from a nice meeting.

9. **Write the killed candidates down with their reasons.** One line each. This costs five minutes and prevents the specific waste of rediscovering the same dead idea in six months, which happens reliably because the reason it died is forgotten faster than the idea itself.

## The generators

Work through these in turn. Each produces a structurally different kind of question, and a shortlist drawn from only one generator is fragile because a single objection kills all of it.

**Setting transfer.** A well-established finding tested where institutions, prices, enforcement, or populations differ enough that the answer could genuinely change. The discipline here is the second clause: say why it could change. "The same thing, in my country" is not a question, and referees say so. "The same thing where enforcement is delegated to municipalities rather than a national inspectorate, which is why compliance may not respond" is.

**Mechanism.** The literature agrees that X affects Y and does not know through which channel. Design the question around distinguishing two named channels that predict different signs somewhere observable. Mechanism questions are strong for a thesis chapter because they are small, defensible, and they cite well.

**Heterogeneity with a reason.** For whom does the effect differ, where theory predicts the difference in advance and the data can measure the moderator. Without the theory clause this is a fishing expedition with subgroups, which is a different and worse activity.

**Policy variation.** A reform, threshold, rollout, quota, lottery, or eligibility cutoff the data can see. This generator gives identification away for free and should always be run when a dataset seed is present, because the strongest questions available to a student with administrative data are usually sitting in the institutional history of the data rather than in the theory.

**Measurement.** An outcome or treatment the literature has proxied badly, where a better measure exists or can be built. These questions are undervalued by students and valued by referees, because a measurement improvement often overturns something.

**Contradiction.** Two credible papers disagree. The question is designed to explain why: different populations, different periods, different estimators, different definitions of the treatment. Reconciliation questions publish well and are unusually safe, because the result is interesting in either direction.

**Negative space.** Something the field assumes and nobody has tested. Rare, high value, and the generator most likely to produce a candidate that is either excellent or already done, so it demands the most careful search.

## The kill round

Score each candidate in this order, one line each, in writing.

**Identification.** What is the source of variation, and what is the biggest threat to it? Name both. If the only honest answer is that treatment is observational and selection is on unobservables, the candidate is descriptive. Descriptive questions can be good chapters in some fields and are not defensible as the empirical core of a quantitative thesis in most. Kill or reclassify.

**Feasibility.** Can the data measure treatment and outcome at the required unit and frequency, over the required period, for enough units? Three specific killers to check every time: the outcome is not observed after treatment for long enough; the treatment variable is only populated in some years; and the number of treated clusters is small enough that inference collapses regardless of how many observations there are.

**Novelty.** Name the closest paper, with a verified reference, and say in one sentence what is different here. "Nobody has done this" without a search is not an answer, and it is usually false. The useful outcome of this line is often not a kill but a sharpening: the closest paper reveals what the question has to add.

**Interest.** Who changes their mind on either answer? Name the reader, the policy, or the debate.

**Fit.** Does it match the constraints on time, method and venue, and does it match this person? A question the student finds boring is a real risk over eighteen months and should be scored honestly rather than treated as unprofessional to mention.

## Estimating feasibility before you have the data

Ten minutes, on paper, for each survivor.

Count the units that will actually identify the effect, which is almost never the number of rows. For a difference-in-differences design it is the number of treated clusters; for a discontinuity it is the number of observations inside a plausible bandwidth; for an instrument it is the variation the instrument actually moves. Students routinely quote a sample of 400,000 observations for a design identified by eleven treated regions.

Then take the effect size the closest paper reports, halve it on the assumption that published effects are optimistic, and ask whether the design could detect that. If not, the question is dead. Say so plainly, and say which change would revive it: more periods, a lower level of aggregation, a different outcome measured with less noise, or a different question.

Where restricted data is involved, add the calendar. An application with a six month queue turns a three month check into a nine month one, and that alone can reorder the ranking of otherwise equal candidates.

## Worked example

**Situation.** A second-year doctoral student at Ashcombe University, working with Professor Ruth Delacroix, arrived at a supervision meeting with the sentence "I want to work on minimum wages and small firms". She had six weeks until a departmental progress review that required a written chapter plan, roughly twenty months to submission, one prior paper using panel fixed effects, and access to a national employer register covering 2014 to 2023 with firm-level payroll, headcount and sector, about 412,000 firm-year observations. She had not searched the recent literature because she had been reading a textbook chapter on the topic.

**Task.** Produce three defensible candidate questions with identification and feasibility already tested, and a recommendation for the progress review, in one working week.

**Action.** The search stage took two hours and changed the shape of the whole exercise. The minimum wage employment literature in that setting was saturated: four papers in the last five years using the same register, two of them by a group at another department with better data access. Any candidate from the topic seed would be competing with them from behind.

The pivot was to re-read the seed as a dataset seed rather than a topic seed. The question became what variation the employer register could see that nobody had used. Running the policy variation generator against the register's own institutional history produced the useful candidate: a training subsidy for firms below fifty employees, introduced across fourteen regions on staggered dates between 2017 and 2020, visible in the register because the subsidy conditioned on a payroll code.

Twelve candidates were generated in total across the seven generators. Eight died in the kill round.

Three died on identification. One asked whether minimum wage increases changed firm survival, where the variation was national and simultaneous and there was no comparison group. Two asked about mechanisms that the register could not distinguish because it recorded headcount but not hours.

Four died on feasibility, and the fourth is the instructive one. It survived the first four checks and died on the minimum detectable effect. The candidate proposed a heterogeneity question comparing subsidy effects across firms above and below a productivity threshold, but productivity had to be constructed from a revenue field populated only from 2019 onwards, leaving one pre-period year for firms treated in 2019 and none for those treated in 2020.

One died on novelty, matching a 2022 paper almost exactly once the closest reference was actually located.

The wrong turn worth recording: the first pass at the surviving subsidy question specified regional clusters as the unit of variation and quoted the full 412,000 observations as the sample. Fourteen treated regions is not enough clusters for conventional inference, and building the design around region-level variation would have produced a chapter whose standard errors nobody would believe. The repair was to move identification to the firm level using the fifty employee eligibility threshold, which the register measures directly, turning a weak staggered difference-in-differences into a discontinuity with a difference-in-differences component around the rollout dates. That change came out of step 6, the ten minute feasibility calculation, and would not have come out of any amount of further reading.

Three survivors went into the memo. The recommendation named one check to do first: count the firms within five employees of the threshold in each region and year, because if that count was in the hundreds the discontinuity was not viable and the whole ranking changed.

**Result.** The threshold count came back at roughly 9,400 firms inside a ten employee bandwidth across the sample, which was ample. The chapter plan went to the progress review with the design already stress-tested and passed without a revision request, which was not the norm in that department. Professor Delacroix's comment was that the memo had done the work the review was supposed to prompt.

The total cost was about eleven hours across five days. The original topic, minimum wages and small firms, was abandoned entirely, and the student's estimate beforehand was that ideation would take an afternoon. It always is.

### A second scenario, where it goes differently

A first-year student in education policy at the same institution, working qualitatively, with no dataset and no reform: the seed was an interest in how school leaders interpret a new national inspection framework. Here the method runs in a different order and two of the generators are useless.

Identification is not the binding constraint, because the question is interpretive rather than causal, and applying the identification kill would destroy every candidate wrongly. Feasibility becomes the first kill instead, and it is dominated by access: how many schools will agree, whether the inspectorate will permit interviews with its own inspectors, and whether the ethics approval for interviewing staff about a body that regulates them will clear. Two candidates died on access alone, one because it required interviewing inspectors during a live inspection cycle and the regulator's own guidance forbade it.

The generators that worked were mechanism, contradiction and negative space. Setting transfer, policy variation and measurement produced nothing usable.

The interest test also changes. For an interpretive question the honest version is not "who changes their mind on either answer" but "what would this let a reader see that they currently cannot", and a candidate that cannot answer that is descriptive in the bad sense whatever its method.

What did not change: the requirement that each candidate be one sentence, the requirement to name the closest existing work, the ten minute feasibility estimate in the form of a recruitment count, and the written record of what was killed and why.

## Output

A memo of one to two pages. The survivors first, because that is what the reader needs.

```
RESEARCH QUESTION MEMO
Seed:            [what was brought, and its type: dataset / policy / topic / contradiction]
Constraints:     [time to submission, target venue, methods available, data access status]
Searched:        [subtopics covered, date of search, number of references reviewed]
Generated:       [n candidates]   Killed: [n]   Survivors: [n]
```

Then the survivors, ranked:

| Rank | Question, one sentence | Source of variation | Closest paper (verified, with DOI) | What is different here | Data required | Rough N that identifies | Main risk |

Then the killed list, which is the part people are tempted to omit and should not:

| Candidate | Killed on | Reason, one line |

Then the recommendation, in this shape:

```
RECOMMENDATION
Take forward:    [question 1]
Check first:     [the single cheapest falsifying check, and how long it takes]
If it fails:     [which survivor is next, and what changes]
Do not revisit:  [any killed candidate the person is likely to bring back, and why it stays dead]
```

## Failure modes

**Generating before searching.** Recognise it because the candidates are all things the person already believed. The literature exists to constrain generation, not to be cited afterwards. Search first, for a bounded time.

**Searching instead of generating.** The opposite failure and the more common one in doctoral work, because reading feels safe. Recognise it when week three arrives with a folder of papers and no candidate sentences. Fix by imposing the ninety minute limit and generating on what you have.

**One generator, eight candidates.** Recognise it because a single objection kills the whole list. Fix by working the generators in order, explicitly, and forcing at least one candidate from each of four of them.

**Scoring interest before identification.** The candidate everyone likes becomes the candidate everyone defends. Fix by scoring in the prescribed order and stopping at the first failure, before anyone has had time to become attached.

**The method in search of a question.** Recognise it when the sentence names an estimator. Ask what the estimator would be estimating and for whom, and if that cannot be answered, the candidate does not exist yet.

**Sample size quoted at the wrong level.** Recognise it when a design identified by a handful of clusters is defended with a six figure observation count. Fix by counting the units that actually vary.

**Novelty asserted rather than checked.** Recognise the phrase "nobody has looked at this". Fix by requiring a named closest paper with a resolving identifier for every candidate, including the ones being killed.

**Kindness at the wrong moment.** Recognise it when a candidate survives because the person is excited about it. The cost of that kindness arrives in month eleven. Name the defect, explain why it sinks the project, then help repair it; the sequence matters and the third part is what makes the first two acceptable.

**No written kill list.** Recognise it six months later when a dead candidate returns as a fresh idea. Fix by spending five minutes on the table.

## Edge cases

**No data and no prospect of any.** The question set is limited to what can be collected, and collection cost becomes the first kill. Do not generate candidates that assume administrative access which has not been applied for. Where primary collection is the only route, hand the surviving question to `survey-and-instrument-design` before the design is fixed, because instrument feasibility will change the question.

**A supervisor's imposed question.** Do not generate alternatives unasked; that is a political act as much as a methodological one. Run the kill round on the imposed question and two near variants, and where it fails, present the failure as a design problem with proposed repairs rather than as a rejection.

**The dataset is excellent and the question is absent.** The most productive case and the one where the policy variation generator earns its place. Read the institutional history of the data before generating: the eligibility rules, the reporting thresholds, the years definitions changed. Discontinuities and rollouts hide in administrative documentation and almost never in the codebook.

**A field with no causal tradition.** Where the discipline does not expect identification, the first kill criterion is replaced by warrant: what makes the evidence adequate to the claim in this tradition, whether that is theoretical saturation, case selection logic, or triangulation. The structure survives; the content of the first check does not.

**The honest answer is that the seed yields nothing.** Say so, in one clear sentence, and propose the nearest adjacent seed rather than ranking three weak candidates to be helpful. Ranking bad options as though they were viable is the most damaging thing this process can do, because it launders a dead end as a decision.

**Three chapters needed from one dataset.** Generate for coherence as well as quality: three questions that share data cleaning and a literature but are separable enough to publish independently. Score an extra line for overlap, and kill any pair that would end up as one paper split in two, which examiners identify immediately.

**The question is already answered but answered badly.** This is viable, and it is a specific kind of contribution that must be stated as such: better data, better identification, or a better measure. State which of the three, in the memo, in those words. A replication framed as a novel question fails; a replication framed as a replication with a stated improvement is publishable.

## Quality bar

- Every candidate is one sentence naming population, variation and outcome.
- Every candidate that survived has a named closest paper with a resolving identifier, and one sentence saying what is different.
- Every survivor has a named source of variation and a named biggest threat to it.
- Every survivor has a rough count of the units that actually identify the effect, at the correct level, and a minimum detectable effect compared against a published effect size.
- Candidates were killed in the prescribed order and the round stopped at first failure.
- The killed list exists in writing with one reason per candidate.
- The recommendation names one cheap check to run first and what happens if it fails.
- Nothing in the memo is cited from memory.

## Adapting this to your context

Written for a quantitative economics doctorate with administrative data, which is why identification is the first kill. Having a first kill is the method; which one depends on your field.

- **What gets killed first.** In interpretive and qualitative work the binding constraint is access, as the second scenario shows. In psychology and health it is often whether a validated instrument exists: a question needing a measure nobody has built is a measurement project in disguise.
- **The search vocabulary.** Search the terms your field indexes on: MeSH in PubMed, PsycINFO thesaurus terms, ERIC descriptors. Search the registries too, PROSPERO, ClinicalTrials.gov, OSF, the AEA RCT Registry, since a question already claimed and unpublished is where "nobody has done this" goes wrong.
- **The feasibility arithmetic.** Minimum detectable effect is the economics phrasing. Elsewhere run a power analysis in G*Power or `simr` against the smallest effect size of interest, remembering that in a multilevel design the binding number is groups, not participants.
- **The generators.** Policy variation is dead where there are no reforms to see. Add one the list lacks: a direct replication of a widely cited result nobody has replicated, publishable in psychology and increasingly elsewhere.
- **What not to change.** One sentence per candidate with population, variation and outcome named, the hard constraint scored before interest, and the killed list written down.

## Related skills

`literature-verification` supplies the citation standard used in the search and novelty steps, and every reference in the memo is verified to it. `research-design` takes the recommended question and turns it into hypotheses, an estimating equation, a sample definition and an exhibit list; this skill deliberately stops short of that. `identification-defense` is where the named threat gets a proper answer once the design exists. `theoretical-framework-review` builds the mechanism argument behind the chosen question. `survey-and-instrument-design` takes over when the data has to be collected rather than found. `research-proposal-and-grant` reuses the memo's survivors and kill list as the core of a proposal's rationale. `thesis-advisor` handles the wider question of whether the chapter portfolio hangs together, which this skill only touches when three questions are needed from one dataset.

