Theoretical Framework Review
A referee asks one question of a theory section and it is not whether the authors have read widely. It is why this effect should exist at all. A results section shows what happened; the framework has to make the reader expect it before they see it, and if it does not, the paper reads as a finding in search of an explanation, which is the single most common reason an otherwise competent empirical paper is rejected on grounds nobody can quite articulate.
The failure has a shape. Four pages naming theories and authors, each paragraph accurate, none of them connected to a hypothesis. Then a list of hypotheses that appear from nowhere, usually numbered, usually stated with confidence, with no argument above them that would have generated them. Readers experience this as a section that is hard to remember, and referees write "the theoretical grounding is weak", which is a diagnosis nobody knows how to act on.
The cost is not only the rejection. A framework that cannot predict the sign of the main coefficient cannot tell the author what a surprising result would be, so when the estimate comes in negative the paper acquires a post hoc explanation, and that explanation was available for the positive result too. The framework was supposed to be what stopped that.
Two rules do most of the work. Every hypothesis is the last sentence of the paragraph that argued for it. Every mechanism is written in plain words before any theory is named.
When to use this, and when not to
Use it when there is a draft framework and someone doubts it, when a referee or supervisor has said the grounding is weak or the hypotheses do not follow, when a chapter has an empirical strategy and no theory section yet, when hypotheses exist but nothing above them explains where they came from, and when a paper has to justify why an effect is expected rather than only report that it was found.
Use it also on inherited text: a coauthor's section, a chapter written a year ago, a proposal being recycled. Frameworks decay in a specific way, which is that the field's account of the mechanism moves on while the citations stay where they were.
Do not use it to search a literature systematically and count what was screened; that is systematic-review-protocol. Do not use it to verify citations as a standalone task, which is literature-verification, although its standard applies here in full and this skill will not proceed without it. Do not use it to write the introduction, which has a different job: the introduction sells the question and states the contribution, and introduction-writer covers it. Do not use it to choose the research question or the identification strategy, which are research-question-ideation and research-design; a framework built before the question is settled will be rebuilt.
Do not use it to add citations to an argument that is already sound in order to make it look more scholarly. That produces the padding this skill exists to remove.
What you need before starting
The hypotheses, or the research question if hypotheses do not exist yet. The framework's structure is determined by what it has to ground. Missing: extract candidate hypotheses from the paper's results or its abstract and confirm them with the author before writing, because a framework grounding the wrong claims is worse than none.
The empirical setting, precisely. Country, sector, period, population, unit, and the institutional detail that makes the mechanism operate or fail. Missing: ask. Mechanisms are setting-dependent, and a framework written without the setting will cite evidence from contexts where the mechanism works differently, which is the error a specialist referee sees first.
The main result, where it exists. Not to reverse-engineer the theory, but to know whether the framework will have to accommodate a surprise. Missing: build the framework on the predicted sign and note that it will need rereading once results exist.
Access to at least one live lookup. A bibliographic database, a DOI resolver, a publisher page, an academic search connector such as Consensus where it is available, or the papers themselves. An academic search connector is useful for one specific thing, which is finding papers by the claim rather than by the keyword, so it is worth reaching for when the mechanism is known and the vocabulary is not. Missing: do not emit formatted citations at all. Write visible placeholders naming the claim that needs a source, and list them so they cannot be missed. Nothing in this skill requires a paid tool; a free bibliographic index and the papers themselves are sufficient, and slower.
The target venue and its conventions. How long theory sections run there, whether hypotheses are stated formally, whether first person is used, which citation style. Missing: match the two nearest published papers in the target and say you have done so.
The word budget. Missing: assume the framework is between fifteen and twenty-five percent of the paper, and say which you assumed. A theory section that runs long is the commonest place a paper exceeds its limit and the commonest place an author cuts the wrong thing.
The method
Decide the mode. Review mode when a draft exists: diagnose before rewriting, because rewriting first destroys the evidence of what was wrong and usually reproduces the same defect in better prose. Build mode when there are hypotheses and no framework: construct from the question outward. A draft that grounds none of its hypotheses is a build, not a review, whatever it looks like on the page.
List every hypothesis and every claim the paper makes that needs theoretical warrant. Number them. This list is the specification for the section, and everything the framework contains has to serve an item on it.
Write each mechanism in plain words, before naming any theory. One or two sentences, with a subject, a verb and a direction. "Devolving hiring authority to principals lets them select for fit with the school, which raises match quality and lowers the probability a teacher leaves within two years" is a mechanism. "Principal-agent theory" is a label, and a label cannot be tested, extended or contradicted. The judgement call: if the mechanism cannot be written in plain words, either it is not understood yet or the hypothesis has no mechanism, and the second case is common and is a finding, not a writing problem.
Name the theory that the mechanism belongs to, second. The order is not stylistic. A mechanism written first constrains which theory is relevant; a theory named first invites the mechanism to be bent to fit it, which is how a paper ends up applying a famous framework to a setting where its assumptions do not hold.
Search per mechanism, as a claim rather than as a topic. Query the mechanism as a statement to be tested, vary the vocabulary because fields name the same construct differently, and read abstracts for the direction of evidence, the setting, the sample size and the estimate. Then find the foundational theoretical sources separately, because search tools that rank on empirical evidence will not surface them. Verify everything to the standard in
literature-verification: nothing cited from memory, nothing cited without a resolving identifier, and nothing cited for a claim whose support was not read.Build the evidence matrix before writing any prose. One row per source: authors and year, theory or mechanism, setting and period, data, method, finding with direction and magnitude, which hypothesis it grounds, and whether it supports, qualifies or contradicts the expected effect. The matrix is what makes the writing fast and the omissions visible, and a section written without one reliably becomes a paragraph per paper rather than a paragraph per idea.
Sort the matrix by mechanism and look at what is thin. A hypothesis with one supporting row is not grounded, it is asserted with a footnote. A hypothesis with a contradicting row closest to your setting is the most important thing on the page and determines what step 9 has to do.
Write each grounding block in the same four-move order. The general theory and its core prediction, with the foundational sources. The mechanism in this setting, with the empirical evidence closest to the setting and its numbers. The contested ground, both sides named. Then the hypothesis, stated formally, as the last sentence of the block. The rule that makes the section coherent: a reader should be able to stop after the third move and predict the fourth.
Engage the contradicting evidence rather than omitting it. Where the closest study to your setting found the opposite sign, the framework says so and does one of three things: explains why this setting differs in a way that predicts a different result, revises the hypothesis to match, or reframes the paper as a reconciliation. All three are respectable. Omission is not, and a specialist referee will know the paper and will assume the omission was deliberate.
Close by restating the hypotheses together, each traceable upward. No summary of the summary, and no new material. The closing list exists so the reader carries the predictions into the results section.
Run the sign test. Give the framework to someone who has not seen the results and ask them to predict the sign and rough size of the main coefficient. If they cannot, the section has failed regardless of how well written it is, and the repair is almost always in step 3.
Cut every source that appears once and never returns. Read the section and mark each citation used in only one sentence with no consequence. These are the padding, they are visible, and they cost credibility rather than adding it.
The review diagnostic
In review mode, score the draft against these six criteria before rewriting anything, quoting the draft where it fails. Deliver it as a table so the author can see the pattern rather than a list of complaints.
Traceability. List every hypothesis and the paragraph that grounds it. Then list every paragraph and the hypothesis it feeds. Both directions matter: an ungrounded hypothesis is a gap, and a paragraph feeding no hypothesis is a cut. This single check usually explains most of what a referee meant by "weak grounding".
Mechanism versus name-dropping. For each theory named, does the draft explain how it produces the expected effect in this setting, or only that someone proposed it? The test is the sign test from step 11, applied paragraph by paragraph.
Currency and proximity. Are the sources the ones the field currently cites for this mechanism, including work from the last five years, and are they from settings close enough to matter? A framework grounded entirely in evidence from a different institutional context needs to say why the mechanism transfers, and usually does not.
Contested ground. Where the literature disagrees, does the draft name both sides or present one as settled? Check specifically for the nearest study to the setting, because that is the one a referee will have read.
Verification. Does every citation resolve to a real record that says what the draft claims? Check each one. Fabricated and misattributed references cluster in theory sections because that is where citation density is highest and where claims are most general, and they are the most damaging single defect a draft can carry.
Structure and budget. Does the section move from general theory to specific setting to hypotheses, or does it wander? Is each subsection there because a hypothesis needs it? Is it the length the venue expects?
Writing standards
Impersonal academic register unless the field or the target journal expects first person. Match the two nearest published papers in the target rather than a general idea of academic style.
Every empirical or theoretical claim carries its citation in the sentence that makes the claim, not at the end of the paragraph where it becomes ambiguous which sentence it supports. This matters more in a framework than anywhere else in a paper, because framework paragraphs make several claims each.
Be dense in specifics. Settings, years, sample sizes, and the effect sizes the cited work reports. "Prior work finds a positive effect" is weaker than "two studies in comparable systems report increases of about four percentage points, while the one study in a setting with centralised pay finds no effect", and the second sentence does work the first does not: it tells the reader what magnitude to expect and it sets up the contrast that motivates the paper.
A framework that could preface any paper on the topic has failed. Test it by asking whether a competitor working on the same question in a different country could use this section unchanged. If yes, the setting has not entered the argument.
Build the reference list in the same session, one verified entry per citation as it is used, so text and bibliography cannot diverge. Keep the identifier in the entry even when the style does not print it.
Worked example
Situation. Rafael Duarte, a fourth-year doctoral student at Brentmoor University, had a paper returned with a major revision. The empirical work was on whether devolving teacher hiring authority from a district office to school principals raised two-year teacher retention, using a staggered reform across 212 schools with about 18,000 teacher-year observations. Two referees liked the identification. The third wrote that the theoretical grounding was weak and that the hypotheses did not follow from the preceding discussion, and recommended rejection. The framework ran 2,800 words and cited 41 sources. The revision was due in eight weeks.
Task. Rebuild the framework so the three hypotheses were each grounded, engage whatever the third referee had actually noticed, and lose no more than the word budget allowed, which was 2,200 words after cuts elsewhere.
Action. The diagnostic came first and took a day. Its findings were not what the author expected.
Traceability was the core problem. Of three hypotheses, only the first had a paragraph that argued for it. The second, on heterogeneity by school size, appeared with no argument above it at all. The third, on effects operating through teacher workload, was grounded in a paragraph that argued for a different mechanism entirely, namely autonomy and job satisfaction, and did not mention workload.
In the other direction, twelve of the 41 sources appeared once, in the first two pages, and never returned. They were a survey of the general literature on school governance, and none of them fed a hypothesis.
Verification found two problems in 41 citations. One paper was attributed to a journal it had not appeared in, which was a real paper misfiled. One was cited for a specific figure, a retention improvement of six percentage points, and the paper reported a change in intention to leave measured on a scale, not a retention rate, so the number in the draft appeared nowhere in the source. That second one had been in the author's notes for two years.
Currency and proximity produced the finding that explained the referee report. The framework cited five empirical studies, four from systems with school-level pay setting. The closest study to the author's setting, a system with centralised pay scales and a comparable devolution reform, reported a small negative effect on retention, of about 1.8 percentage points, and was not cited anywhere in the draft. The author had read it and had not connected it, because it did not fit.
The wrong turn came next and cost about four days. The first repair attempt was to find more citations for hypotheses two and three, on the theory that the referee meant the grounding was thin. The section grew to 3,400 words, the additional sources were real and relevant, and the second reader in the department said it was worse. Adding evidence does not ground a hypothesis; an argument does, and there was still no argument connecting school size to the effect.
The actual repair started at step 3. Writing each mechanism in plain words, before any theory, took two hours and resolved the problem by elimination. Hypothesis one had a clean mechanism: devolved hiring lets principals select for fit, raising match quality and lowering exit. Hypothesis two, on school size, had a mechanism once it was written down, which was that principals in larger schools have more applicants and therefore more scope to select, so the effect should be larger where the applicant pool is deeper. That was writable and testable and the framework had simply never said it. Hypothesis three could not be written at all. Every attempt produced a sentence about workload that did not follow from devolved hiring, and after the third attempt the honest conclusion was that there was no mechanism. Hypothesis three was cut, along with the two tables that tested it.
The contested evidence became the spine of the section rather than an omission. The rewritten framework set the transaction-cost account of hiring devolution against a constraint-based account in which devolution without control over pay gives principals authority they cannot use, and named the centralised-pay study as the evidence for the second. The paper's own setting had partial pay flexibility, which is why a positive effect was still predicted, and saying so gave the framework something to argue rather than something to assert.
Result. The rewritten section ran 1,940 words with 23 sources, of which 19 had survived from the original 41. Three hypotheses became two. The sign test was run on a colleague from another subfield who had not seen the results, and she predicted a positive effect of a few percentage points, larger in bigger schools, which was correct on both counts.
The revision was accepted after one further round. The third referee's second report said the theoretical development was now clear and made no mention of the hypothesis that had been cut, which is the usual pattern: referees notice that something does not follow long before they can name what.
The misquoted six percentage point figure had already appeared in two conference presentations and in a draft of another chapter. Correcting it in all three took an afternoon and was the least pleasant part of the exercise.
A second scenario, where it goes differently
A second-year student at the same university had a question, a dataset and no draft: whether a simplified business registration regime raised formal firm entry, using a reform that reduced registration from an average of 31 days to 5 in a staggered rollout across 94 municipalities. Build mode, and the framework's job turns out to be different.
Step 3 produced two mechanisms rather than one, and they predicted different signs. A transaction-cost account says registration cost is the binding constraint on formalisation, so reducing it raises entry. A second account says the binding constraint is expected enforcement and ongoing tax liability rather than the one-off registration cost, so reducing registration time changes little and may change nothing at all. Both accounts are well established, both have empirical support in different settings, and the field has not settled between them.
That changes what the framework is for. Rather than arguing towards one prediction, it sets up a discriminating test. The four-move structure still holds, but the third move, contested ground, becomes the longest rather than the shortest, and the fourth move produces a pair of hypotheses whose joint pattern separates the accounts: entry rises overall under the transaction-cost account, and under the enforcement account any rise is concentrated among firms with low expected tax liability, with no effect among firms above the threshold where the ongoing obligations bite.
Two consequences follow that would not have arisen in review mode. First, the framework changed the empirical design, because the discriminating test required a firm-level tax liability proxy that had not been in the original data request, and it was added before the request was submitted. This is the strongest argument for building the framework before the data work rather than after. Second, the paper cannot produce an uninformative result: both accounts predict something specific and the data can tell them apart, which is what makes a null publishable here.
What did not change: mechanisms written in plain words before theories were named, an evidence matrix before prose, contested evidence engaged rather than omitted, every hypothesis as the last sentence of the block that argued for it, and the sign test at the end. The sign test in this case asked the reader to predict the pattern rather than a single sign, which is the correct adaptation and not a weakening.
Output
In review mode, the diagnostic first, then the rewrite.
| # | Criterion | Finding | Evidence from the draft | Fix | | 1 | Traceability | | quoted line or paragraph number | | | 2 | Mechanism versus naming | | | | | 3 | Currency and proximity | | | | | 4 | Contested ground | | | | | 5 | Verification | | | | | 6 | Structure and budget | | | |
With two summary lines beneath it: hypotheses grounded, out of total; and sources retained, out of total.
In both modes, the evidence matrix:
| Authors (year) | Theory or mechanism | Setting and period | Data | Method | Finding, direction and size | Grounds which hypothesis | Supports, qualifies or contradicts | DOI |
And the section itself, in this structure:
THEORETICAL FRAMEWORK
[General theory and its core prediction. Foundational sources.]
For each hypothesis:
Move 1: the theoretical prediction, in general terms
Move 2: the mechanism in this setting, with the closest empirical evidence and its numbers
Move 3: the contested ground, both sides named and cited
Move 4: H1, stated formally, as the last sentence of the block
[Closing restatement of all hypotheses together, each traceable upward. No new material.]
Followed by the verified reference list, built in the same session, in the target's style, with identifiers retained.
Failure modes
Hypotheses that appear from nowhere. Recognise it by reading the sentence immediately before each hypothesis and asking whether it argues for it. Fix at step 3, not by adding citations.
Theory named instead of mechanism explained. Recognise it when a paragraph could be summarised as "X proposed the theory of Y". Fix by writing the mechanism in plain words and rebuilding the paragraph around it, with the theory named second.
Padding with famous tangential work. Recognise it as a source cited once, early, that never returns. Cut it. This is the cheapest improvement available and authors resist it because the sources are real and relevant, which they are, and irrelevant to the hypotheses, which is what matters.
The contradicting study that is quietly absent. Recognise it by finding the closest study to your setting and checking whether it is cited. Fix by engaging it in move 3 and adjusting the prediction, the framing, or the hypothesis.
Adding evidence in place of argument. Recognise it when the section grows and the traceability count does not improve. Fix by returning to step 3; more rows in the matrix cannot ground a hypothesis that has no mechanism.
Effect sizes recalled rather than read. Recognise it when a number in the draft is round and its source is not open on the desk. Every number attributed to a paper must have been seen in that paper. This is the failure that propagates, because the next author cites you.
Framework written after the results. Recognise it when every hypothesis is confirmed and each one is stated in exactly the direction found. Fix by writing what the framework would have predicted before, and by labelling anything that cannot survive that as a post hoc interpretation belonging in the discussion.
The interchangeable framework. Recognise it because a researcher in another country could use the section unchanged. Fix by putting the institutional detail of the setting into move 2, where it belongs.
Length by accretion. Recognise it in a section that has grown across drafts and lost none. Fix by rebuilding from the matrix rather than editing the prose, which is faster and produces a shorter result.
Edge cases
No hypotheses, and none appropriate. In exploratory or interpretive work the framework grounds a set of expectations or sensitising concepts rather than testable predictions. The structure holds and move 4 changes: the block ends with the concept the analysis will use and what it directs attention to. The sign test is replaced by asking a reader what the analysis will look for.
A field with a single dominant theory. The risk is that the framework becomes an application exercise. The repair is to state what the theory would have to be wrong about for the effect to be absent, which converts a restatement into a test.
Two theories predicting opposite signs. Treat it as an opportunity rather than a problem. Build the framework as a discriminating test and derive a pattern of predictions that separates the accounts, as in the second scenario above. Check early whether the data can support the discriminating test, because it often requires a variable nobody planned for.
The result contradicts the framework. Do not rewrite the framework to predict what was found. Keep the prediction, report the contradiction, and use the discussion to explain it. A framework revised to match the result is undetectable to the reader and worthless to the author, and it is the mechanism by which fields accumulate results that never replicate.
Theory from another discipline. Legitimate and frequently valuable, and it requires the assumptions to be stated rather than imported. Say what the theory assumes about the actors, and say whether those assumptions hold in this setting. Referees from the borrowing discipline are unforgiving about frameworks that transplant a construct without its conditions.
Non-English literature central to the setting. Cite it in the original, verify it in the original, and give a translated title where the style requires. A framework on a national setting that cites only English-language work will be read as not knowing the field, and often correctly.
A very short venue. Where the target allows eight hundred words of theory, the four-move structure compresses rather than disappears: one sentence of general theory, two of setting-specific mechanism with the closest evidence, one naming the contested alternative, and the hypothesis. The move that must not be dropped is the third.
No lookup available at all. Do not emit formatted citations. Write the framework's argument with visible placeholders naming the claim that requires a source, list those claims separately, and hand the section on unverified with that status stated.
Quality bar
- Every hypothesis is the last sentence of a block that argued for it, and every block feeds a hypothesis.
- A reader who has seen only the framework can predict the sign, and roughly the size, of the main coefficient.
- Every mechanism is stated in plain words with a direction, and named theories appear after the mechanism rather than before.
- The study closest to this setting is cited, including when it contradicts the expected result, and the framework says what follows from that.
- Every citation was verified against a live record in this session, and every number attributed to a source was seen in that source.
- No source is cited once and never used again.
- The setting's institutional detail appears in the argument, so the section could not preface a paper on the same topic elsewhere.
- The section is within the target venue's expected length.
Adapting this to your context
Built for an economics paper where theory is fifteen to twenty-five percent of the words and the payoff is a predicted sign. The two rules are general; the proportion is not.
- The word budget. Management, psychology and education journals commonly give theory and hypothesis development a third or more of the paper, with hypotheses numbered H1a and H1b. Take the proportion from two recent papers in your target.
- The sign test. It assumes a directional quantitative prediction. For an interpretive study, ask instead what observation would disconfirm the account. Where the design specifies mediation or moderation, argue each path separately: a hypothesis about a total effect does not ground a claim about a mediator.
- Inductive designs. A framework fixed before the data contradicts the design. Write the sensitising concepts instead, say what they do and exclude, and keep the requirement that a reader can predict what counts as evidence.
- Searching per mechanism. Query the mechanism as a claim, in your field's controlled vocabulary: PsycINFO thesaurus terms, MeSH, ERIC descriptors. Canonical theories differ by field too, so name the one your reviewers expect.
- What not to change. Write the mechanism in plain words before naming any theory, make every hypothesis the last sentence of the paragraph that argued for it, and engage the study closest to your setting when it disagrees.
Related skills
literature-verification supplies the citation standard this skill applies without exception, and runs as its own pass when the task is checking references rather than building an argument. research-question-ideation and research-design fix the question and the identification the framework has to ground; a framework built before them will be rebuilt. introduction-writer uses the framework's contribution claim but does a different job, which is selling the question. systematic-review-protocol is the right skill when coverage must be reproducible and counted rather than argued. econometric-model-writer translates the hypotheses into the estimating equation, and the mapping between them should be one to one. discussion-and-conclusion is where a result that contradicts the framework is explained, rather than in the framework itself. references-and-bibliography handles style conversion once the sources are verified. peer-review-simulator tests the finished section against the referee this skill is written to anticipate.