# Refereeing For A Journal

> Refereeing for a Journal

- Skill: `ingridleiria/refereeing-for-a-journal` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ingridleiria/refereeing-for-a-journal`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ingridleiria/refereeing-for-a-journal/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Research & Search
- Author: ingridleiria (https://skillmd.com/u/ingridleiria)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ingridleiria/refereeing-for-a-journal

---


# Refereeing for a Journal

The report you write is read by three people with different needs. An editor deciding whether to invest another round in the paper, who needs your judgement stated rather than inferred from the length of your list. An author revising, who needs to know what is fatal, what is fixable, and what is optional, and who cannot tell those apart from a numbered list of twenty-eight points. And frequently a second referee, whose own report the editor is calibrating against yours.

Two failures dominate. The first is vagueness: "the identification is unconvincing" and "the contribution is unclear" tell the author nothing they can act on, cost them a full revision cycle discovering what you meant, and leave the editor unable to tell whether the paper is fixable. The second is scope creep: a list of individually reasonable additional analyses which together amount to eighteen months of work, requested by a referee who has not asked what any of them would change. Scope creep is the more damaging of the two because it looks like rigour, and editors, who are also under pressure, frequently pass it through.

There is a third failure that is rarer and worse. A report written to demonstrate the referee's own cleverness, or in irritation, addressed to the authors rather than to the work. Anonymity is a procedural device, not a licence, and the field is smaller than it feels.

## When to use this, and when not to

Use it from the moment a review invitation arrives, because the first decision, whether to accept at all, is part of the job and is the one most often made badly. Use it while reading, to impose the order that reaches a recommendation fastest. Use it while drafting, and again before submitting the report, for the calibration and tone passes.

Use it also for adjacent refereeing: a conference programme committee review, a grant panel review, an editorial board's assessment, a review for a book series. The structure holds; only the criteria change.

Do not use it to review your own manuscript before submission. That is `peer-review-simulator`, and the difference matters in both directions: that skill should be exhaustive and this one must not be, because exhaustiveness on somebody else's paper is how scope creep happens, and a real referee has duties about conflicts and proportionality that do not apply to attacking your own work.

Do not use it to read a thesis chapter as an examiner; that is `thesis-chapter-review`, where the question is whether a candidate has demonstrated independent research capability rather than whether an article is publishable. Do not use it to answer referees on your own paper; that is `response-to-reviewers`. Do not use it to verify a manuscript's numbers, which would require the authors' data and is not a referee's job; if you suspect an error, say what you suspect and why, and leave the verification to the data editor.

## What you need before starting

**The invitation, with its deadline and the journal's referee guidelines.** Journals differ on whether they want a recommendation in the report, how they handle confidential comments, and whether they use structured forms. Missing: fetch the guidelines from the journal's site, or ask the editor. Guessing produces a report in the wrong shape that the editor has to reformat.

**An honest estimate of your own availability in the window offered.** Not your intention, your availability. Missing: assume the review will take between four and eight hours of concentrated time for an empirical paper, and decide against that number rather than against a vague sense of willingness.

**A conflict check performed before reading the paper.** Once you have read it, you cannot unread it, and a conflict discovered on page twelve is awkward for everyone. Missing: run the check listed below on the title, abstract and author list, which is all you need.

**The manuscript in full, with appendices.** A review of a manuscript whose appendix was not supplied is a review of a different paper. Missing: ask the editor for the missing material and pause the clock.

**The journal's standard, in your own calibration.** A design adequate for this journal may not be adequate for the one you last published in. Missing: read two or three recent articles in the journal on adjacent topics. Ten minutes, and it prevents the most common miscalibration, which is refereeing every paper to the standard of the best journal you know.

**A view on whether you are the right referee.** Missing: decide it explicitly. Being close but not expert is common and is workable if declared, and declaring it is more useful to the editor than a confident review of something outside your competence.

## The method

1. **Decide on the invitation within two or three days, and decline promptly rather than slowly.** A decline on day two costs the editor nothing; a decline on day eighteen costs the paper three weeks. Accept if the paper is inside your competence, you have the hours in the window, and you have no conflict. Where you can judge part of the paper but not all of it, accept and tell the editor precisely which part, in one sentence: this is more useful to them than either a decline or an overreaching review.

2. **Run the conflict check.** A current or recent co-author, a supervisor or supervisee relationship, the same department or institution, a personal relationship, a competing paper of your own in progress on the same question, a financial interest, or a role in the work being evaluated. Where any of these holds, disclose to the editor and let them decide rather than deciding yourself. Where you recognise the authors despite anonymisation, say so to the editor and continue only if you are confident the review would read the same either way. The test is whether you would be comfortable if the author learned of the relationship after publication.

3. **Read in the order that reaches the recommendation fastest: abstract, conclusion, tables, method, introduction.** This order is deliberate and it is not how you read a paper you are learning from. The claims are in the abstract and conclusion; the evidence is in the tables; whether the evidence supports the claims is the only question that determines the recommendation, and it can often be answered in twenty minutes. The introduction goes last because it is the most persuasive part of any paper and reading it first recruits you into the authors' framing.

4. **Read the whole paper once before writing a word.** Points you draft on page six are frequently answered on page nineteen, and a report containing a concern the paper already addresses is the single fastest way to lose an author's trust in the rest of it.

5. **Answer the five questions in order, and stop at the first firm no.** Is the question worth answering? Is it new? Does the design answer it? Do the results support the claims made? Could a reader reproduce this? A paper that fails at question one or two does not need a detailed critique of its standard errors, and providing one is unkind rather than thorough. The earlier the failure, the shorter the report.

6. **Write the summary of the paper in your own words, first, before any criticism.** Three to five sentences: question, data, design, finding, claimed contribution. This is not a formality. It shows the authors whether they were understood, and an author who discovers they were misread has learned something more valuable than most of your points. It also disciplines you: a summary you cannot write is a paper you have not understood, and you should reread rather than review.

7. **Write the assessment paragraph.** Your overall view, stated. Is the question important, is the contribution real, does the evidence support the claim. Editors should not have to infer your position from the number of items in your list, and authors should not have to either.

8. **Write the major points, and apply the change test to each one before it goes in.** For each: what the problem is, why it matters for the claim, and what would resolve it. Then ask the test question: if the authors did this and it came out against them, would the paper's conclusion change? If no, it is not a major point. Move it to minors or delete it. This single test is the strongest available guard against scope creep and it takes ten seconds per item.

9. **Write the minor points, clearly separated and genuinely minor.** Unclear sentences, undefined notation, a table that needs a note, an inconsistency between a number in the text and the table, a missing citation. Do not pad this section to look diligent. A report with four majors and six minors reads as considered; the same report with four majors and thirty-one minors reads as a referee performing effort.

10. **Write the confidential comments to the editor.** Your recommendation, and anything you cannot say to the author: a suspected overlap with published work, a conflict noticed while reading, doubt about your own fitness to judge part of it, a view on fit with the journal. Do not contradict your report here. Editors notice, and a report that is gentle to the author and damning to the editor is the least useful kind, because the author revises against the gentle version and is then rejected on the damning one.

11. **Calibrate the recommendation against the report**, using the rule below, and change one or the other until they match.

12. **Run the tone pass.** Read the report once looking only for sentences about the authors rather than the work, for sarcasm, and for anything you would not write under your own name. Delete or rewrite each one. This takes five minutes and it is the difference between a report that helps and a report that is remembered.

## What a referee is responsible for, and what they are not

You are judging whether the paper's claims are supported by its evidence, and whether the contribution is real relative to what exists. That is the whole remit.

You are not rewriting the paper into the one you would have written. You are not requiring your own work to be cited; if your paper genuinely bears on the argument, describe the finding and let the editor infer, or name it and accept that this is visible. You are not demanding an analysis that constitutes a different paper. You are not enforcing a preference about estimator, software, framing or prose style where the authors' choice is defensible, and defensible is a lower bar than optimal.

You are also not the copy editor. Three or four representative examples of unclear writing, with a note that the problem recurs, is more useful than forty line edits and takes a tenth of the time.

Where a request is worth making but is not a condition of publication, say so explicitly: "This is a suggestion rather than a requirement." Authors cannot tell otherwise, and they will do it, at cost, because it came from a referee.

## Calibrating the recommendation

Match the recommendation to what you actually wrote. Eight major points is not a minor revision. One clarifying question is not a rejection.

**Reject** when the question cannot be answered with this design, or when the contribution is not there. Say which of the two it is, because they carry completely different instructions: a design problem may be fixable with different data or a narrower claim, while an absent contribution means the paper should be abandoned or rethought entirely. Recommend rejection quickly and briefly. A fourteen-page report attached to a rejection reads as contempt and helps nobody.

**Reject and resubmit** when the paper as submitted is not publishable but a defensible paper exists in this material, and the work required exceeds a normal revision cycle. Name the paper you think is there.

**Major revision** when the claims exceed the evidence but the gap can be closed, either by bringing the evidence up or by bringing the claims down. Say which route you think is available; authors default to the first and the second is often cheaper and equally honest.

**Minor revision** when the paper is sound and the remaining issues are presentational or bounded.

**Accept as is** happens rarely and is worth saying plainly when it does. Referees who never recommend acceptance are not maintaining a standard, they are avoiding a judgement.

## Tone

Write about the paper, never about the authors. "The paper does not establish X" rather than "the authors have failed to establish X"; the difference costs nothing and changes how the whole report reads.

Avoid sarcasm completely. It survives anonymity badly, it is remembered longer than any substantive point, and it is the only thing in a report that cannot be defended if it becomes attributable.

Where something is genuinely well done, say what and why. This is not politeness. During a revision an author will change things that were not broken unless they know what to protect, and a report with no positive statement gives them no way to tell.

Assume competence and good faith, including when you are certain the paper is wrong. The most common cause of an unsupported claim is not dishonesty; it is that the authors have been inside the project for three years and can no longer see the step they are skipping.

Write as though your name were on it.

## Worked example

**Situation.** An associate professor, Lena Hofmann, was invited by a mid-tier development journal to review a 38-page manuscript on whether a rural electrification programme raised small-business formation. Household panel data across three survey waves, a difference-in-differences design using programme rollout, headline effect of 2.4 percentage points on business ownership with a standard error of 0.7. The deadline was three weeks. Lena had published on electrification but not on business formation, and she did not know the authors.

**Task.** A report the editor could decide on and the authors could act on, inside about six hours of work.

**Action.** The conflict check ran on the title page alone and cleared. Lena accepted on day two with one sentence to the editor: she could judge the electrification setting and the identification, but not the small-business literature, and the editor should ensure another referee covered that. The editor replied that the second referee was a business formation specialist, which is exactly the outcome that sentence is for.

Reading in order took forty minutes to reach a provisional view. The abstract claimed a causal effect on entrepreneurship. The conclusion extended that to a policy recommendation about programme expansion. Table 4 showed the main estimate, and Table 5 showed heterogeneity by baseline household wealth in which the effect was concentrated entirely in the top wealth tercile, at 6.1 points, with the bottom two terciles indistinguishable from zero. The conclusion's policy recommendation did not mention this.

The full read then found what the fast read had missed: the appendix contained an event study with a clear pre-trend in the wave before rollout, presented without comment.

Three major points went into the report. First, the pre-trend, with what it implied for the parallel trends assumption, and what would resolve it: either a specification robust to differential pre-trends, or a restriction to the subsample where the pre-trend was absent, or an honest downgrade of the causal claim. Second, the heterogeneity result and the policy conclusion, which were inconsistent: the paper's own Table 5 said the programme did not increase business formation among the poorer two-thirds of households, and the conclusion recommended expansion as a poverty policy. Third, clustering at the household level where rollout was assigned at the village level, with 61 villages, which was a straightforward and fixable error.

Seven minor points, including two representative examples of an undefined variable and a note that the same problem occurred elsewhere.

**The wrong turn.** The first draft had six major points. Three of them were additional analyses Lena wanted to see: a triple difference using a neighbouring programme, an alternative outcome definition, and a placebo on a group she thought would be informative. Applying the change test killed all three. If the triple difference came out against the authors, would the conclusion change? Not really; it would add a robustness table. Same for the other two. All three moved to a clearly labelled paragraph headed "Suggestions, not requirements", and the major list dropped to three. The report went from a document demanding roughly a year of work to one demanding about six weeks, without losing anything that determined whether the paper's claims were supported.

Lena's note to herself afterwards: the three cut items were the ones she had found most interesting, which is exactly why they had felt major.

**Result.** Recommendation: major revision, with the confidential comment that the heterogeneity and policy inconsistency was the point she would want to see addressed first, and that she was not competent to judge the business formation literature. The report ran four and a half pages and took five hours including the reading.

The revision came back four months later. The authors had reclustered at the village level, standard error rose to 1.1 and the effect survived; had restricted the sample to villages without a detectable pre-trend, where the estimate was 2.1 with a wider interval; and had rewritten the conclusion around the wealth heterogeneity, which became the paper's most interesting contribution rather than a buried table. Lena recommended acceptance with minor revisions. Two of the three suggestions had been implemented anyway, in the appendix.

### A second scenario, where it goes differently

The same referee, a paper that fails at question two: the design is competent, the execution is careful, and the finding replicates a result established in three prior papers in a slightly different setting, with no argument for why the setting changes anything.

Here almost everything in the method shortens. There is no point listing methodological concerns, because none of them determine the outcome. The report is one page: the summary in the referee's own words, one paragraph of assessment naming the three prior papers and what they established, one paragraph on the specific question the paper does not answer, which is what about this setting would make the result different, and one paragraph saying what would make this a publishable paper, if anything would.

The recommendation is reject, delivered in a page, with the confidential comment stating clearly that this is a contribution problem and not an execution problem, because the editor needs to know the authors did nothing wrong. And the report says so to the authors too, in one sentence, because an author who receives a rejection with no signal about whether the work was competent will assume it was not.

The temptation to resist here is the long report. It is easy to write two thousand words of methodological observation on a competent paper, and it would be worse than useless: it would signal that the problem was fixable when it is not, and the authors would spend three months fixing the wrong thing.

## Output

```
REFEREE REPORT
Manuscript: [title] | Journal: [journal] | Reference: [MS number] | Date:

Summary of the paper
[Three to five sentences in the referee's own words: question, data, design,
finding, claimed contribution.]

Assessment
[One paragraph: is the question important, is the contribution real, does the
evidence support the claim. State the overall view here.]

Major points
1. [What is wrong, with location]
   Why it matters for the claim: [...]
   What would resolve it: [...]
2. ...

Minor points
1. [Location] [Issue]
2. ...

Suggestions, not requirements
[Anything worth doing that is not a condition of publication. Say so plainly.]
```

Separately, the confidential comments:

```
CONFIDENTIAL COMMENTS TO THE EDITOR
Recommendation: [accept / minor revision / major revision / reject and resubmit / reject]
Competence: [which parts of this paper I can and cannot judge]
Conflicts: [any, or none]
The issue I would want addressed first: [one line]
Anything I could not say to the author: [...]
```

## Failure modes

**Slow decline.** Recognise it when the invitation has been sitting for a week while you decide. Decline on day two or accept on day two.

**Scope creep.** Recognise it by counting the major points and asking the change test on each. If a request would not change the conclusion whichever way it came out, it is not a condition of publication.

**The report that is a copy edit.** Recognise it when minors outnumber majors by more than about five to one. Give three examples and a note that the problem recurs.

**Citation self-dealing.** Recognise it when a requested citation is your own and the paper functions without it. If in doubt, describe the finding rather than naming the paper, and let the editor decide.

**Refereeing to the wrong standard.** Recognise it when you cannot say what this journal usually publishes. Read two recent articles first.

**Vague criticism.** Recognise it in any sentence that would apply unchanged to a different paper. Name the assumption, the setting-specific reason to doubt it, and the evidence that would settle it.

**The mismatched recommendation.** Recognise it when the report and the recommendation would surprise each other. Change whichever one is wrong.

**Contradicting yourself in the confidential comments.** Recognise it by reading the two documents back to back as the editor will. If your real view is not in the report, the report is not honest.

**The concern the paper already answers.** Recognise it after the fact, which is too late; prevent it by reading the whole paper before writing.

**Reviewing under irritation.** Recognise it by the presence of any sentence beginning "The authors apparently". Put the report down for a day and run the tone pass.

## Edge cases

**You suspect plagiarism, data fabrication, or undisclosed overlap with a published paper.** Do not accuse in the report. State the observation and the evidence in the confidential comments, factually, with the source you compared against, and let the editor use the journal's process. Referees are not investigators and an accusation in a report can do damage that cannot be undone.

**You realise mid-review that you have a conflict.** Stop and tell the editor immediately, including what you have already read. Editors handle this routinely and would far rather know.

**You have refereed this paper before, for another journal.** Tell the editor. Some journals want a fresh referee, others value the continuity. Do not simply resubmit the earlier report; the paper may have changed, and if it has not, that is itself information the editor wants.

**The paper is in your competence but the method is not.** Review what you can, state the boundary explicitly in both the report and the confidential comments, and recommend that the editor find a methodologist. Do not attempt a confident judgement on an estimator you would not use yourself.

**Second-round review.** Read the response letter first, then the tracked-changes version, then check only that your own points were addressed. Do not raise new major points in a second round unless the revision itself created them; a referee who introduces new conditions after a compliant revision is the reason revisions take years, and editors notice.

**The invitation is from a journal you suspect is predatory.** Decline and do not review. A referee report lends legitimacy to the venue, and an editorial board membership or a review record can be cited by the publisher.

**The deadline will be missed.** Tell the editor before it passes, with a specific new date. Editors can hold a paper for a week they know about and cannot hold one for a week they do not.

**The paper is very poor and clearly written by an inexperienced author.** Keep the rejection short and name the one or two things that would most improve the next paper. This is the highest-value thing a referee ever does and it costs about fifteen minutes.

## Quality bar

- The invitation was answered within two or three days and any conflict was disclosed rather than reasoned away.
- The summary of the paper is written in the referee's own words and precedes any criticism.
- Every major point states what it changes about the conclusion, and every one survives the change test.
- Major and minor points are separated, and requests that are not conditions of publication are labelled as suggestions.
- The recommendation matches the severity of the report, and the confidential comments do not contradict it.
- The boundary of the referee's own competence is stated to the editor.
- No sentence addresses the authors rather than the work, and nothing in the report would embarrass the referee if attributed.
- Anything genuinely well done is named specifically, so the authors know what to protect.

## Adapting this to your context

These defaults come from refereeing empirical economics papers, where reports are long prose and the recommendation ladder has five rungs. Adjust the surface, not the tests.

- **The time budget.** Four to eight hours assumes a forty-page paper with tables. A short psychology report runs two to three; a qualitative manuscript runs longer, because the evidence is the extracts.
- **The reading order.** Abstract, conclusion, tables, method, introduction assumes evidence sits in tables. For qualitative work, read the extracts and the analytic procedure in the tables' place. For a trial or review, read the registration first and compare it to what was reported.
- **Reporting guidelines.** Economics has none. Health, psychology and education journals often expect CONSORT, STROBE, PRISMA, or COREQ for qualitative work. Where the journal names one, checking against it is a legitimate minor point.
- **Report shape.** Some journals want a structured form with scored subratings rather than prose, and some publish reviews signed. Map the five recommendation categories onto the options your form gives before you write.
- **What not to change.** The change test: a major point must name what it would change about the conclusion if it came out against the authors. Anything failing it is a suggestion. And read the whole paper before writing a word.

## Related skills

`peer-review-simulator` is the same activity performed on your own manuscript before submission, where exhaustiveness is a virtue rather than a fault. `identification-defense` supplies the design-specific attack list that makes a methodological major point specific rather than vague. `literature-verification` is how you check a citation you doubt before writing that you doubt it. `analysis-audit` is what a data editor does with the numbers, which is not a referee's job but is worth naming in confidential comments when you suspect a computational error. `response-to-reviewers` is how the author on the other end will handle your report, and reading it explains why an unevidenced request produces an unevidenced refusal. `thesis-chapter-review` applies an examiner's criteria rather than a journal's. `journal-targeting` is the decision that put this paper in front of you.

