Document pN67GqxLQ4q68Q1378B82Y27
Online article and related content current as of July 28, 2010.
Correction Citations Topic collections
What Makes a Good Reviewer and a Good Review for a General Medical Journal?
Nick Black; Susan van Rooyen; Fiona Godlee; et al.
JAMA. 1998;280(3):231-233 (doi:10.1001/jama.280.3.231) http://jama.ama-assn.org/cgi/content/full/280/3/231
Contact me if this article is corrected. This article has been cited 76 times. Contact me when this article is cited. Journalology/ Peer Review/ Authorship; Quality of Care; Quality of Care, Other Contact me when new articles are published in these topic areas.
Subscribe http://jama.com/subscribe
Permissions permissions@ama-assn.orc http://pubs.ama-assn.org/misc/permissions.dtl
Email Alerts http://jamaarchives.com/alerts
Reprints/E-prints reprints@ama-assn.org
Downloaded from www.jama.com by guest on July 28, 2010
3. Wilkes MS, Kravitz RL. Policies, practices, and attitudes of North American medical journal edi tors. J Gen Intern Med. 1995;10:443-450. 4. Schulman K, Sulmasy DP, Roney D. Ethics, eco nomics, and the publication policies ofmajormedical journals. JAMA. 1994;272:154-156. 5. FeurerID, BeckerGJ, Picus D, RamirezE,Darcy MD, Hicks ME. Evaluating peer reviews: pilot test ing of a grading instrument. JAMA. 1994;272:98100. 6. Evans AT, McNutt RA, Fletcher SW, Fletcher
RH. The characteristics of peer reviewers who pro duce good-quality reviews. J Gen Intern Med. 1993; 8:422-428. 7. Oxman AD, Guyatt GH. Validation of an index of the quality of review articles. J Clin Epidemiol. 1991;44:1271-1278. 8. Oxman AD, Guyatt GH, Singer J, et al. Agree ment among reviewers of review articles. J Clin Epidemiol. 1991;44:91-98. 9. Strayhorn J Jr, McDermott JF Jr, Tanguay P. An intervention to improve the reliabilityofmanuscript
reviews for the Journal ofthe American Academy ofChild andAdolescent Psychiatry. Am J Psychia try. 1993;150:947-952. 10. Cicchetti D. The reliability of peer review for manuscript and grant submissions: a cross-dis ciplinary investigation. Behav Brain Sci. 1991;14: 119-186. 11. Nylenna M, Riis P, Karlsson Y. Multiple blinded reviews of the same two manuscripts: effects of ref eree characteristics and publication language. JAMA. 1994;272:149-151.
What Makes a Good Reviewer and a Good Review for a General Medical Journal?
Nick Black, MD; Susan van Rooyen, BSc; Fiona Godlee, MRCP; Richard Smith, FRCP; Stephen Evans, MSc
Context.--Selecting peer reviewers who will provide high-quality reviews is a central task of editors of biomedical journals.
Objectives.--To determine the characteristics of reviewers for a general medi cal journal who produce high-quality reviews and to describe the characteristics of a good review, particularly in terms ofthe time spent reviewing and turnaround time.
Design, Setting, and Participants.--Surveys of reviewers of the 420 manu scripts submitted to BMJ between January and June 1997.
Main Outcome Measures.--Review quality was assessed independently by 2 editors and by the corresponding author using a newly developed 7-item review quality instrument.
Results.--Of the 420 manuscripts, 345 (82%) had 2 reviews completed, for a total of 690 reviews. Authors' assessments of review quality were available for 507 reviews. The characteristics of reviewers had little association with the quality ofthe reviews they produced (explaining only 8% of the variation), regardless of whether editors or authors defined the quality ofthe review. In a logistic regression analysis, the only significant factor associated with higher-quality ratings by both editors and authors was reviewers trained in epidemiology or statistics. Younger age also was an independent predictor for editors' quality assessments, while reviews performed by reviewers who were members of an editorial board were rated of poorer quality by authors. Review quality increased with time spent on a review, up to 3 hours but not beyond.
Conclusions.--The characteristics of reviewers we studied did not identify those who performed high-quality reviews. Reviewers might be advised that spending longer than 3 hours on a review on average did not appear to increase review qual ity as rated by editors and authors.
JAMA. 1998;280:231-233
From theLondon SchoolofHygiene&TropicalMedicine (Dr Black and Mr Evans) and BMJ (Ms Rooyen and Drs Godlee and Smith), London, England.
Presented at the International Congress on Peer Review in Biomedical Publication, Prague, Czech Republic, September 18, 1998.
Reprints: Nick Black, MD, Health Services Research Unit, London School of Hygiene & Tropical Medicine, Keppel Street, London WC1E 7HT, England (e-mail: n.black@lshtm.ac.uk).
ALTHOUGH all editors would like to know how to select good reviewers, there have been only 3 attempts to iden tify their characteristics.1-3 Two ofthese studies found that the best-quality re ports were provided by reviewers who were young1,2 and, therefore, of junior academic status,1 particularly if they-
were working at a top academic institu tion or were known to the editors.2 The third study demonstrated that younger reviewers with considerable refereeing experience provided stricter assess ments ofmanuscripts than other reviewers.3 None of the other characteristics examined (such as research training and postgraduate qualifications) were asso ciated with review quality.
The process ofpeer reviewing has also received little attention. While 3 studies have reported on the time spent by re viewers on the task,4-6 none examined the relationship between time spent and review quality.
Our principal objective was to deter mine the characteristics of reviewers who produce high-quality reviews. In addition, we considered the characteris tics of good reviews in terms of the time spent by the reviewer and time taken to deliver it to the journal.
METHODS
Consecutive manuscripts (research papers) submitted to BMJ and sent for review between January and June 1997 were eligible for inclusion. Each manu script was sent to 2 reviewers (selected from the existing reviewer database as having an interest in and knowledge of the subject matter ofthe manuscript) as part ofa randomized trial ofblinding and unmasking to coreviewer.7 Reviewers were supplied with the journal's stan dard advice and asked to return their
JAMA, July 15, 1998--Vol 280, No. 3
What Makes a Good Reviewer for Medical Journals?--Black et al 231
1998 American Medical Association. All rights reserved.
Downloaded from www.jama.com by guest on July 28, 2010
following 2 ways: the mean of the 2 edi tors' total scores and the author's total score. Initially, the relationship between the total quality score and each reviewer characteristic was assessed using linear regression analysis. Characteristics that were statistically significant (P<.01) were then entered stepwise using mul tiple regression (SPSS [Statistical Pack age for the Social Sciences] for Microsoft Windows, Release 6.1, SPSS, Inc, Chi cago, Ill). The significance of compari sons of categorical variables was as sessed using x2 tests.
Scatterplot of review quality by age of reviewer. The review quality was scored on a 5-point Likert scale (1 = poor, 5 = excellent).
Reviewer Characteristics Associated With Higher-Quality Reviews Based on Editors' (n = 670) and Author's Assessments (n = 507)
Variable
p
Editors' Assessment
Age, y >60
--.104
40-60
.001
Postgraduate training in epidemiology and/or statistics
.201
North American resident
.229
Author's Assessment
Postgraduate training in epidemiology and/or statistics
.180
Member of editorial board
--.166
SE p
.03 .0003 .05 .10
.07 .07
P Value
<.001 <.001 <.001
.02
.01 .02
review within 3 weeks. In addition, two thirds of the reviewers (the third in cluded in the "uninformed" portion ofthe trial had to be omitted to determine any Hawthorne effect) were asked to report how long they spent carrying out the re view (including reading the manuscript, making notes, and writing the review). The time taken to return the review to the journal was recorded.
The quality of each review was as sessed independently by 2 editors and by the corresponding author of the manuscript, using a new review-quality instrument. It considers the following 7 aspects of a review: the extent to which the reviewer addressed the importance of the research question, the originality ofthe question, the strengths and weak nesses of the method, the presentation ofthe paper (writing, organization, illus trations), the interpretation of the re sults, and the extent to which the re viewer provided constructive comments and substantiated the comments. Each item is scored on a 5-point Likert scale (1 = poor, 5 = excellent). The total score is the mean of the 7 item scores. In ad dition, there is a global item seeking an
overall assessment of the quality of the review. The internal consistency of the instrument is high (Cronbach a = .84) as is the interrater reliability of the total score (Kendall coefficient, t = 0.83). Full details of its psychometric properties is available from the authors.
Information on the characteristics of reviewers was obtained from a mailed questionnaire survey carried out in 1996. This covered demographic characteris tics (sex, age, country ofresidence), edu cation and familiarity with research (age when received first degree, postgradu ate qualifications, postgraduate training in epidemiology or statistics, current academic appointment, appointment in university or teaching center, current research investigator), publication ex perience (number of peer-reviewed re search publications in past 5 years, num ber of papers reviewed in past year, number of journals where participated as a reviewer, member of a journal edi torial board, member of a research fund ing body), and willingness to review blinded papers and to have identity re vealed to authors.
Review quality was measured in the
RESULTS
Recruitment and Response Rate
An estimated 420 eligible manuscripts were submitted to the journal during the recruitment period of which 2 reviews were obtained for 345 (82%). Ofthese 690
reviews, information on the characteris tics of reviewers was available for 670. The analyses presented here are based on all670 reviewsfor the editors'assessment of quality and on the first 507 reviews for which the corresponding author's assess ment was available. The only exception is the analysis oftime spent carrying out the review, for which the sample was re stricted to the 438 (editors' assessment) and 342 (author's assessment) reviewers who answered that question.
Characteristics of a Good Reviewer
Editors' Assessments.--Four of the 16 characteristics of reviewers were sig nificantly associated (P< .05) withthe edi tors' assessment of review quality: age (r2 = 0.03), resident in North America (mean score, 3.22 vs 2.88; r2 = 0.02),train-
ing in epidemiology or statistics (3.03 vs 2.74; r2 = 0.04), and current research in vestigator (2.92 vs 2.74; r2 = 0.006). Age showed a quadratic relationship in which younger reviewers up to about 60 years were more likely to produce higher-qual ity reviews (Figure), beyond which there was no statistically significant change in review quality.
When entered in a multiple regression model, only 2 characteristics remained significantly associated (P<.01) with higher-quality reviews: younger age and having training in epidemiology or statis tics (Table). Together with the character istic "resident in North America," these 3 were, however, ofvery limited predictive power (r2 = 0.08).
Author's Assessments.--Four of the 16 characteristics ofreviewers were sig nificantly associated (P<.05) with the author's assessment of review quality: age (r2 = 0.008), training in epidemiol ogy or statistics (2.94 vs 2.77; r2 = 0.009), not a member ofajournal editorial board (2.96 vs 2.81; r2 = 0.011), and resident in
232 JAMA, July 15, 1998--Vol 280, No. 3
What Makes a Good Reviewer for Medical Journals?--Black et al
1998 American Medical Association. All rights reserved.
Downloaded from www.jama.com by guest on July 28, 2010
North America (3.13 vs 2.85; r2 = 0.014). When entered in a multiple regression model, 2 characteristics remained sig nificantly associated with higher-qual ity reviews (Table). As above, these were, however, of very limited predic tive power (r2 = 0.02).
Characteristics of a Good Review
There was no association between the editors' assessment of review quality and the time taken by reviewers to re turn their reviews. There was, in con trast, a clear nonlinear relationship with the time spent by reviewers on their re views. Review qualityimproved with in creasing time up to about 3 hours, but not beyond.
COMMENT
The characteristics of reviewers con sidered in this study had little associa tion with the quality ofthe reviews they produced. This was true regardless of whether editors or authors defined the quality of the review. The only consis tent finding was that reviewers trained in epidemiology or statistics were more likely to produce good reviews. While age was associated with review quality according to editors' assessments (con sistent with previous studies1,2), it was not when reviews were assessed by au thors. The converse was true for the characteristic "not a member of an edi torial board." In regard to what makes a good review, the longer time spent on the task (up to about 3 hours), the better the review.
The lack of association between review quality and certain reviewer characteristics was surprising. We had expected that those actively involved in research, those occupying academic po sitions, and members of research fund ing bodies would have made better re viewers than others. This was not so. We did not seek to categorize the prestige of
academic institutions so were unable to investigate the previously reported as sociation between high-quality reviews and highly prestigious institutions. The association with North American resi dency was probably confounded given that such reviewers are more highly se lected by a British journal than their British counterparts.
Before discussing the implications of these findings, 2 potential methodologi cal limitations need to be considered. First, two thirds of the reviewers knew theywere participating in a study, which may have affected the quality of their review, and the time spent carrying it out. Concern about a Hawthorne effect, however, appears to be unfounded as data presented elsewhere demonstrate.7 The mean total score for the unblinded, masked reviewers was 2.79 and for the uninformed reviewers was 2.87 (differ ence, 0.08; 95% confidence interval, -0.06 to 0.22). Second, our findings depend cru cially on the review-quality instrument. Full details of its development and vali dation is available from the authors. It has good internal consistency and inter rater reliability, and we believe it was sufficiently accurate and robust for the purposes of this study. However, it should be noted that the instrument can only assess review quality in terms of content and completeness, not in terms ofwhether the reviewer's judgment was correct.
So, what makes a good reviewer and a goodreview? Our failure to explain more than 8% of the characteristics of a good reviewer is either because we did not measure the relevant factors or no con sistent pattern exists. In other words, there are almost as many types of good reviewers as there are good reviews. If true, the implication for editors is that they simply have to try new reviewers, assess their performance, and decide whether to continue to use them. This
course of action raises the question of how new reviewers might learn their trade. Itmaybetimeforjournalsto start training their reviewers--though this assumes peer review is worthwhile and that people can be trained. Meanwhile, one suggestion that we can offer editors is to recruit reviewers with training in epidemiology or statistics, and prob ably to enlist people nearer 40 than 60 years of age. Reviewers might also be advised to spend no longer than 4 hours on their task.
Finally, it is unclear whether these findings are applicable to the vast ma jority of biomedical journals that have a specialized rather than a general focus and some ofwhich may not provide guid ance to their reviewers. Further re search in this area might usefully include both types of journal.
We thank all the authors, reviewers, and editors who participated so willingly; Sue Minns and Marita Batten for any disruption we caused; and the Na tional Health Service Executive North Thames Re search & Development Responsive FundingGroup, London,England,fortheirvisioninsupportingpeer review research.
References
1. Stossel TP. Reviewer status and review quality: experience of the Journal ofClinical Investigation. N Engl J Med. 1985;312:658-659. 2. Evans AT, McNutt RA, Fletcher SW, Fletcher RH. The characteristics of peerreviewers who pro duce good quality reviews. J Gen Intern Med. 1993; 8:422-428. 3. Nylenna M, Riis P, Karlsson Y. Multiple blinded reviews of the same two manuscripts: effects of ref eree characteristics and publication language. JAMA. 1994;272:149-151. 4. Yankauer A. Who are the peer reviewers and how much do they review? JAMA. 1990;263:13381340. 5. Lock S, Smith J. What do peer reviewers do? JAMA. 1990;263:1341-1343. 6. McNutt RA, Evans AT, Fletcher RH, Fletcher SW. The effects of blinding on the quality of peer review: a randomized trial. JAMA. 1990;263:13711376. 7. van Rooyen S, Godlee F, Evans S, Smith R, Black N. Effect of blinding and unmasking on the quality of peer review: a randomized trial. JAMA. 1998;280: 234-237.
JAMA, July 15, 1998--Vol 280, No. 3
What Makes a Good Reviewer for Medical Journals?--Black et al 233
1998 American Medical Association. All rights reserved.
Downloaded from www.jama.com by guest on July 28, 2010