Reviewer Guidelines & Evaluation Criteria
This reference documents how reviewers evaluate papers at major ML/AI conferences, helping authors anticipate and address reviewer concerns.
Contents
- - Universal Evaluation Dimensions
- - NeurIPS Reviewer Guidelines
- - ICML Reviewer Guidelines
- - ICLR Reviewer Guidelines
- - ACL Reviewer Guidelines
- - What Makes Reviews Strong
- - Common Reviewer Concerns
- - How to Address Reviewer Feedback
- - Are claims well-supported by theoretical analysis or experimental results?
- - Are the proofs correct? Are the experiments properly controlled?
- - Are baselines appropriate and fairly compared?
- - Is the methodology sound?
- - Include complete proofs (main paper or appendix with sketches)
- - Use appropriate baselines (not strawmen)
- - Report variance/error bars with methodology
- - Document hyperparameter selection process
- - Is the paper clearly written and well organized?
- - Can an expert in the field reproduce the results?
- - Is notation consistent? Are terms defined?
- - Is the paper self-contained?
- - Use consistent terminology throughout
- - Define all notation at first use
- - Include reproducibility details (appendix acceptable)
- - Have non-authors read before submission
- - Are the results impactful for the community?
- - Will others build upon this work?
- - Does it address an important problem?
- - What is the potential for real-world impact?
- - Clearly articulate the problem's importance
- - Connect to broader research themes
- - Discuss potential applications
- - Compare to existing approaches meaningfully
- - Does this provide new insights?
- - How does it differ from prior work?
- - Is the contribution non-trivial?
- - Superficial, uninformed reviews
- - Demanding unreasonable additional experiments
- - Penalizing authors for honest limitation acknowledgment
- - Rejecting for missing citations to reviewer's own work
- - Bidding: May 17-21
- - Reviewing period: May 29 - July 2
- - Author rebuttals: July 24-30
- - Discussion period: July 31 - August 13
- - Final notifications: September 18
- - Top 25% of accepted papers: Score 5-6
- - Typical accepted paper: Score 4-5
- - Borderline: Score 3-4
- - Clear reject: Score 1-2
- - Public reviews (after acceptance decisions)
- - Author responses visible to reviewers
- - Discussion between reviewers and ACs
- - Soundness: 1-4 scale
- - Presentation: 1-4 scale
- - Contribution: 1-4 scale
- - Overall: 1-10 scale
- - Confidence: 1-5 scale
- - Are limitations honest and comprehensive?
- - Do limitations undermine core claims?
- - Are potential negative impacts addressed?
- - Dual-use concerns
- - Data privacy issues
- - Bias and fairness implications
- - Broader AI scope: AAAI covers all of AI, not just ML. Papers on planning, reasoning, knowledge representation, NLP, vision, robotics, and multi-agent systems are all in scope. Reviewers may not be deep ML specialists.
- - Formatting strictness: AAAI reviewers are instructed to flag formatting violations. Non-compliant papers may be desk-rejected before review.
- - Application papers: AAAI is more receptive to application-focused work than NeurIPS/ICML. Framing a strong application contribution is viable.
- - Senior Program Committee: AAAI uses SPCs (Senior Program Committee members) who mediate between reviewers and make accept/reject recommendations.
- - Strong Accept: Clearly above threshold, excellent contribution
- - Accept: Above threshold, good contribution with minor issues
- - Weak Accept: Borderline, merits outweigh concerns
- - Weak Reject: Borderline, concerns outweigh merits
- - Reject: Below threshold, significant issues
- - Strong Reject: Well below threshold
- - Language model focus: Reviewers will assess whether the contribution advances understanding of language models. General ML contributions need explicit LM framing.
- - Newer venue norms: COLM is newer than NeurIPS/ICML, so reviewer calibration varies more. Write more defensively — anticipate a wider range of reviewer expertise.
- - ICLR-derived process: Review process is modeled on ICLR (open reviews, author response period, discussion among reviewers).
- - Broad interpretation of "language modeling": Includes training, evaluation, alignment, safety, efficiency, applications, theory, multimodality (if language is central), and social impact of LMs.
- - 8-10: Strong accept (top papers)
- - 6-7: Weak accept (solid contribution)
- - 5: Borderline
- - 3-4: Weak reject (below threshold)
- - 1-2: Strong reject
- - What the paper does
- - Main contribution claimed
- - Specific positive aspects
- - Why these matter
- - Specific concerns
- - Why these matter
- - Suggestions for addressing
- - Clarifications needed
- - Things that would change assessment
- - Typos, unclear sentences
- - Formatting issues
- - Clear recommendation with reasoning
- - Thank reviewers for their time
- - Address each concern specifically
- - Provide evidence (new experiments if possible)
- - Be concise—reviewers are busy
- - Acknowledge valid criticisms
- - Be defensive or dismissive
- - Make promises you can't keep
- - Ignore difficult criticisms
- - Write excessively long rebuttals
- - Argue about subjective assessments
- - Valid technical errors
- - Missing important related work
- - Unclear explanations
- - Missing experimental details
- - Reviewer misunderstood the paper
- - Requested experiments are out of scope
- - Criticism is factually incorrect
- - [ ] Would I trust these results if I saw them?
- - [ ] Are all claims supported by evidence?
- - [ ] Are baselines fair and recent?
- - [ ] Can someone reproduce this from the paper?
- - [ ] Is the writing clear to non-experts in this subfield?
- - [ ] Are all terms and notation defined?
- - [ ] Why should the community care about this?
- - [ ] What can people do with this work?
- - [ ] Is the problem important?
- - [ ] What specifically is new here?
- - [ ] How does this differ from closest related work?
- - [ ] Is the contribution non-trivial?
Universal Evaluation Dimensions
All major ML conferences assess papers across four core dimensions:
1. Quality (Technical Soundness)
What reviewers ask:
How to ensure high quality:
2. Clarity (Writing & Organization)
What reviewers ask:
How to ensure clarity:
3. Significance (Impact & Importance)
What reviewers ask:
How to demonstrate significance:
4. Originality (Novelty & Contribution)
What reviewers ask:
Key insight from NeurIPS guidelines:
> "Originality does not necessarily require introducing an entirely new method. Papers that provide novel insights from evaluating existing approaches or shed light on why methods succeed can also be highly original."
NeurIPS Reviewer Guidelines
Scoring System (1-6 Scale)
| Score | Label | Description |
| ------- | ------- | ------------- |
| 6 | Strong Accept | Groundbreaking, flawless work; top 2-3% of submissions |
| 5 | Accept | Technically solid, high impact; would benefit the community |
| 4 | Borderline Accept | Solid work with limited evaluation; leans accept |
| 3 | Borderline Reject | Solid but weaknesses outweigh strengths; leans reject |
| 2 | Reject | Technical flaws or weak evaluation |
| 1 | Strong Reject | Well-known results or unaddressed ethics concerns |
| Criterion | Weight | Notes |
| ----------- | -------- | ------- |
| Technical quality | High | Soundness of approach, correctness of results |
| Significance | High | Importance of the problem and contribution |
| Novelty | Medium-High | New ideas, methods, or insights |
| Clarity | Medium | Clear writing, well-organized presentation |
| Reproducibility | Medium | Sufficient detail to reproduce results |
| Criterion | Weight | Notes |
| ----------- | -------- | ------- |
| Relevance | High | Must be relevant to language modeling community |
| Technical quality | High | Sound methodology, well-supported claims |
| Novelty | Medium-High | New insights about language models |
| Clarity | Medium | Clear presentation, reproducible |
| Significance | Medium-High | Impact on LM research and practice |
| Concern | How to Pre-empt | |
| --------- | ----------------- | |
| "Baselines too weak" | Use state-of-the-art baselines, cite recent work | |
| "Missing ablations" | Include systematic ablation study | |
| "No error bars" | Report std dev/error, multiple runs | |
| "Hyperparameters not tuned" | Document tuning process, search ranges | |
| "Claims not supported" | Ensure every claim has evidence | |
| Concern | How to Pre-empt | |
| --------- | ----------------- | |
| "Incremental contribution" | Clearly articulate what's new vs prior work | |
| "Similar to [paper X]" | Explicitly compare to X in Related Work | |
| "Straightforward extension" | Highlight non-obvious aspects | |
| Concern | How to Pre-empt | |
| --------- | ----------------- | |
| "Hard to follow" | Use clear structure, signposting | |
| "Notation inconsistent" | Review all notation, create notation table | |
| "Missing details" | Include reproducibility appendix | |
| "Figures unclear" | Self-contained captions, proper sizing | |
| Concern | How to Pre-empt | |
| --------- | ----------------- | |
| "Limited impact" | Discuss broader implications | |
| "Narrow evaluation" | Evaluate on multiple benchmarks | |
| "Only works in restricted setting" | Acknowledge scope, explain why still valuable |
How to Address Reviewer Feedback
Rebuttal Best Practices
Do:
Don't:
Rebuttal Template
`markdown
We thank the reviewers for their thoughtful feedback.
Reviewer 1
R1-Q1: [Quoted concern]
[Direct response with evidence]
R1-Q2: [Quoted concern]
[Direct response with evidence]
Reviewer 2
...
Summary of Changes
If accepted, we will:
1. [Specific change]
2. [Specific change]
3. [Specific change]
`
When to Accept Criticism
Some reviewer feedback should simply be accepted:
Acknowledge these gracefully: "The reviewer is correct that... We will revise to..."
When to Push Back
You can respectfully disagree when:
Frame disagreements constructively: "We appreciate this perspective. However, [explanation]..."
Pre-Submission Reviewer Simulation
Before submitting, ask yourself:
Quality:
Clarity:
Significance:
Originality: