Updated 2 hours ago
ICML authors report little policy effect—but widespread rule-breaking

AI Policy

ICML authors report little policy effect—but widespread rule-breaking

A September 16 preprint from ICML 2026 reports little estimated effect from assigning limited LLM‑review assistance, while many survey respondents admitted breaking their assigned rules.

The policy experiment separated outcomes from compliance

A [preprint posted September 16](https://arxiv.org/abs/2609.19420) used randomized policy assignments and an anonymous post‑conference survey to study LLM use in ICML 2026 peer review. The conference involved more than 24,000 papers and 17,000 reviewers, while the post‑survey received 1,486 responses. Randomization covered a subset of main‑track papers and reviewers rather than the entire conference. A [September 28 AI Understanding digest](https://aiunderstanding.org/news/study‑finds‑many‑reviewers‑ignore‑ai‑bans‑at‑icml‑conference) linked the study to fresh New Scientist coverage while explicitly saying its source report had not been independently verified. This analysis therefore relies on the paper and ICML's own records for factual claims. The authors report two distinct findings. First, they estimate that assigning the conservative or permissive policy had near‑zero effects on final paper decisions, paper scores and reviewer confidence. Second, self‑reported compliance was weak: 22.5% of surveyed respondents assigned the prohibition said they used an LLM, while 36.5% of surveyed respondents assigned the permissive policy reported at least one use that policy still disallowed. Those percentages describe survey respondents, not all reviews or all reviewers.

Limited permission did not measurably change acceptances

ICML's [two‑policy framework](https://icml.cc/Conferences/2026/LLM‑Policy) did not compare human review with fully automated review. Policy A prohibited substantive LLM use at any stage, while allowing incidental AI built into ordinary search and spelling or grammar checks. Policy B allowed privacy‑compliant tools for limited tasks such as understanding a paper or polishing language, but prohibited using an LLM to judge quality, generate strengths or weaknesses, outline the review, draft it or propose author questions. Reviewers remained responsible for the submitted text. Within that design, the paper reports nearly identical final outcomes in its paper‑level comparison: 27.0% of papers under the switched‑to‑conservative condition were accepted, compared with 26.5% retained under the permissive policy; mean paper scores were 3.31 and 3.32. Its reviewer‑level comparison estimated a score difference of 0.002 with a 95% confidence interval from -0.067 to 0.071, and a confidence difference of 0.004 with a 95% interval from -0.064 to 0.072. The authors also report that reviews under the permissive assignment were roughly 5.5% to 7% longer. Length is an observable writing effect, not evidence that the reviews were better. These results support the authors' “near‑zero” description for the measured assignments and outcomes; they do not prove equivalence, and unrestricted AI review was not the treatment.

Survey and watermark detections measure different behavior

The survey and ICML's earlier watermark enforcement describe different populations and cannot be compared as if they were two estimates of the same rate. The survey result says 22.5% of respondents assigned the prohibition self‑reported some LLM use. In its [March enforcement post](https://blog.icml.cc/2026/03/18/on‑violations‑of‑llm‑review‑policies/), ICML said hidden phrases flagged 795 reviews—about 1% of all reviews—written by 506 unique Policy A reviewers. The denominator, unit and detection method all differ. The watermark counted detected reviews across the full review pool and targeted a specific behavior: a planted phrase traveling from a submission PDF into a review. The later survey counted responding reviewers and captured broader self‑reported use, including assistance that would not reproduce a planted phrase. Neither measurement proves that every AI‑assisted review was poor, and the anonymous survey can be affected by recall, nonresponse and selection bias.

The study cannot isolate a clean all‑human control group

Noncompliance complicates the causal interpretation. Random assignment supports a comparison of policy regimes, but it does not guarantee that the prohibited‑policy group actually avoided LLMs. The strongest causal statement is therefore about the effect of assigning the two policies in the real conference environment—not the effect of AI assistance itself. The survey is also a subset of the full reviewer population. Its 1,486 responses provide direct behavioral evidence but may differ from the thousands who did not answer. A reviewer willing to disclose prohibited use anonymously may not represent a typical reviewer, just as a compliant reviewer may be more likely to respond. The paper’s near‑zero decision effect and its noncompliance estimates should be read with those limits together.

Future review rules need enforcement and measurable task boundaries

The practical lesson for conference organizers is that policy wording alone does not create a clean evaluation condition. A limited‑use policy needs task boundaries reviewers can understand, privacy rules tools can satisfy and auditing methods that measure more than copied hidden phrases. A prohibition needs credible enforcement if organizers expect it to approximate an all‑human comparison group. For authors, the preprint provides one reassuring estimate and one unresolved concern. Its authors report little estimated movement in scores or final outcomes from the policy assignment, but many survey respondents disclosed LLM uses outside their assigned rules. The next useful evidence would replicate the result across venues, compare specific allowed tasks, rate review accuracy and usefulness, and measure actual tool use directly. ICML 2026 therefore supplies evidence about governance and compliance; it does not establish that unrestricted AI review is equivalent to human review. *Figure: Author‑reported outcomes and self‑reported noncompliance in the ICML 2026 preprint. Sources: [study preprint](https://arxiv.org/abs/2609.19420) and [ICML policy](https://icml.cc/Conferences/2026/LLM‑Policy). Figure by OpenTools Team; no third‑party expressive material used.*

Share this article

PostShare

Related News