Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer review, yet a key unresolved question is whether AI can generate technically sound reviews at real-world conference scale. Here we report the first large-scale field deployment of AI-assisted peer review: every main-track submission at AAAI-26 received one clearly identified AI review from a state-of-the-art system. The system combined frontier models, tool use, and safeguards in a multi-stage process to generate reviews for all 22,977 full-review papers in less than a day. A large-scale survey of AAAI-26 authors and program committee members showed that participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions. We also introduce a novel benchmark and find that our system substantially outperforms a simple LLM-generated review baseline at detecting a variety of scientific weaknesses. Together, these results show that state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research.
@article{Biswas2026AIAssistedPR,author={Joydeep Biswas and
Sheila Schoepp and
Gautham Vasan and
Anthony Opipari and
Arthur Zhang and
Zichao Hu and
Sebastian Joseph and
Matthew Lease and
Junyi Jessy Li and
Peter Stone and
Kiri L. Wagstaff and
Matthew E. Taylor and
Odest Chadwicke Jenkins},title={AI-Assisted Peer Review at Scale: The {AAAI-26} {AI} Review Pilot},journal={CoRR},volume={abs/2604.13940},year={2026},url={https://doi.org/10.48550/arXiv.2604.13940},doi={10.48550/ARXIV.2604.13940},eprinttype={arXiv},eprint={2604.13940},timestamp={Mon, 11 May 2026 17:08:47 +0200},biburl={https://dblp.org/rec/journals/corr/abs-2604-13940.bib},bibsource={dblp computer science bibliography, https://dblp.org},month=apr}
2025
IJCAI
The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning
Reinforcement learning (RL) has shown impressive results in sequential decision-making tasks. Meanwhile, Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged, exhibiting impressive capabilities in multimodal understanding and reasoning. These advances have led to a surge of research integrating LLMs and VLMs into RL. In this survey, we review representative works in which LLMs and VLMs are used to overcome key challenges in RL, such as lack of prior knowledge, long-horizon planning, and reward design. We present a taxonomy that categorizes these LLM/VLM-assisted RL approaches into three roles: agent, planner, and reward. We conclude by exploring open problems, including grounding, bias mitigation, improved representations, and action advice. By consolidating existing research and identifying future directions, this survey establishes a framework for integrating LLMs and VLMs into RL, advancing approaches that unify natural language and visual understanding with sequential decision-making.
@inproceedings{Schoepp2025TheEL,author={Sheila Schoepp and
Masoud Jafaripour and
Yingyue Cao and
Tianpei Yang and
Fatemeh Abdollahi and
Shadan Golestan and
Zahin Sufiyan and
Osmar R. Za{\"{\i}}ane and
Matthew E. Taylor},title={The Evolving Landscape of {LLM-} and VLM-Integrated Reinforcement Learning},booktitle={Proceedings of the International Joint Conference on Artificial Intelligence ({IJCAI})},pages={10641--10649},publisher={ijcai.org},year={2025},url={https://doi.org/10.24963/ijcai.2025/1181},doi={10.24963/IJCAI.2025/1181},timestamp={Wed, 24 Sep 2025 17:45:28 +0200},biburl={https://dblp.org/rec/conf/ijcai/SchoeppJCYAGSZT25.bib},bibsource={dblp computer science bibliography, https://dblp.org}}