- Scalene
- Posts
- Scalene 61: Ranking / Trust / Preprints (1)
Scalene 61: Ranking / Trust / Preprints (1)

Humans | AI | Peer review. The triangle is changing.
Oh man! It’s been hot, I’ve been busy, there’s been a World Cup, and I’ve neglected the world of AI and peer review for so long I now have around 20 stories I want to share with you all. Rather than string them out for 2-3 newsletters, I’ve decided to break with my own rule and clear the decks in one fell swoop. My rule is no more than 5 stories and 5 links per newsletter, but that’s going to be broken for this week only. A special issue of sorts. So, stand by your beds, I’m about to switch the firehose on. There’s bound to be something in here you like!
22 July 2026
An Alignment Journal: Adaptation to AI
Great case study of how a real journal has surveyed the landscape of AI use in peer review, and what they are going to do about it according to their own philosophy. A welcome & mature approach in a sea of despair and panic.https://www.nabu.science
Nabu claims to be a tool which reads a paper like a thoughtful reviewer and also provides trust signals so a reader knows whether to trust it or not. It also shows the reasoning for each score. Their Methodology page is particularly enlightening.Towards an Interactive Evidence-RAG Peer-Review Workspace for the Journal of Digital History Preliminary Version of a Full Paper
The Evidence-RAG workflow makes AI-assisted peer-review support more transparent and auditable. By connecting reviewer comments to retrieved manuscript evidence, model labels, confidence components, and editor decisions, the workspace aims to support fairer and more reproducible editorial practice.LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges
This study examined the use of LLMs for automated peer review generation across several key dimensions, including modeling approaches, benchmark resources, security risks, deployment considerations, and practical application scenarios. Our analysis highlights that although LLMs can produce fluent, well-structured reviews, substantial challenges remain in calibrated scoring, deep methodological reasoning, cross-domain generalization, robustness, and fairness. Moreover, the integration of automated systems into peer review introduces broader institutional concerns related to transparency, accountability, and confidentiality.ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact
The authors introduce ReviewGuard, a novel LLM that accurately predicts long-term citation impact for rejected papers. ReviewGuard offers editors a deployable decision-support tool that identifies high-impact papers human reviewers often undervalue, providing a “complementary signal for fairer, more forward-looking editorial decisions”.FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes
Curating 3,668 complete multi-round peer-review dialogues from Nature Communications across five scientific domains (biology, chemistry, neuroscience, physics, and earth science) - the authors describe FirstPass. Deployed in the pre-submission loop as an anticipatory scientific co-author, FirstPass simulates expert critique and predicts revision cycle outcomes before submission, giving authors the judgment a trusted colleague would provide, with consistent cross-domain performance across five disciplines.No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
AI peer reviewers can be gamed by changing only presentation elements like abstracts and related work. No methods, experiments, figures, or numbers need to change to raise scores. This is an approach they term ‘adversarial repackaging’ - which made me think of my last Vinted purchase.Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
I love this paper. LLMs, by their nature, tend to focus on language and text, but scientific papers often have figures, graphs, photos, which convey core evidence, but can also be used for nefarious purposes. PaperGuard is intended to be a toolkit to systematically evaluate and defend AI-generated peer-review against domain-specific, cross-modal attacks.
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
Similar to the ‘adversarial repackaging’ paper above, this paper highlights how changes to title and abstract alone can improve AI-assisted peer reviews. But, in the bigger picture, a concluding insight was worth highlighting: “Any new metric introduced will quickly become a target for optimization, so it is crucial to carefully consider what goals we want the community to pursue and to avoid distorting scientific incentives. We hope this work motivates sustained, collective effort to ensure that AI-driven evaluation strengthens, rather than undermines, the integrity of scientific progress”.
Inclusivity and peer-review: the role of unprotected characteristics
Included here as a reminder that, even when being monitored, human peer review can be incredibly biased. Quality reviews were quite uniform, but impact reviews varied much more.How to Curate and Share Scientific Knowledge Better
This blogpost, from the Solving For team, describes the Discovery Stack Pilot - where people were asked to review papers for either quality or impact. They propose that Quality is easy to assess and Impact is easier to manipulate, and separating the two may be in science’s best interests.How Moderation Makes Science: Precautionary Valuation and Boundary-Making in the Early Circulation of Research
Essentially an examination of moderation practices at preprint servers. Moderators act as filters for unpublished research papers, but don’t consider their work akin to peer review - more like a cross between presubmission checks and rudimentary peer review. It’s a more involved process than I imagined, but also the ideal mix of skills that could be assisted by AI.Experimenting with peer review
Considering lessons learnt in funding peer review and how they can be transplanted to journals: namely lotteries and distributed review.Editorial and peer review dynamics at elite general science journals
The authors analyzed 110,303 submissions to Science and Science Advances from 2015–2020. They linked editorial decisions to author and manuscript features like prestige, topic, team size, gender, and region. Editorial review is a much stronger gatekeeper than peer review. Institutional prestige, topic, geography, and team size strongly predict editorial success.
Well, that wasn’t so bad was it? I will try not to leave it so long next time though.
PS: In writing the last update on my phone, stupid iOS got involved and decided to autocorrect a link I didn’t want autocorrecting. As a result, the good people at unpeer didn’t get their link - and people who clicked on it must have wondered what I was on about. Reproducing again here with the right links:
As well as QED and AIPR, unpeer.org has also quietly been analysing and assessing preprints with its own lenses and monthly summaries - recommended.
Let’s chat
I’m in London a fair bit this coming month. Always happy to chat over coffee. Or preferably an ice-cream in this weather.
Curated by me, Chris Leonard.
If you want to get in touch, please simply reply to this email.