- Scalene
- Posts
- Scalene 61: Ranking / Trust / Preprints
Scalene 61: Ranking / Trust / Preprints

Humans | AI | Peer review. The triangle is changing.
Cactus held an event in central London last week where one of the speakers - a senior leader from pure tech company - gave me reason to pause and reflect on some things. In particular these 2 things; 1 - AI in general, and LLMs in particular are way ahead of what the general public has access to (think Mythos and the like) and some limitations we see now are likely to disappear soon, and 2 - we are on the end of drug dealer economics, where we are getting hooked on AI at a very low cost, and at some point, once we’re hooked, the AI companies are going to start to have to charge something more like their operating costs + profit, which will be significantly more than $20/month. So we’re going to get much better, but much more expensive AI at some point soon. Maybe humans still have a role to play after all?
30th June 2026
1//
How novel is that research paper? Competition to quantify concept crowns winner
Science - 08 Jun 2026 - 4 min read
A contest run by UKRI aimed to find the best method for the automated assessment of novelty in research papers - one of the current elements of peer review where LLMs are weakest (although improving with connectors). The baseline was the evaluation of 37,480 papers by 41,600 researchers to use their own judgement and terms in grading novelty. These human judgements were then placed against machine-derived novelty scores, and the one in closest agreement (LENS) was declared the winner:
Their tool—called LLM-Evaluated Novelty and Significance, or LENS—uses cues including a paper’s cited references to summarize the state of knowledge on which the paper builds. It then evaluates whether a paper solved important problems, introduced valuable methods or concepts, or provided surprising findings that advanced the field. From there the algorithm generates a single novelty score for the paper, as well as forecasting a range of scores that predict what multiple human evaluators might give; papers that humans would likely consider a slam dunk yield smaller predicted ranges, whereas a bigger range indicates the machine expects more disagreement. It also summarizes arguments for and against the paper’s novelty.
2//
AI in peer review: the elephant in the editorial room
Evidence-based Dentistry - 04 Jun 2026 - 7 min read
Referring to the Frontiers report from the end of last year, this opens with a reminder that 53% of peer reviewers have used AI in their work. We can assume the number is somewhat higher six months later. And yet publishers are still telling reviewers what they can’t do with AI, rather than what they can - or even better - providing them with an AI workspace where reviewers can use LLMs in a monitored environment. This overview of policies from major publishers cuts a slightly exasperated air too:
These policies represent a genuine effort at standard-setting. Yet they share a structural limitation: none specifies how compliance will be verified, what constitutes a detectable violation, or what consequences follow from non-adherence. The statement exists. The governance does not.
The real risk is that prohibition prevents us from building the structures needed to make AI use safe, transparent, and accountable. Furthermore, the absence of a non-judgmental framework for acknowledging LLM use limits our understanding of the reasons why researchers rely on these tools.
Keep reading though - the 4-point framework for responsible AI use (governed acknowledgement) in peer review is a welcome suggestion for how the future could look.
3//
The future of peer review: A system-wide perspective
RORI - 22 June 2026 - 38 min read
This thought-provoking new report from RORI makes a bold claim: peer review's problems—capacity overload, conservatism, bias, eroding trust, all now turbocharged by AI—are systemic and not fixable by process tweaks alone. Their proposal to this situation reframes the whole research system. On the input side, funders shift from project grants to strategic funding for institutions, acting as "system shapers." On the output side, they propose moving beyond universal peer review toward three mechanisms: trust markers, peer engagement, and evidence synthesis. It’s ambitious system-wide thinking, but such things are notoriously difficult to implement.
4//
AIPR.pub
AIPR.pub - 14 Jun 2026 - 2/44 min read
One of the nice things about sending this newsletter out is that occasionally people contact me out of the blue and last week I had a great call with Costas Georgantas - developer of AIPR.pub.
Semi-automated solutions are needed if we want to maintain the quality of published work without arbitrarily discarding strong submissions or disadvantaging smaller institutions. My goal is to show that LLMs can enable better review quality, faster feedback, and greater transparency in any scientific field. The aim is to augment human insight, not replace it.
The reviews themselves look great and one spin-off of this has been the ability to grade and review preprints - assigning scores for novelty, rigor, applicability, clarity, and citations. A weekly digest of the best preprints is also helps readers stay on top of the firehose of new papers.
https://aipr.pub/
You can also read the paper behind AIPR here: https://arxiv.org/pdf/2606.15887v1
5//
Building an editorial AI assistant to support peer review with AWS Generative AI Innovation Center
AWS blog - 22 May 2026 - 6 min read
An enlightening tale of how AWS and the BMJ Group put agents to work to improve the lives of their editors in two key ways: Firstly, by better screening submissions for signals of papers that are likely to be rejected before they enter the peer review process - and secondly, by placing manuscripts into the correct journal across the BMJ portfolio (smarter cascading). I won’t reveal all in here, but some of their learnings are valuable and applicable to other endeavours in this space:
Co-create using best scholarly practices
Decompose complex workflows into focused tasks
[Manuscript] Structure drives quality
Predictable workflows build editorial trust
Presentation shapes trust
And finally…
Apologies for the lack of images in stories this week. I’m writing this on my phone and frankly, it’s too much effort. I’m just amazed I can do this on my phone at all.
QED launched ‘The 1%’ - an automated evaluation of 57,455 life science preprints to highlight the most compelling 574. It wasn’t without controversy.
As well as QED and AIPR, under.org has also quietly been analysing and assessing preprints with its own lenses and monthly summaries - recommended.
F1000 launch trust badges for preprints on VeriXiv. Hard to find substantive details on this - but it appears they are human-derived checks.
BioRxiv also partnered with Reviewer3 this week to bring their automated review platform to the system for interested authors, with the intention of pre-reviewing their work before it is submitted.
The Gu/Topol/Poon paper in Nature looks fascinating - but I can’t access it. If you can, click here.
This write-up of the purepub.ai online conference, by Tony Alves, is a must-read
Let’s chat
I’m in London a fair bit this coming month. Always happy to chat over coffee. Or preferably an ice-cream in this weather.
Curated by me, Chris Leonard.
If you want to get in touch, please simply reply to this email.