A short roundup of live AI work in scholarly publishing and libraries — Science on review bottlenecks, Digital Science’s “workflows you can trust” grant, metadata enrichment, and why MARC/ONIX still carry the load.
Scholarly AI chatter on X this week keeps landing on the same operational theme: not speed, but auditability. For publishers and libraries that already live in MARC, ONIX, KBART and JATS, that should sound familiar — it is catalogue discipline applied to generative tools.
Science editor-in-chief H. Holden Thorp’s editorial argues that AI is currently making scientific publishing slower, worse, and more expensive: more submissions and more AI-induced errors mean more human checking, not less (Thorp / Science; abstract also on PubMed). Disclosure rules at Science already distinguish light editing help from drafting that must be declared — a policy problem that quickly becomes a metadata problem (what was assisted, by which system, under which human responsibility).
That framing is useful for DataMercs clients: treat GenAI in the research and publishing stack as workload that must be logged, not as a free productivity dial.
Digital Science’s 2026 Catalyst Grant theme is explicit: Agentic Workflows You Can Trust — up to £25,000 equity-free for early-stage tools that plan, execute and review multi-step research work with provenance, governance and audit trails (press release; grant page; applications close 5 October 2026, 17:00 BST). Their definition of trust is quality and showing the working — not only post-hoc detection of bad actors.
For metadata and product-data teams, that maps cleanly onto existing practice: bounded automation, review queues, field-level change history, and human sign-off before discovery feeds leave the house.
Clarivate’s academic-library messaging continues to push AI for day-to-day metadata enhancement and workflow change (example discussion via @ClarivateAG). Separately, public-heritage projects (e.g. Boston Public Library / Harvard / OpenAI digitisation and metadata enhancement, reported via Berkman Klein) underline the same requirement: OCR and enrichment only help if correction, source linkage, rights and provenance travel with the record.
Cataloguers already know the punchline: machines learn from intellectual work; they do not replace authority control (librarian/cataloguer discourse — follow for the practitioner frame).
JMIR and related integrity voices are naming the threat model aloud: AI-generated research, paper mills, citation manipulation, compromised peer review (@jmirpub). Nature’s researcher polling on when AI use is acceptable and what must be disclosed keeps disclosure on the industry agenda (@Nature).
If the scientific record is a knowledge graph of claims, then trust signals must be queryable — in article metadata, in repository deposits, and in library catalogues.
An old publishing-conference jab still applies: talk of AI is cheap if product feeds are not (ONIX discipline, evergreen). New agent layers still sit on identifiers, controlled fields and partner formats. We wrote the practical cataloguing side of that for AI-related works in Ghost in the MARC (ONIX + MARC21 provenance patterns).
Working rule: put the agent behind the standard — MARC, ONIX, persistent IDs, disclosure fields — not the standard behind the agent.
Consulting and teaching remain available for publishers and library partners who want those rails explicit — without the hype cycle.
For attribution, please cite this work as
Schmalfuß (2026, Sept. 8). OS DataMercs: Industry notes: AI in scholarly publishing is a provenance problem. Retrieved from https://www.datamercs.net/posts/2026-09-08-industry-notes-ai-in-scholarly-publishing/
BibTeX citation
@misc{schmalfuß2026industry,
author = {Schmalfuß, Olaf},
title = {OS DataMercs: Industry notes: AI in scholarly publishing is a provenance problem},
url = {https://www.datamercs.net/posts/2026-09-08-industry-notes-ai-in-scholarly-publishing/},
year = {2026}
}