The Marker Drift Problem: Measuring LLM-Assisted Writing with Unstable Lexical Indicators

Nazarovets, Serhii The Marker Drift Problem: Measuring LLM-Assisted Writing with Unstable Lexical Indicators., 2026 [Preprint]

[thumbnail of The Marker Drift Problem.pdf]
Preview
Text
The Marker Drift Problem.pdf - Draft version
Available under License Creative Commons Attribution.

Download (112kB) | Preview

English abstract

Lexical markers are increasingly used to estimate the prevalence of large language model (LLM)-assisted writing across large publication corpora. However, longitudinal applications of this approach rely on an important assumption: that the relationship between these markers and LLM use remains sufficiently stable over time. Recent research suggests that once publicly identified, individual markers may lose their diagnostic value due to adaptation by users and the evolution of language models. Furthermore, lexical markers may reflect not only LLM use, but also linguistic competence, AI-assisted editing, and broader forms of stylistic convergence. This paper introduces the Marker Drift Problem to describe a methodological challenge in which indicators change alongside the phenomenon they are intended to measure. Marker drift raises questions about temporal comparability, as observed changes in marker prevalence may reflect changes in LLM use, changes in the indicators themselves, or both. Without attention to the temporal stability and construct validity of lexical indicators, increasingly precise estimates of LLM-assisted writing may be based on unstable measurement instruments.

Item type: Preprint
Keywords: large language models; scientific writing; lexical markers; marker drift; text detection; GenAI; LLM
Subjects: G. Industry, profession and education. > GB. Software industry.
L. Information technology and library technology > LL. Automated language processing.
Depositing user: Serhii Nazarovets
Date deposited: 16 Sep 2026 18:43
Last modified: 16 Sep 2026 18:43
URI: http://hdl.handle.net/10760/49030

References

Geng, M., & Poibeau, T. (2025). On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text? http://arxiv.org/abs/2510.20810

Geng, M., & Trotta, R. (2025). Human-LLM Coevolution: Evidence from Academic Writing. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguistics: ACL 2025 (pp. 12689–12696). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.657

Hashemi, A., Shi, W., & Corriveau, J.-P. (2025). AI-generated or AI touch-up? Identifying AI contribution in text data. International Journal of Data Science and Analytics, 20(4), 3759–3770. https://doi.org/10.1007/s41060-024-00693-9

Kehkashan, T., Riaz, R. A., Al-Shamayleh, A. S., Akhunzada, A., Ali, N., Hamza, M., & Akbar, F. (2025). AI-generated text detection: A comprehensive review of methods, datasets, and applications. Computer Science Review, 58, 100793. https://doi.org/10.1016/j.cosrev.2025.100793

Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27). https://doi.org/10.1126/sciadv.adt3813

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779

Liu, X., Li, Y., & Li, K. (2025). Enhancing the Robustness of AI-Generated Text Detectors: A Survey. Mathematics, 13(13), 2145. https://doi.org/10.3390/math13132145

McCreery, Z., & Naser, M. Z. (2026). LLM-Assisted writing in engineering is associated with higher citations as evidenced from 1.17 million papers. Journal of Informetrics, 20(3), 101842. https://doi.org/10.1016/j.joi.2026.101842

Májovský, M., Černý, M., Netuka, D., & Mikolov, T. (2024). Perfect detection of computer-generated text faces fundamental challenges. Cell Reports Physical Science, 5(1), 101769. https://doi.org/10.1016/j.xcrp.2023.101769

Nazarovets, S., & Teixeira da Silva, J. A. (2026). Can Generative AI Assist Post-publication Peer Review? A Pilot Study on LLM-assisted Triage. Journal of Academic Ethics, 24(3), 86. https://doi.org/10.1007/s10805-026-09761-0

Nazarovets, S., & Teixeira da Silva, J. A. (2026). Ethical opportunities and risks of using ChatGPT for open science: evidence from a three-year mixed-methods study. Journal of Ethics in Entrepreneurship and Technology. https://doi.org/10.1108/JEET-09-2025-0059

Suchikova, Y., Tsybuliak, N., Teixeira da Silva, J. A., & Nazarovets, S. (2026). GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishing. Accountability in Research, 33(3). https://doi.org/10.1080/08989621.2025.2544331

Wang, Z., Yao, A., & Ren, M. (2026). AI-generated text detection via multilevel Zipf’s law. Journal of Informetrics, 20(2), 101799. https://doi.org/10.1016/j.joi.2026.101799


Downloads

Downloads per month over past year

Actions (login required)

View Item View Item