Evaluating Retrieval-Augmented Generation for personal collections: architecture, models and criteria

La Gorga, Angelo and Verna, Lorenzo Evaluating Retrieval-Augmented Generation for personal collections: architecture, models and criteria. JLIS.it, 2026, vol. 17, n. 3, pp. 130-149. [Journal article (Paginated)]

[thumbnail of jlis_735_1.pdf]
Preview
Text
jlis_735_1.pdf - Published version
Available under License Creative Commons Attribution.

Download (276kB) | Preview

English abstract

This article presents an ongoing experiment on the use of Retrieval-Augmented Generation (RAG) architectures in library contexts, considering them not as straightforward extensions of information retrieval techniques but as document mediation devices that require explicit methodological reflection. The contribution focuses in particular on the design and analysis of an integrated evaluation framework, conceived as a structural component of system development rather than as an ex post performance check. Using the personal archive of Emanuele Artom as a case study, characterized by bibliographic and archival heterogeneity, the article examines how the quality of generated responses emerges from the interaction between knowledge base modeling, retrieval strategies and evaluation criteria. The proposed framework combines automatic metrics, LLM-as-a-judge approaches, and human-in-the-loop processes, treating divergences between automated evaluation and expert judgment as diagnostic tools for analyzing system behaviour. Preliminary results indicate that core categories of document mediation, such as relevance, completeness, citability and transparency, cannot be fully reduced to computational parameters but, instead, require continuous negotiation between automated models and disciplinary expertise. From this perspective, RAG is framed as an epistemically unstable research object, whose reliability and governability depend on the robustness and reflexivity of the evaluation processes embedded in its development.

Item type: Journal article (Paginated)
Keywords: Retrieval-Augmented Generation (RAG); Academic Libraries; Digital collection; Personal collection; AI evaluation.
Date deposited: 16 Sep 2026 10:16
Last modified: 16 Sep 2026 10:16
URI: http://hdl.handle.net/10760/49022

References

AIB (Associazione italiana biblioteche). 2019. ‘Linee guida sul trattamento dei fondi personali.’ https://www.aib.it/documenti/linee-guida-sul-trattamento-dei-fondi-personali/.

Amirizaniani, Maryam, Jihan Yao, Adrian Lavergne, Elizabeth Snell Okada, Aman Chadha, Tanya Roosta, and Chirag Shah. ‘LLMAuditor: A Framework for Auditing Large Language Models Using Human-in-the-Loop.’ Preprint, submitted February 14, 2024. https://doi.org/10.48550/arXiv.2402.09346.

Barry, Carol L., and Linda Schamber. 1998. ‘Users’ Criteria for Relevance Evaluation: A Cross-Situational Comparison.’ Information Processing & Management 34 (2): 219–36. https://doi.org/10.1016/S0306-4573(97)00078-2.

Bassani, Elias, and Ignacio Sanchez. 2024. ‘GuardBench: A Large-Scale Benchmark for Guardrail Models.’ In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, edited by Yaser Al-Onaizan, Mohit Bansal and Yun-Nung Chen. Miami (USA): Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.1022.

Cassella, Maria. 2020. Biblioteche Accademiche e Terza Missione. Milano: Editrice Bibliografica.

Chandra, Joydeep, Satyam Kumar Navneet, and Yong Zhang. ‘The Hybrid Multimodal Graph Index (HMGI): A Comprehensive Framework for Integrated Relational and Vector Search.’ Preprint, submitted October 11, 2025. https://doi.org/10.48550/arXiv.2510.10123.

Chang, Yapei, Kyle Lo, Tanya Goyal, and Mohit Iyyer. ‘BooookScore: A Systematic Exploration of Book-Length Summarization in the Era of LLMs.’ Preprint,submitted October 1, 2023. https://doi.org/10.48550/arXiv.2310.00785.

Cox, Andrew M. 2024. ‘Artificial Intelligence and the Academic Library.’ The Journal of Academic Librarianship 50 (6): 102965. https://doi.org/10.1016/j.acalib.2024.102965.

Cox, Andrew M., and Suvodeep Mazumdar. 2024. ‘Defining Artificial Intelligence for Librarians.’ Journal of Librarianship and Information Science 56 (2): 330–40. https://doi.org/10.1177/09610006221142029.

Cox, Andrew M., Stephen Pinfield, and Sophie Rutter. 2018. ‘The Intelligent Library: Thought Leaders. Views on the Likely Impact of Artificial Intelligence on Academic Libraries.’ Library Hi Tech 37 (3): 418–35. https://doi.org/10.1108/LHT-08-2018-0105.

Gao, Mingqi, Xinyu Hu, Xunjian Yin, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2025. ‘LLM-Based NLG Evaluation: Current Status and Challenges.’ Computational Linguistics 51 (2): 661–87. https://doi.org/10.1162/coli_a_00561.

Gao, Yunfan, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, Haofen Wang. 2024. ‘Retrieval-Augmented Generation for Large Language Models: A Survey’. arXiv:2312.10997. Preprint, submitted December 18, 2023. https://doi.org/10.48550/arXiv.2312.10997.

Harisanty, Dessy, Nove E. Variant Anna, Tesa Eranti Putri, Aji Akbar Firdaus, and Nurul Aida Noor Azizi. 2025. ‘Is Adopting Artificial Intelligence in Libraries Urgency or a Buzzword? A Systematic Literature Review.’ Journal of Information Science 51 (2): 511–22. https://doi.org/10.1177/01655515221141034.

Huang, Hui, Xingyuan Bu, Hongli Zhou, Yingqi Qu, Jing Liu, Muyun Yang, Bing Xu, Tiejun Zhao. ‘An Empirical Study of Llm-as-a-Judge for Llm Evaluation: Fine-Tuned Judge Model Is Not a General Substitute for GPT-4’. Preprint, submitted March 5, 2024. https://doi.org/10.48550/arXiv.2403.02839.

Huang, Yizheng, and Jimmy Huang. ‘A Survey on Retrieval-Augmented Text Generation for Large Language Models’. Preprint, submitted April 17, 2024. https://doi.org/10.48550/arXiv.2404.10981.

IFLA (International Federation of Library Associations and Institutions). 2020. ‘IFLA Statement on Libraries and Artificial Intelligence.’ https://repository.ifla.org/handle/20.500.14598/1646.

ISO (International Organization for Standardization). 2023. ISO 11620:2023. https://www.iso.org/standard/83126.html.

IPSOS. 2024. Ipsos AI Monitor 2024: opinioni e atteggiamenti sull’Intelligenza Artificiale | Ipsos. https://www.ipsos.com/it-it/ipsos-ai-monitor-2024-opinioni-atteggiamenti-intelligenza-artificiale.

Ji, Ziwei, Nayeon Lee, Rita Frieske, Yu Tiezheng, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. ‘Survey of Hallucination in Natural Language Generation.’ ACM Computing Surveys 55 (12): 1–38. https://doi.org/10.1145/3571730.

Karpukhin, Vladimir, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih. 2020. ‘Dense Passage Retrieval for Open-Domain Question Answering.’ In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), edited by Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu. https://doi.org/10.18653/v1/2020.emnlp-main.550.

La Gorga, Angelo, Roberto Testa, and Lorenzo Verna. 2025. ‘LLM e Retrieval Augmented Generation (RAG) per le biblioteche. Sperimentazioni, prospettive e valorizzazione del patrimonio.’ DigitCult - Scientific Journal on Digital Cultures 10 (1): 59–73. https://doi.org/10.36158/97912566920714.

Lana, Maurizio. 2023. ‘Leggere l’IFLA Statement on Libraries and Artificial Intelligence al tempo di ChatGPT’. Biblioteche Oggi Trends 9 (1): 4–12. https://doi.org/10.3302/2421-3810-202301-006-1.

Ma, Rui, Kai Zhang, Zhenying He, Yinan Jing, X. Sean Wang, and Zhenqiang Chen. ‘CHASE: A Native Relational Database for Hybrid Queries on Structured and Unstructured Data.’ Preprint, submitted January 9, 2025. https://doi.org/10.48550/arXiv.2501.05006.

Marzal, Miguel Ángel, Angelo La Gorga, and Maurizio Vivarelli. 2025. ‘Enriquecimiento del valor de los fondos personales en las bibliotecas universitarias: el fondo “Emanuele Artom” entre archivos, bibliotecas y repositorios digitales.’ Revista Española de Documentación Científica 48 (1): 1653. https://doi.org/10.3989/redc.2025.1.1653.

Natale, Simone, Bruno Surace, Enrico Mensa, and Luca Befera. 2025. ‘ChatGPT for Cultural Heritage and the Customization of Generative AI: A Talkthrough Analysis of the Luigi Einaudi Chatbot.’ New Media & Society. https://doi.org/10.1177/14614448251384258.

Peikos, Georgios, and Gabriella Pasi. 2024. ‘A Systematic Review of Multidimensional Relevance Estimation in Information Retrieval.’ WIREs Data Mining and Knowledge Discovery 14 (5). https://doi.org/10.1002/widm.1541.

Poll, Roswitha. 2007. Measuring Quality: Performance Measurement in Libraries. Munich: K.G. Saur. https://doi.org/10.1515/9783598440281.

Pradhan, Alisha, and Amanda Lazar. 2021. ‘Hey Google, Do You Have a Personality? Designing Personality and Personas for Conversational Agents.’ In CUI ‘21: Proceedings of the 3rd Conference on Conversational User Interfaces, New York, USA, July 27: 1–4. https://doi.org/10.1145/3469595.3469607.

Prandi, Elena. 2014. Emanuele Artom e i suoi libri. Analisi bibliografica del fondo conservato presso la Biblioteca ‘Arturo Graf’ di Torino. https://unitesi.unito.it/handle/20.500.14240/60978.

Raisch, Sebastian, and Kateryna Fomina. 2025. ‘Hybrid Problem-Solving with Large Language Models: A Reply to “Iterative Alternative Evaluation” and “an Assemblage Perspective”.’ Academy of Management Review 50 (2): 482–84. https://doi.org/10.5465/amr.2024.0300.

Reimers, Nils, and Iryna Gurevych. 2019. ‘Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks.’ In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), edited by Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan. https://doi.org/10.18653/v1/D19-1410.

Roncaglia, Gino. 2023a. ‘Intelligenze artificiali generative e mediazione informativa: una introduzione.’ Biblioteche oggi Trends 9 (1): 13. https://doi.org/10.3302/2421-3810-202301-013-1.

Roncaglia, Gino. 2023b. L’architetto e l’oracolo: Forme digitali del sapere da Wikipedia a ChatGPT. Bari; Roma: Laterza.

Roncaglia, Gino. 2024. ‘L’intelligenza artificiale generativa multimodale in ambito umanistico.’ @DigitCult 8 (2): 127-37. https://dx.doi.org/10.36158/97888929589208.

Russell, Stuart, and Peter Norvig. 2022. Artificial Intelligence: A Modern Approach. Hoboken: Pearson India Education Services Private Limited.

Sabba, Fiammetta. 2019. ‘Third Mission, Communication, and Academic Libraries.’ Bibliothecae.It 8 (2): 219–54. https://doi.org/10.6092/ISSN.2283-9364/10368.

Saracevic, Tefko. 2007. ‘Relevance: A Review of the Literature and a Framework for Thinking on the Notion in Information Science. Part II: Nature and Manifestations of Relevance.’ Journal of the American Society for Information Science and Technology 58 (13): 1915–33. https://doi.org/10.1002/asi.20682.

Sullutrone, Giovanni, Riccardo Amerigo Vigliermo, Luca Sala, and Sonia Bergamaschi. 2024. ‘Sensitive Topics Retrieval in Digital Libraries: A Case Study of Ḥadīṯ Collections.’ In Linking Theory and Practice of Digital Libraries, edited by Apostolos Antonacopoulos, Annika Hinze, Benjamin Piwowarski, Mickaël Coustaty, Giorgio Maria Di Nunzio, Francesco Gelati, and Nicholas Vanderschantz. Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-72440-4_5.

Tammaro, Anna Maria. 2024. ‘Intelligenza artificiale e ruolo dei bibliotecari: riflessioni su due seminari di Jesus Lau.’ Bibelot: notizie dalle biblioteche toscane 30 (1). https://riviste.aib.it/bibelot/article/view/14035.

Valacchi, Federico. 2022. ‘The Parts and the Whole. Integrate Knowledge.’ JLIS.It 13 (3): 1–11. https://doi.org/10.36253/jlis.it-477.

Voorhees, Ellen M. 2000. ‘Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness.’ Information Processing & Management 36 (5): 697–716. https://doi.org/10.1016/S0306-4573(00)00010-8.

Yu, Hao, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2025. ‘Evaluation of Retrieval-Augmented Generation: A Survey.’ In Big Data, edited by Wenwu Zhu, Hui Xiong, Xiuzhen Cheng, Lizhen Cui, Zhicheng Dou, Junyu Dong, Shanchen Pang, Li Wang, Lanju Kong, and Zhenxiang Chen. Cham: Springer Nature. https://doi.org/10.1007/978-981-96-1024-2_8.


Downloads

Downloads per month over past year

Actions (login required)

View Item View Item