Development of a semi-automated framework for analytical metadata for comics

De Luce, Valerio Development of a semi-automated framework for analytical metadata for comics. JLIS.it, 2026, vol. 17, n. 3, pp. 91-108. [Journal article (Paginated)]

[thumbnail of jlis_726_1.pdf]
Preview
Text
jlis_726_1.pdf - Published version
Available under License Creative Commons Attribution.

Download (2MB) | Preview

English abstract

The study addresses the inadequacy of traditional cataloguing standards applied to comics, this historically marginalised medium, proposing a practical solution consistent with the ongoing debate. The framework developed here, based on the principles of Augmented Humanities, integrates GPT-4o for the automatic extraction of metadata from a sample of 100 digital comics. The results show high accuracy in extraction, filling in numerous fields of an analytical bibliographic record proposed and validated for the experiment. The main objective is to demonstrate the potential of AI in improving the efficiency of cataloguing practices. However, the study highlights persistent problems such as hallucinations, difficulties in handling incomplete historical data on the material, and the need for human supervision for ambiguous fields. Finally, the work outlines a roadmap for integrating AI into descriptive processes, promoting the bibliographic recognition of comics as a cultural resource.

Italian abstract

Lo studio affronta l’inadeguatezza degli standard catalografici tradizionali applicati al fumetto, questo medium storicamente marginalizzato, proponendo una soluzione pratica coerente con il dibattito in corso. Il framework qui sviluppato, fondato sui principi delle Augmented Humanities, integra GPT-4o per l’estrazione automatica di metadati da un campione di 100 fumetti digitali. I risultati mostrano un’elevata accuratezza nell’estrazione, compilando numerosi campi di una scheda bibliografica analitica proposta e validata per l’esperimento. L’obiettivo principale è dimostrare il potenziale dell’IA nel migliorare l’efficienza delle pratiche catalografiche. Tuttavia, lo studio evidenzia problemi persistenti come le allucinazioni, le difficoltà nella gestione di dati storici incompleti sul materiale e la necessità di supervisione umana per campi ambigui. Infine, il lavoro delinea una roadmap per l’integrazione dell’IA nei processi descrittivi, promuovendo il riconoscimento bibliografico dei fumetti come risorsa culturale.

Item type: Journal article (Paginated)
Keywords: Generative artificial intelligence; Bibliographic description of comics; Automated metadata generation; Library and information science; Intelligenza artificiale generativa; Descrizione bibliografica del fumetto; Generazione automatica di metadati; Biblioteconomia e scienze dell’informazione
Date deposited: 16 Sep 2026 10:16
Last modified: 20 Sep 2026 22:24
URI: http://hdl.handle.net/10760/49021

References

Amato, Maria Concetta, "Il mercato del fumetto in Italia: analisi e prospettive". Tesi, Università degli Studi di Genova, 2024). https://unire.unige.it/bitstream/handle/123456789/9749/tesi30240623.pdf

Bailund, Allison, Steven W. Holloway, Kayla Kuni, e Deborah Tomaras. 2023. "Reprints, Reboots and Retcons: Standardizing Comics Cataloging with the Best Practices Guide from the GNCRT." Art Libraries Journal 48 (3): 62-8. https://doi.org/10.1017/alj.2023.11.

Bryant, Peter T. 2021. "Augmented Humanity: Being and Remaining Agentic in a Digitalized World." In Augmented Humanity, Springer Nature. https://link.springer.com/book/10.1007/978-3-030-76445-6.

De Martin, Chiara, Barbara Leporini, e Gregorio Pellegrino. 2021. "Verso la descrizione automatica delle immagini nell'editoria digitale accessibile: proposta di una tassonomia di immagini per gli algoritmi di IA." In DH per la società: eguaglianza, partecipazione, diritti e valori nell'era digitale. Raccolta degli abstract estesi della 10ª conferenza nazionale, 480-83. Roma: Associazione per l'Informatica Umanistica e la Cultura Digitale. https://arpi.unipi.it/handle/11568/1219148.

Dutta, Arpita, Samit Biswas, e Amit Kumar Das. 2021. "CNN-Based Segmentation of Speech Balloons and Narrative Text Boxes from Comic Book Page Images." International Journal on Document Analysis and Recognition (IJDAR) 24 (1): 49-62. https://doi.org/10.1007/s10032-021-00366-4.

Floridi, Luciano, Josh Cowls, Monica Beltrametti, Raja Chatila, Patrice Chazerand, Virginia Dignum, Christoph Luetge, Robert Madelin, Ugo Pagallo, Francesca Rossi, Burkhard Schafer, Peggy Valcke, e Effy Vayena. 2018. "AI4People-An Ethical Framework for a Good AI Society: Opportunities, Risks, Principles, and Recommendations." Minds and Machines 28: 689-707. https://doi.org/10.1007/s11023-018-9482-5.

Formiga, Federica. 2024. "Il fumetto nelle biblioteche: un percorso distributivo in movimento." Sistema Editoria. Rivista internazionale di studi sulla contemporaneità 2 (2): 25-45. https://doi.org/10.14672/se.v2i2.2703.

Gamage, Ruwan, e Priyanwada Wanigasooriya. 2024. "Using Generative AI for Bibliographic Description: A Study with ChatGPT-4." Journal of the University Librarians Association of Sri Lanka 27 (2): 257-84. https://doi.org/10.4038/jula.v27i2.8083.

Guerrini, Mauro. 2020. Dalla catalogazione alla metadatazione. Tracce di un percorso. Roma: Associazione Italiana Biblioteche.

Guerrini, Mauro. 2025. "La rappresentazione dell'universo bibliografico e delle collezioni in era digitale: l'entity modeling." Biblioteche Oggi Trends 11 (1): 23-6. https://www.bibliotecheoggitrends.it/it/articolo/4120/la-rappresentazione-dell-universo-bibliografico.

Gupta, Asmita. 2025. "Common Issues and Uncommon Solutions to Cataloguing Comic Books in Libraries." In Proceedings of the Annual Conference of CAIS / Actes du congrès annuel de l'ACSI. https://doi.org/10.29173/cais1891.

Jackson, Amy S., Myung-Ja K. Han, Kurt Groetsch, Megan Mustafoff, e Timothy W. Cole. 2008. "Dublin Core Metadata Harvested Through OAI-PMH." Journal of Library Metadata 8 (1): 5-21. https://doi.org/10.1300/J517v08n01_02.

Jobin, Anna, Marcello Ienca, e Effy Vayena. 2019. "The Global Landscape of AI Ethics Guidelines." Nature Machine Intelligence 1: 389-399. https://doi.org/10.1038/s42256-019-0088-2.

Lenadora, Damitha, Rakhitha Ranathunge, Chamath Samarawickrama, Yumantha De Silva, Indika Perera, e Anuradha Welivita. 2020. "Extraction of Semantic Content and Styles in Comic Books." International Journal on Advances in ICT for Emerging Regions (ICTer) 13 (1). https://doi.org/10.4038/icter.v13i1.7212.

Lincoln, Matthew, Julia Corrin, Emily Davis, e Scott B. Weingart. 2020. "CAMPI: Computer-Aided Metadata Generation for Photo Archives Initiative." Library of Congress. https://doi.org/10.1184/R1/12791807.

Maggi, Roberta, Tiziana Pasciuto, Martina Mazzoleni, Maria Teresa Artese, Isabella Gagliardi, e Riccardo Albertoni. 2023. "GECA 3.0-A New Tool for Cataloguing and Enjoying Cultural Heritage." In La memoria digitale: forme del testo e organizzazione della conoscenza. Atti del XII Convegno Annuale AIUCD, 373-79. Siena: Università degli Studi di Siena. https://doi.org/10.6092/unibo/amsacta/7721.

Nguyen, Trinh, e Amany Elbanna. 2025. "Understanding Human-AI Augmentation in the Workplace: A Review and a Future Research Agenda." Information Systems Frontiers. https://doi.org/10.1007/s10796-025-10591-5.

Ponsard, Christophe, Ravi Ramdoyal, e Daniel Dziamski. 2012. "An OCR-Enabled Digital Comic Books Viewer." In Computers Helping People with Special Needs: 13th International Conference, ICCHP 2012, Linz, Austria, July 11-13, 471-78. Berlin: Springer. https://doi.org/10.1007/978-3-642-31522-0_71.

Rigaud, Christophe. 2014. "Segmentation and Indexation of Complex Objects in Comic Book Images". PhD dissertation, Université de La Rochelle, 2014. https://theses.hal.science/tel-01221308v1.

Sferruzza, Marco. 2023. "I fumetti in biblioteca: una palestra per la catalogazione." AIB Studi 63 (2): 297-311. https://doi.org/10.2426/aibstudi-13878.

Soykan, Gürkan, Deniz Yuret, e Tevfik Metin Sezgin. 2024. "ComicBERT: A Transformer Model and Pre-training Strategy for Contextual Understanding in Comics." In Document Analysis and Recognition-ICDAR 2024 Workshops, edited by Harold Mouchère e Anqi Zhu. Cham: Springer. https://doi.org/10.1007/978-3-031-70645-5_16.

Sugie, Noriko, Akiko Hashizume, Djoke Dam, e Mari Agata. 2025. "Are Japanese Manga Held in Dutch Libraries? An Analysis of Collection Trends Using Library Metadata." Journal of Library Metadata 25 (2): 69-97. https://doi.org/10.1080/19386389.2025.2469993.

Vivoli, Emanuele, Mohamed Ali Souibgui, Andrey Barsky, Artemis LLabrés, Marco Bertini, e Dimosthenis Karatzas. 2024. "One Missing Piece in Vision and Language: A Survey on Comics Understanding." Preprint, submitted September 14, 2024. https://doi.org/10.48550/arXiv.2409.09502.

Wang, Xuezhi, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, e Denny Zhou. 2023. "Self-Consistency Improves Chain of Thought Reasoning in Language Models." In International Conference on Learning Representations (ICLR 2023). https://arxiv.org/abs/2203.11171.

Walsh, John A. 2012. "Comic Book Markup Language: An Introduction and Rationale." Digital Humanities Quarterly 6 (1). https://dhq.digitalhumanities.org/vol/6/1/000117/000117.html.

Zhang, Yunlong, e Seiji Hotta. 2023. "Automatic Reading Order Detection of Comic Panels." In Pattern Recognition, Computer Vision, and Image Processing. ICPR 2022 International Workshops and Challenges, 76-90. Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-37742-6_6.

Zhu, Deyao, Jun Chen, Xiaoqian Shen, Xiang Li, e Mohamed Elhoseiny. 2024. "MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models." In Proceedings of the 12th International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.2304.10592.


Downloads

Downloads per month over past year

Actions (login required)

View Item View Item