Open this publication in new window or tab >>2025 (English)Conference paper, Oral presentation only (Refereed)
Abstract [en]
This paper is a report from my creation of a vector database from all Ancient Greek texts in the Perseus project, in order to facilitate semantic and thematic similarity searches. As opposed to keyword searches, this approach is not dependent on exact wordings or phrases to identify relevant texts. Instead, the database can be searched for passages with thematic similarities, by searching with whole sentences or paragraphs in Greek or modern languages. I present the database, discuss the choice of embedding model, evaluate how the database performs for different kinds of searches, and suggest different ways in which vector databases based on embedding models trained on ancient languages can be useful for historians searching for relevant primary texts on a certain topic.
Keywords
AI, GPT, embeddings, vector database, semantic vectors, search
National Category
Religious Studies
Research subject
Biblical Studies, New testament
Identifiers
urn:nbn:se:ths:diva-2863 (URN)
Conference
EABS/ISBL in Uppsala,Sweden, June 23–27, 2025.
2025-06-272025-06-272025-10-03Bibliographically approved