- Revue : Journal of Open Humanities Data (9)
- Pages : 4
Résumé
This paper present a novel segmentation and handwritten text recognition dataset for Medieval Latin, from the 11 th to the 16 th century. It connects with Medieval French dataset as well as ealier Latin dataset, by enforcing common guidelines. We provide our own addition to Ariane Pinche's Old French guidelines to deal with specific Latin case. We also offer an overview of how we addressed this dataset compilation through the use of pre-existing resources. With a higher abbreviation ratio and a better representation of abbreviating marks, we offer new models that outperform the base Old French model on Latin dataset, reaching readability levels on unknown manuscripts.
Partager sur les réseaux sociaux
Publications de chercheur
Publication de chercheur
CATMuS-Medieval: Consistent Approaches to Transcribing ManuScripts
Communication dans un congrès
- Date de parution : 2024
Publication de chercheur
Layout Analysis Dataset with SegmOnto
Communication dans un congrès
- Date de parution : 2024
Publication de chercheur
Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts
Communication dans un congrès Nouveau
- Date de parution : 2024