Dataset Open Access

Mini-dataset for VL-Models fine-tuning (VL-Tune-dataset-mini)

Hussein Mohammed

A minimal dataset of 125 image-text pairs and 10 text queries for fine-tuning vision-language models on manuscript images. It is dedicated to the task of text-based image retrieval, and splited into "train" and "test" sets. The train set consists of 100 image-text pairs, while the test set consists of 25 image-text pairs. This dataset is constructed from the following sources:

- images from the DocExplore dataset of medieval manuscripts.

- images from two manuscripts from Al-Ḥarīrī, Maqāmāt, © Paris, Bibliothèque nationale de France. Département des manuscrits, namely MS arabe 3929 and MS arabe 5847.

- the descriptions in the text files are prepared by Martina Dinelli

The research for this work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany's Excellence Strategy – EXC 2176 ‘Understanding Written Artefacts: Material, Interaction and Transmission in Manuscript Cultures', project no. 390893796. The research was conducted within the scope of the Centre for the Study of Manuscript Cultures (CSMC) at Universität Hamburg.

Files (42.3 MB)
Name Size
VL_Drawing_Dataset.zip
md5:4c79a7d397ab6ad09f14828c75ca2fa9
42.3 MB Download

Cite record as