Dataset Open Access
<?xml version='1.0' encoding='utf-8'?>
<resource xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://datacite.org/schema/kernel-3" xsi:schemaLocation="http://datacite.org/schema/kernel-3 http://schema.datacite.org/meta/kernel-3/metadata.xsd">
<identifier identifierType="DOI">10.25592/uhhfdm.12671</identifier>
<creators>
<creator>
<creatorName>Hussein Mohammed</creatorName>
<nameIdentifier nameIdentifierScheme="ORCID" schemeURI="http://orcid.org/">0000-0001-5020-3592</nameIdentifier>
<affiliation>Universität Hamburg</affiliation>
</creator>
</creators>
<titles>
<title>Mini-dataset for VL-Models fine-tuning (VL-Tune-dataset-mini)</title>
</titles>
<publisher>Universität Hamburg</publisher>
<publicationYear>2023</publicationYear>
<subjects>
<subject>Vision-Language Models, Dataset</subject>
</subjects>
<dates>
<date dateType="Issued">2023-06-30</date>
</dates>
<language>en</language>
<resourceType resourceTypeGeneral="Dataset"/>
<alternateIdentifiers>
<alternateIdentifier alternateIdentifierType="url">https://www.fdr.uni-hamburg.de/record/12671</alternateIdentifier>
</alternateIdentifiers>
<relatedIdentifiers>
<relatedIdentifier relatedIdentifierType="DOI" relationType="IsPartOf">10.25592/uhhfdm.12670</relatedIdentifier>
</relatedIdentifiers>
<version>1.0</version>
<rightsList>
<rights rightsURI="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</rights>
<rights rightsURI="info:eu-repo/semantics/openAccess">Open Access</rights>
</rightsList>
<descriptions>
<description descriptionType="Abstract"><p>A minimal dataset of 125 image-text pairs&nbsp;and 10&nbsp;text queries&nbsp;for fine-tuning vision-language models on manuscript images.&nbsp;It is dedicated to the task of text-based image retrieval, and splited into &quot;train&quot; and &quot;test&quot; sets. The train set&nbsp;consists of 100 image-text pairs, while the test set consists of 25 image-text pairs. This dataset is constructed from the following sources:</p>
<p>- images from&nbsp;the <a href="http://spotting.univ-rouen.fr/">DocExplore</a> dataset&nbsp;of medieval manuscripts.</p>
<p>- images from&nbsp;two manuscripts from Al-Ḥarīrī,&nbsp;<em>Maqāmāt</em>,&nbsp;&copy; Paris,&nbsp;Biblioth&egrave;que nationale de France. D&eacute;partement des manuscrits,&nbsp;namely MS arabe 3929 and MS arabe 5847.</p>
<p>- the descriptions in the text files are prepared by&nbsp;Martina Dinelli</p>
<p>The research for this work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany&#39;s Excellence Strategy &ndash; EXC 2176 &lsquo;Understanding Written Artefacts: Material, Interaction and Transmission in Manuscript Cultures&#39;, project no. 390893796. The research was conducted within the scope of the Centre for the Study of Manuscript Cultures (CSMC) at Universit&auml;t Hamburg.</p></description>
</descriptions>
</resource>