Dataset Open Access
Hamburger Zentrum für Sprachkorpora
Schmidt, Thomas;
Watzke, Franziska;
M., Peter;
Rolle, Andrea;
Hedeland, Hanna;
Al-Jaraf, Tara;
Schnieder, Annette;
Merkel, Silke;
Yusun, Secil;
Wörner, Kai;
Stäwen, Nicole;
Lehmberg, Timm
EXMARaLDA Demo Corpus 1.2
A selection of short audio and video recordings in various languages to be used for instruction or demonstration of the EXMARaLDA system.
The EXMARaLDA Demo Corpus is a small corpus which you can use to try out the functionality of the EXMARaLDA system. Please note that this corpus is for demonstration purposes only and will be changed occasionally. Further information can be found in a PDF document that describes the online and offline use of the EXMARaLDA Demo Corpus.
CLARIN Metadata summary for EXMARaLDA Demo corpus (CMDI-based)
================================================================================
EXMARaLDA Demo corpus 1.2
================================================================================
METADATA
--------
Created by: Universität Duisburg-Essen
Creation date: 2026-09-18
Self link: https://www.fdr.uni-hamburg.de/record/21551/files/EXMARaLDA-DemoKorpus.cmdi
Profile: clarin.eu:cr1:p_1422885449343
Collection: Hamburger Zentrum für Sprachkorpora (HZSK)
GENERAL INFORMATION
-------------------
Name: EXMARaLDA Demo corpus 1.2
Title: EXMARaLDA Demo corpus 1.2
PID: 10.25592/uhhfdm.21551
Resource class: corpus
Publication date: 2026
Life cycle status: released
Legal owner: Hamburger Zentrum für Sprachkorpora, Max-Brauer-Allee 60 / D-22765 Hamburg, corpora@uni-hamburg.de
Availability: Free to use for research and teaching purposes.
Version: 1.2
DESCRIPTION
-----------
The EXMARaLDA Demo Corpus is a small corpus originally designed for
demonstrating the functionality of the EXMARaLDA system. It consists of a
selection of short transcribed audio and video recordings in various
languages. Version 1.2 is a subset of version 1.1, curated for the
Transcription+ project, with the main aim of demonstrating the use of ISO
24624:2016 Language resource management — Transcription of spoken language.
KEYWORDS
--------
- L1 data
- EXMARaLDA
- ISO 24624:2016
CONTACT
-------
Organisation: Hamburg Centre for Language Corpora
Address: Max-Brauer-Allee 60, D-22765 Hamburg
Email: corpora@uni-hamburg.de
Website: https://www.slm.uni-hamburg.de/hzsk/
LICENSE
-------
Name: CC BY-NC-SA 4.0
URL: https://creativecommons.org/licenses/by-nc-sa/4.0/
Distribution: public
Non-commercial only: true
PROJECTS
--------
1. Z2 "Computer Assisted Methods for the creation and analysis of multilingual data"
Funder: German Research Foundation (DFG)
Website: https://www-archiv.fdm.uni-hamburg.de/sfb538/www.uni-hamburg.de/sfb538/forschungsprogramm/computergestuetze-methoden.html
2. ISO 24624:2016 – Transcription of spoken language: Resources, Documentation and multilingual demo corpus (Transcription+)
Funder: German Research Foundation (DFG)
Website: https://transcription-plus.awhamburg.de/
CREATORS
--------
- Hamburger Zentrum für Sprachkorpora (Max-Brauer-Allee 60 / D-22765 Hamburg)
- Thomas Schmidt (thomas@linguisticbits.de)
CONTRIBUTORS
------------
compiler:
- Hamburger Zentrum für Sprachkorpora, Max-Brauer-Allee 60 / D-22765 Hamburg, corpora@uni-hamburg.de
- Transcription+ project (Thomas Schmidt)
- Wilfried Schütte
data_inputter:
- Secil Yusun
- Annette Schnieder
- Andrea Rolle
- Silke Merkel
- Thomas Schmidt
- Martina Schwalm
- Peter M. Fischer
- Kim-Chi Hamze
- Franziska Watzke
- Roman Stachowicz
- Karolina Kaminska
- Nicole Stäwen
- Maria Görlich
- Tara Al-Jaraf
- Florian Fuchs
- Heidemarie Sambale
depositor:
- Hamburger Zentrum für Sprachkorpora, Max-Brauer-Allee 60 / D-22765 Hamburg, corpora@uni-hamburg.de
developer:
- Hamburger Zentrum für Sprachkorpora, Max-Brauer-Allee 60 / D-22765 Hamburg, corpora@uni-hamburg.de
- Transcription+ project
researcher:
- Thomas Schmidt
- Kai Wörner
- Hanna Hedeland
- Timm Lehmberg
sponsor:
- Deutsche Forschungsgemeinschaft (DFG)
- NFDI consortium Text+
CORPUS
------
Type: general corpus
Temporal class.: synchronic
Time coverage: 1964-01-07/2011-09-16
Countries: DE, PT, GB, ES, IT, US, RU, PL, FR
Multilinguality: Multilingual
Genre: discourse
Modalities: spoken
Sub-genres:
- television debate
- television sketch
- television interview
- telephone interview
- telephone interview in television
- interview
- television conversation
- radio interview
SUBJECT LANGUAGES
-----------------
- German (deu)
- English (eng)
- French (fra)
- Spanish (spa)
- Polish (pol)
- Italian (ita)
- Russian (rus)
- Portuguese (por)
ANNOTATIONS
-----------
- transcription (manual): HIAT (simplified); HIAT
- de: German translation
- en: English translation
- k: free comment
- akz: accentuation/stress
- nv: non-verbal
- sup: suprasegmental information
- pos: part of speech
- norm: orthographic normalisation
- lemma: lemmatisation
- phon: phonetic transcription
- speech-rate: speech rate
SIZE
----
- 12 communications
- 56 minutes
- 12 audio recordings
- 8 video recordings
- 12 transcriptions
- 10769 words
SPEECH CORPUS
-------------
Duration: 0.93 h
Number of speakers: 38
Demographics: 10 female, 28 male
DOCUMENTATION
-------------
- website: https://exmaralda.org/en/exmaralda-demo-corpus/
RESOURCES
---------
- LandingPage: https://exmaralda.org/en/exmaralda-demo-corpus/
| Name | Size | |
|---|---|---|
|
EXMARaLDA-Demokorpus-1.2.zip
md5:7278c193a599ee6d714cb542e6f32b79 |
1.2 GB | Download |
|
EXMARaLDA-DemoKorpus.cmdi
md5:92ccc763f12206f6c116e0e63018396d |
28.2 kB | Download |
|
EXMARaLDA-DemoKorpus.coma
md5:0e030d570bf4b2039bc41a8505a7ddc9 |
221.8 kB | Download |
|
EXMARaLDA-DemoKorpus.xml
md5:0c1c419fc902031950d1b04ff6c275cc |
3.4 kB | Download |
|
readme_de.txt
md5:bdf4519d12645b5217358a8649597cc8 |
5.3 kB | Download |
|
readme_en.txt
md5:5bcf0b82c77a473cafe967448ddcac36 |
5.2 kB | Download |