BerufsSchulKorpus (BerSchKo) (EN)

The BerSchKo (BerufsSchulKorpus) documents the longitudinal development of writing skills among newly arrived adolescents enrolled in the InteA program at a vocational school in Hesse over two academic years. It comprises of written texts addressing the operators instructing, reporting, and arguing, as well as extensive metadata on the participating students.
BerSchKo is being developed within the framework of the Center of Excellence for German as a Second Language at Justus Liebig University in Giessen. Current and former project members include:
Prof. Dr. Jana Gamper (Principal Investigator)
Yee Cheng Foo (Research Associate; Corpus Design, Data Preprocessing and Quality Management, Documentation)
Dominik Kainz (Research Associate; Data Cleaning and Documentation)
Lara Hilbert (Data Collection, Transcription and Data Cleaning)
Lisa Schlothane (Data Collection, Transcription, Data Cleaning and Segmentation)
Malin Brinkmann (Data Collection, Transcription, Data Cleaning and Segmentation)
Data Preparation and Annotation
The collected data was manually transcribed and segmented (based on Analysis of Speech Units (AS-Units), Foster et al. 2000). The EXMARaLDA Partitur Editor (Schmidt/Wörner 2014) was used for multilevel annotation.
The annotation includes:
- automatic lemmatization (manually corrected)
- automatic POS tagging with TreeTagger (Schmid 1995, manually corrected)
The XML-based EXMARaLDA files are converted into an ANNIS corpus using an Annatto pipeline (Schlauch et al. 2026).
Annotation Tiers
| Tier | Description |
| tok_part | Tokenised Speaker-Tier |
| ctok_part | Cleaned ctok_part-Tier |
| pos | Part-Of-Speech tagging automatically generated with TreeTagger and manually corrected |
| lemma | Lemmatisation automatically generated with TreeTagger and manually corrected |
| unit | Segmented AS-Units (Independent/Matrix Clauses) |
| subclause | Dependent Clauses |
| coord | Coordinated AS-Units on [unit]-Tier |
| coordsc | Coordinated AS-Units on [subclause]-Tier |
Corpus Data and Design
BerSchKo contains 880 texts from n = 139 InteA students, which were elicited over two school years over eight measurement points. Identical material-based tasks related to the operators instructing and reporting were used, adapted to the text types description and explaination according to Feilke/Rezat (2019). The primary goal is to collect text samples at the intersection of school and vocational education. In addition, argumentative texts were collected from the model exam of the DSD-I-PRO.
Extensive language-biographical data were collected (including age at the start of acquisition, educational background prior to migration to Germany, existing language skills, and a comprehensive record of text-related receptive and productive activities in the languages of origin and target languages), in order to conduct a nuanced investigation of the influence of individual prior experiences on the development of schooling language competencies in German (Schleppegrell 2004).
Corpus Handbook
The first version of the corpus and corpus handbook will follow.
References
Feilke, H. & Rezat, S. (2019). Operatoren 'to go'. Prozedurenorientierter Schreibunterricht. Praxis Deutsch 274, 4–13.
Foo, Y. C., & Gamper, J. (2025). Neuzugewanderte an beruflichen Schulen. Zur Rolle von Schreibkompetenzen im Kontext separater Beschulungsmodelle. New Immigrants in Vocational Schools. The Role of Writing Competencies in Separate School Integration Models. Sprache Im Beruf, 8(1), 53–72. https://doi.org/10.25162/sprib-2025-0004
Krause, T. (with Berlin, H.-U. Z., & Berlin, H.-U. Z.). (2019). ANNIS: A graph-based query system for deeply annotated text corpora. Humboldt-Universität zu Berlin. https://edoc.hu-berlin.de/handle/18452/20436
Schlauch, Julia; Braunewell, Aylin; Gamper, Jana; Klotz, Martin (2026). SeiKo: Ein Lernerkorpus neu zugewanderter Schüler:innen (Seiteneinsteiger:innen) in Vorbereitungsklassen (Version 1.0-preview). Zenodo. DOI: 10.5281/zenodo.21476831
Schmid, H. (1995). Improvements in Part-of-Speech Tagging with an Application to Germa. Proceedings of the ACL SIGDAT-Workshop. Dublin, Ireland.
Schmidt, T., & Wörner, K. (2014). EXMARaLDA. In Handbook on Corpus Phonology (pp. 402–419). Oxford University Press.