Inhaltspezifische Aktionen

BerufsSchulKorpus (BerSchKo) (EN)

 

German Version

The BerSchKo (BerufsSchulKorpus) documents the longitudinal development of writing skills among newly arrived adolescents enrolled in the InteA program at a vocational school in Hesse over two academic years. It comprises of written texts addressing the operators instructing, reporting, and arguing, as well as extensive metadata on the participating students.

 

BerSchKo is being developed within the framework of the Center of Excellence for German as a Second Language at Justus Liebig University in Giessen. Current and former project members include:

Prof. Dr. Jana Gamper (Principal Investigator)
Yee Cheng Foo (Research Associate; Corpus Design, Data Preprocessing and Quality Management, Documentation)
Dominik Kainz (Research Associate; Data Cleaning and Documentation)
Lara Hilbert (Data Collection, Transcription and Data Cleaning)
Lisa Schlothane (Data Collection, Transcription, Data Cleaning and Segmentation)
Malin Brinkmann (Data Collection, Transcription, Data Cleaning and Segmentation)

 

Data Preparation and Annotation

The collected data was manually transcribed and segmented (based on Analysis of Speech Units (AS-Units), Foster et al. 2000). The EXMARaLDA Partitur Editor (Schmidt/Wörner 2014) was used for multilevel annotation.

 

The annotation includes:

  • automatic lemmatization (manually corrected)
  • automatic POS tagging with TreeTagger (Schmid 1995, manually corrected)

The XML-based EXMARaLDA files are converted into an ANNIS corpus using an Annatto pipeline (Schlauch et al. 2026). 

 

Annotation Tiers

Tier Description
tok_part Tokenised Speaker-Tier
ctok_part Cleaned ctok_part-Tier
pos Part-Of-Speech tagging automatically generated with TreeTagger and manually corrected 
lemma Lemmatisation automatically generated with TreeTagger and manually corrected
unit Segmented AS-Units (Independent/Matrix Clauses)
subclause Dependent Clauses
coord Coordinated AS-Units on [unit]-Tier
coordsc Coordinated AS-Units on [subclause]-Tier     

 

Corpus Data and Design 

BerSchKo contains 880 texts from n = 139 InteA students, which were elicited over two school years over eight measurement points. Identical material-based tasks related to the operators instructing and reporting were used, adapted to the text types description and explaination according to Feilke/Rezat (2019). The primary goal is to collect text samples at the intersection of school and vocational education. In addition, argumentative texts were collected from the model exam of the DSD-I-PRO.

Extensive language-biographical data were collected (including age at the start of acquisition, educational background prior to migration to Germany, existing language skills, and a comprehensive record of text-related receptive and productive activities in the languages of origin and target languages), in order to conduct a nuanced investigation of the influence of individual prior experiences on the development of schooling language competencies in German (Schleppegrell 2004).

 

Corpus Handbook

The first version of the corpus and corpus handbook will follow. 

 

References

Feilke, H. & Rezat, S. (2019). Operatoren 'to go'. Prozedurenorientierter Schreibunterricht. Praxis Deutsch 274, 4–13.

Foo, Y. C., & Gamper, J. (2025). Neuzugewanderte an beruflichen Schulen. Zur Rolle von Schreibkompetenzen im Kontext separater Beschulungsmodelle. New Immigrants in Vocational Schools. The Role of Writing Competencies in Separate School Integration Models. Sprache Im Beruf, 8(1), 53–72. https://doi.org/10.25162/sprib-2025-0004

Foster, P., Tonkyn, A., & Wigglesworth, G. (2000). Measuring spoken language: A unit for all reasons. Applied Linguistics, 21(3), 354–375. https://doi.org/10.1093/applin/21.3.354
 

Krause, T. (with Berlin, H.-U. Z., & Berlin, H.-U. Z.). (2019). ANNIS: A graph-based query system for deeply annotated text corpora. Humboldt-Universität zu Berlin. https://edoc.hu-berlin.de/handle/18452/20436

Krause, T., & Klotz, M. (2026). Annatto (Version 1.2.0) [Computer software]. Humboldt-Universität zu Berlin. https://github.com/korpling/annatto/
 
Schleppegrell, M. J. (2004). The language of schooling: A functional linguistics perspective. Lawrence Erlbaum.
 

Schlauch, Julia; Braunewell, Aylin; Gamper, Jana; Klotz, Martin (2026). SeiKo: Ein Lernerkorpus neu zugewanderter Schüler:innen (Seiteneinsteiger:innen) in Vorbereitungsklassen (Version 1.0-preview). Zenodo. DOI: 10.5281/zenodo.21476831

Schmid, H. (1995). Improvements in Part-of-Speech Tagging with an Application to Germa. Proceedings of the ACL SIGDAT-Workshop. Dublin, Ireland.

Schmidt, T., & Wörner, K. (2014). EXMARaLDA. In Handbook on Corpus Phonology (pp. 402–419). Oxford University Press.