Building a corpus of student academic writing in EMI contexts: Challenges in corpus design and data collection across international higher education settings

Dana Gablasova*, Luke Harding, Raffaella Bottini, Vaclav Brezina, Haoshan (Sally) Ren, Giovanni Iamartino, Yingyu Li, Tanjun Liu, Laura Poggesi, Kristof Savski, Anuchit Toomaneejinda, Angela Zottola

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

The article discusses methodological procedures and challenges in a project requiring multi-site, transnational data collection for the construction of a corpus of academic writing in EMI higher education contexts. Drawing on our decision-making experiences as a research team, together with empirical data generated through data collection logs recorded by a network of researchers involved in the project, we reflect on key issues in conducting the project and the solutions we found to address specific challenges. After describing the background to the project and the current status of the corpus, we focus on four broad challenges: (1) selecting partners and managing a multi-site project; (2) defining a working construct of academic writing; (3) categorising data according to disciplinary areas; and (4) managing data collection “on the ground”. Throughout, we provide descriptions of our solutions to the challenges identified, and we conclude with a call for further publication of corpus construction records to provide greater transparency and detail around decisions and judgements made at all stages of a corpus construction project.

Original languageEnglish
Article number100140
JournalResearch Methods in Applied Linguistics
Volume3
Issue number3
DOIs
Publication statusPublished - Dec 2024
Externally publishedYes

Keywords

  • Corpus construction
  • Corpus data collection
  • Corpus design
  • EMI
  • English as a Medium of Instruction
  • Written academic English

Fingerprint

Dive into the research topics of 'Building a corpus of student academic writing in EMI contexts: Challenges in corpus design and data collection across international higher education settings'. Together they form a unique fingerprint.

Cite this