TY - JOUR
T1 - DICO-JALF v.1.0: Diponegoro Corpus of Japanese Learners as a Foreign Language in Indonesia with AI Error Annotation and Human Supervision
AU - Prihantoro,
AU - Ishikawa, Shin’Ichiro
AU - Liu, Tanjun
AU - Fadli, Zaki Ainul
AU - Rini, Elizabeth Ika Hesti Aprilia Nindia
AU - Kepirianto, Catur
N1 - Publisher Copyright:
© 2025, Fakultas Ilmu Budaya Universitas Andalas. All rights reserved.
PY - 2025/9/30
Y1 - 2025/9/30
N2 - There is a growing body of research in using AI for corrective feedback in foreign language teaching. However, few studies have specifically addressed the accuracy of AI analysis in learner corpus research. This study aims to create an AI-annotated corpus whose data were obtained from learners of Japanese as a Foreign Language (JFL) in Indonesia with human supervision; branded it as DICO-JALF v.1.0. The aim is to measure to what extent ChatGPT accurately annotates errors. A task was first administered to collect corpus data and metadata to build the corpus. The corpus was error-annotated using ChatGPT 4.0. Human annotators manually supervised the accuracy of AI-generated annotations. Regarding errors committed by learners, it is observed that incorrect lexical choices and forms dominate the cause of errors, while underuse and overuse are minimal. It can be concluded that ChatGPT demonstrated an average accuracy of 70% correct identification of errors. Regarding error rate, the verb is the category where errors are most frequent, which maybe driven by its conjugation, a feature absent in Indonesian, the L1 of the students. This suggests that Indonesian learners’ acquisition of Japanese verbs needs greater emphasis. As compared to other similar studies, this is relatively low. However, it can be argued that one factor determining the accuracy of ChatGPT annotations, or any other LLM-based tool, is the complexity of the annotation scheme they adhere to. The corpus have been made available for download. The annotations shall be readable by a corpus query system that reads XML tags. This corpus serves as a foundational resource for future research on AI-assisted error analysis in JFL learning contexts in Indonesia.
AB - There is a growing body of research in using AI for corrective feedback in foreign language teaching. However, few studies have specifically addressed the accuracy of AI analysis in learner corpus research. This study aims to create an AI-annotated corpus whose data were obtained from learners of Japanese as a Foreign Language (JFL) in Indonesia with human supervision; branded it as DICO-JALF v.1.0. The aim is to measure to what extent ChatGPT accurately annotates errors. A task was first administered to collect corpus data and metadata to build the corpus. The corpus was error-annotated using ChatGPT 4.0. Human annotators manually supervised the accuracy of AI-generated annotations. Regarding errors committed by learners, it is observed that incorrect lexical choices and forms dominate the cause of errors, while underuse and overuse are minimal. It can be concluded that ChatGPT demonstrated an average accuracy of 70% correct identification of errors. Regarding error rate, the verb is the category where errors are most frequent, which maybe driven by its conjugation, a feature absent in Indonesian, the L1 of the students. This suggests that Indonesian learners’ acquisition of Japanese verbs needs greater emphasis. As compared to other similar studies, this is relatively low. However, it can be argued that one factor determining the accuracy of ChatGPT annotations, or any other LLM-based tool, is the complexity of the annotation scheme they adhere to. The corpus have been made available for download. The annotations shall be readable by a corpus query system that reads XML tags. This corpus serves as a foundational resource for future research on AI-assisted error analysis in JFL learning contexts in Indonesia.
KW - AI
KW - ChatGPT
KW - Corpus linguistics
KW - Indonesia
KW - JFL
KW - error annotation
UR - https://www.scopus.com/pages/publications/105018758782
U2 - 10.25077/ar.12.3.274-288.2025
DO - 10.25077/ar.12.3.274-288.2025
M3 - Article
AN - SCOPUS:105018758782
SN - 2339-1162
VL - 12
SP - 274
EP - 288
JO - Jurnal Arbitrer
JF - Jurnal Arbitrer
IS - 3
ER -