Data di Pubblicazione:
2024
Abstract:
This paper presents a model for the automatic classification of writing proficiency in Italian as a second language (L2) according to the Common European Framework of Reference (CEFR) for languages. The proposed method integrates lexical and morphosyntactic quantitative analysis with phraseological dimensions. Phraseological aspects include the ability to use and understand fixed expressions, idioms, and other multiword units that are common in a language and reflect the depth of language comprehension typically manifested by native speakers. Specific techniques for encoding phraseological features have been introduced, and basic phraseological statistics, previously unavailable for Italy, have been extracted from an Italian corpus. The proposed model was experimentally compared with widely used machine-learning models using a dataset of written texts produced by non-native speakers for official Italian CEFR certification exams. The experimental results outperformed previous work on the CEFR classification of Italian L2 proficiency in terms of accuracy and all relevant prediction metrics, demonstrating the effectiveness of the proposed approach, which integrates morphosyntactic and phraseological features.
Tipologia CRIS:
1.1 Articolo in rivista
Keywords:
Machine learning; Classification algorithms; Complexity measures; Text complexity; Language proficiency; L2 learners; NLP; Complexity theory; Syntactics; Linguistics; Feature extraction; Accuracy; Standards; Machine learning; Vocabulary; Europe; Current measurement
Elenco autori:
Franzoni, Valentina; Biondi, Giulio; Milani, Alfredo
Link alla scheda completa:
Pubblicato in: