From Prompting to Fine-Tuning: LLM Strategies for CEFR-Based Readability Assessment
| dc.contributor.author | Comelles Pujadas, Elisabet | |
| dc.contributor.author | Alonso Alemany, Laura | |
| dc.contributor.author | Oviedo Ferreyra, Juan Cruz | |
| dc.date.accessioned | 2026-09-30T17:35:09Z | |
| dc.date.available | 2026-09-30T17:35:09Z | |
| dc.date.issued | 2026-09-24 | |
| dc.date.updated | 2026-09-30T17:35:12Z | |
| dc.description.abstract | This paper explores the use of generative open-source large language models (LLMs) to identify the CEFR level of texts for learners of English, a critical step in readability-controlled text adaptation for pedagogical purposes. To enable a systematic and cost-effective evaluation, we assess model performance on a manually annotated CEFR classification dataset, rather than relying on the more complex and less stable evaluation of generated adaptations. We evaluate several prompting strategies, including prompts enriched with CEFR descriptors or vocabulary lists, across multiple LLM families, and compare their performance to traditional readability tools and proprietary LLMs. Results show that prompting alone yields moderate performance and high variability, with models below 8B parameters behaving inconsistently, while fine-tuning markedly improves classification accuracy. Our experiments also show the importance of balanced, manually annotated training data, particularly for low-resource CEFR levels such as A1. | |
| dc.format.extent | 16 p. | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.idgrec | 772350 | |
| dc.identifier.issn | 1135-5948 | |
| dc.identifier.uri | https://hdl.handle.net/2445/231812 | |
| dc.language.iso | eng | |
| dc.publisher | Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN) | |
| dc.relation.isformatof | Reproducció del document publicat a: https://doi.org/10.2634/2-2026-77-34 | |
| dc.relation.ispartof | Procesamiento del lenguaje natural, 2026, vol. 77, p. 485-500 | |
| dc.relation.uri | https://doi.org/10.2634/2-2026-77-34 | |
| dc.rights | (c) Comelles Pujadas, E. et al., 2026 | |
| dc.rights.accessRights | info:eu-repo/semantics/openAccess | |
| dc.subject.classification | Tractament del llenguatge natural (Informàtica) | |
| dc.subject.classification | Models lingüístics | |
| dc.subject.classification | Ensenyament de llengües estrangeres | |
| dc.subject.other | Natural language processing (Computer science) | |
| dc.subject.other | Linguistic models | |
| dc.subject.other | Foreign language teaching | |
| dc.title | From Prompting to Fine-Tuning: LLM Strategies for CEFR-Based Readability Assessment | |
| dc.type | info:eu-repo/semantics/article | |
| dc.type | info:eu-repo/semantics/publishedVersion |
Fitxers
Paquet original
1 - 1 de 1