Tipus de document

Article

Versió

Versió publicada

Data de publicació

Tots els drets reservats

Si us plau utilitzeu sempre aquest identificador per citar o enllaçar aquest document: https://hdl.handle.net/2445/231812

From Prompting to Fine-Tuning: LLM Strategies for CEFR-Based Readability Assessment

Títol de la revista

Director/Tutor

Contribució addicional

ISSN de la revista

Títol del volum

Resum

This paper explores the use of generative open-source large language models (LLMs) to identify the CEFR level of texts for learners of English, a critical step in readability-controlled text adaptation for pedagogical purposes. To enable a systematic and cost-effective evaluation, we assess model performance on a manually annotated CEFR classification dataset, rather than relying on the more complex and less stable evaluation of generated adaptations. We evaluate several prompting strategies, including prompts enriched with CEFR descriptors or vocabulary lists, across multiple LLM families, and compare their performance to traditional readability tools and proprietary LLMs. Results show that prompting alone yields moderate performance and high variability, with models below 8B parameters behaving inconsistently, while fine-tuning markedly improves classification accuracy. Our experiments also show the importance of balanced, manually annotated training data, particularly for low-resource CEFR levels such as A1.

Citació

Citació

COMELLES PUJADAS, Elisabet, ALONSO ALEMANY, Laura and OVIEDO FERREYRA, Juan Cruz. From Prompting to Fine-Tuning: LLM Strategies for CEFR-Based Readability Assessment. Procesamiento del lenguaje natural. 2026. Vol. 77, num. 485-500. ISSN 1135-5948. [consulted: 7 of October of 2026]. Available at: https://hdl.handle.net/2445/231812

Exportar metadades

JSON - METS

Compartir registre