Large language models for personalized feedback and communicative skills in english learning: a systematic review

Abstract

Artificial intelligence is increasingly embedded in English language learning, but evidence on generative AI and large language models (LLMs) remains fragmented across skills, learner groups, and tool designs. This systematic review synthesized evidence on personalized learning, automated feedback, speaking, writing, vocabulary, assessment, English for specific purposes (ESP), and physical education. Reporting followed PRISMA 2020. Scopus, PubMed, and ScienceDirect returned 403 records; 14 duplicates were removed, 389 were screened, 54 reports were sought, 48 full texts were assessed, and 17 studies were included. The corpus spanned experimental, quasi-experimental, comparative, mixed-methods, qualitative, survey, review, and system-development designs. Where reported, participant samples included 40 undergraduates, 79 graduate students, and 327 primary pupils, while several design and qualitative studies did not report a single numeric sample. Narrative/thematic synthesis identified four themes: feedback and writing support; speaking and personalized tutoring; changing learner-teacher roles; and affective and ethical conditions. The evidence is recent and heterogeneous, limiting quantitative pooling. FICO was used as an internal appraisal rubric; reviewer-level logs were not retained, so Cohen’s kappa and aggregate FICO scores were not reconstructed retrospectively. Future research should prioritize controlled longitudinal designs, validated outcome measures, transparent reporting, data privacy, and under-represented ESP and sport-science populations.

How to Cite
Burhan, K., & Desman, M. A. (2026). Large language models for personalized feedback and communicative skills in english learning: a systematic review. Lentera Negeri, 7(1), 1337–1352. https://doi.org/10.29210/993470