Journal of English for Specific Purposes Praxis

Journal of English for Specific Purposes Praxis

A Comparative Error Analysis of Iranian EFL Learners’ Writing Compositions by Human Evaluators Vs. Perplexity AI Platform

Document Type : Original Article

Authors
1 Department of English Language and Literature, Faculty of Persian Literature and Foreign Languages, Allameh Tabataba’i University, Tehran, Iran
2 Department of English language, Faculty of Humanities, Imam Khomeini International University, Qazvin, Iran
Abstract
Abstract
The advent of the AI as a supplementary tool for error analysis and how it is different from human error analysis seems to be an underexplored and enchanting research area. This study sought to examine the types of errors found in the written compositions of 16 intermediate-level Iranian students— male and female learners included—selected through convenience sampling whose data was gleaned from two language academies in Rasht. All participants were using American English File 3 (2nd Edition) serving as their main course book, with the researchers also being their instructor. The students were tasked with writing a response to a letter, based on a model provided in their textbook. A total of 16 writing samples were collected and analyzed using both qualitative and quantitative methods, guided by Keshavarz’s (2013) linguistic error classification framework. The evaluation was conducted by a human rater and the AI tool Perplexity, which was specifically prompted to identify errors according to the same classification system. The results revealed a range of error types with varying frequencies. Morphosyntactic errors tended to be the most common, followed by orthographic, lexicosemantic, and phonological errors, respectively. Moreover, Inter-rater reliability was calculated via Cohen’s Kappa, which indicated substantial level of agreement between human and AI raters (κ = 0.78, p < 0.001). Teachers can combine the precision and consistency of AI with the subjective and interpretive depth of human assessment to provide more responsive feedback that supports linguistic accuracy and learner agency. Ultimately, this study advocates for integrative approaches to language assessment to use technological development under the supervision of the human agents. Future research can utilize different learner characteristics (e.g., different proficiency and cultural levels, with more varied written assignments) in different cultures, compared with other AI tools, to study the longitudinal consistency of the AI and human raters under different situations.
Keywords

Article Title Persian

تحلیل مقایسه‌ای خطاهای نگارشی زبان‌آموزان ایرانی زبان انگلیسی توسط ارزیابان انسانی در مقایسه با پلتفرم هوش مصنوعی پرپلکسیتی

Authors Persian

امید استاد 1
هادی حیدری 2
1 گروه زبان و ادبیات انگلیسی، دانشکده ادبیات فارسی و زبان‌های خارجی، دانشگاه علامه طباطبایی، تهران، ایران
2 گروه زبان انگلیسی، دانشکده علوم انسانی، دانشگاه بین الملل امام خمینی (ره)
Abstract Persian

هدف این مطالعه بررسی خطاهای نگارشی ۱۶ دانش‌آموز ایرانی، شامل دختر و پسر در سطح متوسط، بود که از دو مؤسسه زبان در شهر رشت انتخاب شدند. این دانش‌آموزان کتاب American English File 3 را به عنوان منبع درسی خود مطالعه می‌کردند و پژوهشگران نیز مدرسین آن‌ها بودند. به آن‌ها مدلی از کتاب درسی داده شد تا پاسخی به یک نامه بنویسند که خود، پاسخ به نامه‌ای دیگر بود. در مجموع ۱۶ نمونه جمع‌آوری شد و این نمونه‌ها به صورت کیفی و کمی بر اساس طبقه‌بندی زبان‌شناختی خطاها معرفی شده توسط دکتر کشاورز (۲۰۱۳) تحلیل شدند. این تحلیل توسط انسان و همچنین ابزار هوش مصنوعی Perplexity انجام شد.

یافته‌ها نشان دادند که دانش‌آموزان انواع مختلفی از خطاها را با درجات متفاوتی از فراوانی مرتکب شدند. خطاهای صرفی-نحوی بیشترین فراوانی را داشتند و پس از آن خطاهای املایی، واژگانی-معنایی و آوایی قرار گرفتند. همچنین، سطح بالایی از همخوانی در نمره‌دهی بین ارزیاب انسانی و Perplexity مشاهده شد (آلفا = ۰٫۹۴). معلمان می‌توانند دقت و ثبات هوش مصنوعی را با عمق تفسیری و ذهنی ارزیابی انسانی ترکیب کنند تا بازخوردی پاسخ‌گوتر ارائه دهند که دقت زبانی و عاملیت یادگیرنده را تقویت کند.

در نهایت، این مطالعه از رویکردهای تلفیقی در ارزیابی زبان حمایت می‌کند تا از توسعه فناوری تحت نظارت عوامل انسانی بهره‌برداری شود. پژوهش‌های آینده می‌توانند ثبات طولی ارزیابی‌های انسانی و هوش مصنوعی را در زمینه‌های اجتماعی-فرهنگی مختلف و در سطوح گوناگون زبانی بررسی کنند.

Keywords Persian

هوش مصنوعی
تحلیل خطا
ارزیابان انسانی
تداخل زبان اول
خطاهای صرفی-نحوی
پرپلکسیتی
انشاهای نوشتاری