INVESTIGATION OF ARTIFICIAL INTELLIGENCE SUPPORTED AUTOMATIC ASSESSMENT AND EVALUATION SYSTEMS IN SCIENCE EDUCATION IN TERMS OF QUALITY
DOI:
https://doi.org/10.22452/Keywords:
Artificial Intelligence, Automated Assessment, Validity and Reliability, Assessment Quality, Deberta-V3Abstract
This study examines the quality dimensions of artificial intelligence-based automated assessment systems in science education in terms of validity, reliability, and conceptual understanding. A DeBERTa-v3-based multiple-choice question answering model was trained using the SciQ dataset. The model was evaluated through accuracy analysis, justification-based semantic similarity, and comparative performance analysis. The findings indicate that the model achieved high accuracy (90.4%) and demonstrated strong alignment between selected answers and supporting texts. Beyond accuracy, the system shows potential to assess students’ conceptual understanding through justification-based analysis. The results suggest that AI-supported systems can enhance the quality of both formative and summative assessment practices by providing meaningful, explanation-based insights into student reasoning. However, considerations related to explainability, pedagogical alignment, and ethical use remain critical for effective classroom implementation. These findings contribute to bridging the gap between technical model performance and pedagogical assessment practices in science education. This study provides a novel contribution by integrating justification-based semantic analysis into AI-supported assessment, offering a multidimensional evaluation framework for science education.







