命令行考试的可扩展且可靠的评分仍然是计算教育, 中的一个挑战,其中入学人数的增加使得手动评分变得困难,并且基于规则的自动评分器无法处理部分学分, 等效解决方案, 或语法变化。本文评估四种前沿大型语言模型 (GPT, Claude Opus, Gemini, 和 GLM) 在对短 Linux/bash 命令响应进行评分时是否可以近似专家判断。该研究采用四级认知分类法,结合了认知复杂性和操作影响,,范围从信息检索(L1)和基本文件操作(L2)到结构操作(L3)和高级系统管理(L4)。这些模型使用两个提示变体,、最小基线和评分增强版本, 对来自二年级计算机工程专业学生的 1200 个真实答案进行了测试,并由三位专家讲师独立评分。采用标题引导提示的 Gemini~3.0 Pro 实现了最高的人类 AI 一致性 (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman 偏差 = -0.014)。随着分类级别增加,,一致性持续下降,较高级别的差异最大。在所有模型中, 标题质量比提供商选择, 具有更大的影响,结构化提示不断提高一致性。这些结果表明,问题复杂性是法学硕士在准确评分, 时面临的难度的可靠预测因素,他们建立了一个原则性的, 基于分类的框架,用于确定哪些问题适合人工智能辅助评分,哪些问题需要人工审核,,同时还提供可转移的评估协议和提示模板。
Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1) and basic file manipulation (L2) to structural operations (L3) and advanced system management (L4). The models were tested with two prompt variants, a minimal baseline and a rubric-enhanced version, on 1200 real responses from second-year Computer Engineering students independently graded by three expert instructors. Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014). Agreement declined consistently as taxonomy level increased, with the largest discrepancies at higher levels. Across all models, rubric quality had a larger effect than provider choice, with structured prompts consistently improving agreement. These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted grading and which require human review, while also providing a transferable evaluation protocol and prompt templates.
科目: 人工智能 (cs.AI); 计算与语言 (cs.CL); 计算机与社会 (cs.CY)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)