生成式人工智能越来越多地支持教育设计任务,,例如,通过大型语言模型(LLMs),展示了设计与教学框架(一致的评估问题的能力(例如, Bloom的分类)。然而, 他们经常依赖于主观或有限的评估方法; 主要关注专有模型; 或很少系统地检查一代, 评估, 或实际教育环境中的部署约束。与此同时,, 小语言模型 (SLMs) 已成为本地替代方案,可以更好地解决隐私和资源限制;,但其评估任务的有效性仍未得到充分探索。为了解决这一差距,,我们系统地比较LLM和SLM以进行评估问题设计;,使用可重复的,教学基础指标;评估Bloom的分类水平的生成质量,并通过分析可靠性和协议模式,根据专家知情的评估进一步评估基于模型的判断。结果表明,SLM 在关键的教学驱动的质量维度上实现了具有竞争力的性能,同时支持本地, 隐私敏感部署。然而,, 基于模型的评估也表现出与专家评级相关的系统不一致和偏差。这些发现提供了证据,将语言模型视为评估工作流程中的有界助手;,强调了人在环; 的必要性,并通过检查质量, 可靠性, 和部署感知权衡来推进自动化教育问题生成领域。
Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom的 taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom的 taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.
科目: 人工智能 (cs.AI); 计算与语言 (cs.CL); 人机交互 (cs.HC)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)