自动口语评估系统越来越多地部署在高风险环境中,以标记第二语言 (L2) 学习者 口语测试,,因此表明他们的分数取决于口语熟练程度而不是不相关的说话者属性(例如第一语言 (L1) 或年龄)至关重要。基于 Transformer 的基础模型提高了这些 L2 口语评分者, 的准确性,但其黑盒表示使公平性和可解释性分析变得更加困难。在之前的工作基础上,我们使用概念激活向量 (CAVs) 来检测基于特征的评分器, 中对不需要的属性 (`concepts') 的偏差,,我们将基于 CAV 的分析扩展到两个神经口语评估系统:、一个基于文本的 BERT 评分器和一个基于 Whisper 的语音和文本多模态评分器。 CAV 将人类可解释的概念表示为模型的 激活空间, 中的方向,使我们能够区分概念是否在模型的 内部表示中编码,以及它是否影响预测分数,,后者使用基于梯度的灵敏度度量进行量化。由于 CAV 依赖于线性可分离性,,这在复杂的神经嵌入空间, 中不太可能出现,因此我们还研究稀疏自动编码器 (SAE) 是否通过在稀疏潜在空间中学习 CAV 并将其映射回激活空间来提供更清晰的概念方向。我们的分析表明,概念的可恢复性在很大程度上取决于所探测的表示和体系结构,,而不是仅仅取决于概念。对概念的敏感性也取决于架构。 SAE 使概念更加线性可恢复,,但削弱了原始激活空间敏感性,,尤其是在低维层中。这些发现强调了在审计口语评估系统中的偏差时区分概念可恢复性和概念影响的必要性。
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model的 activation space, allowing us to distinguish between whether a concept is encoded in a model的 internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)