视觉语言模型 (VLMs) 的最新进展在许多任务中取得了令人印象深刻的性能,,但之前的研究报告称,在应用大型语言或多模态模型来查找序列数据中的异常模式时,性能并不令人满意。公共异常检测基准通常提供区间注释,但不提供自然语言原理,,因此很难微调 VLM 以生成有依据的, 可解释决策。为了解决这个差距,,我们构建了VisAnomBench,,这是一个根据公共时间序列数据集构建的策划基准,并使用细粒度,特定任务奖励从多个大型VLM中选择的高质量异常解释进行了增强。通过在此基准, 上进行微调,我们开发了 VisAnomReasoner, 一个用于时间序列异常检测的参数高效 VLM。 VisAnomBench 上的实验结果表明,VisAnomReasoner 实现了更准确的异常定位,并且始终优于所有基线,,精度和 F1, 分别提高了至少 21.23 和 23.87 个百分点。 TSB-AD-U 基准的其他实验表明,VisAnomReasoner 具有强大的跨基准泛化,,精度和 F1 分别提高了 9.57 和 13.39 个百分点,。
Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory performance when applying large language or multimodal models to finding abnormal patterns in sequential data. Public anomaly detection benchmarks typically provide interval annotations but not natural-language rationales, making it difficult to fine-tune VLMs to produce grounded, interpretable decisions. To address this gap, we construct VisAnomBench, a curated benchmark built from public time-series datasets and augmented with high-quality anomaly explanations selected from multiple large VLMs using fine-grained, task-specific rewards. Through fine-tuning on this benchmark, we develop VisAnomReasoner, a parameter-efficient VLM for time-series anomaly detection. Experimental results on VisAnomBench show that VisAnomReasoner achieves more accurate anomaly localization and consistently outperforms all baselines, with improvements of at least 21.23 and 23.87 percentage points in precision and F1, respectively. Additional experiments on the TSB-AD-U benchmark demonstrate strong cross-benchmark generalization, with VisAnomReasoner improving precision and F1 by 9.57 and 13.39 percentage points, respectively.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)