Abstract
Multimodal traffic anomaly detection is affected by differences in visual and audio score ranges, temporal fluctuations, and unstable decision thresholds. This paper proposes a calibration-guided score fusion (CGSF) framework that processes video frames and audio spectrograms through separate reconstruction-based models. The resulting anomaly scores are temporally smoothed, normalized using validation data, and combined at the score level. A percentile estimated from normal validation samples is then used as the decision threshold. The framework was evaluated on the MAVD and DADA2000 datasets. On MAVD, CGSF achiev,,,,,,ed a ROC-AUC of 0.553, a PR-AUC of 0.082, and an F1-score of 0.129. It outperformed direct fusion in precision, recall, and F1-score, although the gain in ROC-AUC was small. Analysis on DADA2000 showed smoother temporal score behaviour after calibration and smoothing. The results indicate that CGSF mainly improves score comparability and threshold consistency rather than producing a large increase in detection accuracy. Its modular design also allows the visual and audio branches to be trained and updated independently.