Abstract
Infrared small target detection (IRSTD) aims to identify dim, tiny targets with extremely low signal-to-noise ratios in complex infrared scenes, where targets are often submerged in heavy clutter and exhibit weak, ambiguous visual characteristics. These challenges are further exacerbated by background noise, low contrast, and limited structural information, making reliable detection particularly difficult. Although recent methods have achieved promising progress, most rely solely on the visual modality, which restricts their ability to capture high-level semantic cues and limits their robustness and generalization in complex environments. To address these limitations, we propose a novel Vision-Language Guided Infrared Detection Network (VLGNet), which introduces auxiliary textual semantics to enhance feature representation for infrared small target detection. By leveraging complementary information from the language modality, the proposed framework compensates for the inherent deficiencies of visual features in infrared imagery. Specifically, we design a Vision-Language Feature Pyramid Network (VL-FPN) to perform multiscale cross-modal feature fusion, enabling hierarchical interaction between visual features and text embeddings. A Cross-level Context Attention Module (CCAM) is further introduced to capture long-range dependencies and strengthen semantic consistency across feature levels, thereby improving the coherence of multimodal representations. In addition, we develop a Spatial-guided Decoupled Refinement Block (SDRB) to dynamically fuse target-associated frequency cues while enhancing small-target energy and suppressing background interference. Furthermore, we construct a multimodal IRSTD benchmark by augmenting three widely used datasets (NUDT-SIRST, SIRST, and IRSTD-1K) with descriptive textual annotations, which provide richer semantic supervision for both training and evaluation. Extensive experiments demonstrate that the proposed method achieves superior detection accuracy and significantly reduces false alarms compared with state-of-the-art approaches. Notably, VLGNet exhibits strong robustness in complex and low-contrast scenarios, validating the effectiveness of integrating vision-language information for infrared small target detection and highlighting its potential for real-world applications.