Abstract
Abstract
Walking, as a form of active and sustainable mobility, plays a critical role in future smart transportation systems. Accurate prediction of pedestrian crossing intentions is essential for preventing collisions, particularly with the increasing deployment of autonomous vehicles. Existing approaches to near-miss prevention typically rely on computationally intensive computer vision and deep learning techniques. In contrast, this work proposes
CSE
, a lightweight contextual stacked ensemble-learning framework to efficiently predict pedestrian crossing intent. Pedestrians are first detected and their visual representation is compressed through skeletonization, and complementary pose, trajectory, and contextual cues are fused using a stacked ensemble model. Unlike prior approaches that fuse all features into a single unified representation, CSE employs a modality-decomposed strategy in which each feature stream is processed independently and integrated via a lightweight meta-classifier. This design preserves complementary decision cues across modalities. Experimental results on multiple datasets demonstrate that the proposed approach achieves performance comparable to state-of-the-art pedestrian intent prediction methods while reducing computational complexity by at least
$$25\times $$
in FLOPs and trainable parameters compared to the most efficient existing baseline. This reduction enables deployment on resource-constrained edge devices without compromising accuracy and while avoiding the latency associated with cloud-based processing.