Abstract
Animal pose estimation typically requires manual annotation of body parts to train neural networks for each new recording system or environmental condition. Here we use a general-purpose vision-language model (VLM) to generate annotations in multi-animal recordings of fruit flies (Drosophila melanogaster). Initial annotations follow written anatomical descriptions and are reviewed by an independent VLM session using neighboring video frames. On 120 held-out fly images, AI annotations had a median distance of 2.28 pixels from human labels, with 84.0% agreement within eight pixels, averaged across landmark types. The comparison included 10 landmarks, with head and eyes analyzed separately because of systematic differences in placement. Review improved agreement primarily at the feet, with the largest remaining discrepancies at positions with limited visual evidence. Visual inspection revealed cases in which the AI annotations were more anatomically accurate than those of an experienced human annotator, particularly at the feet. A convolutional network trained using SLEAP on 300 AI-annotated fly images showed similar agreement with human labels and closely reproduced the AI annotations on held-out images, with a median distance of 1.56 pixels. We also applied the procedure to recordings from another laboratory, using new AI annotations to train a separate pose network. Although VLM annotation required minutes per fly, the trained networks enabled pose estimation throughout more than 700,000 frames. These results demonstrate that VLM annotations can support fast animal pose estimation without manually labeled training images.