Abstract
Abstract
The icon of Leonardo da Vinci’s
Mona Lisa
(1503-1519) is an image that, for centuries, has been termed “the enigmatic smile,” or the smile that is “there but not there.” - it is sometimes said to smile when you look away from it, but not when you look at it. Neurobiologist Margaret Livingstone suggested in 2000 that this visual ambiguity is due to the differential spatial frequency processing of human vision: a smile can be seen by the coarse magnocellular visual pathway but is not visible when the smile is reconstructed from fine (foveal) parvocellular spatial frequencies. Previous computational studies of the portrait have sought to assess monolithic facial expression recognition (FER) models that are tested using static un-manipulated reproductions of a portrait; such studies have been carried out to classify the facial expressions of a portrait taking a single classification without considering the different dynamics and scales of the visual presentation of the portrait. In this case, we release to the first systematic computational testing of Livingstone’s hypothesis with up-to-date self-attention Vision Transformers (ViTs) and frequency domain decomposition. We apply 2D Butterworth bandpass filtering (0–4, 4–8, 8–16, 16–32, and 32–64 cycles per face-width, cpf) alongside a cumulative low-pass series (2
,
4
,
8
,
16
,
32
,
64
,
128 cpf) simulating progressive viewing distance. We find that our empirical results are in line with the spatial frequency hypothesis, with the probability of detection of happiness being more than twice as high in the peripheral/distance regime (14
.
06% for 4 cpf, 13
.
51% for 8 cpf, and 13
.
00% for 16 cpf) as compared to the high-resolution foveal regime, where the probability of happiness detection sharply drops by 54% to 6
.
44% for 32 cpf, while the probability of detecting neutral faces increases from 0% to 57
.
34%. Beyond this, outputs from the full-spectrum analysis show that, at the unmanipulated level, the portrait is overwhelmingly classified as both neutral (55%) and sad (26%); when its spectrum is manipulated by using frequency bandpass filtering, however, it reveals a latent expression of affective ambiguity among anger, surprise and happiness. We also find a localised peak in happiness in the considerable range of 16–32 cpf, that we ascribe to the combination of patch-level self-attention and the sparse gradient boundaries found in Leonardo’s
sfumato
micro-glazes. The results provide a reliable computational connection between psychophysics, visual neurobiology, and computer vision in cultural heritage based on a transformer network.