vlms see better when squinting, kinda
ever looked at a pixelated image and couldn't recognize it, then stepped back and suddenly could see it? it seems like vlms follow this same pattern.
ai models can't back up to look at an image, but you can inset the object in the image so it is "smaller" i guess. i put 100 album covers pixellated and increasingly inset in a 1024x1024 image, then asked a vllm to guess what album it is. here's how the model accuracy degrades on 100 different 95x95 covers (4 shown).
"what album cover is this?":

i don't know why they fail so early; it's strange. fwiw, control case with full 1024x1024 has 98% accuracy, so i don't think something funny is happening here.
![]()
humans seem to fail much later. take this fleetwood mac image. if i open this on my laptop screen full page, my own ability to recognize the image is inconsistent 16x16px row. but the model gets confused at 95x95!