Like its sibling 2DCLIP, CLIP-MAP turns OpenAI’s multimodal CLIP model against itself

Description

CLIP-MAP compares images obtained from Google Street View to CLIP’s para-visual concept of that city.

Background

Leonardo Impett and Fabian Offert argue in their paper “There is a Digital Art History” that the advent of multimodal embedding networks ushers in a new era of digital art history, by “facilitat

\[ing\]

a completely new form of image retrieval … based on para-visual concepts of arbitrary complexity” (192–3). Moreover, the discipline of digital art history appears to be uniquely well-positioned to address the entanglement of model and dataset, or ‘way of seeing’ and object: “in the age of foundation models, we can never quite isolate the neural network from the object of study” (202).

Use

CLIP-MAP probes this entanglement of model and dataset by effectively turning it on itself. It features the urban geographies of six cities, sampling 10.000 images from each city through Google Street View. For each of these images, it plots the strength of its association to CLIP when prompted “A photo of <city>”; the resulting scores are overlaid as a heatmap on the city map. Click around the map to develop an idea of the vectorised imaginaries that CLIP builds its para-visual concepts out of.

Locations

  • Johannesburg

  • London

  • New York

  • Paris

  • Rome

  • Tokyo

    CLIP’s model of anything is always a vector imaginary, necessarily deeply culturally situated, linked to the image-economies of the internet: in this case, a visual imaginary of Paris linked both to tourism and to national identity (and state power). (Impett & Offert 2022, 197)

Technical

CLIP-MAP runs entirely in the browser—that is, locally—consuming minimal compute.

CLIP’s model of anything is always a vector imaginary, necessarily deeply culturally situated, linked to the image-economies of the internet: in this case, a visual imaginary of Paris linked both to tourism and to national identity (and state power).

CLIP has a concept, a vector imaginary, of Paris. A CLIP-search for “Paris” in the Museum of Modern Art collection, for instance, turns up images of the Tour Eiffel, of Notre Dame, and of generally “French-looking” streets. Paris, for CLIP, is bound up with a set of particular visual, we might say symbolic, properties; but how do they relate to the material reality of Paris as it exists today? Where in Paris looks (to CLIP) like Paris – and what does that tell us about CLIP?