Introduction to
Computer Vision & AI

Artificial Nouveau

Outline

1. Foundations : Who Am I / What is AI / How CV Works

2. Images & Consent : Scraping, Datasets & "Is My Face In There?"

3. Language & Power : Captions, Labels & the Feedback Loop

4. Surveillance & Imagination : Control, Beauty & Global Perspectives

5. Failure & Resistance : When Machines Break & Artists Push Back

Foundations

Who Am I

Ex-Academic

Focusing on Medical Surveillance

AI Engineer (Computer Vision)

AI Engineer (Media Researcher)

Computational Artist

What is AI

What is AI

Artificial intelligence is the name of a whole knowledge field in which machines perform tasks that mimic human intelligence.

Why AI

Why AI Now

Introduction to Computer Vision

Computer Vision Prep

Feature Extraction

CV extracts information such as edges, textures, colors, etc.

Training the AI

Model → Predictions → Mario Probability / Luigi Probability

CV Applications: Seeing

Object Recognition

Identifying and locating objects in images

Pose Estimation

Tracking body position and movement

Self-Driving Cars

Navigating roads, detecting pedestrians

CV Applications: Describing

Image Captioning

Generating text descriptions of images

Emotion Detection

Classifying facial expressions into emotional states

Medical Imaging

Detecting disease in scans and pathology

CV Applications: Judging

Facial Recognition

Identifying or verifying people by their face

Content Moderation

Deciding what is acceptable to show

Deepfake Detection

Deciding what is real and what is fabricated

Images & Consent

Where Do the Images Come From?

The Scale Problem

Modern CV models need millions to billions of images to train. No one photographs that many consenting subjects. Instead, images are scraped from the open internet without consent.

Where They Scrape From

  • Flickr (Creative Commons loophole)
  • Social media (Instagram, Facebook, X)
  • News sites and press archives
  • Personal websites and blogs
  • CCTV and webcam feeds

Scraping at Scale

Clearview AI

Scraped 30+ billion photos from Facebook, Instagram, LinkedIn, and news sites. Built a facial recognition tool and sold it to law enforcement, without a single person's consent.

Fined in the EU, Australia, and Canada. Still operating in the US.

LAION-5B

5.8 billion image-text pairs scraped from the internet. Used to train Stable Diffusion and other generative AI.

In 2023, researchers found CSAM (child sexual abuse material) in the dataset, including images of children scraped without parental knowledge. LAION was forced to take it offline.

The Weight of Seeing

CV models are small

A MobileNet is 17MB. CLIP is 400MB. Compare that to GPT-4 at hundreds of gigabytes.

Training a CV model uses a fraction of the energy of training a large language model.

But deployment is massive

The cost isn't in one model. It's in millions of cameras running 24/7, data centers storing billions of images, and real-time processing at global scale.

The cloud is not weightless. It's buildings full of machines that consume electricity and water.

Is Your Face in a Dataset?

If you have ever posted a photo online, on social media, a personal website, a news article, a university page, your face may already be in a training dataset.

You were never asked. You were never told. And in most countries, you have no legal right to have it removed.

We'll find out in the workshop using pimeyes.com

Language & Power

Whose Language?

"A woman posing for a picture"

Their vocabulary is inherited from millions of image-caption pairs scraped from the internet. The biases, assumptions, and blind spots of those captions become the model's worldview.

The Caption is the Training Data

CLIP (OpenAI, 2021)

Pulls images and text into the same mathematical space. Trained on 400 million image-caption pairs scraped from the internet. Most captions come from HTML alt-text, written for SEO, not for teaching AI.

What the machine learned from

alt="": empty, invisible
alt="IMG_0392.jpg": meaningless
alt="exotic woman": stereotyping
alt="normal family": whose normal?
alt="scientist at work": rare and specific

Labels Are Not Neutral

ImageNet's "Person" Categories

ImageNet, the dataset that powered the deep learning revolution, contained 2,832 subcategories under "person." Many were racial slurs, sexual orientations, and derogatory terms.

The act of labeling is an act of power

  • Who decides the categories?
  • What gets collapsed into a single label?
  • What is made invisible by the taxonomy?

Same Image, Different Captions

One photograph. Three captions. Three different machines.

"A woman walking home at night"

"A suspicious person in a dark alley"

"A model in an urban photoshoot"

The Feedback Loop

AI generates captions

Captions get posted online

Future AI trains on those captions

The machine's language becomes self-reinforcing

Those who control the data control the narrative

They control Representation

What gets collected and what is ignored?

They control The Language

Define the labels, categories, and the taxonomies

They control The Outcome

They control what gets optimized. Who and what gets to use the data.

A meaningful data is a big dataset. Once datasets meet a certain scale, they become harder to audit, and therefore have no opacity.

The Invisible Workers

Who labels the data?

Every labeled image was labeled by a person. Outsourced through Amazon Mechanical Turk, Sama, Scale AI, Appen, and dozens of smaller contractors to workers in Nairobi, Manila, Hyderabad, Caracas, and Accra.

Often less than $2/hour to categorize violence, abuse, and hate so the model knows what to filter.

You've done it too

Every time you solve a CAPTCHA ("select all traffic lights"), you are labeling training data for free. Google's reCAPTCHA was used to digitize books and train self-driving car models.

Surveillance & Imagination

Whose Imagination Built This System?

Ruha Benjamin, Imagination: A Manifesto (2024)

  • Imagination is not a luxury, it is a contested resource. Those who control how we imagine the future control who gets to live in it.
  • When AI systems encode only the imaginations of their creators, they foreclose alternative futures before they can be articulated.

Every CV dataset encodes an imagination of what the world looks like.

Surveillance as World Building

Data Collection

Czech artist Jakub Geltner

Machine Vision as Frozen Imagination

Violence

Frozen Imagination is Profitable

  • 1.5 million Uyghurs and Turkic Muslims
  • European & American Predictive Policing Software

Profit in 3 ways:

  • State contracts
  • Advanced Software
  • Cheap Labour

Same Tool, Different Context

Object detection counts visitors in a gallery

The same model identifies targets for a drone strike

Facial recognition unlocks your phone

The same model identifies protesters at a demonstration

A portrait in a gallery vs the same portrait in a police database.
The image is identical. The power relation is completely different.

CV Surveillance is Global

India: Aadhaar

The world's largest biometric database, 1.4 billion people enrolled. Iris scans, fingerprints, and facial data linked to identity, banking, and welfare.

Exclusion errors have locked millions out of food rations and pensions. The system fails most for those who need it most.

Hong Kong: Counter-Surveillance

During the 2019–2020 protests, demonstrators developed real-time counter-surveillance: umbrellas over faces, laser pointers to blind cameras, and coordinated destruction of smart lampposts.

Resistance to CV is itself a form of political expression.

Frozen Imagination: Health

Self-Surveillance & Beauty

Training a Cosmetic AI: World Building

Criteria

  • No Textured Skin
  • Lighter Skin
  • Smaller nose
  • Bigger Eyes
  • Slimmer or Wider Chin
  • Fuller Lips

Training Photos

  • Celebrities
  • Pre/Post Surgeries

Training a Cosmetic AI

Frozen Imagination: The Algorithmic Ideal Face

Living in the Mirror World

Failure & Resistance

When Computers Fail

Is AI better than pigeons?

Is AI better than pigeons?

When Computer Vision Fails

When Computer Vision Fails

Grok

Who Becomes Invisible?

Invisibility as danger

If the system can't see you, you don't exist to it. If that system controls access, mobility, or safety, invisibility becomes danger.

The error rate isn't evenly distributed. It clusters around the people the training data didn't include.

When Machines Accuse

Robert Williams (Detroit, 2020): arrested at home in front of his family. Facial recognition matched him to a shoplifting suspect. Held for 30 hours. The match was wrong.

Porcha Woodruff (Detroit, 2023): arrested while 8 months pregnant on a carjacking charge. Detained for 11 hours. Had contractions from stress. The match was wrong.

Netherlands (2022): Dutch Data Protection Authority fined Clearview AI €20 million for illegally scraping Dutch citizens' faces. Dutch police use the CATCH facial recognition system, but accuracy disparities across racial groups remain undisclosed.

In the US cases: no human verified the match before arrest. The algorithm was treated as evidence. Nobody was held accountable.

Creative Responses

CV Dazzle

CV Dazzle

Toko Kihara

How Not to Get Hit By A Self-Driving Car

Watch on Vimeo if embed is unavailable

Toko Kihara: Is this Violence? Am I too sexy?

Dries Depoorter

The Follower, 2023–2026

Capture

Paolo Cirio

Trevor Paglen & Kate Crawford

ImageNet Roulette, 2019

A web app that classified visitors' faces using ImageNet's "person" categories, exposing labels like "rape suspect", "alcoholic", and racial slurs that the dataset had quietly been using for a decade.

  • Went viral, millions of people classified themselves
  • Forced ImageNet to remove 600,000+ images from the "person" subtree
  • Proved that making the system visible is itself a form of resistance

Conclusion

What We Covered

Machines see math
Pixels, matrices, probability. Not meaning.

Images are taken without consent
Scraped at scale, with real environmental and human cost.

Language shapes vision
Captions become training data, labeled by invisible, underpaid workers.

The same tool serves different masters
A gallery counter or a military targeting system.

Failure is unevenly distributed
People are arrested. Nobody is accountable.

Artists are pushing back
Making the invisible visible, refusing the default.

What Can You Do?

As a Citizen

  • Search yourself: use PimEyes, HaveIBeenTrained to check where your face appears
  • Opt out: request removal from datasets where possible (GDPR, CCPA)
  • Advocate: support legislation that requires consent for biometric data

As a Creative

  • Caption with care: your alt-text and metadata become training data
  • Question the tools: where did the training data come from?
  • Create counter-narratives: use your work to make the invisible visible
  • Refuse the default: whose gaze are you reproducing?

Questions?

Thank You

Artificial Nouveau

Workshop

Try It Yourself

How Does the Machine See Your Photo?

tinyurl.com/workcv00

1. Pixels

See your photo as the machine does: pixel grids, colour channels, edge detection, and classification heatmaps.

2. The Machine Speaks

AI captions your photo and classifies it using categories you define: emotions, gender, beauty, age.

3. Who Is Watching?

Live webcam surveillance: face detection, demographics, pose, objects, NSFW scanning, all at once.

4. Mirror

See how the machine maps your face: 468 landmarks, symmetry, and proportion analysis.

5. Clone

Upload two faces and measure how similar the machine thinks they are.

6. Prompt the Machine

Ask Google's Gemini anything about your photo. See how framing changes the answer.