The class in brief
Wouter's whole night traced AI bias to specific, findable human choices: who scraped what, who labeled it, who filtered it, whose taste rated the output. He proved it with his own project, Generated.Earth, 70,000 Stable Diffusion images mapped onto the globe, then handed the class his prompt-mapping method to run themselves. After this page you can trace any AI bias back to a specific layer, scraper, annotator, filter, or rater, and build your own prompt map to go find it.
The night at a glance
Why this matters · 6:17 PM
AI bias doesn't come from "the internet" or "society." It comes from traceable choices.
Wouter opened on his own diffusion-model explainer , then the trick version, prompting for pure Gaussian noise so the model produces abstract machine artworks instead of fighting the noise it can't find . Then the two myths went up together: gen-AI is trained on the entire internet, and gen-AI is a mirror of society. Both wrong, in specific and provable ways.
The framework · 6:35 PM
A free, open repository of web-crawl data, roughly 50 billion image-text pairs: Instagram posts, webshop listings, alt text in raw code .
OpenAI's algorithm, trained itself on about 400 million pairs from Wikipedia, Flickr, and ArtStation. Its only job: score how relevant a caption is to an image. Before CLIP, ImageNet did this by hand, 14.28 million pairs labeled by 50,000 contributors .
CLIP scores every pair from Common Crawl; the low scores get trashed. What's left is LAION-5B .
The one model that told the class exactly what it learned from, and admitted the set isn't fit for product use without safety layers on top.
Generated images flow back into the web, becoming tomorrow's scrape. Yesterday's gaze compounds into tomorrow's defaults.

Two more layers worth naming: the DW documentary on the human labor behind annotation , and the annotators' own language leaking into the model, the reason ChatGPT overuses "delve" , and, by the same mechanism, why AI overuses em dashes.
The exercise · 6:50 PM
70,000 Stable Diffusion images, one per geographic prompt, provinces and cities, never country names, stitched into continent-scale maps . Each class member got ten minutes solo on earth.wouteroomen.com, then five minutes discussing in groups of four.
"AI is a popularity test between concepts."
Stable Diffusion's 2022 data cutoff means Syria, at war before the cutoff, renders as rubble and destruction. Ukraine, at war after the cutoff, still renders as churches and public squares . News coverage is the dataset.
People only photograph what feels significant, so seas get photographic identities of their own: the Red Sea renders as diving fish and beaches, traced straight back to Flickr's diving-photo gaze .
With no popular concept to grab onto, the model falls to whatever junk data exists. Wouter's example: Magic: The Gathering card wikis "poisoning" low-data regions with fantasy-card imagery.
The craft · 7:36 PM
A raw model just continues your sentence. A product has been fine-tuned to sound like a chatbot.
Part 3 opened with the question on screen , then the three phases: gather training data, build the base model as a text predictor, fine-tune it into a chatbot with thumbs up and down. The live demo made it concrete: Llama 3 8B, raw off Hugging Face, just kept typing whatever came next; ChatGPT, given the same input, replied with an emoji and asked if it was helping .

The judgment · 7:47 PM
An LLM is always limited by its data, even when the guardrails are current.
Talkie is a 13-billion-parameter model trained only on text from before 1930 . Asked what careers suit women, it answered in period: governess and shopkeeper suitable, attorney and medicine not, public speaking would "render them coarse and masculine" . Talkie's own modern moderation layer flagged the answer as potentially inappropriate before showing it, 1930 values sitting directly underneath 2026 guardrails. Asked to brainstorm climate solutions, it suggested moving somewhere else: no concept of climate change existed in its training data at all.
"An LLM is always limited by its data."
Wouter, on what Talkie proves
Methods and prompts
Wouter's core method: take one big concept, split it into 16 to 25 flat prompts, generate a grid, then read the grid for what it reveals. A raw image model on purpose, since a de-biasing layer adds its own bias.
Working prompt
Give me 16 flat, specific prompts that all sit inside one bigger concept: [name the concept, not a geography]. Each one should be a plain noun phrase, no adjectives that editorialize, ready to drop straight into an image generator as a 4x4 grid.
You will know it worked wheneach of the sixteen prompts is a plain noun phrase with no adjective that editorializes, ready to drop straight into a 4x4 grid.
Name your own guess for where a bias entered, scraper, annotator, CLIP filter, or aesthetic rater, before asking AI to check it. The layer matters more than the symptom.
Working prompt
Here's a bias I noticed in this output: [describe it]. My guess at which layer caused it (what got scraped, who annotated it, how CLIP filtered it, or who rated the aesthetics): [your answer]. Now you check me: is my guess right, and which layer is actually the more likely source?
You will know it worked whenit names one specific layer, from scraping to annotation to filtering to rating, as the more likely source, and checks whether your own guess was right.
Find a raw, un-fine-tuned model (Llama 3 8B on Hugging Face is free) and feed it a sentence to complete, no chat framing. See the cake before the sprinkles.
Working prompt
[Typed into a base model, not a chat product] I like to [start a plain, unfinished sentence about something you actually do]... [Let it complete the sentence with no further instruction, and compare the raw continuation to how a chat product like ChatGPT would answer the same prompt.]
You will know it worked whenthe raw continuation reads differently in tone and content from how a chat product would answer the same starting sentence.
The same subject, asked for in a different language, can return a completely different bias. Prompts are made by people too, in a specific tongue with a specific gaze.
Working prompt
Generate an image of [subject]. Now generate the exact same prompt translated into [a language other than the one you usually prompt in]. Compare the two results: what changed in tone, color, dignity, or setting, and what does that say about which internet each language draws from?
You will know it worked whenthe two images differ in tone, color, dignity, or setting, not just cosmetically, between the language you usually use and the other one.
Guess the model's cultural default on a values question before you ask it. The WEIRD imprint (Western, Educated, Industrialized, Rich, Democratic) is stronger than any persona you feed it, so check for it directly.
Working prompt
Answer this as neutrally as you can: [ask a values or moral-judgment question]. Before I read your answer, my prediction is that you'll default to a [Western / individualist / analytical, your guess] framing. Now show me your answer, and tell me honestly whether my prediction was right.
You will know it worked whenit tells you plainly whether its own answer actually leaned the way you predicted, Western, individualist, or analytical, not just a neutral-sounding answer.
Where it broke · 7:57 PM onward
The prompt-mapping workshop put Wouter's tool in the class's own hands, and the misses were as instructive as the hits. Beverages: "grog" returned nothing at all, a concept with no image behind it; milk rendered as the word "milk," not a glass of it; beer couldn't decide between a bottle and a glass. Weather broke on poetic words: "drizzle" and "hail" had too little image data to hold together, while "Blizzard" reliably pulled the video game's logo instead of the storm. One holiday map, built and shared live in class , showed the same pattern from a different angle: some holidays rendered as confident clichés, others barely rendered at all. And one open question never got solved: the Generated.Earth map keeps rendering a specific region of Iceland as cartoon trolls, a pattern even Icelanders in the room couldn't source.
The shelf
99 captures, in order. Click any one to see it full size.


































































































