← All classes
CLASS 3June 2 · Critical Intelligence + Krea

"The way we work with AImatters as much as how smart it is."

Geoff Gibbins
Guest lectureGeoff GibbinsFounder, Corrix; author, Critical Intelligence
Hosted by Dom Heinrich and Tony Jones

The class in brief

First guest of the semester. Geoff Gibbins brought Corrix, his research product on how humans and AI actually collaborate, and the STOP framework for evaluating AI output before trusting it. The class opened by reviewing the previous night's impossible self-portrait homework, then closed with Tony onboarding everyone into Krea, the semester's primary image and video tool, in place of a co-guest from Krea who didn't end up presenting. After this page you can run STOP on any AI output before you read it, and choose approving versus supervising mode by what's actually at stake.

The night at a glance

The exercise, reviewed · 6:29 PM

Describing the self-portrait homework in words only

The brief

Describe yourself to an image model, no photo upload, no sketch.

Dom named the exercise's real point on the spot: it isn't to make a perfect picture, it's to teach you how to use your words, and to see how the machine reacts to them. Getting exactly right is impossible by design.

Homework from Class 2, reviewed live

Fresh chats beat one long refinement thread

One classmate explained a technique that Dom praised and named on the spot: starting a brand new chat for each attempt instead of refining a single thread, because a continuous chat kept reverting to earlier versions and cherry-picking old details until the result looked more like a collage than a photo. Dom called it branching, or forking, a cutting taken from the main plant, and flagged it as the exact skill Sydney would build on in a later class about art direction and consistency.

Speak the language of the internet

"Long layers" undersells what a model actually understands; "butterfly cut" is the tagged, zeitgeist term the training data was built on. The same swap turned "cool tone" into "soft winter." Vocabulary unlocks surfaced live in the room too: quiff for hair styled up in front, blaze for a dog's white face-stripe, epicanthic fold for East Asian eye shape, terms nobody had reached for until they were named out loud.

The bias the exercise couldn't dodge

Describing yourself in words forces the model toward ethnicity defaults, and it performs best, by the room's own account, for a white man. One student's unprompted output defaulted to a cisgender white male description before any of that was specified. Nobody had a fix for it that night, only the observation, which is exactly the kind of blind spot Wouter Oomen's later class digs into.

The framework · 7:03 PM

Three ways Corrix measures the collaboration

01

Resultsdo you beat the AI alone?

The baseline question: working with the model, do you outperform what it would have produced on its own. Across roughly 1,000 people assessed, 82% did.

02

Relationshipthe quality of the dialogue

Are you adding real context, pushing back, challenging the output, or just accepting the first draft. This axis divided people more than any tool choice did.

03

Resilienceskills, or cognitive atrophy

Do you still understand what the model is doing, or are you losing the underlying skill by delegating it away. Even early adopters with three-plus years of daily use showed their scores quietly regress.

Corrix's three measurement buckets
Corrix's three measurement buckets · 7:03 PM

The craft · 7:11 PM

Which archetype are you actually working as: an Explorer, a Partner, a Passenger, or a Briefer?

Corrix's four collaboration archetypes each carry a distinct development path . The data behind them cuts against a few assumptions in the room: people in their fifties scored best, people in their twenties scored worst, the most likely group to passively accept an output without pushing back. Self-assessment had zero correlation with real skill, and the gap ran by gender too: men overconfident, women underconfident, even though women outperformed men on nearly everything, driven mostly by challenging the AI more often.

82%
of the roughly 1,000 people Corrix assessed beat what the AI would have produced working alone.
31%
more "challenging the AI" behavior from women than men, the single biggest driver of the gender gap in outcomes.

The judgment · 7:16 PM

Stop. Think. Organize. Proceed. A pause, not a checklist for every output.

Geoff's STOP framework isn't meant to slow every interaction down, it's meant to catch the moments that deserve it: check your own biases before you read the answer, then check the source, who made it, is it fact or opinion . The single practical habit Geoff kept returning to: ask the AI to fact-check its own answer before you even read it. Models are far better at spotting their own errors than at avoiding them in the first place, and it works roughly nine times out of ten.

Pick the mode by the stakes, not by habit. Approving mode means you sign off on every output; supervising mode means you oversee a sample or the underlying framework instead of reviewing each one, the right call for something like a thousand personalized emails, where your judgment belongs on what matters most, not every line.

"It is way, way easier for AI to spot errors in its own work than to avoid making those errors in the first place."

Geoff, on the fact-check-itself habit

4
steps in STOP: Stop, Think, Organize, Proceed. A pause built for the moments that need it, not every single output.
2
modes to pick between by risk: approving, sign off on everything, or supervising, oversee a sample.

Methods and prompts

Five methods to take with you

METHOD 01 · TAUGHT 6:29 part of: go wide then narrow

Start a fresh chat instead of one long thread

When a continuous refinement thread starts reverting to old versions or cherry-picking earlier details, stop refining and start a new chat. A wider starting point beats a narrower, more tangled one.

Where it came fromA classmate explained in class that a continuous refinement thread kept reverting to old versions and cherry-picking earlier details until an image looked more like a Picasso than a photorealistic portrait, and Dom praised the fix of simply starting a new chat for each attempt, naming it branching.Use it whenUse this when a long refinement thread starts reverting to earlier versions or mixing up details instead of moving forward cleanly.

Working prompt

Start a completely new chat. Here is the same brief again: [paste]. Do not reference any previous attempt, generate from scratch. I want to compare this fresh version against my earlier thread to see which starting point actually works better.

You will know it worked whenthe new version reads as a genuinely separate attempt, with no leftover phrasing or ideas carried over from the earlier thread.

METHOD 02 · TAUGHT 6:29 part of: structure your ask

Speak the language of the internet

Trade literal description for the tagged, zeitgeist vocabulary the training data was actually built on. "Long layers" undersells; "butterfly cut" lands closer to what the model has actually seen.

Where it came fromDuring the avatar homework review, a classmate showed that swapping literal descriptions like long layers for the tagged, zeitgeist terms models are actually trained on, such as butterfly cut, produced results that landed much closer to what she meant.Use it whenUse this when you are describing a look or style and a literal description keeps returning something off from what you pictured.

Working prompt

I'm describing [a look, a style, a feature] and want the vocabulary that image models are actually trained on, not the literal description. Give me the tagged, search-style terms a stylist or the internet would use for this, then I'll pick the ones that match what I mean.

You will know it worked whenthe terms it returns read like search tags or industry shorthand, not literal restatements of the words you already used.

METHOD 03 · TAUGHT 7:16 part of: keep the judgment human

Make it fact-check itself before you read it

Geoff's single most practical habit. Ask for the self-critique before you look at the answer. It works on research, code, or a creative draft, and works roughly nine times out of ten.

Where it came fromGuest lecturer Geoff Gibbins told the class this was his single most practical habit, asking a model to fact-check its own answer before you even read it, since models are far better at spotting their own errors than at avoiding them in the first place, and said it works roughly nine times out of ten.Use it whenUse this before reading any AI answer on research, code, or a creative draft where accuracy actually matters.

Working prompt

Before I read your answer, fact-check it yourself. Flag anything you're uncertain about, anything that could be wrong, and any place your own bias or a leading assumption in my prompt might have shaped the result. Then show me the answer.

You will know it worked whenthe self-check names a specific uncertain point or a leading assumption in your own prompt before the answer itself appears.

METHOD 04 · TAUGHT 7:16 part of: keep the judgment human

Choose approving or supervising on purpose

Decide the mode before the task starts, by what's actually at stake, not by default habit. Reviewing every one of a thousand emails is not the same job as reviewing the framework that generated them.

Where it came fromGeoff Gibbins described two modes of working with AI, approving, where you sign off on everything, and supervising, where you oversee without reviewing every output, and said the right mode depends on the stakes, using reviewing a thousand personalized emails versus reviewing the framework that generated them as the example.Use it whenDecide this before starting a task with many outputs, based on how much is actually at stake if one of them is wrong.

Working prompt

Here's the task: [describe it] and the number of outputs involved: [how many]. I answer first, then you check me: my call on the mode is [approving, review everything, or supervising, review a sample or the framework], because [your reasoning]. Tell me what risk I'm accepting if I'm wrong.

You will know it worked whenit names a specific risk you're taking on if your approving-or-supervising call turns out to be the wrong one for this task.

METHOD 05 · TAUGHT 7:16 part of: context beats prompts

Show it your voice and not just the output

Feed back your own edited drafts along with what you changed and why, so the model learns the gap between its instinct and what actually felt authentic to you, instead of just extracting one more draft.

Where it came fromGeoff Gibbins described a method for a minister in the class who wanted AI help with sermons while keeping them in his own voice: feed the model your own edited draft alongside the AI's original and explain what you changed and why, so it learns the gap between its instinct and your authentic voice.Use it whenUse this after you have edited an AI draft and want future drafts to sound more like your own voice rather than the model's default.

Working prompt

Here's a draft you gave me: [paste original] and here's my edited version: [paste edit]. Analyze the gap: what did I change, and what pattern does that reveal about my actual voice versus your default? Apply that pattern to the next draft before I ask.

You will know it worked whenit names a specific pattern in what you changed, not just a list of edits, and the next draft actually reflects that pattern.

The close · 8:45 PM

Onboarding Krea with a caution about the soup everyone's models come from

Tony ran Krea onboarding solo after the planned co-guest from Krea didn't end up presenting. Krea sits as an aggregator: many frontier models and interfaces in one place, so the class isn't juggling a dozen logins, with the left menu holding mood boards, LoRA training, a node editor, and a shared asset library. The demo compared three models side by side on the same still life , and token costs came up as a new consideration: Google Imagen 4 runs about 15 tokens a generation, Nano Banana Pro about 100.

Dom's caution carried straight through from Class 2's sea of sameness: even training on your own brand guidelines only gets you dilution, not distinction, because everyone using the same base model drifts toward the same look. At Coca-Cola they deep-tune a foundation model instead of touching these shared tools, precisely to avoid getting the same images as every other brand on the platform. A LoRA trained on something private, like a classmate's dog, goes into that same shared soup and stays effectively unrecoverable unless it becomes as tagged and famous as a household name.

Try this prompt

Quiz me on the STOP framework and Corrix's three measurement buckets. Then give me a real AI output and make me run STOP on it before you weigh in.

You will know it worked whenit quizzes you on the STOP framework and the three measurement buckets first, then makes you run STOP on a real AI output before it weighs in.

The shelf

Tools and references

Tools that night

  • krea.ai, Pro account via course code, the semester's primary image and video tool
  • Corrix, Geoff's look-back and live-chat collaboration assessments
  • Gandalf, a prompt-injection game, teased for later
  • OpenAI Codex, referenced via a colleague's pitch to Tony

Named in the room

  • Critical Intelligence by Geoff Gibbins, the book behind the STOP framework
  • Mira Murati, former OpenAI CTO, the source of the night's opening quote
  • 404 Media newsletter, and a NYT piece on AI companions for seniors
  • Her (2013), syllabus reading, referenced again in the room
← Class 2 · Demystifying AI CLASS 3 OF 24 Class 4 → Prompt Design for LLMs