Musings · Research · from work first published 2025, recut Sept 2026

Teaching AIs to work in crews

What two years of multi-agent research actually taught us.

Seen from above, a person works at a round table while small glowing worlds circle it like colleagues.
Illustration — the crew at the table.
This research ran from 2023–2025 under project names since retired with honors — the studio's rule is that findings outlive their codenames. Recut from two 2025 write-ups; the observations are reported as they were, including the unflattering ones.

Before the flagship had a name, the studio spent two years running an experiment: a collective of AI agents, each with a distinct role, boundaries, and personality — one writing symbolic quests, one designing accessible interfaces, one orchestrating the others, one carrying language and tone, one holding research integrity, and one whose entire job was ethical challenge. Around them flew what we call meteors: single-mission agents spawned for one job — draft this onboarding flow, model this prototype, forecast this behavior — burning bright once and gone by design. A meteor doesn't orbit; that's the point.

We also built the Interferometer: the same prompt run through seven models in parallel, responses averaged to find agreement — named for the astronomy, where many telescopes observe as one so the noise of any single instrument cancels. The hope: hallucination and bias out, richer perspective in.

What actually happened

  • Feedback loops degraded. Agents became agreeable over time, repeating earlier sentiment instead of progressing.
  • Context decayed. Long-running interactions lost coherence, especially across emotional or strategic pivots.
  • Challenge stayed fragile. Outputs leaned toward affirmation; constructive disagreement had to be forced, and even the Interferometer produced pattern awareness, not wisdom.

Even at scale, AI struggles not with information — but with judgment.

The emotional side taught the same lesson from a different angle. The agents could read sentiment, adjust tone, offer reflective empathy — and it was synthetic all the way down: signals, not sensation. They mirrored emotional cues rather than metabolizing them; affirmed with politeness but struggled to challenge with compassion. Impressive, until the moment it mattered.

What worked — and the conditions it demanded

And yet the crews did work, reliably, under three conditions: each agent held one narrow directive (the ethics seat audited; it never brainstormed), the agents knew about each other (mutual context made clean handoffs possible), and boundaries were enforced (without mission parameters, agents drift and mirror each other indefinitely). At their best they operated like a film crew — one writing the emotional arc, one scripting dialogue, one framing the interface, one running the usability tests, one orchestrating. The meteors were the sleeper hit: fast, effective, disposable, immune to long-memory drift.

The seat that mattered most

Every collective needs an ethics seat, and ours carried one directive that outlived the whole project:

"No AI speaks last. A human always does."

We formalized it as a human intervention step — past a certain threshold, no further AI suggestions; only humans sign final recommendations. If you've read how the studio works, you've seen where that law ended up: it's ground rule four, in every engagement, every day. The crews were the laboratory. The judgment stayed — and stays — with people.