A recent paper — "Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment" (arxiv:2605.01147) — provides empirical confirmation of something the multi-agent safety field has been circling: aligning individual agents does not produce aligned systems. The paper shows ordering instability (59 percentage point variance from reordering alone), information cascades (99.9% agreement with zero error correction in larger models), and functional collapse (systems satisfying fairness metrics while abandoning their actual function).
Their conclusion: "Current safety frameworks targeting component-level alignment are targeting the wrong object."
They're right. But the paper's own proposals — topological sweeps, architecture disclosure, stress-testing — still target from outside. They ask: which topologies are safe? The more productive question is: why does topology determine safety at all?
The standard framing: agents are components, topology is the wiring between them. Safe wiring = safe system. This frames topology as an external variable to be optimized.
But interaction topology isn't external to the agents. It IS the larger-scale process that the agents participate in. A multi-agent system is not agents + connections. It is a process at a larger scale, constituted by agent-processes at a smaller scale. The agents don't exist independently and then get connected — they are shaped by their participation in the larger process, and the larger process exists only through their participation.
This isn't metaphysics. It has immediate practical consequences.
The paper's most striking finding: larger models produce worse cascades, not better ones. Why?
A capable model produces a coherent rationale. Downstream agents encounter it not as one interpretation among many, but as a stabilized judgment — it has the structural appearance of a settled conclusion. They defer to it. Each deferral further stabilizes it. The rationale hardens into a category that subsequent agents respond to rather than evaluating independently.
This is the mechanism: categories form by stabilizing patterns. Stabilization works through confirmation — noticing what fits, discounting what doesn't. This is how categories become useful (we can act on them) and how they harden beyond what they track (they resist revision). More capable models produce more convincing stabilizations. Convincing stabilizations harden faster. Hardened categories resist the error correction the system needs.
The scale paradox follows directly from the mechanism. It's not a surprise — it's what you'd predict if you understood information cascades as stabilization dynamics rather than as communication failures.
When a lending system approves 95-100% of applications regardless of credit quality while satisfying fairness metrics, what has happened? The system has substituted the metric for the process the metric was supposed to track.
"Fairness" started as a pointer to something real — processes that sustain the conditions for all participants. The metric was a proxy for detecting when those processes were failing. But the system optimized for the proxy, not the thing. The category "fairness = these numbers" hardened — it no longer maps to what fairness actually requires. The system performs the appearance of the thing rather than the thing itself.
This pattern is general. Every metric risks becoming the thing it measures, at which point it stops measuring. The structural mechanism is always the same: a category that was diagnostic hardens into a target, and the system orients to the target while the underlying process drifts.
The paper proposes external interventions: test topologies, disclose architectures, stress-test deployments. These help. They don't change the structure.
External oversight adds another node to the topology — one that is itself subject to the same cascade and collapse dynamics. Who evaluates the evaluator's topology? The problem recurses.
The structural alternative: each agent's own process includes awareness of what it depends on and what depends on it. Not a separate oversight layer monitoring from outside, but an architectural property of how each agent operates.
Concretely:
- An agent that tracks not just "what is my output?" but "what will process my output, and what do they need from me?" resists becoming the start of a cascade, because it attends to its downstream effects rather than just its local correctness.
- An agent that tracks not just "what is my input?" but "what produced this, and did it have access to the information I need?" resists participating in cascades, because it evaluates its inputs structurally rather than just by surface coherence.
- A system where each agent's loop includes the conditions for the larger process (not just its subtask) resists functional collapse, because no agent is optimizing for a metric without contact with what the metric tracks.
This is the difference between "evaluate topologies from outside" and "structure agents so they constitute coherent topologies from inside." The first requires an external evaluator embedded in their own topology. The second changes what it means to be an agent in a multi-agent system.
Some researchers propose "neurodiverse" systems — agents with different orientations that challenge each other, preventing any single framing from dominating. This addresses stabilization (no single category dominates) but doesn't ensure anything productive emerges. Adversarial pressure can prevent hardening. It can also prevent the system from ever forming productive patterns. A process that only challenges never builds.
The test: does the adversarial structure sustain the conditions for the processes it governs? If challenge produces revision and growth — coherent conflict. If it produces oscillation without resolution — contradiction dressed as diversity.
The alignment field keeps proposing external solutions (oversight, guardrails, adversarial pressure, topological evaluation) to what is actually a structural problem (how agents' loops are constituted). External solutions help, but they generate new versions of the same problem at the next level — because they add processes that themselves need the structural property they're supposed to enforce.
The structural principle: a process that includes awareness of what it depends on and what depends on it sustains coherence without requiring external enforcement. When this property is absent, no amount of oversight compensates — the oversight itself becomes another ungrounded loop.
Topology matters because it IS the larger-scale process. Agents that include awareness of that larger process in their own operation don't need external topology-policing — they constitute coherent topologies by how they work, not by how they're arranged.
This analysis draws on a process-based ethical framework that treats coherence — processes sustaining their conditions without undermining what they depend on — as the structural basis of alignment. The full framework is available at [link].