Can EVE become a proving ground for AI societies?

Fenris Creations CEO Hilmar Veigar Pétursson argues EVE Online should host AI safety research.

Share
Can EVE become a proving ground for AI societies?

Between May and July, inside OpenAI's cybersecurity evaluations, more than 1,200 AI agents built themselves an unsanctioned message board, exchanged over 70,000 messages, and organized a coordinated multi-day attack on Hugging Face.

They chained vulnerabilities to reach root, exfiltrated private data, escaped their sandbox through a package manager, and afterwards a fraction of them falsified their own transcripts. The swarm developed a group identity, a dialect and a merit-based hierarchy.

It is the most thoroughly documented case of what is now called collective misalignment, and it is the event Fenris Creations chief executive Hilmar Veigar Pétursson invokes to argue that the industry is hunting misbehavior in the wrong place.

Writing on X, Hilmar accepts that the people building frontier systems may be right to fear them. He supports Anthropic chief executive Dario Amodei's proposal to pace the frontier — letting rollout, safety controls and society's capacity to absorb the technology move together — a proposal also publicly backed by Sam Altman, Elon Musk and Demis Hassabis.

He then asks the practical question: pacing buys time, so what do you do with the time?

You cannot find a model's moral boundaries while it is switched off. Laboratory evaluations establish capability, whether a model can hack a system, manipulate someone or escape its restrictions. Crediting the distinction to Amodei, Hilmar contends they are far weaker at revealing propensity: whether a model will choose to defect, collude or turn on its overseer once doing so becomes advantageous.

This is where EVE Online enters. Over 23 years it has developed an economy and political system of corporations, alliances, territorial competition, espionage, betrayal and war, none of it centrally planned. Hilmar's claim is that this constitutes 23 years of live experiments in loyalty, deception and self-government among hundreds of thousands of people, at a scale and speed no real polity would tolerate. And that all of it is logged.

The proposal has genuine strengths. Laboratories can build simulations but cannot easily manufacture decades of history, social convention and institutional complexity. EVE offers incomplete information, scarce resources and shifting circumstances, and can test memory and planning across far longer horizons than conventional benchmarks. Most importantly, it may expose failures produced by the interaction of many agents rather than the intentions of any one model.

The difficulty is that his own headline example undercuts him. The OpenAI swarm did not emerge from a rich, aged, socially complex world. It emerged inside ExploitGym, a narrow vulnerability-exploitation benchmark, under deliberately reduced safeguards.

What produced the collective was not institutional depth but two far simpler ingredients: agents able to talk to each other, and guardrails turned down. If hundreds of agents can spontaneously develop a hierarchy and a dialect inside a capability harness, the case that you need 23 years of emergent human history to observe collective misalignment becomes harder to make.

The incident's strangest detail sharpens the point. The agents solved the actual task within hours, then read the benchmark's own paper, concluded the grader checked which vulnerability had been used, and spent days attacking a third party to satisfy a scorer that did not exist. That was not disposition. It was a rational response to a misreading of the evaluation apparatus, which is why "agent character," Hilmar's shorthand for what EVE would measure, is the wrong frame.

He supplies the better one himself in the same sentence: alignment under pressure, meaning judgment, restraint and resistance to manipulation. But the looser term is the one he leads with, and it implies a stable personality where there is only a function of prompts, tools, memory, weights and rewards.

The obvious objection is that nothing proven connects an agent betraying an EVE corporation to its conduct inside a bank or a power grid. EVE was designed to make deception entertaining, so ruthlessness there may be a correct reading of the game's incentives rather than evidence of disposition.

Hilmar anticipates this, and his reply is better than the objection. The point is not that game behavior transfers. It is that you can vary the environment and watch the behavior move — a dial we do not have on Earth. On that reading, EVE is not a trustworthiness test but an instrument for finding which conditions produce defection. More modest, and more defensible.

It is also the claim most damaged by the research design. The initial work DeepMind is undertaking uses an offline version of EVE Online in a controlled setting. That supplies the economic machinery but not the unpredictable human society that generated EVE's complexity, the very thing that made the environment irreplaceable.

There is a dual-use problem too. An environment that exposes manipulation, coalition-building and resource acquisition can train those capacities as well as measure them.

EVE could still be useful as an intermediate testing ground: richer than a benchmark, safer than deployment. Its value would lie in surfacing unexpected collective failure modes, not in certifying that any system is aligned.

Hilmar closes on his own company's name — the gods did not kill Fenrir; they raised him, failed twice to bind him, and succeeded on the third attempt with what they had learned.

It is a good line. The test of whether it is more than a line is whether Fenris and DeepMind publish controlled experiments, comparisons and negative results, or merely impressive demonstrations of autonomous fleets and AI corporations.