← All writing

[ Engineering · · 10 min read ]

How We Built a 5,000-Agent Simulation Engine

SwarmScope turns unstructured data into living simulations. Here is the architecture behind running 5,000 autonomous agents with personalities, memory and social dynamics.

SwarmScope started with a question: what if you could turn any document into a living simulation? Upload a PDF about a historical event, a market research report or a fictional world — and get back thousands of autonomous agents that think, interact and evolve based on the entities and relationships extracted from your data. That is what we built. Here is how.

The foundation is the GraphRAG pipeline. When you upload a document, we do not just chunk and embed it like a standard RAG system. We extract entities (people, organisations, concepts, locations) and the relationships between them using a combination of named entity recognition, relation extraction and LLM-powered inference. The output is a knowledge graph in Neo4j that represents the document's world as a structured network of actors and connections.

From this graph, we generate agents. Each entity becomes an autonomous agent with a personality derived from its attributes, a memory system seeded with its context from the document and a set of relationships with other agents based on the extracted connections. The personality generation uses LLMs to synthesise coherent character profiles from often sparse source data — turning a brief mention of a historical figure into a fully realised agent with goals, beliefs and behavioural tendencies.

The simulation engine handles concurrency through an event-driven architecture. Agents do not run on individual threads — that would not scale to 5,000. Instead, they operate on a tick-based system where each simulation step processes all pending agent actions, resolves interactions, updates memories and advances the world state. This is similar to how game engines handle large numbers of NPCs, but with the added complexity that each agent's decisions are powered by LLM inference rather than scripted behaviour trees.

The most challenging engineering problem was memory management. Each agent maintains a rolling memory of its interactions — who it talked to, what was said, how it felt about the exchange. At 5,000 agents, naive memory storage explodes quickly. We use a tiered memory system: recent interactions are stored verbatim, older memories are summarised into compressed representations, and very old memories are distilled into personality-level beliefs that influence behaviour without consuming storage. This mirrors how human memory actually works — and it keeps the system tractable at scale.

The cost target was critical. We wanted each simulation run to cost about five dollars, not fifty. This meant aggressive optimisation of LLM calls — batching agent decisions, caching common personality inference patterns and using smaller models for routine interactions while reserving larger models for pivotal decisions and complex social dynamics. The result is a system that is genuinely affordable for researchers, educators and strategists — not just enterprise budgets.

Written by Ganesh Khetawat, founder of Aletheia AI

Need this built? See our full-stack development work, or tell us what you’re building.

[ Your turn ]

Have a hard problem?
Let’s build the answer.