The ghost in the machine: Simulation and the future of AI-powered determination
Enterprise decision-making is scaling more accurately and efficiently with AI-simulated populations.
Consider this: A CMO at a Fortune 500 brand ran three packaging concepts against 10,000 consumers overnight before committing to one. A political consultant pressure-tested a Senate campaign's messaging against voters in three swing districts before spending a dollar on media. A NATO analyst modeled how 5.1 million people across an entire European country would respond to a foreign disinformation campaign, then tested countermeasures against the same population.
Incredibly, none of these studies conducted interviews, surveys, or focus groups. These were all queried against a synthetically generated population through a process that took hours rather than weeks, cost a fraction of what a human panel would, and was on par with (or even more representative than) legacy methodology.
Behavioral simulation may be one of the most consequential applied AI use cases that we’ve seen emerge. Today, it's already a commercial category with paying enterprise customers in wedges like consumer market research. If the technology works the way its best builders intend it to, this technology has the potential to become the layer that enterprises run every meaningful decision through. Here’s an exploration of what’s real today, who's building it, and what will separate the winners from the rest.
The growing demand for synthetic research
The bottleneck behind almost every consequential decision organizations make is the same: not knowing how people will actually behave before you commit. Market research is a $140 billion industry that has been broken in painfully obvious ways. A single concept test from a large research firm can start in the high five-figures and take weeks, while deeper projects can cross into millions and take months, forcing organizations to selectively limit their studies to just a few key decisions.
The deeper problem goes beyond just cost and speed: the source data is often inaccurate. Respondents are paid a nominal amount to take a survey or sit in a focus group, and some studies find their answers correlate only about 40% with their subsequent behavior. Part of the problem is unrepresentative panels and tight budgets, but the bigger issue is that interviewers can introduce bias and respondents can be unreliable narrators. Conventional studies rest almost entirely on what people say rather than what they actually do. The hardest populations to reach, including policy insiders, CISOs, or orphan populations, are nearly impossible to find and study at all.
A related category of AI-native platforms, like Strella and Listen Labs, emerged to automate qualitative interviewing itself by conducting and synthesizing conversations with real respondents at a speed and cost previously impossible. While groundbreaking, population simulation takes a different route: modeling behavior (with data or human interviews as inputs), then building synthetic populations to simulate outcomes, find emergent causal links, and drive decision making.
Once LLMs made it possible to simulate coherent human reasoning affordably enough to commercialize in late 2023, next came a wave of academic papers exploring the possibilities. By 2024, the techniques had reached commercial research, and by 2025, real enterprise revenue was flowing. In our conversations with buyers, many have already affirmed the massive cost and time benefits. As a result, some leading enterprises are now spending seven figures annually with these emerging companies. When decisions that would have never justified a strategy consulting engagement can now be queried against a synthetic population in an afternoon, the surface area of what deserves to be researched grows exponentially and is fundamentally changed.
Two camps of simulation

Emerging methodologies for building a synthetic population
Market players are coalescing around two primary camps when it comes to how they build synthetic populations. The first builds based on real human interviews, then enriches them with layers of behavioral data to query repeatedly across future projects. Simile is a Stanford spinout which grew out of the seminal "Smallville" generative agents study, in which AI agents populated a simulated town and formed routines and interactions without human scripting. Their deep research focus has led them away from prompting off-the-shelf models and towards training proprietary simulation and confidence models, incorporating collected behavioral and interview data into the model weights themselves rather than only the context window. Simile has an individual model that maintains the internal logic to govern individual preferences, biases, and worldviews, and also a population model that applies these learnings at scale. Rehearsals is another player in this first camp that also interviews humans, but they approach with a psychologically-grounded core and multi-modal cloning as the entry point.
The second camp focuses on rebuilding the population from ingested data sources rather than individual interviews. These data sources can include census and demographic studies, behavioral traces, historical research panels, and social listening. Aaru is an expression of this philosophy. They believe relying only on non-interview data gives customers faster time to value and also avoids letting the interview itself become a contaminant that introduces human bias. This camp argues that self-reporting inherits the biases that can make traditional research unreliable, so grounding a persona in interviews could import those same errors that simulation is meant to avoid. Artificial Societies is another commercial example that assembles populations from synthetically modeled personas rather than 1:1 interviews.
The interview-based methodology rebuts that these “biases” can be valuable in simulating how real populations will act, using only data and behavioral traces as inputs makes validation and backtesting almost impossible, and calibrating against real, named individuals is what separates a believable twin from a generic persona. That being said, the line can get blurry between camps. Interview-first approaches often still fold in public data, while data-only approaches can also absorb a client's first-party records (including prior surveys, focus groups, and interviews) to improve their outputs. Players like Electric Twin have already emerged to occupy a more hybrid view.
Whichever approach a company starts with reveals its founders' beliefs about where the value lives, and that belief is what will shape their initial ICP and future roadmap.
The enterprise buyer perspective
While startups have varied across both camps, buyers have largely taken a more pragmatic approach. Enterprise users frequently evaluate products from both camps, and their purchasing decisions are more often dictated by specific use cases and technical requirements rather than the provider’s philosophy around building the population. Accuracy is the first checkpoint. If outputs are unverifiable or untraceable, buyers won’t have confidence in the results and the category may not survive serious procurement or renewal.
However, even directional accuracy can be valuable for narrowing choice sets and pre-filtering ideas before more expensive human research, and often represent a better option than the rough guesses decision-makers are forced to rely upon today.
Those who are building for the enterprise in this space will need to successfully navigate three tiers of accuracy:
|
Attitude replication |
Does the simulation answer align with what people say? |
This is what most published figures measure. The ceiling is set by the surveys and benchmarks used, which can correlate weakly with actual behavior. |
|
Behavioral replication |
Does the simulation align with what a real person does and how that would change under different conditions? |
This is the harder and more meaningful bar. The category has scattered evidence here but rigorous out-of-sample evaluation is rare. Buyer benchmarks suggest the best products land within 10–15% of actual campaign outcomes, good enough to drive real commercial decisions. |
|
Emergent simulation |
Does a simulated population remain accurate in the face of variable inputs not in the training set or interview data? |
This requires the simulation to reproduce macro phenomena that no individual agent could, such as sentiment indices, retail trends, or electoral outcomes. The test of accuracy here is whether agent-to-agent interaction produces genuine emergent behavior. |
Separately, while we have focused on simulation of human behavior here, there is a related and emerging category of companies like Mantic or Preseen that aim to predict event outcomes such as election results, probability of supply chain disruptions, or regulatory decisions. Many get their start competing on Metaculus leaderboards or betting in prediction markets. Both predictive forecasting and enterprise simulation have uncapped real world applications if they can prove accurate or sufficiently informative.
What the winners will need to get right
Given the pace of change, we don't view synthetic personas, the data pipelines, or the query engine as where the moats lie, since foundation models and public data will likely commoditize all three in the coming years. We believe the real moats will be: the proprietary first-party data that calibrates each population, the workflows the product embeds into, and the institutional trust that comes from rigorous validation.
The data flywheel
An enterprise engagement should integrate a customer’s own records, transaction logs, and prior survey waves to calibrate the population. Each deployment feeds response data back, sharpening the next prediction and building a proprietary signal. Constructing that population is where the craft lives. A query run against a poorly assembled population is fast and wrong, so composition and calibration are where accuracy is won or lost.
The product workflow
The product is stickiest once it stops being something customers consult for one-off questions and becomes a fixed step in how decisions get made, running on every launch, pricing call, and campaign. At that point, removing the engine means rebuilding part of the organization's decision-making process.
Institutional trust
Trust compounds as predictions hold up across engagements, as customers can defend a choice by pointing to how the system reached it, and as the tool becomes the one they instinctively reach for on the next high-stakes decision. Institutional trust is slow to build, but once established, it makes a provider difficult to dislodge.
What’s next for simulation
The “ghost in the machine" is no longer a figure of speech. With enough tokens and data, a recognizable model of a real person takes shape, including their tastes, instincts, and how they weigh a decision. What was a thought experiment a few years ago is now a working product, accurate enough for enterprises to route real decisions through.
As technical capabilities evolve, the need for commercial and ethical boundaries grows. Nothing stops a state from running these simulations for political targeting or propaganda optimization, and no one has answered who owns the responses of a live person's digital twin once it becomes accurate enough to query in their absence.
The category is still in its early innings, and companies need time to earn enterprise trust and test their simulations against real world outcomes. The winners will need to get the most difficult parts right: real accuracy, compounding data flywheels and workflows, the discipline to validate what they sell, and a clear view of where this technology should and shouldn’t be directed. At the frontier, building fast and building thoughtfully have to go hand in hand.
If you're a founder building in this space, get in touch with Elliott Robinson, Andrew Ren, or Caty Rea to discuss more.






