Why Not Just Use ChatGPT?
Why simulation is an order-of-operations problem, not an AI problem
Author’s Note: This is Part 2 of 3 on simulation and Project Brainstorm. In Part 1, I explained why simulation can be useful for exploring questions before they’re ready to be measured. In this article, I explore a more fundamental question: if modern AI models can already simulate conversations, why build an entire platform around them at all?
Project Brainstorm is a simulation-based market research and consumer intelligence platform designed to help users better understand how people evaluate products, concepts, and decisions. Its role is to help users explore questions, assumptions, tensions, and decision dynamics before committing significant resources to validation, experimentation, or execution.
In my previous article, I argued that market research simulation is useful because it helps us – whether we’re managers, investors, marketers, analysts, strategists, or entrepreneurs - explore questions before they can be measured.
The obvious response to my argument, then, would be:
"If AI models can already simulate conversations, why bother building an entire system around them?"
When I first started thinking about building an AI-powered market research platform, I asked myself the same question. Why bother building a system around ChatGPT or Claude? After all, I know that our Managing Partner at Oakseed Ventures – Chee-We Ng – loves nothing more than to throw shade on start-up ideas that turn out to be little more than AI wrappers. I respect Chee-We immensely and have no desire to be ridiculed by him.
At first glance, Project Brainstorm looks dangerously close to being one of those.
When I was planning the architecture, I realised something important. The value wouldn’t come from the models themselves. Don’t get me wrong, AI models are a necessary component of any simulation system. Without them, you’d have a car with no engine. But just as a Porsche is more than its motor, a simulation platform needs to be more than just the AI model.
What matters as much, if not more, are the rules you build around it: how information is gathered, how context is constructed, how people are represented, and how the simulation is designed to reason. In other words, the most important part of this system wouldn’t be the large language model - it would be the structure.
It took me a few months to understand that distinction, but once I did the project took on a clear direction. Let’s deal head on with the elephant in the room.
Why Not Just Use ChatGPT?
If modern AI models can already simulate conversations, why bother building an entire system around them?
I think that's a fair question. After all, the Project Brainstorm platform uses a mix of the models from OpenAI, Anthropic, DeepSeek, Baidu - plus a revolving cast of a few others - to drive its simulations. So, why bother?
You bother because, as I alluded to earlier, the models aren't really the key parts. What I found is the key to a realistic simulation is the structure that you build around them.
A modern chatbot is designed to respond to prompts. You ask a question and it gives you an answer. It's fast, flexible, and capable - like Google Search on performance-enhancing substances. In many situations that's exactly what you want.
But when you're trying to understand how real people in some place think, this approach starts to run into problems. Just as we concluded that context influences reasoning, prompts influence outputs.
Any model can only work from the information and framing it receives. Change the framing and the reasoning changes with it. Change the framing again and the answer changes again. Emphasise one aspect of a problem and the reasoning shifts with it. Ask a different question and you're effectively starting over from scratch.
There's nothing inherently wrong with that process. The model is simply responding to the information you've given it.
The challenge, then, isn’t that the model answers incorrectly, the challenge is that the model can only reason from the assumptions you’ve provided. If you're exploring a new market, testing a product idea, or trying to understand how people might interpret a concept, you quickly run into problems that you’re not even aware exist. The danger isn't getting a bad answer, it's getting an answer that quietly inherits the assumptions you started with.
The result is that you’re returning a result that feels persuasive, but only because it confirms your inherent biases.
From that starting point, I aimed to build something different.
Instead of going directly from question to answer, I wanted a system that established who existed in a market first, what context those people operated within, and what information might shape their thinking.
I wanted individuals whose identities had been formed independently of whatever product, idea, or message was being tested, before asking them to evaluate anything. People first, product second.
In other words, I wasn't looking for a better response engine. I was looking for a structured environment from which responses could naturally emerge. That distinction sounds subtle, but what emerges is completely different.
Why Most AI Simulations Are Shallow
Moving on from the chatbot problem, I looked at the growing number of AI simulation platforms appearing, from both academic research and the private sector.
Researchers are increasingly using large language models as stand-ins for human populations, with synthetic personas, agent-based simulations, and AI-driven focus groups as ways to understand attitudes, behaviours, decision-making, even psychiatric disorders.
Many of them are impressive. They generate personas. They hold conversations. They produce reports filled with opinions, objections, recommendations, and behavioural insights. On the surface, they look remarkably human.
The problem is that realism and simulation are not the same thing. A system can generate highly convincing dialogue without representing how a person arrived at that conclusion.
That distinction is paramount. Most simulation platforms focus on generating responses. I am interested in generating people.
It sounds like semantics, but I don’t think it is. If the people only exist to react to a particular product, campaign, or idea, then they are effectively part of the prompt. They are props within the exercise rather than participants in it.
The result may look realistic, but the simulation has very little independence. The personas exist because the product exists. The result is a subtle form of circularity where a simulation effectively begins from the answer space. The product, message, or concept is established before the people who are supposed to evaluate it have even been defined.
I wanted the opposite. I wanted the people to exist before the product arrived.
The Structure That Makes Simulation Useful
The more I thought about solving this problem, the more I realised that the quality of a simulation depends less on the sophistication of the model and more on the order in which information enters the system.
In retrospect, I’d even go as far as to say that this is less of an AI problem and more of an order-of-operations problem.
Most simulation systems start with the thing being tested. A product, concept, message, or idea is introduced first. Personas are then created around it and asked to react.
At first glance, that seems perfectly reasonable.
But the issue is that the thing being evaluated immediately starts shaping the simulation itself. The people are defined in relation to the product. They are shaped by the thing they are supposed to be evaluating. The evaluation begins before the evaluator has even been established.
That’s completely backwards to me.
In real life, people do not spring into existence the moment they encounter a product. They already have lives, experiences, constraints, habits, beliefs, and sources of information. They have friends and families. They live in communities. They consume media. They navigate economic and social realities that influence how they interpret everything around them.
That all exists before the product arrives. So, I insisted the system to reflect that sequencing.
It’s a surprisingly simple sequence, too:
Population → Personas → Context → Reasoning
First, understand the market. Who lives there? What demographics exist? What pressures, opportunities, and constraints shape daily life?
Next, create individuals who plausibly emerge from that environment.
Then we establish the context surrounding them. The information they encounter, the news they consume, the narratives they’re exposed to, and the events shaping public perception.
And only after all those pieces are in place is the product, idea, message, or decision introduced.
What emerges from that isn't just a generic response. It's an individualized response generated by a particular person, operating in a particular situation, with a particular set of constraints.
It means the simulation isn't trying to answer:
"What do people think about this?"
Instead, it's trying to answer:
"How would that person in that market make sense of this?"
That's a subtle difference, but I think it's paramount to realism. The first assumes opinions already exist. The second focuses on how opinions formed in the first place. And the latter is where I believe most of the useful signal sits.
Grounding, Context, and Simulation Are Different Things
People tend to lump everything together under the label of "AI" nowadays.
In practice, there are multiple separate processes taking place inside the system and treating them as the same thing creates confusion.
The first is grounding.
Grounding aims to answer a simple question:
“What kind of world does this person live in?”
Who lives there? What circumstances shape their lives? What social, economic, cultural, and geographic factors influence their day-to-day reality?
The second is context.
Context answers a difference question than grounding:
“What information exists within that world?”
News stories, local narratives, cultural references, public debates, media coverage, recent events and countless other signals all influence how people interpret their surroundings.
The third is simulation.
Simulation asks:
“How might a particular person reason within that world?”
This is the stage where people encounter an idea, a product, a message, or a decision and attempt to make sense of it based on their circumstances and the information available to them.
These concepts are closely related, but they're not interchangeable. Grounding establishes the environment. Context shapes interpretation. Simulation explores reasoning.
This distinction matters not because it produces better-looking outputs. It matters because it allows you to see where those outputs come from.
Rather than treating every response as the product of some mysterious AI process, you can examine the environment, the information available within the environment, and the reasoning that follows.
Beyond the Architecture
At first, separating grounding, context, and simulation was about preserving the integrity of the process.
If these were all collapsed together, it would become difficult to distinguish between the world people live in, the information they encounter, and the reasoning that emerges from the interaction between the two. I viewed it as an important architectural decision.
The full implications of that separation only became apparent later.
Once the environment, information people are exposed to, and the reasoning that emerges from both are treated as separate layers, conclusions no longer appear as if they came out of nowhere.
Instead, it becomes possible to understand not just what the system concluded, but how it reached that conclusion.
A reaction can be traced back to the context that influenced it. That context can be linked back to the market conditions, narratives, and circumstances that shaped it. Rather than receiving a conclusion as a black-box output, you can follow the chain of reasoning that connects one stage to the next.
The more I tested and worked with the system, the more I realised this was where much of the value sat.
And that leads directly to the next question:
How do you make simulated reasoning traceable?
That’s what I’ll explore in the next article.
This article was originally published in Substack (https://josephlo.substack.com/p/why-not-just-use-chatgpt?r=7fr2) on July 17, 2026.