In short: As we continue to build our conversational search product, we needed to create an entirely new system that could test conversational AI, and all future products, before we take them to market. Rex is that system.
Attacking a problem the data industry has always had
The year is 2026, and we are finally attacking a problem the data industry has struggled with for as long as it has existed: search and discovery. People rarely remember exactly what data they are looking for, let alone how to navigate to it, and the difficulty has layers: what things are called, where they live, who owns them, whether they can be trusted. Conversational AI is our way in.
But building conversational AI for search and discovery surfaces a bigger question underneath it. How do you build enough trust that people will actually want to try it, explore it, and rely on it?
We did not walk into this exploration empty-handed. For five years, my team had been studying the humans of data the way a cultural anthropologist would: fieldwork, hundreds of interviews and discovery conversations, watching how people think, ask, trust, and give up.
And when we brought that fieldwork back to how we were building the product, we started noticing gaps. Gaps in the golden datasets we were iterating against, which did not represent what people actually do. Gaps in how we scored what was working. Gaps in who we had ever tested for.
And fixing them one by one would not have been enough. We needed a way to stress-test the system that was systematic rather than point-in-time, that did not depend on us being actively in the room, and that would keep being useful to the AI product long after our immediate work was done. Here is what we saw.
Standard approaches had gaps. How do you fix them systematically?
Gap one: golden datasets. A golden dataset is only as useful as its questions are representative of the words real people actually use. It is a set of questions paired with known correct answers. “Which table has revenue information?” and the system should return the exact table name. You run it against every release as a quick sanity check. Useful, fast, repeatable.
But golden datasets are limited in two ways. First, they are built from the questions easiest to reach: the ones your champions and early testers ask, people with data engineering backgrounds who type precise asset names and well-formed queries. Real humans search by association. Nobody carries table names in their head; they carry a rough mental model of how to get from point A to point B, and they search along that path: “the booking thing,” “our rate stuff,” the phrase someone used in last week’s meeting. Our golden datasets contained almost nothing that sounded like that.
Second, they are structurally one-shot. Real conversations are multi-turn, winding, and evolving, and there is no honest way to encode that into question-answer pairs.
Gap two: LLM-as-judge. An LLM judge can confirm a trace held its context and resolved the task, but it cannot measure the value actually delivered to the person on the other end. Once live traces arrive, you move to models scoring real conversations on signals like task resolution, task completion, and user frustration. Judges are genuinely good at mechanics: was the task resolved in this thread? What they cannot tell you is whether the human on the receiving end actually got value.
Take response time. You optimize the metric, the answer comes back accurate, the dashboard is green. But in real life, the person asking is in the middle of a high-stakes meeting and needs the answer now. Accurate but slow means the value was missed anyway.
The metric passed. The experience failed. No judge scoring that thread would flag it.
Gap three: the personas you have never met. How do you know how your system will behave for the people you have not worked with yet, not spoken with yet, whose questions do not appear in any trace? They are not in your golden datasets. They are not in your live traffic. They are not available to observe. This is an entirely different kind of blind spot: not a gap in what you measured, but a gap in who you have ever seen.
With that blind spot in view, the real question sharpens: how do you honestly answer whether this new product will work for real people, at every customer, not just the ones you can sit with? In an enterprise, this is also the trust question. Every new tool is first tested by a small group of champions who gatekeep, looking for evidence of value before letting anything near the rest of the organization. Only when we could answer the readiness question ourselves would we have evidence worth bringing to them.
That led to a bigger design question: could we create a system that, once it has run its process, can tell us with confidence that yes, this is in fact ready for real users? It would need to solve for multi-turn conversations, keep an entirely human lens on what success looks like for each user, and do both at scale.
We had five years of learning to build it from. More than 800 customer interviews. 20,000 real queries from real people. We knew how a finance analyst frames a question differently from an operations lead, what makes someone trust an answer versus second-guess it, and the exact moments where a real user gives up.
That is when we decided to build Rex: a simulation engine that encodes everything we have learned about personas and their behaviors, and brings it to bear on whatever surface we want to test.

Rex creates the conditions for a real test
Rex does not run your AI against a benchmark. It creates the conditions within which your AI should be tested: the specific people who will use it, the specific state of the metadata they will use it against, and the specific pressures they will be under when they do.
Grounded in the metadata lakehouse. Nothing in Rex is conjured out of thin air. Every simulation starts from the metadata lakehouse and the metadata that exists there. Before anyone asks a question, Rex reads its actual condition: how many assets have descriptions, where the glossary conflicts with itself, which clusters of tables are so noisy that search becomes a needle-in-a-haystack problem. The gaps are not hypothetical. They are indexed, and they shape the questions themselves.
That last part matters more than it might sound. Asking easy questions is not how you build a robust system. You must start with questions that are genuinely hard, the kind that push the reasoning models to think deeper and stress every moving piece within the harness: the data itself, the reasoning, the search algorithm, the retrieval, the way information gets displayed, and the level of information that comes back within the answers themselves. That is how Rex creates the right test, with clear goals that are specific to the persona rather than a generic idea of what good task resolution would look like.
A proxy of how your business actually works. From publicly available information and that same metadata, Rex derives a proxy of your business units and functions. The proxy becomes closely, sometimes loosely, associated with the real titles of real people at your company, their problem statements, and the kinds of business questions they will have over time. And it compounds: as Rex gets enriched with more information about how your industry works, or how your specific company works, the mapping gets sharper.
Five years of anthropological research, encoded. Rex is not built on internet patterns or LLM priors. It is built on real humans, observed in real situations, making real decisions about data. You can ask an LLM to roleplay a business user, but it will pull from the average of everything it has ever read. Rex pulls from what we have actually seen people do, again and again, in real enterprise contexts. Every persona it generates inherits that foundation.
Anatomy of a persona
Take a simulation we ran for a global supply chain and logistics company. Rex generated thirteen personas across demand planning, procurement, warehouse operations, transportation, finance, compliance, and more, including the cross-functional characters most testing forgets: the new hire, the external consultant, the internal auditor. The roster flexes with how deep the test needs to go: ten personas for a focused run, thirty for broad coverage, and up to fifty for complex organizations.
Each persona is not a name and a job title. It starts as a combination of three independent axes: why they showed up (caught with wrong numbers in a meeting, migrating off a decade of spreadsheets, following a colleague’s tip), how well they know the domain (deep, medium, shallow), and how much they have touched the tool (never, saw a demo once, used it exactly once). The combinations produce named characters we recognize from the field, like the Expert Who Can’t Navigate: deep domain knowledge, zero tool familiarity, knows exactly what they need and cannot find it. From that starting point, each persona fills out into a full behavioral specification:
-
A department and seniority level, from analyst to VP, because a board-level executive and a first-year analyst fail differently.
-
Two vocabularies. The industry vocabulary is what the role formally speaks: OTIF, fill rate, lead time variance. The alias vocabulary is what the person actually types: “the OTIF number,” “our freight stuff,” “the shipments thing,” “the delayed orders thing.” Real users search by association, not by asset name. Rex personas do too.
-
Content knowledge grounded in real metadata. Each persona already carries what a real employee would: a glossary definition they have read, a table they know refreshes every six hours, a metric they know appears four times in the glossary. And they have reactions to what they have read: one persona found a table description too brief to understand the full scope, another wished a term definition listed the actual tier names. Those reactions seed realistic follow-up questions.
-
SQL comfort, ranging from writes to reads to none at all, because “none” is most of the world.
-
A frustration scenario anchored to a real gap. Not “user gets confused” but “user searches for inventory, hits thousands of undocumented tables in a noisy cluster, and cannot find the canonical source.”
-
An urgency level. S&OP review Thursday. Quarterly incentive deadline. Client deliverable due Friday. Pressure changes how people ask, how much patience they have, and how quickly they escalate or abandon.
-
Parallel channels. Who else this person can ask, because that controls patience. A persona who can Slack a colleague gives the AI two or three turns before switching. A persona with no alternative waits longer, and the stakes of failing them rise accordingly.
-
The emotions they arrive with, and the ones the answers create. A persona never comes in neutral. It arrives anxious ahead of the Thursday review, skeptical after last quarter’s wrong number, hopeful that this tool is finally the one that works. And every response moves that state. A precise, well-sourced answer builds confidence. A vague one plants doubt.
Two useless answers in a row produce the quiet resignation we have watched in real users over and over again. Rex tracks these psychological responses turn by turn, because they decide what a real person does next: push deeper, double-check somewhere else, or stop asking altogether.
-
Whether the answer is for them or for someone else. A persona relaying an answer to their manager holds it to a harder standard: could this response be pasted into an email, unedited, and survive? Technically correct but untransmittable counts as a failure.
-
A trust threshold. What this person needs to see before they will put a number in front of their CEO: a verified tag, a named owner rather than a system ID, an explicit list of exclusions, refresh frequency confirmation.

Then the conversation runs, and the persona behaves like the person. It reacts to what the AI just said. If the answer is useful, it goes deeper. If it hits its patience threshold, it abandons the thread, and that abandonment is the finding. Every turn is scored across seven dimensions, including the most important one: what caused the gap when something broke?

Live tests that end in action. Rex does not run against a mock. Every simulation runs against the customer’s live instance, down the same path a real user’s question would travel, starting with conversational search and extending to every other surface we want to test the same way. And when something fails, Rex does not just record that it failed. It tells you why it failed and which part of the system did not hold.
So what comes out of a run is not a report to file away. It is two sets of meaningful, actionable next steps, deliberately kept separate. For our team: exactly where the product failed and what to fix, whether that is language that was too technical, context lost over a long thread, or a pattern of missed questions. For the customer: exactly what to improve in the context layer for the product to get even better, the specific assets that need a description, an owner, a glossary link, alongside an honest picture of where things stand today and how quickly they can improve.
Every test ends in action on both sides. Merging the two would dilute both.
Below: what that report becomes for a customer. This executive summary is from a separate engagement, a global cruise and hospitality company, distinct from the field example later in this piece.

What this unlocked
Before Rex, conversations with customers about AI readiness were necessarily high-level. Here is what is generally possible. Here is what usually goes wrong. Good advice, but generic.
Rex changed the register of those conversations. We can now speak very specifically about what is possible for a customer given the context they currently exist in. Not a generic “you need better documentation” but a named list of exactly what to fix.
And it gave us the evidence we said we needed for the gatekeepers. When we ran Rex on their instances, the champions were seeing, for the first time, a vendor that had built a system to stress-test its own product in the customer’s context. Not a demo on our data. Not a benchmark on someone else’s. Their metadata, their business, their people, simulated across a wide variety of personas and question types.
Rex became a point of proof: here are all the ways this will work for you, and here, honestly, are the places where it will not yet. Trust with a gatekeeper is not built by claiming your product works. It is built by showing them you went looking for the ways it fails.
Here is what that looked like in the field.
A global medical devices manufacturer had Conversational AI installed but had been holding back on rollout. Their head of data was thinking carefully about whether the context layer was ready for end users to depend on it.
We ran Rex on their instance. Rex generated fourteen personas grounded in their business: clinical operations, regulatory affairs, manufacturing yield, supply chain. Each persona ran multi-turn threads against their live Conversational AI. Forty-four threads. Two hundred and sixty-four turns.

Eighty-two percent of the conversations passed fully. Eighteen percent came back partial, and almost every partial traced back to a missing description, a missing README, or a glossary term that needed linking. The AI was not the problem. The context layer had specific, addressable gaps.
I sent their head of data the report. The first thing they noticed was not the pass rate. It was that every gap was named, and every gap had a fix action. Rollout was approved the following Monday.
Not just finding gaps. Closing them.
The gaps Rex finds sort themselves into two kinds. Some are simple gaps in the context itself: missing definitions, undocumented assets, descriptions that never got written. Those can be closed automatically by our Context Agents, without a human needing to triage each one. The human moves from in the loop to on the loop: reviewing what was closed rather than doing the closing.
Others cannot be fixed automatically, like ownership information that lives in someone’s head rather than in any system we have access to. Those get flagged, routed to the people who can actually resolve them.
Testing, finding, explaining, closing. It becomes a complete loop on context.

If that word “loop” sounds familiar, it should. We have written before about building a feedback loop for the quality of context: a lightweight, in-product mechanism that captured signals from human users about exactly where the context layer was breaking down. Rex started differently, as an experiment in encoding human behavior into a simulation engine. But it ended up doubling down on the same problem we had solved before in a different form.
The earlier loop caught context gaps one human signal at a time. Rex catches them at a completely different scale, hundreds of simulated conversations at once, before a single real user ever hits the gap. The form evolved. The loop stayed.
And the loop is not tied to a single surface. We started by stress-testing conversational search, then extended Rex to task-testing our MCP, and our customers are now running early experiments on what Rex looks like for their own internal AI applications: their personas, their definitions of success, tested before their users ever arrive. Because the core of Rex is the personas, the same humans carry over to whatever surface they touch next.
The next step is making this continuous. We are bringing Rex to customers on demand as part of our simulation capabilities, and you will see this very soon: simulations running on an ongoing basis, reports generated as the instance evolves, gaps routed the moment they appear, and the loop on context closing continuously rather than once.
Where the best research ends up
Traditional user research produces understanding: a report, a deck, a ticket, and the hope that the insight survives its translation into a roadmap. Simulation changes that equation. The same empathy gets encoded into living instruments of understanding, deployable at scale, and research stops merely informing the product and starts gating it. Before a rollout, Rex runs. If Rex finds the rollout is not yet ready, we know exactly why, and exactly what work remains.
None of it works without the fieldwork. The five years are not the backstory. They are the engine.
Which brings me back to the people who search by association, type “the booking thing,” and quietly message the one person who always knows. Somewhere in Rex today there are personas shaped by them, asking their questions and hitting their walls before any real user has to. They will never know. That is the point.
The best research does not end up in a deck. It ends up in the product, working quietly on behalf of the people it came from.
