Blog
What should a user research agent do?
What we're learning at Bolt about building an AI user research agent on an LLM wiki, and about making shared knowledge work for a team.

The art and craft of creating LLM wikis feels like gardening.
A long time ago (centuries in AI terms, I think), in early 2023, we built our first gen AI chatbot to help teams reach insights faster. We used it with about 30 distributed innovation teams at Deutsche Bahn, each running a customer discovery project during an innovation sprint. It helped them bring their research together and work on a shared synthesis. To us, it felt like a new way of putting atomic research into practice.
Looking back, it was a simple RAG system on top of a voice-to-text pipeline, where people could chat with the interview transcripts. But the executive product manager who had set the challenge took time to explore the teams' growing research findings through the chatbot, asking questions across the different sprint projects. That experience shaped the vision we still work towards today: bringing customer insights to where decisions are made.
We are back at that same question today, working with Bolt to build a user research agent. At its centre is an LLM wiki built from the team's research studies and customer call transcripts. With RAG, the agent searches the raw transcripts again for every question. With the wiki, knowledge compounds, because each new study is worked into what is already there. People ask questions about customer insights in Slack, or via MCP from their own AI tools. A maintenance routine ingests new studies and customer call transcripts, keeps the wiki searchable and cleans it up.
Today these are separate building blocks. We call the whole thing an agent because we are aiming for a loop in which use, ingestion and maintenance feed each other, and the wiki maintains and extends itself as part of the researchers' daily work.
Working on the wiki feels like gardening to me. We start with a structure, observe how the agent uses it, and experiment with small changes. More on building and maintaining an LLM wiki in Everyone can become a knowledge builder.
I think there are three levels here:
- Getting your own personal knowledge system up and running. That's pretty straightforward.
- Making the knowledge compound over time while keeping the system from drifting and rotting.
- Getting this to work for a team. That's a completely different story.
With Bolt, we are working on the third.
Which role should the agent play?
We spend a lot of time thinking about wiki structures and agentic context engineering. I find this work fascinating. But the question for our user research agent is whether it helps people make better product decisions.
Jonny Longden makes a related point about experimentation: "Experimentation was never the point. Making good decisions was."
For us, that raises questions about how the agent responds. Can it push back when a product manager is only looking for evidence to support an idea? Can it say it doesn't know and suggest further research? Can it judge how far the available evidence supports a particular decision?
That leads to a difficult question about the agent's role. Should it help an experienced researcher analyse evidence and reach their own conclusions? Should it formulate research insights directly for product managers? Or should it act as a product sparring partner, questioning ideas and suggesting solutions?
We see requests for all three in our experiments. Deciding where to position the system is proving difficult.
Jeff Gothelf asks: "The real question isn't 'can we?', it's 'should we?'" We see that question in the effects on researchers' influence, motivation and relationships with colleagues.
When a researcher's judgment becomes part of an agent's instructions, who retains influence over it, and who benefits? I explore that tension, starting from Garry Tan's talk, in my new article on knowledge ownership (in German).
What we're learning at Bolt
Knowledge is not an island. Understanding what drivers say requires context about the product and how it works in a specific market. For now, most of that context lives in people's heads, or in planning and strategy documents outside the research team. The wiki holds studies and transcripts, so the agent cannot make these connections on its own yet. We need someone who can tell when context is missing or out of date.
Automated transcripts and meeting notes have not been accurate enough for our purposes. The data is there, but the tools we tried, from off-the-shelf AI interview platforms to meeting transcription, promise more than they deliver. We now refine the material before it enters the wiki, checking it against the source. Cleaning up a sentence must not turn uncertainty into something the participant never said. Getting from speech to text a team can rely on is a harder problem than it looks, and an interesting one.
Writing evals is demanding work: researchers need reference answers drawn from their research and domain knowledge independently of the agent and its wiki. These help judge its responses, including when it should say it doesn't know. Their own research keeps that judgment grounded. So far, our evals have focused on getting the facts right. They do not yet cover how the agent delivers them, and that brings back the question about its role. More on this contribution through evals, skills and shared work with agents in The future of customer research.
Then there is the learning loop, which we have not solved yet. In my personal system, I ask the agent a question, inspect its answer and review the session. I can see where the agent took a wrong turn, collect the problem, and improve the wiki structure or the checks used to maintain it.
In a team, the loop is longer. Research informs an answer. A product manager uses it in a decision. A feature gets built, and customers use it, struggle with it or ignore it. For learning to accumulate, some of that experience needs to find its way back into the shared knowledge. We need to know which assumptions to revisit, who notices, and who takes responsibility for the update. A well-supported answer is only one part of that work.