Introducing Insights: Finding What Spot Checks Miss
Voice agents in production generate more conversations than any team can review. Dashboards answer the questions you thought to ask when you built them. Spot-checks find what you happen to look at. Neither is particularly good at finding the behavior nobody knew to search for. Liberate built Insights to close that gap: It’s an investigation agent that answers plain-language questions about an agent's conversations and returns a cited report, without requiring a bespoke evaluation pipeline for each question.
The gap
Liberate runs roughly a thousand-plus calls an hour across deployments. Each call leaves a transcript. Behind each transcript sits the configuration that produced it: system prompts across multi-agent setups, tool definitions, prompt and tool revisions, evaluation results, and the API logs from the customer integrations the agent called during the conversation.
The questions we get asked about this data vary. "What were the common mistakes last week?" "Is the agent re-prompting callers too often?" "Did it give any unsolicited insurance advice?" "Why did the completion rate drop after Tuesday's revision?" Each is a reasonable question. Each, until now, prompted one of two actions: someone reading a sample of transcripts by hand, or someone building a bespoke analytics pipeline for that one question.
The first option does not scale. The second does not keep up. The hard part is not volume alone. It is that the evidence is large, unstructured, and spread across artifacts with different shapes. Answering "why did the agent do that?" requires joining a line in a transcript to the prompt revision live at that moment, to the tool the agent had available, to the response the customer's API returned. That join is what a human engineer does when they investigate an incident.
The investigation is slow, but it is also how we found the 91 conversations that made the problem concrete. During a compliance review, we found our agent telling callers their claim had been "filed" before a confirmed submission number existed. Not once. In 91 conversations. Fortunately, the agent was still in test mode, and not in production yet. Nothing in our evaluations flagged it, because no evaluation had been written for it. Somebody looked, and only then did it exist.
Where the idea came from
Coding agents are already good at a version of this problem. Point an agent at an unfamiliar codebase and it does not read every file. It searches for candidates with grep, opens the ones that matter, follows references across files, and, when it works well, only claims what it can point back to in source.
That behavior is exactly what a quality investigation over conversations needs: locate candidates, read selectively, follow references from behavior to configuration, and cite the evidence.
So we stopped treating voice agent analytics as a dashboard problem and started treating it as a codebase investigation problem. If conversation data is structured in a form these agents already handle well (a filesystem of readable text with clear naming and cross-references), an agent equipped with bash, grep, and file read tools can investigate a week of calls the same way it investigates a repository.
The complication is size. A codebase fits on disk. A quarter of conversations across a hundred agents does not fit in a working set an agent can grep efficiently, and most of it is irrelevant to any given question. That drove the two-stage design: semantic search runs first against the entire call dataset, narrowing the agent’s task to just the relevant samples.
What we built
A user asks a plain-language question about an agent's conversations. Insights returns a report with citations back to the specific conversations and configuration that support each claim.
Data path. Conversation data sits encrypted in S3. An ingest step decrypts, extracts, and formats it, then loads it into a vector store alongside its metadata: Agent, revision, timestamp, outcome, tools called, evaluation scores.
Stage one: retrieval. At query time, the system uses semantic search plus metadata filters to pull the conversations relevant to the question. This is deliberately a wide net. Recall matters more than precision here, because stage two is cheap to run over a few hundred extra files and expensive to run over a missing one.
Stage two: materialization. The hit set is written to a local filesystem as a structured workspace. The layout is designed for an agent to navigate:
Agent identity and configuration, with prompts and tool definitions stored by revision, so the agent can see exactly what was live when a given call happened.
One directory per conversation, holding metadata and the full transcript.
Evaluation artifacts attached to the conversations they scored.
Customer-integration API logs where available, so a claim the agent made can be checked against what the backend returned.
Stage three: investigation. An LLM agent runs over that workspace with file tools such as read, grep, and glob, plus bash execution. It works in a loop: search, verify against the source files, then write. It greps for candidate patterns, opens the transcripts that match, checks the claim against the prompt revision and the API logs, and only then writes the finding into the report with the conversation IDs attached.
No per-question pipeline. No new dashboard. The question is the specification.
What we found
Behavior has to be joined to configuration, or the answer is a score and not a cause. Seeing completion rate fall between revision 14 and revision 15 is a signal. Seeing that revision 15 changed a tool description and that the agent started calling the wrong tool in the conversations that failed is a diagnosis. The revisioned prompts and tools in the workspace are what make the second kind of answer possible.
The compliance failures are the ones nobody wrote an eval for. The "filed without a submission number" case is representative. It was not a hallucination in the usual sense. The agent said something reasonable that happened to be false in a regulated context. No metric existed for it because no one had anticipated it. An open-ended question ("Where does the agent assert something it has not confirmed?") is what surfaced it, and that question would not have justified a two-week pipeline build.
Wide retrieval, then narrow reading, is the right split. Semantic search alone gives you a ranked list and no verification. Grep alone over the whole corpus is too slow and too noisy. Retrieval to a working set, then grep and read within it, gets the coverage of the first and the rigor of the second.
Citations change how the reports get used. A finding with conversation IDs attached gets acted on. A finding without them gets debated. Requiring the agent to point back to source before it writes a claim made the reports shorter and more trusted, and it cut down on the agent asserting patterns it had inferred rather than observed.
Development and testing are where this pays off first. The 91-conversation case, and others like it, were caught in test and staging traffic before they reached production callers. The cost of finding a compliance issue early is a prompt change. The cost of finding it late is a regulator.
On-demand questions replace a backlog. Customers ask questions we have no prebuilt view for. Before, those went into a queue for custom analytics work measured in weeks. Now the answer comes back in minutes, which means the customer asks the follow-up question, and the one after that.
The production reality
Three limits are worth stating plainly.
Retrieval recall is the ceiling. If stage one does not pull a relevant conversation into the working set, stage two never sees it. Semantic search misses things, especially behaviors described differently from how they appear in transcripts. We compensate with wide nets and metadata filters, and we treat "the agent found nothing" as a weaker claim than "the agent found this."
The investigator is itself an LLM. It can over-generalize from a handful of examples or misread a transcript. The citation requirement is the control: a claim without a pointer to source does not go in the report, and reviewers check the pointers. Our Insights tool finds candidates and evidence. A human still decides what it means and what to change.
Investigation costs more than a dashboard query. An agent that greps, reads, and verifies across hundreds of files spends real tokens and real minutes. That is the right trade for a question asked once, or for a question whose shape changes every time. For a metric you will watch every day, build the dashboard. Insights is for the questions that come before you know which dashboard to build.
What comes next
The immediate uses are the ones you might guess: comparing agent revisions and explaining regressions by cause rather than score, answering bespoke customer questions on demand, and running compliance sweeps over development traffic before a change ships.
The use we expect to matter most is the loop. Run Insights over a large set of conversations. Change the agent based on what it finds. Run it again. Each pass tightens the agent against behaviors that no eval anticipated, because the question set is not fixed in advance.
Quality at scale has never been a problem of reading harder. It is a problem of being able to ask a new question and get a cited answer before the answer stops mattering. Coding agents learned how to do that over source code. It turns out that when conversations, configurations and system evidence are structured the right way, they can be investigated much the same way.
And that changes the quality loop from test what you know to look for to investigate what you didn't know to ask.
Key takeways
At production scale, manual review is a sampling method, not a quality method. It finds the failures you already suspect and misses the rest.
Coding agents already know how to investigate large, unstructured corpora with search tools and cite what they find. Conversation data can be laid out so the same skills apply.
A two-stage design (semantic retrieval to cast a wide net, then a materialized filesystem the agent greps and reads) scales the approach to very large corpora.
The result is a closed loop: ask, find, change the agent, ask again.
By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.