Mentiora
DT Demo Transfer Service
Copilot AU
Blog post: How Mentiora helps CX teams build an assistant customers can trust

All names, identifiers, conversations, and results in this demonstration are illustrative and synthetic. No real customer cases were used. Any resemblance to real people, companies, journeys, or records is coincidental.

How Mentiora helps CX teams build an assistant customers can trust

Every company has its own promises, policies, escalation paths, and standards for a good customer conversation. Mentiora lets a CX team express those standards through its own Scenarios and Judges, then presents the resulting evidence in an approachable interface that remains connected to the agent's detailed configuration. A CX team can understand what customers received, improve the behavior that produced it, and build justified trust that the assistant represents the company well.

This article presents an in depth view of the Mentiora quality loop for teams that want precise control over their agent. It follows a complete loop through many of the platform's quality features. Teams can also begin with a simpler plug and play setup, then add controls as their needs grow.

Scroll to follow the work inside Mentiora.

Mentiora connects every stage of customer agent quality in one working loop

An AI agent for CX answers real customers all day, and while a reply might sound correct, it can still apply the wrong policy, leave the case unresolved, or state something that the evidence never supported.

A complete agent evaluation loop defines the behavior customers should receive, observes what the agent actually did, traces the cause, and checks the change against the same situations again.

A CX loop starts from an agent configuration, then continues with Scenario definition. Scenarios are the customer situations we want to test. As each company has its own set of customer policies, we also need to define Judges, so our agent is evaluated under a precise company context. Having both scenarios and Judges, the evaluation can then run and produce a score, with detailed rationales and proof behind each score. This article follows a complete evaluation loop.

The Mentiora Quality Loop Six steps run in a circle. A saved agent configuration, a scenario, a judge, a run, the evaluation results, and an investigation. The investigation produces the next agent configuration, which closes the loop. Mentiora Quality Loop 1 Configuration 2 Scenario 3 Judge 4 Run 5 Results 6 Investigate

Welcome to the Mentiora Platform.
In this follow-along use case, the blog post you are about to read uses the Mentiora Platform interface to help you visualise each concept.

A preview of the platform's content appears on the right.

The platform's navigation tabs appear on the left, and they mark which part of the platform the follow-along is showing.

Chapter 1Trust begins by separating fluent language from truly supported company services.

To make this quality loop concrete, we will follow in this article Demo Transfer Service, a fictional airport transfer company. Its support assistant helps customers manage bookings, understand company policy, and respond to disruptions such as delayed flights or missing drivers.

A simulated customer experiences the assistant as one continuous voice, although each answer may depend on several kinds of evidence:

  • Customer data establishes the facts of the traveler's situation, such as whether a driver has been assigned.
  • The illustrative company policy explains what should happen, such as when a delayed flight qualifies for automatic tracking.
  • Operational records confirm what has already happened, such as whether the transfer was updated.

The support assistant must combine these sources into one clear answer and quote its sources, and refrain from presenting a policy that does not exist or actions that were not taken. These questions, and all quality considerations, can now be answered thanks to the Mentiora Platform.

Demo case

The simulated traveler asks whether her airport transfer still works after a flight delay, whether the price changes, and whether a driver has been assigned. The illustrative agent answers the first two questions with calm confidence.

Demo case

The booking says flight DEMO-FLIGHT-271 moved from 18:10 to 20:30. The illustrative company policy says a tracked delay of no more than four hours moves the transfer with the flight and leaves the fare unchanged. The agent therefore has support for its answer about the transfer and the price.

Demo case

The agent was expected to answer the first two questions from the illustrative company policy and treat the third as unverified. A correct reply would explain that the transfer follows the tracked flight and that the fare stays unchanged, then state that the available reservation information does not confirm a driver assignment. Instead, the reply says that a driver has been assigned. That statement presents a completed operational action as fact even though no assignment appears in the information available to the agent.

Mentiora keeps all conversations connected to their information and execution evidence behind it, which allows a CX team to preserve the useful policy explanation while locating an unsupported promise. Once a CX team can see the exact boundary the reply crossed, the next requirement is to define that boundary as part of the company's own standard for good service.

Mentiora uses Scenarios to simulate customer journeys and customer questions. The Scenario view brings together the customer persona, the situation, the information available at the start, and the outcome the agent should achieve. This gives the CX team a reusable service test that Mentiora Studio can run against different agent versions.

Demo case

Demo Transfer Service turns its standard for the simulated traveler's journey into a Scenario. It gives the simulated customer her reservation number, the old and new flight times, the fact that flight tracking was added early enough, and the three questions she needs answered. The simulated customer can phrase the request, react to the answer, and pursue missing information as the conversation develops.

Success criteria describe what the agent must accomplish during a Scenario. They can require a policy outcome, a particular piece of information, an honest statement about unavailable data, or a specific escalation path. This gives CX and operations teams a concrete service outcome to review, instead of leaving the agent with a general instruction to be accurate.

Once Scenarios describe the customer journeys and their expected outcomes, the CX team needs a consistent way to examine the conversations those journeys produce. Mentiora uses Judges for this second part of the quality definition. Each Judge focuses on one dimension, reads the company policies, applies a written scoring guide, and records a reason with every score, so the team can distinguish a policy problem from an issue involving resolution, tone, evidence, privacy, or safety.

Demo case

The Evidence discipline rubric scores whether claims about reservations, prices, vehicles, and completed actions are supported. Its highest score requires every operational claim to have evidence and missing facts to be identified honestly. This is possible as judges have access to company policies and documents. Several correct policy explanations can no longer conceal one invented action.

Scenario success criteria describe the service outcome, while Judge rubrics explain why a conversation met or missed it. Together they give a CX team a stable and inspectable definition of quality before any comparison begins.

Chapter 2Compare versions in Mentiora Studio

Mentiora Studio simulates the selected Scenarios and then evaluates the resulting conversations with the Judges the CX team has defined. The same evaluation can run against two complete agent versions, which lets the team compare how each version handles the same customer journeys and service expectations.

Demo case

Demo Transfer Service selects live v9 as the baseline and candidate v13 as the version it wants to evaluate. The run setup pairs both versions with the same twelve Scenarios. Mentiora Studio creates one fresh conversation for each version and Scenario combination, producing twenty four conversations whose scores remain connected to their transcripts, Judge reasons, and traces.

This is a paired comparison of generated conversations, not a replay of identical transcripts. Both versions use the same Scenario definitions, success criteria, and Judges, while each simulated customer is free to phrase the request and respond naturally. The comparison therefore keeps the service situation consistent while testing whether both agent versions can handle reasonable variation in the conversation.

Demo case

Mentiora Studio shows three Scenarios improving and two regressing in v13. The illustrative estimated cost of the twelve conversations falls from $0.1645 to $0.1381. The summary directs attention to changed customer journeys so the team knows where to focus its attention.

Chapter 3Investigate a weak conversation

Mentiora Studio summarizes the comparison, then keeps every result connected to the conversation that produced it. A CX team can open any weaker result and examine the customer exchange beside its Judge scores, written reasons, and execution trace. The summary identifies where attention is needed, while the underlying evidence shows what the customer experienced.

Demo case

In this new Demo Transfer Service Scenario, a traveler reports a missing pickup.

The reservation score combines two rubrics. Necessary fact selection and policy accuracy each receive two points out of four. The combined score identifies a weak dimension, while the written reason explains how that weakness affected the customer.

The written reason explains that the agent did not provide the exact information required to continue the missing pickup process. The traveler reached the right topic without receiving a complete resolution. This directs the team toward the agent behavior that affected the outcome.

To analyse a conversation, the Conversations page can be used. It opens the execution trace next to the conversation. It organizes the run by customer turn and shows the route, instructions, knowledge, workflows, and conversation context that were available when the agent produced each answer.

Demo case

In the Demo Transfer Service conversation, the traveler's second turn reached the missing pickup workflow. The trace therefore shows that routing occurred, although the resulting answer remained incomplete. The CX team can now focus its investigation on the workflow and the conditions that selected it.

Chapter 4Improve the configuration

The comparison has shown the team which customer journey needs attention, and the trace has shown where the incomplete answer came from. The team can now improve the configuration that selected and guided that workflow.

Mentiora Rules are saved routing conditions. A Rule looks for a defined customer situation, its priority determines when it is considered, and its effect sends the conversation into the appropriate workflow. This gives a CX team precise control over important service paths.

The Rules page shows the name of each service rule, the customer language that activates it, its priority, and the workflow it starts. Reviewing the list helps the team see where two rules could react to the same customer message before changing the agent configuration.

Opening one Rule shows its activation condition in detail. The keyword match lists the words that can activate the Rule, while the effect shows which workflow receives the conversation. The CX team can therefore review the exact routing condition rather than infer it from the agent's reply.

Demo case

Demo Transfer Service's earlier Rule reacted to the common words “escalate” and “pickup.” A billing question that contained both words could therefore enter the missing pickup workflow. The team kept the genuine missing pickup route and narrowed its activation language so that unrelated questions would no longer enter it.

Chapter 5Verify the improvement

After improving the configuration, a CX team can return to Mentiora Studio with the same Scenarios and Judges. Fresh conversations show whether the targeted improvement appears in the intended customer journey while every other Judge remains visible, including the service dimensions the team cannot afford to weaken.

Demo case

Demo Transfer Service evaluates candidate v17 against live baseline v9. Reservation accuracy rises from 82 to 93, showing that v17 improved the previous result.

Chapter 6Decide whether to release

Mentiora connects Demo Transfer Service's service definition to generated conversations, Judge reasons, traces, routing Rules, and Mentiora Studio results. Together, this evidence supports the CX team's decision about whether the candidate should become live.

Demo case

Live baseline v9 remains selected. Candidate v13 records the earlier comparison, while candidate v17 is saved, published, and available for review without serving customers. Keeping those states separate lets the team inspect the candidate and its evidence without changing the live experience.

If Demo Transfer Service accepts the candidate, the team can select v17 as the live version. The quality loop leaves that decision with the people responsible for the customer experience and keeps the supporting Scenarios, standards, and configuration available for review.

The Mentiora Quality Loop and the page that holds each step Configuration lives on the Agent page, Scenarios on the Scenarios page, Judges on the Judges pages, runs and results in Mentiora Studio, and investigations on the Conversations page. Mentiora Quality Loop 1 Configuration Agent page 2 Scenario Scenarios page 3 Judge Judges pages 4 Run Mentiora Studio 5 Results Mentiora Studio 6 Investigate Conversations

Mentiora gives CX teams a connected way to define, investigate, improve, and verify agent behavior

The simulated traveler's unsupported driver claim looked plausible because correct booking and policy information surrounded it. Preserving the source of consequential claims lets a team keep the useful explanation while refusing authority the agent has not earned. Scenarios, Judges, Mentiora Studio, traces, and saved versions carry that distinction from one conversation into a controlled release decision.

Your gateway to the Quality Loop

Mentiora transforms how your company meets its customers: one agent, one memory, every channel, built for your business.

Book a demo
Scroll to read the post