When your product stops having a single path, how do you know it's working?

Measuring agentic AI products means following the full path from a customer's intent through execution, correction, and outcome. I call that sequence the agentic product loop. It is not the same as scoring a single chat turn or a step in a funnel you designed.

If you have spent your career building traditional software, you are used to a comforting idea: you designed the path.

Step one. Step two. Step ten.

You mapped the milestones and defined the route customers would take. You conducted qualitative research, shipped a baseline, measured the funnel, and tested whether changing step three improved conversion at step four.

The path was largely deterministic, and therefore measurable.

Agentic products change that contract. The customer tells the system what they want to accomplish and why it matters. The agent decides how to get there, working within the tools and context available to it.

For one customer, that might produce a ten-step workflow. For another, the agent might jump ahead, backtrack, or pause to ask for clarification. Both paths could succeed. Or a workflow that appears successful at turn three could fail when the customer inspects the result at turn six.

That changes the unit of measurement. We can no longer evaluate only a click, screen, or individual response. We have to consider the sequence from the customer's original intent through execution, correction, and the eventual outcome.

We encounter this problem every day with GoDaddy Airo AI Builder. Customers do not follow one prescribed path. They state a goal, and the system adapts to help them reach it.

That is powerful. It is also a measurement problem.

This post explains why traditional product measurement no longer tells the whole story. In a follow-up, I will share how we are approaching the problem with Product Loop in Airo AI Builder.

When the path is no longer fixed

In traditional product development, success often means progressing from one predefined step to the next. Product teams know the expected route because they built it.

In an agentic product, the workflow is not fully specified in advance. The customer supplies the intent, while the agent selects, and sometimes revises, the steps required to fulfill it.

Two customers can begin with the same goal, take different paths, and both succeed. Two others can follow similar paths but end with very different results.

You are no longer measuring progress through a route you designed. You are trying to determine whether a dynamic trajectory produced the outcome the customer wanted.

That is non-deterministic software: the path is not fixed in advance. Measuring agentic AI products means measuring outcomes and trajectories, not just clicks on a screen you designed.

The GPS analogy

The analogy I keep coming back to is GPS navigation.

The driver provides a destination. The navigation system determines a route based on the available context. Two drivers heading to the same place might take different routes and both arrive successfully. If one misses a turn or encounters a road closure, the system can recalculate the path.

You would not measure the success of navigation by asking whether every driver passed through the same intersection. You would ask whether they reached the intended destination, and how much delay, confusion, or unnecessary rerouting occurred along the way.

Agentic software works in much the same way. The customer states a goal, while the agent determines and revises the steps required to reach it. A deviation does not always mean failure, and completing an intermediate step does not mean the customer's job is done.

That makes the questions product teams rely on harder to answer:

  • Did the product help the customer reach the intended outcome?
  • Did the requested change actually happen, or did the agent merely say it happened?
  • Was the agent adapting productively or repeatedly recalculating without making progress?
  • Did the agent recover from a problem, or did the customer give up?
  • Did we complete the customer's job, or just complete one turn?

In a button-and-screen product, the possible routes are relatively constrained. In a chat-first agentic product, customers can ask for things you did not explicitly wire into a flow.

That is part of the promise: customers can pursue what they need without waiting for a product team to ship a dedicated workflow for every variation.

It is also why traditional dashboards can feel increasingly incomplete.

A response is not an outcome

Consider a common pattern in GoDaddy Airo AI Builder.

A customer uploads images and asks the agent to place them on a website. The agent responds that the work is complete.

If the analysis stops there, the interaction might be counted as a success.

But then the customer says they cannot find the images where they expected them. The agent tries again. After another correction, the customer confirms that the images are now in the right place.

So, was the interaction successful?

The first response appeared successful, but the product state did not match the customer's intent. The workflow contained a failure, additional customer effort, and an eventual recovery.

Evaluating only the first response would record a false success. Evaluating only the failure would miss the recovery. Evaluating only the final response would hide the friction required to get there.

The meaningful unit is the complete loop: what the customer asked for, what the agent did, what changed in the product, how the customer responded, and whether the intended outcome was eventually achieved.

Agentic product management has to compare outcomes and trajectories across that loop, not assume everyone hit step four of your funnel.

Three questions traditional dashboards struggle to answer

1. Did the customer's job complete, or did one turn merely look successful?

Agent responses are not reliable evidence that the requested work occurred. An agent may produce a confident confirmation while the resulting product state is incomplete, incorrect, or different from what the customer intended.

Success often becomes visible only across multiple turns and through what actually happened in the product.

2. Was the path productive, or was the customer stuck?

A longer conversation is not necessarily a failure. Some goals genuinely require clarification and iteration.

But additional turns can also indicate repeated attempts, misunderstood intent, unnecessary corrections, or a customer trying different language because the agent is not making progress.

The same number of turns can represent healthy collaboration in one workflow and serious friction in another.

3. What kind of problem occurred?

A frustrated customer does not automatically indicate a product bug.

The underlying cause might be a missing capability, an agent reasoning error, insufficient context, an unsuccessful tool call, an ambiguous request, or an edge case behaving as designed.

Without enough product context, analytics teams may report gaps in features that were never intended to exist. Engineering teams may investigate bugs that are actually failures of interpretation. Product and go-to-market teams may see the same conversation and reach different conclusions.

Everyone has part of the picture, but no one can confidently say whether the agent completed the job.

Why chat traces are necessary, but not sufficient

Conversation traces are essential. They show what the customer asked, how the agent interpreted the request, and how the interaction developed.

But a transcript primarily tells you what the customer and agent said. It does not necessarily tell you what the system did.

An agent can say that it added an image, published a site, updated a page, or changed a setting. The conversation alone cannot always confirm that the action occurred correctly, or that the resulting state matched the customer's intent.

The reverse can also happen. An operation may succeed technically, but the customer may not recognise the result, find it useful, or consider the original goal complete.

Understanding the full workflow requires the conversation to be interpreted alongside the product context in which it occurred. Without that connection, apparent success can hide failure, apparent failure can hide recovery, and customer friction can be difficult to diagnose.

That gap is why we are building Product Loop for GoDaddy Airo AI Builder with Ashish, Ankur, and our agent reliability team. I will go deeper on how it works in a follow-up post. This one is about the problem it is trying to solve.

The scope of the problem matters

Once you begin looking beyond individual turns, it is tempting to evaluate everything the agent might possibly do.

But broader scope does not automatically produce better understanding. As the evaluation boundary expands, context becomes noisier, workflows become harder to compare, and it becomes more difficult to determine why something succeeded or failed.

The challenge is not simply collecting more AI-generated analysis. It is identifying a meaningful product capability, understanding the context around it, and evaluating the customer's complete workflow within that boundary.

For internal systems like Product Loop, that means focusing on specific capabilities and workflows rather than trying to explain every possible thing a customer might do in Airo AI Builder. Get the unit of analysis right, and the system becomes useful. Try to boil the ocean, and you spend more time debating outputs than acting on them.

That discipline matters for customer-facing agents as well as the internal systems teams build to understand them.

A new unit of measurement

When the product no longer has one fixed path, its unit of measurement has to change.

A response can look correct while the workflow fails. A workflow can contain failures and still recover successfully. Two customers can pursue the same goal through different trajectories and both reach the right outcome.

Clicks, funnels, and individual turns still provide useful signals. They just do not tell the complete story.

To understand whether an agentic product is working, we have to follow the customer's full loop, from intent to execution, correction, and outcome.

The first challenge is recognising why the old measurement model is incomplete. The next is building one that can account for the way agentic products actually work.

In the next post, I will walk through what Product Loop does today: how we infer scenarios from real usage, how we look past a single turn to see if a workflow succeeded, how we connect conversation signals to what the product can and cannot do, and how that feeds engineering priority.

If you are a PM or engineering leader shipping agentic experiences, the shift is not optional. The product stops being a fixed path. Your measurement model has to catch up.

Start building with GoDaddy Airo AI Builder — describe your idea and watch it come to life. 50 free AI credits to get started.

FAQ

What is an agentic product loop?

It is the full sequence from a customer's original intent through what the agent did, what changed in the product, how the customer responded, and whether the intended outcome was achieved. Measuring agentic AI products means evaluating that loop, not a single chat response or funnel step.

What is agentic product measurement?

It is how you evaluate whether an AI agent completed the customer's job across a full workflow, including follow-up turns where they confirm success or report a problem. The agentic product loop is the unit you measure.

Why don't traditional funnel metrics work for agentic AI?

Funnel metrics assume a designed path. Agentic products let customers and agents co-create the path. The same goal can succeed through different step counts, so step-based conversion alone undercounts failure and overcounts single-turn wins.

How do you know if an AI agent completed the task?

Look past the final message in a single turn. Walk the conversation forward: what did the customer ask for, what did the agent do, what changed in the product, and what happened in the next few turns? Task completion shows up in that sequence, not in one reply.

What should product managers measure for agentic AI?

Prioritise task completion across the full loop, multi-turn satisfaction, recovery after failure, and whether failures cluster around specific capabilities. Pair conversation traces with product context so you can separate bugs from prompt clarity issues.

What is agentic product management?

It is product management when the interface is a goal-directed agent, not a fixed UI. You optimise for outcomes and trajectories customers actually take, not only the flows you designed in a wireframe.