# How GoDaddy Builds AI Agent Evaluation Datasets from Real Customer Workflows

In [Part 2](/discover/learn/product-loop-part-2-finding-bugs-nobody-reported), we explored how Product Loop helps our GoDaddy Airo AI Builder team find and diagnose customer problems. This post looks at another use of the same production signals: extending our **AI agent evaluation datasets** with scenarios drawn from how customers actually use Airo.

If you are new to the series, [Part 1](/discover/learn/measuring-agentic-ai-products) explains why agentic products break traditional measurement. The meaningful unit is the full loop from intent through correction to outcome, not a single chat turn.

> **Key takeaway:** Specification-based evals define expected behavior. Production conversation patterns reveal how customers actually phrase requests, combine steps, and refer back to earlier turns — turning real workflows into regression tests and deeper eval coverage.

## In this article

- [What are AI agent evaluation datasets?](#what-are-ai-agent-evaluation-datasets)
- [Start with the specification, then learn from usage](#start-with-the-specification-then-learn-from-usage)
- [Textbook exercises and everyday conversations](#textbook-exercises-and-everyday-conversations)
- [Turning customer interactions into evaluation scenarios](#turning-customer-interactions-into-evaluation-scenarios)
- [Learning from failures](#learning-from-failures)
- [Protecting the workflows customers use most](#protecting-the-workflows-customers-use-most)
- [Protecting customer privacy](#protecting-customer-privacy)
- [What this changes for us at GoDaddy](#what-this-changes-for-us-at-godaddy)
- [Product Loop series](#product-loop-series)
- [Frequently Asked Questions](#frequently-asked-questions)

---

## What are AI agent evaluation datasets?

**AI agent evaluation datasets** are structured collections of test scenarios used to check whether an AI agent behaves as intended across a range of real-world tasks. Each scenario typically includes a starting product state, a sequence of user requests, and a clear definition of what a successful outcome looks like.

Unlike single-turn benchmarks, agent eval datasets must account for multi-turn conversations, context carried across steps, and the gap between what a user says and what they mean. For agentic products like [GoDaddy Airo AI Builder](https://www.godaddy.com/airo/ai-builder?utm_source=brgd_blog&utm_medium=blog&utm_campaign=product-loop-part-3-evals-customer-workflows), a robust eval dataset is the primary mechanism for catching regressions before they reach customers.

---

## Start with the specification, then learn from usage

Our initial evaluation dataset starts with product specifications: what a feature should do and what a successful result looks like. These requirements give us a deliberate foundation for testing the product before customers use it.

Specification-based evaluations remain essential after launch. They define the behavior we expect and provide a consistent way to check it as the product changes. Production usage adds another perspective: how customers express their goals, combine requests, and respond to the agent along the way.

That distinction matters for [GoDaddy Airo AI Builder](https://www.godaddy.com/airo/ai-builder?utm_source=brgd_blog&utm_medium=blog&utm_campaign=product-loop-part-3-evals-customer-workflows). For instance, a spec can define what it means to replace an image correctly. Customer conversations reveal many ways people ask for that change, including references to earlier messages, follow-up adjustments, and details they expect the agent to remember.

Product Loop helps us bring those patterns into our evaluation dataset. It extends the coverage we build from specifications. It does not replace that foundation.

## Textbook exercises and everyday conversations

Think about learning a language. Textbook exercises teach vocabulary, grammar, and how to put a sentence together. Listening to everyday conversations adds practice with how people use those skills in different situations. Both serve a purpose.

A textbook might teach someone to order coffee by saying, "I would like a large iced coffee, please." At a familiar café, they might say, "Same as yesterday, but iced," then change the size. Understanding that request depends on knowing what came before.

Customers use Airo in much the same way. One test might ask, "Replace the homepage image." A customer might instead say, "Use the other one," followed by, "Keep the crop the way it was." Airo needs to understand which image they mean and which earlier choices should stay the same.

We can design tests for conversational context from the outset. Product Loop adds evidence about the specific patterns customers use, helping us expand those tests and decide where deeper coverage would be most valuable. That builds on the multi-turn measurement frame from [Part 1](/discover/learn/measuring-agentic-ai-products): a response is not the same as an outcome.

## Turning customer interactions into evaluation scenarios

Product Loop reads customer conversations alongside telemetry and product context. It can group interactions by what customers are trying to accomplish, then identify common ways they phrase a request and the steps that follow.

The useful unit is often a sequence, not a single prompt. "Use the other one" makes little sense without the earlier image choices. A meaningful scenario needs the relevant conversation and starting product state, followed by the customer's request and a clear definition of the expected result.

That last part matters. A production conversation tells us what happened. It does not automatically tell us what should have happened. The expected result still needs to be checked against product requirements and customer intent. For image editing, the agent saying "Done" is not enough. The resulting page should reflect the requested change.

Product Loop can generate candidate scenarios from these patterns and create tickets for gaps in coverage. Before a candidate becomes a reusable test, it needs a reviewable setup, a clear success criterion, and appropriate handling of customer data. Private details should be removed or replaced without losing the context that makes the scenario useful.

This connects what we learn in production to concrete evaluation work, rather than leaving the findings in a report. It is the evals counterpart to the diagnostic workflow in [Part 2](/discover/learn/product-loop-part-2-finding-bugs-nobody-reported), where production signals feed engineering tickets instead of test cases.

## Learning from failures

One use is turning a recurring failure into a **regression test**: a test designed to detect whether a previously fixed problem returns.

Suppose customers repeatedly need to correct an image replacement because an earlier crop setting is lost. Product Loop can identify that pattern and generate a candidate scenario that preserves the relevant sequence of requests. The team can then check that the test reproduces the problem and that its expected result matches the intended behavior.

Once the issue is fixed, that scenario provides a repeatable check for future changes. It does not guarantee the problem can never return, but it gives us a targeted way to detect a recurrence before a release reaches customers.

The value is that the investigation contributes both a fix and a lasting addition to our evaluation coverage.

## Protecting the workflows customers use most

Failures are only one reason to add coverage. Product Loop also helps us understand which capabilities customers rely on and the variations they use most often, even when those interactions succeed.

Take image editing as an example. Our specification-based dataset can cover the defined editing behaviors. If production usage shows that image editing is especially common, Product Loop can help us deepen that coverage around actual workflows: replacing an image while preserving its crop, adjusting it after changing a layout, or returning to a previous choice later in the conversation.

Those are illustrative examples, but the principle is practical. We want to protect the combinations customers depend on, not wait for each one to fail before adding a test. As customer behavior changes, Product Loop helps us identify further additions. The goal is not to regenerate the entire dataset every day. A stable set of tests remains valuable for comparing releases, while reviewed additions keep our coverage connected to current usage.

## Protecting customer privacy

Protecting customer information is an integral part of how we build eval scenarios. During scenario generation, we mask personally identifiable information (PII) while preserving the relevant wording and context needed to test the agent's behavior. The focus is on understanding the customer's task, not identifying the person behind it.

## What this changes for us at GoDaddy

For our Airo AI Builder team, Product Loop connects two kinds of knowledge: what we designed the product to do and how customers put it to work. Our specifications establish the expected behavior. Production patterns help us see where more variations, conversation context, or workflow combinations deserve coverage.

This gives our engineers and product team a clearer basis for deciding where to invest evaluation effort. A recurring problem can lead to a targeted regression test. A heavily used capability can receive deeper coverage before a failure brings it to our attention. The evidence comes with the proposed work, reducing the effort needed to reconstruct each scenario.

That is the direction we want to take evaluation at GoDaddy: a strong foundation in product requirements, extended by what we learn from customers. As AI products support more varied workflows, both perspectives become more important, not less.

In the next and final post, we will explore how Product Loop helps us identify missing capabilities and new product opportunities.

---

## Product Loop series

1. [When Your Product Stops Having a Single Path: Measuring Agentic AI Products](/discover/learn/measuring-agentic-ai-products) — why funnel analytics and single-turn scoring fall short for agentic products
2. [Product Loop, Part 2: Finding the Bugs Nobody Reported](/discover/learn/product-loop-part-2-finding-bugs-nobody-reported) — how production signals surface recurring customer problems
3. **Product Loop, Part 3: Extending Our Evals with Real Customer Workflows** (this post)
4. Part 4 (coming soon) — missing capabilities and new product opportunities

---

## Frequently Asked Questions

### What are AI agent evaluation datasets?

AI agent evaluation datasets are structured collections of test scenarios that check whether an AI agent behaves as intended. For agentic products, each scenario typically includes a starting product state, a sequence of user requests, and a clear success criterion — not just a single prompt and reply.

### Why are specification-based evals not enough after launch?

Specs define expected behavior, but customers phrase goals in many ways, combine requests, and refer back to earlier turns. Production usage reveals those real patterns so the eval dataset stays aligned with how people actually work.

### How does Product Loop turn conversations into eval scenarios?

Product Loop groups production interactions by task, surfaces common phrasing and step sequences, and proposes candidate scenarios with setup and success criteria. Each candidate is reviewed, scrubbed for PII, and checked against product requirements before it joins the reusable dataset.

### What is a regression test in the context of AI agent evaluation?

A regression test reproduces a previously fixed failure so future releases can be checked for the same problem. Product Loop can propose these from recurring correction patterns in customer conversations.

### How do you protect customer privacy when building eval datasets?

PII is masked or replaced during scenario generation while keeping the wording and context needed to test agent behavior. The goal is to preserve the task structure, not identify the person behind it.

### How does this relate to measuring agentic AI products?

[Measuring agentic AI products](/discover/learn/measuring-agentic-ai-products) focuses on whether the full workflow succeeded across multiple turns. Extending evals with real customer workflows applies that same unit of analysis to the tests you run before each release.

---

**Ready to see agentic building in practice?** [Start building with GoDaddy Airo AI Builder](https://www.godaddy.com/airo/ai-builder?utm_source=brgd_blog&utm_medium=blog&utm_campaign=product-loop-part-3-evals-customer-workflows). Describe your idea and watch it come to life. 50 free AI credits to get started.
