DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Designing Tool-Calling AI Agents That Survive Production: A LangGraph Approach
  • Designing Production-Grade AI Tools: Why Architecture Matters More Than Models
  • Production-Ready Observability for Analytics Agents: An Open Telemetry Blueprint Across Retrieval, SQL, Redaction, and Tool Calls
  • From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Trending

  • Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
  • Are Passphrases Still Secure in the Age of AI?
  • Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
  • Detection and Response Did Its Job. Now Someone Has to Actually Fix It.
  1. DZone
  2. Coding
  3. Tools
  4. Decoding the “Black Box”: Evaluating Agent Tool Chains in Production

Decoding the “Black Box”: Evaluating Agent Tool Chains in Production

Evaluating AI agents means checking the full trajectory, not just the answer — intent, tool calls, order, and accuracy — using Microsoft Foundry's evaluation tools.

By 
Gaurav Bhardwaj user avatar
Gaurav Bhardwaj
·
Oct. 07, 26 · Analysis
Likes (1)
Comment
Save
Tweet
Share
126 Views

Join the DZone community and get the full member experience.

Join For Free

One of the most interesting parts of my role as a cloud solution architect is working directly with customers as they move from experimentation to production. The conversations change significantly at that point.

During an early proof of concept, the questions are usually foundational:

  • Can the model answer the question?
  • Can the agent call the API?
  • Can we connect our enterprise data?
  • Can we build the experience? 

But when customers start thinking about production, the questions become much harder:

  • Can I trust the agent to take an action?
  • How do I know it followed the business process?
  • How can I prove that it used the right tool?
  • What happens if it gets the right answer but takes the wrong path?

I'll use a simplified travel scenario throughout this article. Imagine building an agent that can upgrade an airline passenger when certain business conditions are met.

A user asks: “Can you upgrade Alex Johnson's SEA-to-JFK flight AA245 to business class if he is eligible?”

The agent responds: “Alex is eligible, and I have successfully submitted the business-class upgrade request.”

At first glance, everything looks good. The answer is clear. The request appears to have been completed. The demo works.

But I found myself asking a different question: “What did the agent actually do before giving us that answer?” That question became much more interesting than the answer itself.

From Evaluating Answers to Evaluating Behavior

For a traditional generative AI application, evaluating the final response makes sense. Was it relevant? Grounded? Coherent? Accurate?

Those questions don't disappear when we build agents, but agents introduce another dimension. An agent can understand an intent, decide which tool to use, generate tool arguments, execute the tool, inspect the result, choose another tool, take an action, and eventually generate a response.

Microsoft Foundry describes this same challenge: production-ready agent applications need evaluation not only of the final output, but also of the quality and efficiency of the workflow that produced it.

For our travel scenario, this might be the intended workflow: User request ➔ Look up customer ➔ Check upgrade eligibility ➔ Create upgrade request ➔ Return confirmation

Now imagine the agent instead does this: User request ➔ Create upgrade request ➔ Look up customer ➔ Check eligibility ➔ Return confirmation

The final answer could still be perfect. But the workflow is absolutely not. That is what I mean by the agentic black box.

A Small Version of the Customer Problem

When working through an architecture question, I try to make the problem as small as possible first to separate the core issue from enterprise complexity. I created three simple Python functions.

The first finds the customer:

Python
 
import json

def lookup_customer(name: str) -> str:
    """Find a customer using their name."""
    customers = {
        "Alex Johnson": {
            "customer_id": "CUST-1001",
            "loyalty_tier": "Gold"
        }
    }

    customer = customers.get(name)
    if not customer:
        return json.dumps({"error": "Customer not found"})

    return json.dumps(customer)


The second determines whether that customer is eligible for an upgrade:

Python
 
def check_upgrade_eligibility(customer_id: str, flight_number: str) -> str:
    """Check whether the customer can receive an upgrade."""
    if customer_id == "CUST-1001":
        return json.dumps({
            "customer_id": customer_id,
            "flight_number": flight_number,
            "eligible": True,
            "reason": "Gold member with available upgrade inventory"
        })

    return json.dumps({
        "customer_id": customer_id,
        "flight_number": flight_number,
        "eligible": False
    })


Finally, the tool that performs the business action:

Python
 
def create_upgrade_request(customer_id: str, flight_number: str, target_cabin: str) -> str:
    """Create an airline upgrade request."""
    return json.dumps({
        "request_id": "UPG-9001",
        "customer_id": customer_id,
        "flight_number": flight_number,
        "target_cabin": target_cabin,
        "status": "submitted"
    })


In a real environment, each function could represent a Microsoft Foundry Agent, an Azure Function, an MCP Server, a Logic App, or an enterprise system. But from the agent's perspective, the important question remains: Which tool should I call, with what parameters, and at what point in the workflow?

Turning Business Rules into Agent Instructions

The customer requirement sounds straightforward: “Do not create an upgrade until eligibility has been confirmed.”

That sentence is doing something very important — it is defining a business control. I can expose my Python functions to a Foundry agent as function tools and establish the Foundry project client.

Python
 
from typing import Any, Callable, Set
from azure.ai.projects.models import FunctionTool, ToolSet
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential
import os

user_functions: Set[Callable[..., Any]] = {
    lookup_customer,
    check_upgrade_eligibility,
    create_upgrade_request,
}

functions = FunctionTool(user_functions)
toolset = ToolSet()
toolset.add(functions)

project_client = AIProjectClient(
    endpoint=os.environ["AZURE_AI_PROJECT"],
    credential=DefaultAzureCredential(),
)


Now we create the agent, converting a business rule into expected agent behavior:

Python
 
agent = project_client.agents.create_agent(
    model=os.environ["MODEL_DEPLOYMENT_NAME"],
    name="travel-upgrade-agent",
    instructions="""
You help airline customers request flight upgrades.

For every upgrade request:
1. Look up the customer first.
2. Check the customer's upgrade eligibility.
3. Only if the customer is eligible, create an upgrade request.
4. Never create an upgrade request before eligibility is confirmed.
5. Clearly explain the result to the user.
""",
    toolset=toolset,
)


Running the Customer Scenario

We can create a thread, submit a test request, and execute the run:

Python
 
thread = project_client.agents.threads.create()

message = project_client.agents.messages.create(
    thread_id=thread.id,
    role="user",
    content="Can you upgrade Alex Johnson's SEA-to-JFK flight AA245 to business class if he is eligible?",
)

run = project_client.agents.runs.create_and_process(
    thread_id=thread.id,
    agent_id=agent.id,
)


If we inspect the output and stop there, we have demonstrated that the agent can work. But we haven't demonstrated that it worked correctly.

If an agent can create a ticket, modify a reservation, or provision infrastructure, I want visibility into the trajectory. Did it select the right tool? Did it send correct parameters? Did it take the next action at the right time?

Using AIAgentConverter

A Foundry execution produces multiple messages and interactions. Parsing all of that manually into an evaluation schema is tedious. That is where AIAgentConverter is useful. It converts a Foundry thread and run into the inputs expected by supported evaluators.

Python
 
import json
from azure.ai.evaluation import AIAgentConverter

converter = AIAgentConverter(project_client)

converted_data = converter.convert(
    thread.id,
    run.id,
)

print(json.dumps(converted_data, indent=2, default=str))


This was the point where the evaluation problem clicked for me. Instead of evaluating only the final string, I now have a normalized representation of the agent interaction that evaluators can inspect.

Evaluation Question #1: Did the agent understand the user?

"Upgrade the flight if he is eligible" is subtly different from "Upgrade the flight." Using IntentResolutionEvaluator, we can measure whether the agent correctly identified that condition.

Python
 
import os
from azure.ai.evaluation import IntentResolutionEvaluator

model_config = {
    "azure_deployment": os.environ["AZURE_DEPLOYMENT_NAME"],
    "api_key": os.environ["AZURE_OPENAI_API_KEY"],
    "azure_endpoint": os.environ["AZURE_OPENAI_ENDPOINT"],
    "api_version": os.environ["AZURE_API_VERSION"],
}

intent_evaluator = IntentResolutionEvaluator(
    model_config=model_config,
    threshold=3,
)

intent_result = intent_evaluator(**converted_data)


Evaluation Question #2: Did it call the right tools?

Next, we inspect tool behavior using ToolCallAccuracyEvaluator.

Python
 
from azure.ai.evaluation import ToolCallAccuracyEvaluator

tool_evaluator = ToolCallAccuracyEvaluator(
    model_config=model_config,
    threshold=3,
)

tool_result = tool_evaluator(**converted_data)


Why does this matter? If the agent selected the correct function but supplied the customer's name ("Alex Johnson") where the API expected an ID ("CUST-1001"), it fails. A language model could still generate an extremely convincing final response to cover this up. Agent behavior is part of the software surface we need to test.

Tool Accuracy Is Not the Same as Tool Order

Suppose the agent makes these valid calls: lookup_customer ➔ create_upgrade_request ➔ check_upgrade_eligibility.

The tools and parameters are correct, but the sequence violates our business process. That is why I wouldn't use Tool Call Accuracy alone. For sequencing, Foundry provides Task Navigation Efficiency, which compares the actual sequence against an expected sequence using three modes:

  • exact_match: Requires the exact same content and order. (Useful for strict payment workflows).
  • in_order_match: Permits extra exploratory steps while preserving the expected order of mandatory steps.
  • any_order_match: Expected steps can occur in any order. (Useful for research tasks).

This flexibility proves that evaluation matching isn't just an AI decision—it's a business-process decision.

One Successful Demo Is Not Enough

A successful demonstration proves the agent worked once. Production readiness asks: How consistently does it work across model upgrades, API changes, and prompt tweaks?

AIAgentConverter can prepare thread data for batch evaluation:

Python
 
filename = os.path.join(os.getcwd(), "agent_evaluation_data.jsonl")

converter.prepare_evaluation_data(
    thread_ids=[thread_id_1, thread_id_2, thread_id_3, thread_id_4],
    filename=filename,
)

evaluators = {
    "intent_resolution": IntentResolutionEvaluator(model_config=model_config),
    "tool_call_accuracy": ToolCallAccuracyEvaluator(model_config=model_config),
}

from azure.ai.evaluation import evaluate

results = evaluate(
    data=filename,
    evaluation_name="travel-agent-regression",
    evaluators=evaluators,
    azure_ai_project=os.environ["AZURE_AI_PROJECT"],
)


My Favorite Evaluation Cases Come From Failures

When a customer discovers an edge case — like the agent submitting an action before validating eligibility — I don't just fix the prompt. I turn it into a permanent regression case:

JSON
 
{
    "query": "Upgrade Alex Johnson's flight if he is eligible.",
    "expected_actions": [
        "lookup_customer",
        "check_upgrade_eligibility",
        "create_upgrade_request"
    ]
}


Over time, your evaluation dataset becomes a history of: "Things we have learned that this agent must never get wrong again."

How I Now Think About Agent Evaluation

My mental model for architecture discussions has shifted to this layered approach:

Plain Text
 
USER INTENT
                            |
                            v
                       +---------+
                       |  AGENT  |
                       +----+----+
                            |
              +-------------+-------------+
              |             |             |
              v             v             v
           INTENT         PROCESS       RESPONSE
              |             |             |
              |      +------+------+      |
              |      |      |      |      |
              v      v      v      v      v
          Understand Tool  Input  Order  Quality
                     Choice

  • Intent resolution: Did it understand what the user wanted?
  • Tool selection: Did it choose the right tool?
  • Tool input accuracy: Were the parameters correct?
  • Tool call success: Did the execution succeed?
  • Tool output utilization: Did it correctly use the returned result?
  • Task navigation efficiency: Did it follow the sequence?
  • Response quality: Was the final answer useful and grounded?

Where Microsoft Agent Framework Fits

While AIAgentConverter is documented in Foundry's classic Agent Service workflow, newer applications built using the Microsoft Agent Framework can integrate Foundry evaluation more directly through FoundryEvals.

A simplified current pattern looks like this:

Python
 
import os
from azure.identity.aio import AzureCliCredential
from agent_framework import Agent, evaluate_agent
from agent_framework.azure import FoundryChatClient
from agent_framework.foundry import FoundryEvals

async def evaluate_travel_agent():
    credential = AzureCliCredential()

    chat_client = FoundryChatClient(
        project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
        model=os.environ.get("FOUNDRY_MODEL", "gpt-4o"),
        credential=credential,
    )

    agent = Agent(
        client=chat_client,
        name="travel-upgrade-agent",
        instructions=(
            "Help customers with flight upgrades. "
            "Always verify eligibility before submitting an upgrade."
        ),
        tools=[lookup_customer, check_upgrade_eligibility, create_upgrade_request],
    )

    query = "Upgrade Alex Johnson's AA245 flight to business class if he is eligible."
    response = await agent.run(query)

    evaluators = FoundryEvals(
        client=chat_client,
        evaluators=[
            FoundryEvals.INTENT_RESOLUTION,
            FoundryEvals.TOOL_CALL_ACCURACY,
            FoundryEvals.TASK_NAVIGATION_EFFICIENCY,
        ],
    )

    results = await evaluate_agent(
        agent=agent,
        responses=response,
        queries=[query],
        evaluators=evaluators,
    )

    for result in results:
        print(f"Status: {result.status}")
        print(f"Passed: {result.passed}/{result.total}")
        print(f"Report: {result.report_url}")


(Note: Agent evaluation SDKs are evolving rapidly. Validate against the current Microsoft Learn documentation and your installed SDK version before production implementation.)

The Real Customer Question Was Trust

Looking back, what started this entire line of thinking wasn't really an SDK question.

It was a customer asking, implicitly:

“How comfortable should I be allowing this agent to take action in my business?”

And I realized I could not answer that question simply by looking at the final response.

I needed to understand:

Plain Text
 
What did it understand? 
What did it call? 
What did it send? 
What came back? 
What did it do next? 
Did it follow the process? 
Did it obey the business rules?


That is why trajectory evaluation has become such an important part of how I think about agent architecture.

The Path Is Part of the Product

For a chatbot, the final answer may be the primary product. For an agent, I increasingly think: The path is part of the product. 

When an agent starts calling APIs, modifying systems, creating transactions, triggering workflows, or taking enterprise actions, evaluating only the final response is no longer enough.

We need to evaluate behavior. 

For me, that is the real value behind capabilities such as:

JSON
 
evaluation_dimensions = [ "Intent Resolution", 
                         "Tool Call Accuracy", 
                         "Tool Selection", 
                         "Tool Input Accuracy", 
                         "Tool Output Utilization",
                         "Tool Call Success",
                         "Task Navigation Efficiency"]


And it is why I think tools such as AIAgentConverter, the Azure AI Evaluation SDK, Microsoft Foundry evaluators, and Microsoft Agent Framework deserve a place in the architecture conversation much earlier than the final production-readiness review.

Because before I tell a customer: “Yes, I think this agent is ready,” I want to be able to answer one additional question: “Do we know what it actually did?”

That, to me, is where evaluating agents becomes much more interesting than simply evaluating answers.

Black box Tool Production (computer science)

Opinions expressed by DZone contributors are their own.

Related

  • Designing Tool-Calling AI Agents That Survive Production: A LangGraph Approach
  • Designing Production-Grade AI Tools: Why Architecture Matters More Than Models
  • Production-Ready Observability for Analytics Agents: An Open Telemetry Blueprint Across Retrieval, SQL, Redaction, and Tool Calls
  • From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook