DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
  • Understand the Sidecar Pattern by Deploying n8n to AWS Fargate
  • Deliberate Decoupling: 6 Architectural Patterns From a Regulated WAS-to-AWS Migration
  • Multi-Account AWS Architecture: Isolating PHI Workloads Without Slowing Down Engineering Teams

Trending

  • Extracting Entities and Relationships From Engineering Documents With spaCy
  • The Request Timed Out, But the Payment Succeeded: Building Retry-Safe Mobile APIs
  • GenAI Isn't Solving the Problem Most Development Teams Actually Have
  • AI Architectures That Drive Real Business ROI
  1. DZone
  2. Software Design and Architecture
  3. Cloud Architecture
  4. Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A

Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A

Multi-agent systems are common. Here we build a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol.

By 
Purnanga Borah user avatar
Purnanga Borah
·
Oct. 09, 26 · Tutorial
Likes (0)
Comment
Save
Tweet
Share
194 Views

Join the DZone community and get the full member experience.

Join For Free

Most agent projects start the same way: one model, one prompt loop, three or four tools. That's the right place to start, and for a lot of workloads it's also the right place to stop.

The trouble shows up later, when the same agent is asked to inspect infrastructure, parse OpenAPI specs, write docs, answer developer questions, and call operational tools. The system prompt turns into a wall of text. Tool selection gets flaky in ways you can't reproduce. Someone tunes the prompt to fix documentation quality and quietly breaks spec analysis, and nobody notices for a week.

Splitting the agent is the obvious fix. But the agents were never the hard part. The hard part is the contract between them: how one agent finds another, how it hands over work, where authentication lives, and how you trace a request that crossed three runtimes before it failed.

This article builds a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol:  an orchestrator, an OpenAPI spec analyzer, and a documentation generator. Three agents are a toy. The decisions are not.

A note on versions. AgentCore and A2A are both moving fast. Everything here was checked against the AWS documentation and AgentCore CLI 0.25.x in August 2026. Pin your CLI and Python dependencies, and re-check the docs before you deploy anything you care about.

Is a Second Agent Actually Worth It?

Multi-agent isn't a maturity level. Every remote handoff you add buys you latency, an extra auth boundary, and a new class of failure that only shows up under load.

Keep one agent when one team owns the workflow, the tools sit inside the same security boundary, and a single prompt can plausibly cover the job. Don't split because the architecture diagram looks better with more boxes in it.

Split when responsibilities genuinely diverge. The signals are usually organizational before they're technical: two teams own two capabilities, the specialists want different models, the release cadences don't match, or one capability touches data the other has no business seeing. At that point the agent stops being a prompt and starts being a service, and it needs a service's contract.

The arithmetic argument for a shared protocol is the familiar one. Pairwise integrations grow with the square of the number of participants, and each one tends to invent its own payload shape, auth handling, timeout, and error format. A standard protocol doesn't make any of those problems disappear. It just means every participant expresses them in the same place, which is most of the benefit.

MCP and A2A Aren't Competing

These two get discussed as alternatives. They aren't: they sit on different edges of the architecture.

  • MCP connects an agent to tools and data.
  • A2A connects an agent (or an application) to another agent.

A spec analyzer might pull an OpenAPI document out of a repo over MCP, then return its analysis to an orchestrator over A2A. Both, same request. In the code below, the OpenAPI document is passed inline so the A2A flow stays readable.

One caveat about "framework-agnostic," since it gets oversold: A2A is framework-agnostic at the protocol level. That doesn't mean two arbitrary SDKs will talk to each other on the first try. Both sides still need a compatible A2A implementation, transport, protocol version, and auth configuration. Once those line up, the internals genuinely don't matter; Strands on one side, LangGraph on the other, fine.

What Runtime Gives You

Configure an AgentCore Runtime for A2A, and it becomes a transparent proxy in front of your container. Your side of the contract is narrow: a stateless, streamable HTTP server on 0.0.0.0:9000, mounted at the root path, speaking JSON-RPC 2.0. AgentCore passes the JSON-RPC payload straight through without touching it, and adds session isolation, authentication, scaling, and observability around it.

Note the port. A2A is 9000 at /, MCP is 8000 at /mcp, plain HTTP is 8080 at /invocations. Getting this wrong is the most common reason a container deploys fine and then never answers.

The A2A vocabulary you need for the rest of this:

  • A2A client: whoever is sending work. Here is the orchestrator.
  • A2A server: the agent endpoint receiving it.
  • Agent Card: metadata at /.well-known/agent-card.json describing identity, endpoint, skills, capabilities, and auth requirements. This is how discovery works.
  • Message: one turn, carrying one or more parts (text, structured data).
  • Task and artifact: a task tracks work at one remote agent; artifacts carry its output. An orchestrator fanning out to three specialists creates three separate tasks.

For inbound calls, Runtime supports IAM SigV4 or OAuth 2.0 bearer tokens (JWT). A runtime version uses one of them at a time, so pick based on who's calling and where the security boundary sits. If a client authenticates with the wrong one, you get a 403 that looks nothing like an agent error.

The Shape of the System

The system has a hub and spoke architecture. The orchestrator is an A2A server to the user and an A2A client to the specialists, which is worth stating explicitly because it's easy to forget when you're debugging.

Hub and Spoke Achitecture

The orchestrator should know what each specialist advertises. It should not know how any of them work. The test: adding a fourth agent a changelog writer, say should mean one new config entry, not edits inside the agents you already shipped.

Building a Specialist

Prerequisites

  • Python 3.10+
  • Node.js 20+ (the AgentCore CLI ships as an npm package; the A2A tutorial page still says 18, but the CLI's own docs say 20)
  • AWS CDK installed, the CLI deploys through it
  • AWS credentials, plus model access enabled in Bedrock
  • npm install -g @aws/agentcore

If you still have the old Python starter toolkit installed, uninstall it first. Both provide an agentcore command, and you will lose an afternoon to that.

Scaffold an A2A project and pick Strands when prompted:

Shell
 
bash
agentcore create --protocol A2A


The generated project gives you the config, the model loader, and the A2A server wiring. The specialist below does one thing: read an OpenAPI document and pull out a compact, honest view of its operations. 

Python
 
# app/SpecAnalyzer/main.py
from typing import Any

import yaml
from strands import Agent, tool
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model

HTTP_METHODS = {
    "get", "put", "post", "delete",
    "options", "head", "patch", "trace",
}


def _auth_mode(security: Any) -> str:
    if security is None:
        return "unspecified"
    if security == []:
        return "none"
    if isinstance(security, list) and any(item == {} for item in security):
        return "optional"
    return "required"


@tool
def summarize_endpoints(openapi_yaml: str) -> list[dict[str, Any]]:
    """Return a compact list of operations from an OpenAPI document."""
    spec = yaml.safe_load(openapi_yaml)

    if not isinstance(spec, dict):
        raise ValueError("The OpenAPI document must contain a YAML object.")

    paths = spec.get("paths", {})
    if not isinstance(paths, dict):
        raise ValueError("The OpenAPI 'paths' field must be an object.")

    global_security = spec.get("security")
    operations: list[dict[str, Any]] = []

    for path, path_item in paths.items():
        if not isinstance(path_item, dict):
            continue

        for method, operation in path_item.items():
            if method.lower() not in HTTP_METHODS:
                continue
            if not isinstance(operation, dict):
                continue

            security = operation.get("security", global_security)
            operations.append(
                {
                    "path": path,
                    "method": method.upper(),
                    "operation_id": operation.get("operationId"),
                    "summary": operation.get("summary", ""),
                    "auth": _auth_mode(security),
                }
            )

    return operations


agent = Agent(
    model=load_model(),
    system_prompt=(
        "You are an API specification analyst. "
        "Call summarize_endpoints before reasoning. "
        "Report endpoints, authentication behavior, and concrete "
        "specification risks. Never invent fields that aren't present."
    ),
    tools=[summarize_endpoints],
)

if __name__ == "__main__":
    serve_a2a(StrandsA2AExecutor(agent))


serve_a2a is the AgentCore SDK helper that does the boring, mandatory parts: /ping, Agent Card serving, AGENTCORE_RUNTIME_URL handling, Bedrock header propagation, and binding to port 9000.

The validation inside that tool is the part I'd argue for hardest. OpenAPI path items legally contain keys that aren't HTTP methods  parameters, servers, $ref, orsummary. Loop over every key as though it were an operation and you'll either crash on a string or, worse, emit a confident summary containing an endpoint called PARAMETERS. Then the doc generator writes it up, and it ships. Demos never surface this because demo specs are clean.

Direct dependencies for both specialists:

Shell
 
bedrock-agentcore[a2a]
strands-agents[a2a]
a2a-sdk
PyYAML
httpx


The a2a extras matter. Without them, serve_a2a and the client classes aren't there. a2a-sdk is what the orchestrator's client code imports from, so it's a direct dependency even though the server side pulls it in transitively.

The documentation generator uses the same wrapper with a narrower job:

Python
 
# app/DocGenerator/main.py
from strands import Agent
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model

agent = Agent(
    model=load_model(),
    system_prompt=(
        "You write developer documentation from an API request and a "
        "specification analysis. Produce Markdown with an overview, "
        "authentication notes, request examples, error handling, and a "
        "short Python quickstart. Never invent endpoints, fields, or "
        "credentials."
    ),
)

if __name__ == "__main__":
    serve_a2a(StrandsA2AExecutor(agent))


That one's prompt-only on purpose. If your docs have to match a house format, this is where you add output validation or template tools instead of asking a model to remember a style guide.

Test It Locally

Shell
 
agentcore dev
# or, without the inspector UI:
python main.py


Then send it a JSON-RPC message at the root path:

JSON
 
curl -X POST http://localhost:9000/ \
  -H "Content-Type: application/json" \
  -d '{
    "jsonrpc": "2.0",
    "id": "req-001",
    "method": "message/send",
    "params": {
      "message": {
        "role": "user",
        "parts": [
          {
            "kind": "text",
            "text": "Analyze the authentication on this API."
          }
        ],
        "messageId": "12345678-1234-1234-1234-123456789012"
      }
    }
  }' | jq .


Pull the Agent Card too, before you deploy anything: 

Shell
 
curl http://localhost:9000/.well-known/agent-card.json | jq .


Thirty seconds of curl catches a missing discovery endpoint or a card whose skills describe an agent you deleted two commits ago. Card problems are miserable to debug from the client side, because what you see is a resolver failure, not a bad card. 

Deploy the Specialists

The CLI defaults to IAM SigV4. Configure bearer-token auth when the caller can't sign requests, the AWS tutorial walks through Cognito for this, though Cognito isn't a requirement; any OIDC provider works.

Deploy each specialist as its own runtime:

Shell
 
agentcore deploy


The CLI packages the code, uploads the artifact to S3, creates the runtime, and hands back an ARN: 

Shell
 
arn:aws:bedrock-agentcore:us-west-2:<account-id>:runtime/SpecAnalyzer-xyz123


Separate runtimes are half the point of doing this at all. Each specialist scales, releases, and rolls back on its own, without a coordinated deploy across the whole system.

The URL Nobody Warns You About

Here's the step that eats the most time. The invocation URL isn't the ARN, and it isn't a friendly hostname. It's the ARN, URL-encoded, embedded in a path: https://bedrock-agentcore.us-west-2.amazonaws.com/runtimes/{url-encoded-arn}/invocations/

Every colon and slash in the ARN gets escaped with quote(arn, safe='') in Python, or let the SDK do it:

Python
 
from bedrock_agentcore.runtime import build_runtime_url

url = build_runtime_url(
    "arn:aws:bedrock-agentcore:us-west-2:123456789012:runtime/SpecAnalyzer-xyz123"
)


The Agent Card sits under that path, at .../invocations/.well-known/agent-card.json. Keep the trailing slash on the base URL. A missing slash or an unescaped ARN both produce a 404, which reads like the runtime doesn't exist. 

Calling a Remote Agent

This client assumes bearer auth. For an IAM runtime, sign with SigV4 or use a client that signs for you.

Python
 
# app/Orchestrator/client.py
import os
from uuid import uuid4

import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, TextPart

DEFAULT_TIMEOUT = 300


def _message(text: str) -> Message:
    return Message(
        kind="message",
        role=Role.user,
        parts=[Part(TextPart(kind="text", text=text))],
        message_id=uuid4().hex,
    )


async def call_remote_agent(
    runtime_url: str,
    text: str,
    session_id: str | None = None,
) -> str:
    headers = {
        "Authorization": f"Bearer {os.environ['BEARER_TOKEN']}",
        "X-Amzn-Bedrock-AgentCore-Runtime-Session-Id":
            session_id or str(uuid4()),
    }

    async with httpx.AsyncClient(
        timeout=DEFAULT_TIMEOUT,
        headers=headers,
    ) as http:
        resolver = A2ACardResolver(httpx_client=http, base_url=runtime_url)
        card = await resolver.get_agent_card()

        config = ClientConfig(httpx_client=http, streaming=False)
        client = ClientFactory(config).create(card)

        async for event in client.send_message(_message(text)):
            if isinstance(event, Message):
                return event.model_dump_json(exclude_none=True)

            if isinstance(event, tuple) and len(event) == 2:
                task, _ = event
                return task.model_dump_json(exclude_none=True)

    raise RuntimeError("The remote agent returned no A2A result.")


With streaming=False, you get exactly one result, which is either a Message or a (Task, UpdateEvent) tuple. Handle both. The tuple form is what trips people up on the first run.

Returning raw JSON keeps the protocol visible for a walkthrough, but don't ship it. Normalize messages and artifacts into your own response type at this boundary. Otherwise, every downstream prompt has to understand the full A2A envelope, and you end up with models reasoning about kind fields instead of about your problem.

Session IDs deserve a real decision, not a uuid4() you stopped thinking about. Fresh ID for an isolated one-shot task. Same ID when you're continuing a multi-turn conversation with the same remote agent. And keep a separate correlation ID of your own, because one user request may span several runtime sessions and you'll want to stitch them back together in the traces.

The Orchestrator

Each remote agent becomes a tool. Endpoints come from config; the client resolves the Agent Card at the endpoint before sending work.

Python
 
# app/Orchestrator/main.py
import os

from strands import Agent, tool
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model

from client import call_remote_agent

REMOTE_AGENTS = {
    "spec_analyzer": os.environ["SPEC_ANALYZER_URL"],
    "doc_generator": os.environ["DOC_GENERATOR_URL"],
}


@tool
async def delegate(agent_name: str, instruction: str) -> str:
    """Send work to a named A2A specialist."""
    runtime_url = REMOTE_AGENTS.get(agent_name)
    if runtime_url is None:
        allowed = ", ".join(sorted(REMOTE_AGENTS))
        return f"Unknown agent '{agent_name}'. Available agents: {allowed}."

    return await call_remote_agent(runtime_url, instruction)


orchestrator = Agent(
    model=load_model(),
    system_prompt=(
        "You coordinate two API specialists. Use spec_analyzer for OpenAPI "
        "parsing and risk analysis. Use doc_generator for documentation and "
        "SDK examples. When both are needed, analyze the specification "
        "first, then include that analysis in the documentation request. "
        "Never assign work outside an agent's advertised capability."
    ),
    tools=[delegate],
)

if __name__ == "__main__":
    serve_a2a(StrandsA2AExecutor(orchestrator))


Two small things in there matter more than they look. The tool is async, which avoids wrapping the remote call in asyncio.run() that blows up when the agent is already running inside an event loop, and the traceback doesn't point anywhere near the cause. And returning a friendly string for an unknown agent, rather than raising, gives the model something it can recover from instead of a dead turn.

The URLs belong in deployment config, a service catalog, or an agent registry. Not in source control.

Let the Model Route, or Hard-Code the Path?

The code above lets the orchestrator model decide who to call and in what order. That's genuinely useful when requests vary. It's also the default choice for the wrong reasons.

If every documentation request must be analyzed before generation, always, no exceptions encode that sequence in Python and use the model only inside each step. Deterministic orchestration is easier to test, retry, and audit, and it doesn't drift when someone edits a prompt.

Save model-driven routing for routes that are actually ambiguous. It shouldn't be there to make a fixed workflow feel more autonomous.

Running It

With all three runtimes deployed, send the user request to the orchestrator. Something like "Review this customer API spec and create a Python quickstart" should cause it to call the analyzer, feed that analysis to the doc generator, and return one answer.

Should. That's expected behavior, not a guarantee the protocol hands you. Whether the right handoffs happen comes down to your orchestrator prompt, your tool descriptions, your eval cases, and your workflow code. A2A standardizes the conversation. It has no opinion about whether your orchestration logic is correct.

Before You Call It Production

Least privilege per runtime. Each runtime gets permissions for its own tools and data, nothing more. The failure mode to watch for is the orchestrator's role slowly accumulating every permission any specialist ever needed.

Distinguish auth failures from agent failures. A2A application errors come back as JSON-RPC error objects, frequently with HTTP 200. Auth and authorization failures come back as native HTTP errors AccessDeniedException is a 403. AgentCore also maps runtime exceptions onto JSON-RPC codes: -32051 for not found (404), -32053 for throttling (429), -32054 for conflict (409), -32055 for a runtime client error (424). Inspect the JSON-RPC shape. A 200 is not a success.

Retry deliberately. 429 and retryable-conflict 409 are worth retrying with backoff. A 424 usually means your container failed and the answer is in CloudWatch, not in a second attempt. Never blind-retry non-idempotent work, and pass a task or request identifier so a specialist can recognize a duplicate.

Trace every handoff. AgentCore Observability gives you sessions, traces, and spans. Propagate your own business correlation data on top, or you'll have three sets of traces and no way to prove they belong to the same user request.

Set lifecycle values on purpose. Defaults are a 15-minute idle timeout and an 8-hour max lifetime, both configurable (60 to 28,800 seconds). Long-running tasks may justify raising them; a bigger number should be a decision, not a leftover. Also watch that your A2A sessions actually go idle — a server that holds a connection open after the client is gone will happily keep a microVM warm on your bill.

Cache Agent Cards carefully. Resolving on every call is simple and costs you a round trip per hop. Cache with a short TTL and invalidate when an agent's version or capabilities change. A stale card advertising a skill that no longer exists is a genuinely confusing bug.

Keep protocol and SDK versions aligned. Pin dependencies, validate cards, and test a mixed-framework setup rather than assuming it works.

Use private connectivity if you need it. Runtime can attach to a VPC for internal resources, and interface VPC endpoints via PrivateLink keep AgentCore API calls off the public internet.

A Failure Test Matrix

Before you call it done, break the handoffs rather than admiring the happy path:

  • The specialist's Agent Card is unreachable or malformed.
  • The bearer token expires between card resolution and invocation.
  • The analyzer returns a valid task with no usable artifact.
  • The doc generator times out after the analyzer succeeded.
  • The orchestrator picks an agent that doesn't exist, or one it isn't allowed to call.
  • A retry re-runs a task that already had an external side effect.

Each of these teaches you something a successful demo can't. They also make the division of labor concrete: Runtime hosts and protects the endpoint, A2A defines the conversation, and workflow correctness is still entirely yours.

Closing

The value here isn't the agent count. It's that each capability gets a clear boundary, a narrow permission set, and its own release cycle. Runtime provides the managed execution boundary; A2A provides a common way to discover agents and pass work between them.

So, start with one agent. Split it when responsibilities, ownership, or security boundaries justify a network hop, and not before. Where the sequence is fixed, keep the orchestration deterministic. Where routing is genuinely variable, let the orchestrator reason over well-described capabilities, and instrument every handoff you add.

References

  • AWS: Deploy A2A servers in AgentCore Runtime
  • AWS: A2A protocol contract
  • AWS: AgentCore Runtime authentication and authorization
  • AWS: LifecycleConfiguration API reference
  • AWS: Configure AgentCore Runtime for VPC access
  • AgentCore CLI
  • Agent2Agent protocol specification


AWS

Opinions expressed by DZone contributors are their own.

Related

  • AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
  • Understand the Sidecar Pattern by Deploying n8n to AWS Fargate
  • Deliberate Decoupling: 6 Architectural Patterns From a Regulated WAS-to-AWS Migration
  • Multi-Account AWS Architecture: Isolating PHI Workloads Without Slowing Down Engineering Teams

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook