Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A
Multi-agent systems are common. Here we build a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol.
Join the DZone community and get the full member experience.
Join For FreeMost agent projects start the same way: one model, one prompt loop, three or four tools. That's the right place to start, and for a lot of workloads it's also the right place to stop.
The trouble shows up later, when the same agent is asked to inspect infrastructure, parse OpenAPI specs, write docs, answer developer questions, and call operational tools. The system prompt turns into a wall of text. Tool selection gets flaky in ways you can't reproduce. Someone tunes the prompt to fix documentation quality and quietly breaks spec analysis, and nobody notices for a week.
Splitting the agent is the obvious fix. But the agents were never the hard part. The hard part is the contract between them: how one agent finds another, how it hands over work, where authentication lives, and how you trace a request that crossed three runtimes before it failed.
This article builds a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol: an orchestrator, an OpenAPI spec analyzer, and a documentation generator. Three agents are a toy. The decisions are not.
A note on versions. AgentCore and A2A are both moving fast. Everything here was checked against the AWS documentation and AgentCore CLI 0.25.x in August 2026. Pin your CLI and Python dependencies, and re-check the docs before you deploy anything you care about.
Is a Second Agent Actually Worth It?
Multi-agent isn't a maturity level. Every remote handoff you add buys you latency, an extra auth boundary, and a new class of failure that only shows up under load.
Keep one agent when one team owns the workflow, the tools sit inside the same security boundary, and a single prompt can plausibly cover the job. Don't split because the architecture diagram looks better with more boxes in it.
Split when responsibilities genuinely diverge. The signals are usually organizational before they're technical: two teams own two capabilities, the specialists want different models, the release cadences don't match, or one capability touches data the other has no business seeing. At that point the agent stops being a prompt and starts being a service, and it needs a service's contract.
The arithmetic argument for a shared protocol is the familiar one. Pairwise integrations grow with the square of the number of participants, and each one tends to invent its own payload shape, auth handling, timeout, and error format. A standard protocol doesn't make any of those problems disappear. It just means every participant expresses them in the same place, which is most of the benefit.
MCP and A2A Aren't Competing
These two get discussed as alternatives. They aren't: they sit on different edges of the architecture.
- MCP connects an agent to tools and data.
- A2A connects an agent (or an application) to another agent.
A spec analyzer might pull an OpenAPI document out of a repo over MCP, then return its analysis to an orchestrator over A2A. Both, same request. In the code below, the OpenAPI document is passed inline so the A2A flow stays readable.
One caveat about "framework-agnostic," since it gets oversold: A2A is framework-agnostic at the protocol level. That doesn't mean two arbitrary SDKs will talk to each other on the first try. Both sides still need a compatible A2A implementation, transport, protocol version, and auth configuration. Once those line up, the internals genuinely don't matter; Strands on one side, LangGraph on the other, fine.
What Runtime Gives You
Configure an AgentCore Runtime for A2A, and it becomes a transparent proxy in front of your container. Your side of the contract is narrow: a stateless, streamable HTTP server on 0.0.0.0:9000, mounted at the root path, speaking JSON-RPC 2.0. AgentCore passes the JSON-RPC payload straight through without touching it, and adds session isolation, authentication, scaling, and observability around it.
Note the port. A2A is 9000 at /, MCP is 8000 at /mcp, plain HTTP is 8080 at /invocations. Getting this wrong is the most common reason a container deploys fine and then never answers.
The A2A vocabulary you need for the rest of this:
- A2A client: whoever is sending work. Here is the orchestrator.
- A2A server: the agent endpoint receiving it.
- Agent Card: metadata at
/.well-known/agent-card.jsondescribing identity, endpoint, skills, capabilities, and auth requirements. This is how discovery works. - Message: one turn, carrying one or more parts (text, structured data).
- Task and artifact: a task tracks work at one remote agent; artifacts carry its output. An orchestrator fanning out to three specialists creates three separate tasks.
For inbound calls, Runtime supports IAM SigV4 or OAuth 2.0 bearer tokens (JWT). A runtime version uses one of them at a time, so pick based on who's calling and where the security boundary sits. If a client authenticates with the wrong one, you get a 403 that looks nothing like an agent error.
The Shape of the System
The system has a hub and spoke architecture. The orchestrator is an A2A server to the user and an A2A client to the specialists, which is worth stating explicitly because it's easy to forget when you're debugging.

The orchestrator should know what each specialist advertises. It should not know how any of them work. The test: adding a fourth agent a changelog writer, say should mean one new config entry, not edits inside the agents you already shipped.
Building a Specialist
Prerequisites
- Python 3.10+
- Node.js 20+ (the AgentCore CLI ships as an npm package; the A2A tutorial page still says 18, but the CLI's own docs say 20)
- AWS CDK installed, the CLI deploys through it
- AWS credentials, plus model access enabled in Bedrock
npm install -g @aws/agentcore
If you still have the old Python starter toolkit installed, uninstall it first. Both provide an agentcore command, and you will lose an afternoon to that.
Scaffold an A2A project and pick Strands when prompted:
bash
agentcore create --protocol A2A
The generated project gives you the config, the model loader, and the A2A server wiring. The specialist below does one thing: read an OpenAPI document and pull out a compact, honest view of its operations.
# app/SpecAnalyzer/main.py
from typing import Any
import yaml
from strands import Agent, tool
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model
HTTP_METHODS = {
"get", "put", "post", "delete",
"options", "head", "patch", "trace",
}
def _auth_mode(security: Any) -> str:
if security is None:
return "unspecified"
if security == []:
return "none"
if isinstance(security, list) and any(item == {} for item in security):
return "optional"
return "required"
@tool
def summarize_endpoints(openapi_yaml: str) -> list[dict[str, Any]]:
"""Return a compact list of operations from an OpenAPI document."""
spec = yaml.safe_load(openapi_yaml)
if not isinstance(spec, dict):
raise ValueError("The OpenAPI document must contain a YAML object.")
paths = spec.get("paths", {})
if not isinstance(paths, dict):
raise ValueError("The OpenAPI 'paths' field must be an object.")
global_security = spec.get("security")
operations: list[dict[str, Any]] = []
for path, path_item in paths.items():
if not isinstance(path_item, dict):
continue
for method, operation in path_item.items():
if method.lower() not in HTTP_METHODS:
continue
if not isinstance(operation, dict):
continue
security = operation.get("security", global_security)
operations.append(
{
"path": path,
"method": method.upper(),
"operation_id": operation.get("operationId"),
"summary": operation.get("summary", ""),
"auth": _auth_mode(security),
}
)
return operations
agent = Agent(
model=load_model(),
system_prompt=(
"You are an API specification analyst. "
"Call summarize_endpoints before reasoning. "
"Report endpoints, authentication behavior, and concrete "
"specification risks. Never invent fields that aren't present."
),
tools=[summarize_endpoints],
)
if __name__ == "__main__":
serve_a2a(StrandsA2AExecutor(agent))
serve_a2a is the AgentCore SDK helper that does the boring, mandatory parts: /ping, Agent Card serving, AGENTCORE_RUNTIME_URL handling, Bedrock header propagation, and binding to port 9000.
The validation inside that tool is the part I'd argue for hardest. OpenAPI path items legally contain keys that aren't HTTP methods parameters, servers, $ref, orsummary. Loop over every key as though it were an operation and you'll either crash on a string or, worse, emit a confident summary containing an endpoint called PARAMETERS. Then the doc generator writes it up, and it ships. Demos never surface this because demo specs are clean.
Direct dependencies for both specialists:
bedrock-agentcore[a2a]
strands-agents[a2a]
a2a-sdk
PyYAML
httpx
The a2a extras matter. Without them, serve_a2a and the client classes aren't there. a2a-sdk is what the orchestrator's client code imports from, so it's a direct dependency even though the server side pulls it in transitively.
The documentation generator uses the same wrapper with a narrower job:
# app/DocGenerator/main.py
from strands import Agent
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model
agent = Agent(
model=load_model(),
system_prompt=(
"You write developer documentation from an API request and a "
"specification analysis. Produce Markdown with an overview, "
"authentication notes, request examples, error handling, and a "
"short Python quickstart. Never invent endpoints, fields, or "
"credentials."
),
)
if __name__ == "__main__":
serve_a2a(StrandsA2AExecutor(agent))
That one's prompt-only on purpose. If your docs have to match a house format, this is where you add output validation or template tools instead of asking a model to remember a style guide.
Test It Locally
agentcore dev
# or, without the inspector UI:
python main.py
Then send it a JSON-RPC message at the root path:
curl -X POST http://localhost:9000/ \
-H "Content-Type: application/json" \
-d '{
"jsonrpc": "2.0",
"id": "req-001",
"method": "message/send",
"params": {
"message": {
"role": "user",
"parts": [
{
"kind": "text",
"text": "Analyze the authentication on this API."
}
],
"messageId": "12345678-1234-1234-1234-123456789012"
}
}
}' | jq .
Pull the Agent Card too, before you deploy anything:
curl http://localhost:9000/.well-known/agent-card.json | jq .
Thirty seconds of curl catches a missing discovery endpoint or a card whose skills describe an agent you deleted two commits ago. Card problems are miserable to debug from the client side, because what you see is a resolver failure, not a bad card.
Deploy the Specialists
The CLI defaults to IAM SigV4. Configure bearer-token auth when the caller can't sign requests, the AWS tutorial walks through Cognito for this, though Cognito isn't a requirement; any OIDC provider works.
Deploy each specialist as its own runtime:
agentcore deploy
The CLI packages the code, uploads the artifact to S3, creates the runtime, and hands back an ARN:
arn:aws:bedrock-agentcore:us-west-2:<account-id>:runtime/SpecAnalyzer-xyz123
Separate runtimes are half the point of doing this at all. Each specialist scales, releases, and rolls back on its own, without a coordinated deploy across the whole system.
The URL Nobody Warns You About
Here's the step that eats the most time. The invocation URL isn't the ARN, and it isn't a friendly hostname. It's the ARN, URL-encoded, embedded in a path: https://bedrock-agentcore.us-west-2.amazonaws.com/runtimes/{url-encoded-arn}/invocations/
Every colon and slash in the ARN gets escaped with quote(arn, safe='') in Python, or let the SDK do it:
from bedrock_agentcore.runtime import build_runtime_url
url = build_runtime_url(
"arn:aws:bedrock-agentcore:us-west-2:123456789012:runtime/SpecAnalyzer-xyz123"
)
The Agent Card sits under that path, at .../invocations/.well-known/agent-card.json. Keep the trailing slash on the base URL. A missing slash or an unescaped ARN both produce a 404, which reads like the runtime doesn't exist.
Calling a Remote Agent
This client assumes bearer auth. For an IAM runtime, sign with SigV4 or use a client that signs for you.
# app/Orchestrator/client.py
import os
from uuid import uuid4
import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, TextPart
DEFAULT_TIMEOUT = 300
def _message(text: str) -> Message:
return Message(
kind="message",
role=Role.user,
parts=[Part(TextPart(kind="text", text=text))],
message_id=uuid4().hex,
)
async def call_remote_agent(
runtime_url: str,
text: str,
session_id: str | None = None,
) -> str:
headers = {
"Authorization": f"Bearer {os.environ['BEARER_TOKEN']}",
"X-Amzn-Bedrock-AgentCore-Runtime-Session-Id":
session_id or str(uuid4()),
}
async with httpx.AsyncClient(
timeout=DEFAULT_TIMEOUT,
headers=headers,
) as http:
resolver = A2ACardResolver(httpx_client=http, base_url=runtime_url)
card = await resolver.get_agent_card()
config = ClientConfig(httpx_client=http, streaming=False)
client = ClientFactory(config).create(card)
async for event in client.send_message(_message(text)):
if isinstance(event, Message):
return event.model_dump_json(exclude_none=True)
if isinstance(event, tuple) and len(event) == 2:
task, _ = event
return task.model_dump_json(exclude_none=True)
raise RuntimeError("The remote agent returned no A2A result.")
With streaming=False, you get exactly one result, which is either a Message or a (Task, UpdateEvent) tuple. Handle both. The tuple form is what trips people up on the first run.
Returning raw JSON keeps the protocol visible for a walkthrough, but don't ship it. Normalize messages and artifacts into your own response type at this boundary. Otherwise, every downstream prompt has to understand the full A2A envelope, and you end up with models reasoning about kind fields instead of about your problem.
Session IDs deserve a real decision, not a uuid4() you stopped thinking about. Fresh ID for an isolated one-shot task. Same ID when you're continuing a multi-turn conversation with the same remote agent. And keep a separate correlation ID of your own, because one user request may span several runtime sessions and you'll want to stitch them back together in the traces.
The Orchestrator
Each remote agent becomes a tool. Endpoints come from config; the client resolves the Agent Card at the endpoint before sending work.
# app/Orchestrator/main.py
import os
from strands import Agent, tool
from strands.multiagent.a2a.executor import StrandsA2AExecutor
from bedrock_agentcore.runtime import serve_a2a
from model.load import load_model
from client import call_remote_agent
REMOTE_AGENTS = {
"spec_analyzer": os.environ["SPEC_ANALYZER_URL"],
"doc_generator": os.environ["DOC_GENERATOR_URL"],
}
@tool
async def delegate(agent_name: str, instruction: str) -> str:
"""Send work to a named A2A specialist."""
runtime_url = REMOTE_AGENTS.get(agent_name)
if runtime_url is None:
allowed = ", ".join(sorted(REMOTE_AGENTS))
return f"Unknown agent '{agent_name}'. Available agents: {allowed}."
return await call_remote_agent(runtime_url, instruction)
orchestrator = Agent(
model=load_model(),
system_prompt=(
"You coordinate two API specialists. Use spec_analyzer for OpenAPI "
"parsing and risk analysis. Use doc_generator for documentation and "
"SDK examples. When both are needed, analyze the specification "
"first, then include that analysis in the documentation request. "
"Never assign work outside an agent's advertised capability."
),
tools=[delegate],
)
if __name__ == "__main__":
serve_a2a(StrandsA2AExecutor(orchestrator))
Two small things in there matter more than they look. The tool is async, which avoids wrapping the remote call in asyncio.run() that blows up when the agent is already running inside an event loop, and the traceback doesn't point anywhere near the cause. And returning a friendly string for an unknown agent, rather than raising, gives the model something it can recover from instead of a dead turn.
The URLs belong in deployment config, a service catalog, or an agent registry. Not in source control.
Let the Model Route, or Hard-Code the Path?
The code above lets the orchestrator model decide who to call and in what order. That's genuinely useful when requests vary. It's also the default choice for the wrong reasons.
If every documentation request must be analyzed before generation, always, no exceptions encode that sequence in Python and use the model only inside each step. Deterministic orchestration is easier to test, retry, and audit, and it doesn't drift when someone edits a prompt.
Save model-driven routing for routes that are actually ambiguous. It shouldn't be there to make a fixed workflow feel more autonomous.
Running It
With all three runtimes deployed, send the user request to the orchestrator. Something like "Review this customer API spec and create a Python quickstart" should cause it to call the analyzer, feed that analysis to the doc generator, and return one answer.
Should. That's expected behavior, not a guarantee the protocol hands you. Whether the right handoffs happen comes down to your orchestrator prompt, your tool descriptions, your eval cases, and your workflow code. A2A standardizes the conversation. It has no opinion about whether your orchestration logic is correct.
Before You Call It Production
Least privilege per runtime. Each runtime gets permissions for its own tools and data, nothing more. The failure mode to watch for is the orchestrator's role slowly accumulating every permission any specialist ever needed.
Distinguish auth failures from agent failures. A2A application errors come back as JSON-RPC error objects, frequently with HTTP 200. Auth and authorization failures come back as native HTTP errors AccessDeniedException is a 403. AgentCore also maps runtime exceptions onto JSON-RPC codes: -32051 for not found (404), -32053 for throttling (429), -32054 for conflict (409), -32055 for a runtime client error (424). Inspect the JSON-RPC shape. A 200 is not a success.
Retry deliberately. 429 and retryable-conflict 409 are worth retrying with backoff. A 424 usually means your container failed and the answer is in CloudWatch, not in a second attempt. Never blind-retry non-idempotent work, and pass a task or request identifier so a specialist can recognize a duplicate.
Trace every handoff. AgentCore Observability gives you sessions, traces, and spans. Propagate your own business correlation data on top, or you'll have three sets of traces and no way to prove they belong to the same user request.
Set lifecycle values on purpose. Defaults are a 15-minute idle timeout and an 8-hour max lifetime, both configurable (60 to 28,800 seconds). Long-running tasks may justify raising them; a bigger number should be a decision, not a leftover. Also watch that your A2A sessions actually go idle — a server that holds a connection open after the client is gone will happily keep a microVM warm on your bill.
Cache Agent Cards carefully. Resolving on every call is simple and costs you a round trip per hop. Cache with a short TTL and invalidate when an agent's version or capabilities change. A stale card advertising a skill that no longer exists is a genuinely confusing bug.
Keep protocol and SDK versions aligned. Pin dependencies, validate cards, and test a mixed-framework setup rather than assuming it works.
Use private connectivity if you need it. Runtime can attach to a VPC for internal resources, and interface VPC endpoints via PrivateLink keep AgentCore API calls off the public internet.
A Failure Test Matrix
Before you call it done, break the handoffs rather than admiring the happy path:
- The specialist's Agent Card is unreachable or malformed.
- The bearer token expires between card resolution and invocation.
- The analyzer returns a valid task with no usable artifact.
- The doc generator times out after the analyzer succeeded.
- The orchestrator picks an agent that doesn't exist, or one it isn't allowed to call.
- A retry re-runs a task that already had an external side effect.
Each of these teaches you something a successful demo can't. They also make the division of labor concrete: Runtime hosts and protects the endpoint, A2A defines the conversation, and workflow correctness is still entirely yours.
Closing
The value here isn't the agent count. It's that each capability gets a clear boundary, a narrow permission set, and its own release cycle. Runtime provides the managed execution boundary; A2A provides a common way to discover agents and pass work between them.
So, start with one agent. Split it when responsibilities, ownership, or security boundaries justify a network hop, and not before. Where the sequence is fixed, keep the orchestration deterministic. Where routing is genuinely variable, let the orchestrator reason over well-described capabilities, and instrument every handoff you add.
References
- AWS: Deploy A2A servers in AgentCore Runtime
- AWS: A2A protocol contract
- AWS: AgentCore Runtime authentication and authorization
- AWS: LifecycleConfiguration API reference
- AWS: Configure AgentCore Runtime for VPC access
- AgentCore CLI
- Agent2Agent protocol specification
Opinions expressed by DZone contributors are their own.
Comments