DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Cloud Architecture

Cloud architecture refers to how technologies and components are built in a cloud environment. A cloud environment comprises a network of servers that are located in various places globally, and each serves a specific purpose. With the growth of cloud computing and cloud-native development, modern development practices are constantly changing to adapt to this rapid evolution. This Zone offers the latest information on cloud architecture, covering topics such as builds and deployments to cloud-native environments, Kubernetes practices, cloud databases, hybrid and multi-cloud environments, cloud computing, and more!

icon
Latest Premium Content
Trend Report
Cloud-Native Foundations
Cloud-Native Foundations
Refcard #370
Data Orchestration on Cloud Essentials
Data Orchestration on Cloud Essentials
Refcard #379
Getting Started With Serverless Application Architecture
Getting Started With Serverless Application Architecture

DZone's Featured Cloud Architecture Resources

Supercharging AI Agents with Azure Context: A Hands-On Guide to Azure MCP

Supercharging AI Agents with Azure Context: A Hands-On Guide to Azure MCP

By Ammar Ekbote
As Large Language Models (LLMs) continue their rapid trajectory of development, software engineers and cloud architects regularly run into two systemic challenges: The Fixed Knowledge Cutoff: Model intelligence is inherently restricted to its training data window, making it blind to real-time changes. The "Air-Gap" Limitation: Out of the box, LLMs cannot securely interact with external systems or private APIs on their own. Historically, developers bypassed these hurdles by writing fragile, ad-hoc API wrappers or custom orchestrators. Enter the Model Context Protocol (MCP): an open standard designed to standardize how AI applications safely connect to external data sources and execution environments.In this article, we will explore the core architecture of MCP, look at why the Azure MCP Server is a game-changer for cloud engineers, and walk through a step-by-step guide to configuring it inside Visual Studio Code. The Core Architecture of MCP At its heart, MCP establishes a uniform "language" that allows AI applications (Hosts) to talk to external resources (Servers). Rather than building custom integrations for every new LLM or tool, developers can rely on a single, clean architecture: MCP Architecture The standard is built around four fundamental building blocks: MCP Host: The runtime environment or user interface where the AI agent operates (e.g., VS Code, Claude Desktop, Cursor).MCP Client: The architectural component within the host that initiates and maintains the active connection.MCP Server: A lightweight, modular helper service that exposes specific resources, prompts, and tools.Transport Layer: The underlying protocol facilitating communication, typically utilizing JSON-RPC 2.0 over standard input/output (stdio) or HTTPS. MCP Components How the MCP Handshake Works Instead of executing raw, unpredictable bash scripts, the interaction is highly structured: Initialization: The client connects to the server and queries its capabilities. Declaration: The server returns a structured schema listing the specific tools it supports. Execution Request: When an LLM determines it needs external data, the host requests a specific tool execution from the server. Context Injection: The server executes the local process, gathers the result, and returns a semantic JSON response to the host, which is then cleanly formatted for the user. MCP in Action Why Use the Azure MCP Server? If you are managing infrastructure on Microsoft Azure, the Azure MCP Server bridges the gap between your local AI assistant and your active cloud resources. Operating as a secure local process, it natively integrates with the Azure command-line context. The server supports over 40 Azure services and more than 170 tools out of the box, spanning critical cloud primitives: Compute & Containers: Azure Container Apps and Azure Kubernetes Service (AKS). Storage & Resource Management: Azure Storage (blobs and containers) and Azure Resource Groups. Infrastructure as Code: Integrated Azure Terraform Best Practices. Key Advantages Over Raw CLI Executions While you could technically let an AI agent execute arbitrary commands in a terminal, using a dedicated MCP server provides several major structural benefits: Rich Semantics: Instead of parsing messy, unstructured terminal stdout text, the MCP server passes rich semantic data blocks directly back to the LLM. Strict Governance & Scope: You can explicitly configure the server to run in read-only mode or expose only a select subset of tools, preventing the AI from accidentally deleting production infrastructure. Interactive Guardrails: The protocol requires explicit user confirmation before executing tools that touch sensitive data or perform mutative actions. Azure MCP Server Setup and Configuration Modes You can run the Azure MCP server locally across multiple development setups using stdio transport. Below are the three most common configuration schemas. NuGet Configuration For .NET teams, the server can be dynamically fetched and started using the dnx toolchain: JSON { "mcpServers": { "Azure MCP Server": { "command": "dnx", "args": [ "Azure.Mcp", "--source", "https://api.nuget.org/v3/index.json", "--yes", "--", "azmcp", "server", "start" ], "type": "stdio" } } } Node.js Configuration If you are developing in a standard JavaScript/TypeScript ecosystem, you can spin up the server dynamically using the latest npm package via npx: JSON { "mcpServers": { "azure-mcp-server": { "command": "npx", "args": [ "-y", "@azure/mcp@latest", "server", "start" ] } } } Docker Configuration For isolated development environments, you can run the server in a container. Note that you must provide a local environment file (.env) containing your Azure Service Principal credentials: JSON { "mcpServers": { "Azure MCP Server": { "comman { d": "docker", "args": [ "run", "-i", "--rm", "--env-file", "/full/path/to/.env", "mcr.microsoft.com/azure-sdk/azure-mcp:latest" ] } } } Step-by-Step Visual Studio Code Integration A great feature for those who want to work within Visual Studio Code is that they can also manage the Azure MCP Server through an exclusive extension available directly in their editor. The following steps walk you through the setup. Step 1: Install the Extension Search for and install the Azure MCP Server extension directly from the Visual Studio Code marketplace. Install Azure MCP Server Extension Step 2: Initialize and Verify Open the Command Palette (Cmd + Shift + P on macOS or Ctrl + Shift + P on Windows) and search for the extension commands to verify the server is active and running. Select an MCP Server Step 3: Configure Your Tool Accessibility Open your integrated AI chat window and select the Tools icon. Here, you will see a list of all active tools provided by the Azure MCP Server. You can toggle specific permissions on or off—for instance, disabling write operations while keeping read operations active. Select Your Tools Step 4: Interact Natively Now, your AI chat assistant can securely call Azure tools in the background. You can ask complex queries like: "Are there any inactive containers running in my resource group?""Upload our local configuration file directly to our Azure storage container blob storage." Any further interactions that reference an Azure account would use the MCP tools. Confirm Tools and Resources Real-World Engineering Use Cases To see the power of this setup, let’s look at how this changes day-to-day operations: Use Case A: Automated AKS Incident Troubleshooting When an incident occurs in an Azure Kubernetes Service (AKS) cluster, engineers typically run dozens of diagnostic commands. With the Azure MCP server connected, you can simply ask the LLM: "Investigate why the pods in our production namespace are crash-looping." The agent will call the relevant AKS tools, inspect the logs, identify the misconfiguration, and suggest the fix - all in seconds. Use Case B: Continuous Terraform and Compliance Audits Before deploying infrastructure, you can point your local AI agent to your code directory. Because the server incorporates Azure Terraform Best Practices, the agent can audit your configuration files, cross-reference them against your live Azure Resource Groups, and warn you if you are violating security compliance rules or generating drift. Conclusion The Model Context Protocol represents a major step forward in AI-assisted development. By standardizing the communication layer, the Azure MCP Server enables software engineers to transform static, isolated LLMs into active, context-aware cloud operators. Whether you are monitoring active Kubernetes clusters, auditing Terraform configurations, or automating file uploads to blob storage, MCP gives your AI assistant safe and highly scoped access to the "live" Azure ecosystem. More
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent

Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent

By Jubin Soni, FBCS DZone Core CORE
The Foundry model catalog covers a lot of ground, but it doesn't cover everything. If you need a model trained on your own proprietary, tabular data, a churn predictor built on your actual customer history, a fraud score trained on your actual transaction patterns, that's not a model catalog problem. That's a real machine learning problem, and on Microsoft's stack it belongs to Azure Machine Learning, a genuinely separate platform from Foundry with its own SDK, its own workspace concept, and its own deployment model. This is a hands-on build of the whole path: train a model with Azure ML's SDK, register it, deploy it behind a managed endpoint, and then wire that endpoint into a Foundry agent as a function tool, so a conversational agent can call your custom model mid-conversation the same way it would call any other tool. The Mental Model First Azure ML and Foundry don't share a runtime. They share an ecosystem, and a genuine amount of engineering effort has gone into making the seam between them small, but it's worth being precise about where that seam actually is: A command job is how training happens. You point Azure ML's SDK at a script, a compute target, and a set of inputs, and it runs your training code as a tracked, reproducible job.The model registry is where a trained model becomes a named, versioned artifact, independent of the job that produced it. This is what lets you promote a specific version to production without recreating the training run.A Managed Online Endpoint is how a registered model actually serves predictions. It's a real, always-on (or autoscaling) piece of infrastructure with its own URL, its own auth, and its own cost, running inside the Azure ML workspace, not inside Foundry.A Foundry agent's FunctionTool is the bridge. Foundry doesn't know or care that the function it's calling happens to invoke an Azure ML endpoint. From the agent's perspective, it's just a tool with a name, a schema, and a result. Prerequisites An Azure ML workspace and a Foundry project, in the same or different resource groups; it doesn't matter which, since nothing about this integration requires them to share infrastructure.Python 3.9+ with both SDKs installed. Python pip install azure-ai-ml azure-ai-projects azure-identity Python from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential ml_client = MLClient( DefaultAzureCredential(), subscription_id="<subscription-id>", resource_group_name="<resource-group>", workspace_name="<aml-workspace-name>", ) Step 1: Train the Model A command job wraps a training script, here a plain scikit-learn classifier, and runs it on managed compute. Python from azure.ai.ml import command, Input, Output train_job = command( code="./src", command="python train.py --data ${{inputs.training_data} --model_output ${{outputs.model_output}", inputs={"training_data": Input(type="uri_folder", path="azureml://datastores/workspaceblobstore/paths/churn-training/")}, outputs={"model_output": Output(type="uri_folder")}, environment="azureml://registries/azureml/environments/sklearn-1.5/labels/latest", compute="cpu-cluster", display_name="churn-model-training", ) returned_job = ml_client.jobs.create_or_update(train_job) ml_client.jobs.stream(returned_job.name) ml_client.jobs.stream blocks and prints logs until the job finishes, which is worth doing in any script you're actually going to run rather than fire-and-forget, since a training job failing silently in the background is a bad way to find out your pipeline is broken. Step 2: Register the Trained Model Registration turns the job's output into a named, versioned artifact you can reference independently of the job. Python from azure.ai.ml.entities import Model from azure.ai.ml.constants import AssetTypes model = ml_client.models.create_or_update( Model( path=f"azureml://jobs/{returned_job.name}/outputs/model_output", name="churn-classifier", type=AssetTypes.MLFLOW_MODEL, description="Customer churn classifier, trained on 18 months of account history.", ) ) Using MLFLOW_MODEL as the type here isn't incidental. If your training script logs the model with MLflow's autologging, Azure ML's managed endpoints can deploy it with a built-in scoring container, no custom score.py inference script required. That's a real time saver worth designing your training script around from the start rather than discovering after the fact. Step 3: Deploy It Behind a Managed Endpoint Python from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment endpoint = ManagedOnlineEndpoint(name="churn-endpoint", auth_mode="key") ml_client.online_endpoints.begin_create_or_update(endpoint).result() deployment = ManagedOnlineDeployment( name="blue", endpoint_name="churn-endpoint", model=model, instance_type="Standard_DS3_v2", instance_count=1, ) ml_client.online_deployments.begin_create_or_update(deployment).result() endpoint.traffic = {"blue": 100} ml_client.online_endpoints.begin_create_or_update(endpoint).result() The blue deployment name isn't a convention you have to follow, but it's worth keeping, since it sets up the pattern you'll want the first time you deploy a new model version: create a green deployment alongside blue, split traffic between them, and shift fully once you trust the new version, rather than replacing blue outright and hoping. Step 4: Confirm It Works Before Anything Else Touches It Python import json test_input = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[14, 89.50, 3]]} response = ml_client.online_endpoints.invoke( endpoint_name="churn-endpoint", request_file=None, deployment_name="blue", input_data=json.dumps(test_input), ) print(response) Get a real prediction back here before wiring anything else into it. Debugging a broken endpoint through the extra layer of an agent's function-calling loop is meaningfully harder than debugging it directly. Step 5: Wrap the Endpoint as a Foundry Agent Function Tool This is the actual bridge. The Foundry agent doesn't call Azure ML directly; your application code does, in response to the agent asking for it. Python import os import requests from azure.ai.projects import AIProjectClient from azure.ai.projects.models import PromptAgentDefinition, Tool, FunctionTool from azure.identity import DefaultAzureCredential def get_churn_score(tenure_months: int, monthly_spend: float, support_tickets: int) -> dict: payload = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[tenure_months, monthly_spend, support_tickets]]} resp = requests.post( "https://churn-endpoint.<region>.inference.ml.azure.com/score", headers={"Authorization": f"Bearer {os.environ['AML_ENDPOINT_KEY']}", "Content-Type": "application/json"}, json=payload, timeout=10, ) resp.raise_for_status() return {"churn_probability": resp.json()[0]} func_tool = FunctionTool( name="get_churn_score", description="Predict churn probability for a customer given tenure, spend, and support ticket history.", parameters={ "type": "object", "properties": { "tenure_months": {"type": "integer", "description": "How many months the customer has been active."}, "monthly_spend": {"type": "number", "description": "Average monthly spend in dollars."}, "support_tickets": {"type": "integer", "description": "Number of support tickets in the last 90 days."}, }, "required": ["tenure_months", "monthly_spend", "support_tickets"], "additionalProperties": False, }, strict=True, ) project = AIProjectClient(endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"], credential=DefaultAzureCredential()) tools: list[Tool] = [func_tool] agent = project.agents.create_version( agent_name="retention-agent", definition=PromptAgentDefinition( model="gpt-4.1-mini", instructions="Help the team assess churn risk. Call get_churn_score whenever specific customer numbers are provided.", tools=tools, ), ) Step 6: Run It End to End Python import json from openai.types.responses.response_input_param import FunctionCallOutput openai_client = project.get_openai_client() conversation = openai_client.conversations.create() response = openai_client.responses.create( input="A customer's been with us 14 months, spends about $90/month, and filed 3 tickets recently. Churn risk?", conversation=conversation.id, extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}, ) for item in response.output: if item.type == "function_call" and item.name == "get_churn_score": result = get_churn_score(**json.loads(item.arguments)) follow_up = openai_client.responses.create( input=[FunctionCallOutput(type="function_call_output", call_id=item.call_id, output=json.dumps(result))], conversation=conversation.id, extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}, ) print(follow_up.output_text) The model decides whether the question warrants calling get_churn_score at all; your code executes the actual HTTP call to Azure ML when it does, and the result flows back into the same conversation for the model to finish answering. Nothing about the agent definition or the Responses API calls changes based on what the function actually does behind the scenes, whether it hits a database, a Lambda function, or, as here, a completely separate Azure ML workspace. Where This Fits in the Bigger Picture Worth being explicit about something the SDKs don't make obvious on their own: Azure ML's managed endpoint and the Foundry project endpoint are two different resources with two different auth boundaries, and nothing about this integration merges them. The Azure ML endpoint has its own key or managed identity, scoped to that workspace. The Foundry project has its own credential, scoped separately. FunctionTool doesn't create a trust relationship between the two; it's just a name and a schema the model uses to ask your code to do something. The actual authentication to Azure ML happens entirely in your own function implementation, same as it would if that function called any other external API. That's a feature, not a limitation. It means the Azure ML side of this can be owned, secured, and rotated independently of the Foundry side, by a different team if that's how your organization is structured, without either side needing write access to the other's resources. Production Considerations Before You Commit Process expires 10 minutes after a function call is issued. Submit the tool's output back to the conversation before that window closes, or the run fails. If get_churn_score's HTTP call to Azure ML is slow, that's a hard deadline, not a soft one; budget for it.Prefer a managed identity over a static endpoint key once this is more than a prototype. A key baked into an environment variable is fine for local testing and a real liability in anything that runs unattended. Grant the identity running your function tool code the AzureML Data Scientist or a more narrowly scoped custom role against just this endpoint, not the whole workspace.Blue/green the endpoint; don't overwrite it. Standing up a new deployment alongside the old one and shifting traffic gradually is the only way to catch a regression in a newly trained model before it's serving 100% of real requests.Treat strict=True on the FunctionTool schema as a correctness feature, not boilerplate. It's what keeps the model from calling get_churn_score with a malformed or partial argument set that would otherwise fail inside your function rather than being caught before the call.Online endpoint compute is billed whether or not it's actively scoring. Unlike a serverless model call, a Managed Online Endpoint with instance_count=1 running around the clock costs money at idle. If call volume is low and bursty, look at scale-to-zero options or batch endpoints instead of defaulting to always-on instance count 1.Model drift doesn't announce itself. Nothing in this pipeline retrains automatically when the real-world distribution shifts away from what the model was trained on. Log the inputs and outputs of every get_churn_score call somewhere you can actually review, and revisit training data on a real cadence, not only when someone notices predictions have gotten worse. Where This Leaves You Foundry's model catalog is the right tool for the overwhelming majority of generative AI work, and for teams that never need a model trained on their own proprietary data, it's genuinely the whole story. But the moment the problem is "predict something specific about our own customers, from our own historical data," that's Azure ML's job, and pretending otherwise means either forcing a language model to approximate a task it was never built for, or building your own training and serving infrastructure by hand. The actual integration effort here is small: a training job, a registered model, a managed endpoint, and one FunctionTool definition. The two platforms don't need to merge for that to work. They just both need to keep doing the one thing each is actually good at. References Microsoft Learn. "Train models with the Python SDK v2." Azure Machine Learning. learn.microsoft.com/en-us/azure/machine-learning/how-to-train-model?view=azureml-api-2Microsoft Learn. "Deploy machine learning models to online endpoints." Azure Machine Learning. learn.microsoft.com/en-us/azure/machine-learning/how-to-deploy-online-endpoints?view=azureml-api-2Microsoft Learn. "Use function calling with Microsoft Foundry agents." learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/function-callingMicrosoft Learn. "Explore Microsoft Foundry Models in Azure Machine Learning." learn.microsoft.com/en-us/azure/machine-learning/foundry-models-overview?view=azureml-api-2Microsoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overview More
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
By Naga Santhosh Reddy Vootukuri DZone Core CORE
Your Cloud Diagram Is Already Out of Date: An Operating Model for Continuous Security Architecture
Your Cloud Diagram Is Already Out of Date: An Operating Model for Continuous Security Architecture
By Avik Mukherjee
The Silent Container Death: A TCP Dial That Never Times Out
The Silent Container Death: A TCP Dial That Never Times Out
By Alexander Fo
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams

Most AWS migration projects don't fail because of technical complexity. They fail because teams treat migration as a single activity rather than a set of distinct strategies applied to different workloads. AWS defines seven migration strategies — the 7Rs — that determine how each application moves to the cloud. The decision of which strategy applies to which workload has more impact on project cost, timeline, and outcome than any architectural choice you'll make after. Yet in practice, most teams default to "lift-and-shift everything" without evaluating whether that's appropriate. This article presents a practitioner's framework for classifying workloads into the 7Rs, based on delivering 50+ AWS migrations across fintech, SaaS, healthcare, and e-commerce. The 7R Strategies Retire Not every workload deserves migration. During discovery, you will invariably find applications that are redundant, unmaintained, or replaceable. In a typical enterprise portfolio of 20–40 applications, 10–20% qualify for retirement. Decision criteria: No active users, duplicate functionality already covered by another system, or maintenance cost exceeds business value. Common mistake: Teams skip this step because retiring applications requires stakeholder conversations. The result is migrating dead applications that consume compute budget indefinitely. Retain Some workloads shouldn't migrate in this wave. Applications with deep hardware dependencies, pending end-of-life within 12 months, or complex regulatory constraints that require legal review before cloud deployment are candidates for retention. Decision criteria: High migration complexity combined with low business urgency, or external constraints that prevent cloud deployment within the project timeline. Retain is not "never migrate." It's "not now." Document these workloads with a future migration path and trigger conditions. Rehost (Lift-and-Shift) Moving applications to EC2 or containers without code modifications. AWS Application Migration Service (MGN) automates this by continuously replicating servers and orchestrating cutover with minutes of downtime. Decision criteria: Application has a short remaining lifespan (1-2 years), speed of migration matters more than optimization, or the application is a black box with no available source code. Timeline: Days to weeks per workload. Trade-off: You gain cloud elasticity and pay-as-you-go pricing immediately, but you inherit all existing architectural inefficiencies. A poorly designed monolith on-premises becomes a poorly designed monolith on EC2. Relocate Hypervisor-level migration, primarily for VMware workloads moving to VMware Cloud on AWS. The OS, application, and configuration remain untouched. Decision criteria: Large VMware estate, tight data center exit deadline, and applications that cannot tolerate any configuration change. Replatform Migration with targeted adaptations to managed services. The application architecture stays intact, but you replace self-managed infrastructure components with AWS equivalents: Self-ManagedAWS ManagedOperational BenefitSelf-hosted PostgreSQLRDS for PostgreSQLAutomated backups, patching, failoverCron jobs on EC2EventBridge + LambdaNo server to maintain, pay-per-invocationSelf-managed RedisElastiCacheAutomatic failover, scalingNginx load balancerApplication Load BalancerManaged TLS termination, WAF integrationSelf-hosted ElasticsearchOpenSearch ServiceManaged cluster scaling, snapshots Decision criteria: Application is well-structured but operationally expensive. The team spends significant time on database maintenance, patching, backup verification, or scaling. Timeline: 2-4 weeks additional per workload compared to rehost. Trade-off: Moderate additional effort (schema compatibility testing, connection string changes) in exchange for a 40-60% reduction in ongoing operational cost. For most mid-complexity applications, replatforming represents the optimal balance between migration effort and long-term benefit. Refactor (Re-Architect) Rebuilding applications for cloud-native patterns: microservices decomposition, containerization (ECS/EKS), serverless (Lambda), event-driven architecture (EventBridge, SQS, SNS, Step Functions). Decision criteria: The application is a core business asset that needs capabilities the current architecture cannot deliver — true horizontal scaling, independent service deployments, multi-region active-active, or zero-downtime deployments. Timeline: Months. Budget accordingly. Trade-off: Highest upfront investment, but delivers the best long-term results in terms of deployment velocity, fault isolation, and scaling capability. Reserve this for 2-3 applications maximum in a migration portfolio. Repurchase Replacing custom-built software with a commercial SaaS product. The application doesn't move to AWS; it moves to a vendor. Decision criteria: The in-house application solves a problem that is not a core competency and commercially available alternatives have matured to cover your requirements. Common candidates: CRM, HR systems, monitoring, project management. The Decision Framework Classification should happen during the assessment phase, before any infrastructure work begins. For each workload, evaluate four dimensions: 1. Business Value How critical is this application to revenue generation or core operations? High: Core product, customer-facing, revenue-generatingMedium: Internal operations, supports core processesLow: Legacy, rarely used, or duplicate functionality 2. Technical Complexity How difficult is it to migrate given current architecture, dependencies, and state management? High: Stateful, tightly coupled, hardware dependencies, proprietary protocolsMedium: Standard web application with database, some external integrationsLow: Stateless, containerizable, standard protocols 3. Team Capacity Does your engineering team have the skills and bandwidth to support a complex migration approach? High capacity: Can support re-architecting alongside other workLimited capacity: Can handle replatforming with some external supportMinimal capacity: Rehost or retain is the only realistic option 4. Time Constraint How quickly must this workload be operational on AWS? Immediate (weeks): Data center exit, contract expiryStandard (1-3 months): Planned migration within a programFlexible (3-6 months): Can wait for deeper optimization Mapping Dimensions to Strategy Business ValueComplexityCapacityTimeRecommended StrategyLowAnyAnyAnyRetire or RepurchaseAnyHighLowImmediateRehost (with future replatform plan)MediumMediumMediumStandardReplatformHighMedium-HighHighFlexibleRefactorAnyAnyAnyBlockedRetain A Practical Example Consider a portfolio of 15 applications for a mid-size SaaS company: Plain Text ┌─────────────────────────────────────────────────────┐ │ RETIRE (3) │ │ - Legacy admin panel (replaced by new one 2024) │ │ - Internal wiki (moved to Confluence) │ │ - Prototype service (never went to production) │ ├─────────────────────────────────────────────────────┤ │ RETAIN (1) │ │ - Hardware security module integration │ │ (requires legal review for cloud deployment) │ ├─────────────────────────────────────────────────────┤ │ REHOST (4) │ │ - Backoffice tools (low traffic, stable) │ │ - Legacy reporting engine (EOL in 18 months) │ │ - Monitoring collector agents │ │ - Staging environment clone │ ├─────────────────────────────────────────────────────┤ │ REPLATFORM (5) │ │ - Main API (PostgreSQL → RDS, cron → Lambda) │ │ - Worker services (EC2 → ECS Fargate) │ │ - File processing pipeline (S3 + Lambda) │ │ - Authentication service (→ElastiCache for sessions│ │ - Notification service (→ SES + SQS) │ ├─────────────────────────────────────────────────────┤ │ REFACTOR (1) │ │ - Core product platform (monolith → microservices) │ ├─────────────────────────────────────────────────────┤ │ REPURCHASE (1) │ │ - Custom CRM (→ HubSpot) │ └─────────────────────────────────────────────────────┘ This distribution — 20% retire, 7% retain, 27% rehost, 33% replatform, 7% refactor, 7% repurchase — is representative of what I see in practice. The replatform bucket is almost always the largest. Migration Tooling Alignment Each strategy maps to specific AWS tooling: StrategyPrimary ToolsRehostAWS Application Migration Service (MGN), Migration HubReplatformDMS (databases), manual adaptation, Terraform/IaCRefactorECS/EKS, Lambda, Step Functions, custom developmentRelocateVMware Cloud on AWS AWS Migration Hub provides a unified tracking dashboard across all strategies. For database migrations specifically, AWS Database Migration Service (DMS) handles both homogeneous and heterogeneous migrations with continuous replication (CDC), enabling near-zero-downtime cutovers. Common Anti-Patterns "Rehost everything, optimize later." Teams that plan to rehost first and replatform in a second phase rarely execute phase two. The urgency disappears once applications are running, and the team moves to other priorities. If replatforming is the right strategy, do it during migration. "Refactor everything for cloud-native." The opposite extreme. Not every application needs microservices. A well-structured monolith running on ECS Fargate can serve thousands of requests per second with simpler operations than a distributed system. "One strategy for all workloads." Every application in the portfolio has different characteristics. The decision framework exists because one size does not fit all. Conclusion The 7R classification exercise takes 3-5 days for a typical portfolio. It requires involvement from engineering leads, product owners, and sometimes finance (for retire/repurchase decisions). The output - a workload-by-workload strategy map - becomes the foundation for accurate timeline estimates, resource planning, and budget allocation. Without it, you're building infrastructure for workloads that might not need to exist. For a comprehensive breakdown of migration costs, the full 6-phase delivery process, and AWS tooling details, see my complete AWS cloud migration guide.

By Jerzy Kopaczewski
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data

Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq

By Jubin Soni, FBCS DZone Core CORE
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults

Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report

By Naga Santhosh Reddy Vootukuri DZone Core CORE
Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction
Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction

Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. After a few years of operating a shared Kubernetes environment, the shift in the center of gravity becomes clear. Cluster provisioning, container scheduling, and upgrades become routine, yet releases still stall over ownership, access, telemetry, and cost allocation. Consider a hypothetical product team adding a stateful order-processing service to a shared platform. The service has an API, a worker, database migrations, and a data store backed by cluster-managed persistent storage, and it must run in staging and production. We will follow that service through its delivery path to examine where mature infrastructure stops helping, how local workflow differences compound, and which operating model decisions restore consistency without stripping teams of useful autonomy. When Infrastructure Maturity Stops Solving the Hard Part At first glance, onboarding the service should be routine. The cluster exists, the CI system can build an image, and infrastructure as code can create the namespace. Then the service reaches production and encounters a different StorageClass, quota profile, network policy, or service account configuration from staging. Each difference may be valid, but the delivery workflow didn’t surface the environment contract early enough. This is the practical limit of infrastructure maturity. Reliable clusters provide capable building blocks, while reliable delivery also requires a shared agreement about how teams use those blocks, what evidence a release produces, where exceptions go, and who owns the outcome. How Cloud Complexity Starts to Compound Follow the service through one release and the pattern becomes clearer: The same components use inconsistent service and environment identifiers across logs and traces.Ownership labels exist in one cluster but not the other.The team copies a pipeline because the shared template cannot sequence migrations.Production access and policy exceptions move through separate ticket queues. The operational cost shows up in the manual coordination required before each deployment. An engineer has to reconstruct which rules apply every time. During an incident, responders can’t move cleanly from an alert to the owning team, deployment record, runbook, and cost center. Finance sees shared-cluster spend that cannot be attributed reliably, while the security team receives evidence in different formats. OpenTelemetry semantic conventions and FinOps allocation practices rely on consistent service, environment, and allocation metadata, so local naming schemes undercut the value of the underlying tools. As the same pattern spreads across clusters and cloud accounts, small differences become a persistent operating burden. Operating Models Set the Terms of Scale The operating model decides who turns those building blocks into a usable delivery system. For our example, the product team owns the order domain, data model, migration safety, scaling behavior, SLOs, and on-call response. The platform team owns the interface through which the service receives a namespace, workload identity, baseline policy, deployment workflow, and telemetry defaults. Security, SRE, and FinOps teams contribute requirements and review the evidence that the workflow produces. This split keeps service-specific decisions close to the people who understand them while centralizing cross-cutting capabilities that every team would otherwise rebuild. CNCF’s platform guidance makes an important distinction here: A platform team is responsible for the interfaces and experience around shared capabilities, even when another team or provider operates the backing service. In practice, a mature platform offers versioned workflows, clear support boundaries, self-service for common requests, and feedback loops based on real usage. In this way, the platform team is an enabler of consistency rather than the operator of every component. Standardization, Autonomy, and Shared Operating Logic The defensible baseline is narrower than a universal application architecture. For this service, shared standards should cover: Service and environment identityWorkload identity and minimum network policyResource requests, quota expectations, and cost-allocation metadataRelease evidence, rollback behavior, and minimum telemetry These rules belong in the shared workflow because inconsistency affects other teams and complicates incident response, security, and cost allocation. The product team still chooses its schema, partitioning strategy, cache design, scaling thresholds, and release timing, and defines SLOs around the behavior users experience. The stateful workload then tests that boundary. A default pipeline built for stateless HTTP services may need a supported hook for migrations and worker rollout. An overly broad standard becomes an approval layer or bottleneck, while an overly narrow one leaves every team maintaining its own release and recovery logic. A bounded extension with an owner, tests, constraints, and review date preserves autonomy without creating an unsupported parallel system. Why Team Friction Turns Into a Scaling Tax Weaknesses in the operating model become most visible in the friction between teams. If a developer must request a namespace, ask another team for credentials, copy a pipeline, and find a production approver, the architecture may be automated while delivery remains ticket-driven. Each handoff adds queue time and loses context. Adding a portal without changing that path gives the developer one more place to check. Effective self-service completes the request, applies policy, records the change, and returns a clear support path. To see whether self-service is reducing friction, track metrics like request-to-environment time, time to first production deployment, exception rate, support demand, and failed-deployment recovery time. CNCF recommends tracking fulfillment and new-service delivery latency; DORA advises applying delivery metrics in the context of a specific service. Together, these measures show whether the workflow reduced coordination overhead or moved it to another queue. When Control Models Backfire The order-processing service example exposes two ways the control model can fail: Overly rigid standardization. A workflow designed only for stateless services forces the team to create a separate migration path, fragmenting release evidence.Unbounded local variation. Unrestricted cluster access allows identity, policy, and resource controls to drift between teams. The scalable approach pairs a narrow baseline, enforced through mechanisms like admission policies, with a documented extension path for legitimate workload-specific behavior, keeping the standard credible without turning each exception into a permanent fork. Operating Assumptions That Fail at Scale Old Assumption Why It Breaks Operating Model Replacement Healthy clusters make a workload portable Storage, identity, policy, and quota profiles differ by environment Versioned environment contract with a shared baseline One shared pipeline can serve every workload Stateful rollout and migration steps don’t fit the default sequence Core workflow with bounded, tested hooks A portal provides self-service Tickets and manual approvals remain behind the interface Workflow that provisions, enforces policy, and records evidence Local conventions remain harmless when teams own their services Metadata and controls drift across services Small enforced baseline with governed exceptions Making Cloud Complexity More Manageable To begin, you don’t need to redesign your entire platform. You can trace one representative delivery workflow and find where coordination breaks. For the order-processing service example, map the path from repository creation to production, including owners, queues, controls, evidence, and exceptions. Improvements should then be tested through adoption and outcomes such as lead time, failed-deployment recovery time, support demand, exception volume, and cost-attribution coverage. This sequence shows whether the platform is reducing operational variation for real workloads before the model expands to more teams and environments. References: Platforms for Cloud-Native Computing, CNCFResource Quotas, KubernetesStorage Classes, KubernetesAdmission Control in Kubernetes, KubernetesResource Semantic Conventions, OpenTelemetryAllocation FinOps Framework Capability, FinOps FoundationService Level Objectives, Google SRESoftware Delivery Performance Metrics, DORA This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report

By Igboanugo David Ugochukwu DZone Core CORE
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control

Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Kubernetes environments can drift, accumulate one-off fixes, and diverge across teams until a routine deploy breaks or a cost spike forces a review. This checklist gives platform, SRE, and engineering teams a way to keep clusters, deployments, and automation manageable as Kubernetes operations scale and teams grow. It covers standards, observability, releases, access, drift, and cost. Review it before promoting a service to production and revisit it as your environments shift. Cluster Standards and Environment Discipline Most teams run more than one Kubernetes cluster, and those clusters diverge over time as they are upgraded and modified independently. At that point, a fix or runbook that works on one cluster can’t be trusted to work on another. Keeping the fleet operable requires every cluster to run a supported Kubernetes version and follow the same approved platform settings and policies. Document centrally controlled settings (e.g., Kubernetes versions, networking, admission policies) separately from service-team settings (e.g., pod resource requests, autoscaling, ConfigMaps)Standardize namespace, labeling, and resource-quota conventions so workloads are identified and bounded consistently across clustersMaintain each cluster’s baseline configuration in version control; use reconciliation to apply it and correct untracked changesMaintain an approved Kubernetes version range across environments; track each cluster against this range and upgrade it before its current version reaches end of supportRecord any cluster setting that differs from the standard baseline, including the justification, approver, expiry date, and whether it must be restored or reapproved control areawhat to standardizeminimum evidence Kubernetes version Supported version range and upgrade cadence Version inventory showing every cluster within the supported range Cluster baseline Networking, ingress, and baseline policies Declarative config in version control, reconciled to live state Namespaces and quotas Naming, labels, and resource quotas Quota and label audit across clusters Exceptions Approved deviations from baseline Record of approved deviations with justification and expiry date Deployment Consistency and Release Safety Kubernetes makes it easy to ship a change to production several ways: a CI/CD pipeline, a Helm upgrade run by hand, or a kubectl apply straight from a laptop. Each runs different checks, but the manual ones skip the tests and approvals that a pipeline would enforce. A repeatable release path applies the same gates every time and provides a reliable way to recover when a deployment fails. Require every service to follow the same approved deployment path from commit to production, with consistent release steps and controls across teams and environmentsPromote the same versioned, immutable artifact through every environment without rebuilding it at each stageRequire every change to clear the same automated gates (e.g., tests, policy checks, health checks) before reaching productionRoll out production changes in stages (e.g., canary release, percentage-based traffic shift); automatically stop or roll back when predefined health criteria are not metFor every production change, require a rollback, feature disablement, or recovery path that has been tested before releaseFor each deployment, assign an owner accountable for monitoring it through release and triggering rollback on failureRecord every production deployment with its artifact version, approver, and timestamp so the active release stays auditable Observability and Operational Readiness A Kubernetes cluster keeps workloads running by restarting and rescheduling them, so a service can keep failing without the failure ever becoming obvious. A pod stuck in CrashLoopBackOff or failing its readiness probe can remain unhealthy for hours, and if it emits no metrics or logs of its own, there’s nothing to tell you what went wrong. Catching that early depends on each service surfacing its own signals rather than waiting for the cluster to show something is wrong. Require every new service to ship with a minimum observability baseline before production: metrics, structured logs, traces, and liveness and readiness probesDefine service health signals (e.g., latency, traffic, errors, saturation), each with a threshold and assigned team that responds when it is breachedStandardize structured logging and trace context so a request can be followed end to endRoute every alert to an on-call rotation or runbook; retire alerts no one acts onMaintain a quarterly reviewed runbook for each service, including known failure modes, escalation contacts, and recovery stepsSet minimum retention periods for metrics, logs, and traces, with documented justification and explicit approval for shorter retention periodsRun a post-incident review after every major outage; apply findings to update runbooks, alerts, and service baselines Access Controls and Automation Guardrails A Kubernetes cluster usually serves many teams and workloads through a single shared control plane. A role with too much access, for example, can affect them all at once. And when the credential is shared, there’s no way to tell later who actually made the change. Access that stays narrow and tied to a single identity keeps a mistake or a compromised account from impacting the whole cluster. Use namespace-scoped RBAC roles with only the required permissions; grant cluster-wide administrator access only through logged, justified, time-limited exceptionsGive each automation its own scoped service account so automated and privileged actions trace to a distinct identity instead of shared credentialsReserve break-glass access for emergency production changes, with time limits and post-use reviewUse admission policies to reject workloads with unsigned images, privileged containers, or settings barred by platform standardsRecord the actor, target, and timestamp for every privileged or automated action in the Kubernetes audit log; regularly review for activity that does not match an approved change or access requestUse short-lived, automatically rotated ServiceAccount tokens for workloads; revoke credentials and RBAC bindings when a person, workload, or automated process is decommissioned Drift and Failure Management Over time, a Kubernetes cluster’s live state can drift from the configuration stored in version control. This could be due to a hotfix applied directly to a live resource during an incident or an incomplete rollout that leaves the cluster partially updated. If those differences are not fixed, a subsequent deployment may conflict with the live state or overwrite a manual change, and version control may no longer accurately reflect what is running in the cluster. Use automated checks to compare live cluster state with the version-controlled baseline at defined intervals; record each mismatch and notify the team responsible for the affected resourceSet risk-based remediation deadlines for detected drift, requiring teams to restore the baseline or approve a time-limited exception for the changed configuration before the deadlineLog every manual production change and resolve it within a defined period by updating the baseline or reverting the live resource to its declared stateSet an SLO and error budget for each service, identify the team tracking budget use, and pause feature work to prioritize reliability fixes when the budget is exhaustedRun root-cause reviews for recurring failures and apply findings to update baselines, policies, and admission checks instead of patching each instanceTest failure scenarios (e.g., pod disruption, node loss, dependency outages) on a defined schedule, confirm services recover as expected, and track remediation for any gaps example drift patternwhat usually reveals it Manual live-resource change Reconciliation diff against declared state Version or baseline skew Scheduled cluster inventory audit Expired break-glass fix Exception register entry past its window Repeated failure patched one service at a time Same root cause across incident reviews Cost Awareness and Resource Discipline In Kubernetes, resource requests for CPU and memory determine how much cluster capacity is reserved for a workload. Teams may size these requests for peak demand and leave them unchanged even when normal usage is much lower. Across many workloads, this unused capacity adds up and can cause the cluster to run more nodes than actual demand requires, increasing infrastructure costs. Set CPU and memory requests based on representative usage data; set limits where appropriate based on workload behavior and reliability requirementsReview workloads whose requests exceed observed use by a defined threshold, accounting for traffic patterns and reliability needsRequire cost-allocation labels for each workload by team and namespace; correct unallocated spend and missing or inaccurate labelsReclaim idle and orphaned resources (e.g., unused volumes, stale namespaces, oversized nodes) on a monthly cadenceSet autoscaling thresholds based on demand and reliability requirements; periodically review settings that fall outside the approved rangeRegularly review sustained overprovisioning or low utilization; reduce excess capacity or record why it must be retained when avoidable cost exceeds a set threshold Closing Run this checklist before a service enters production and at regular intervals afterward. Repeat it when clusters are upgraded, team responsibilities change, or services are added or retired. Resolve failed checks and revisit approved exceptions before they expire. Unresolved configuration drift can accumulate across environments until teams begin to treat it as the intended baseline. This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report

By Abhishek Gupta DZone Core CORE
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.

The first warning sign wasn't an outage. It was a boring pull request. We changed one App Service setting. It was the sort of change that should have resulted in a small plan and a quick review. Instead, Terraform refreshed networking, private endpoints, DNS, Key Vaults, storage accounts, app services, and monitoring before showing what would actually change. Nothing was broken; that was the point. Terraform did exactly what it was designed to do: account for everything represented in state before calculating change. The problem was that our Terraform state had become a single, platform-sized boundary that every small change had to pass through, and one no team could fully own. If you have run a landing zone as a single Terraform configuration, you have probably had a version of that pull request. The instinct afterward is to blame size: the configuration has grown too large, so break it up. That instinct is wrong, or at least incomplete. Size is uncomfortable, but coupling is what actually hurts. Nothing in the change touched networking, DNS, or those key vaults. They were dragged into the plan because everything was bound together through one state. At first, that coupling just means slow plans and noisy reviews. Later, it raises a harder question: who actually owns this? Where the Coupling Shows Up Start with the plan. In a monolith, Terraform has to account for everything represented in the state before it can tell you what changed. You can target a single resource, but that is an escape hatch, not a way to run a platform. So the wait scales with the size of the estate, not your change. Both a one-line edit and a fifty-resource migration get stuck behind the same refresh before the diff appears. Provider upgrades show the same problem. A single root configuration pins one set of provider versions, so you cannot move networking to a newer azurerm version and leave everything else behind. Every upgrade becomes all-or-nothing, which means it keeps losing to smaller, safer priorities. Ours sat on azurerm 2.97 and only moved to the 4.x line once the upgrade could no longer be put off. The monolith had made the jump too big to schedule any sooner. The bigger concern is blast radius. One state file, one lock, one plan. A bad apply, a corrupted state, a destroy that catches more than you aimed at: whatever goes wrong can reach more of the platform than the change was ever meant to touch, because nothing in the layout is there to contain it. The dependency graph suffers too. Unrelated resources get sequenced together just because they share a graph. A network change might wait on unrelated compute, DNS on policy. The graph ends up reflecting accidental grouping rather than real dependencies. The result is clear. There is no small change. You cannot ship a DNS record or a new Key Vault without running the entire configuration through plan and apply. Every change is a platform change, carrying platform risk and requiring review, no matter how minor. These look like separate problems, but all come from the same design choice: too many unrelated concerns tied into one Terraform boundary. Where Coupling Becomes Ownership It is easy to call these operational annoyances: slow plans, awkward upgrades, risky applies, the tax you pay for a big configuration. But the same coupling appears in review and approval, where it stops being just an operational problem. Once too many concerns share the same state, pipeline, and approval path, the question is no longer only "how long did the plan take?" It becomes "who is accountable for the boundary this change is crossing?" Take private connectivity. A single private endpoint on Azure isn't handled by just one team. The application team owns the service behind it. The platform team manages the landing zone, subnet, and endpoint placement. Private DNS zones might be managed centrally or by another team. Security or governance may require the service to be private. How these map to teams varies, but in a monolith, everything ends up in the same state, pipeline, and plan. So "who owns this?" rarely has a clear answer. However you split teams, they are coupled through a single configuration that none can truly own. When the application team changes its service, the same config still carries platform connectivity and governance controls. You cannot draw ownership along your real organizational boundaries, because the code does not have them. Both slow plans and unclear ownership trace back to the same issue: shared concerns treated as if they belong to just one team. Figure 1: When Terraform boundaries stop matching ownership boundaries. The monolith gives Terraform one boundary. Organizations have several. The pain comes when small changes have to cross boundaries that no team fully owns. Reach for the Coupling, Not the Size The reflex now is to split the state and move on. But splitting a landing zone poorly can be worse than leaving it alone. If you split along the wrong lines, you trade one blast radius for tangled cross-state dependencies. You also lose the single plan that at least showed the whole graph in one place. For example, splitting private endpoints into one state and private DNS zones into another may look clean on paper. But if different teams deploy them without a clear agreement, every new endpoint becomes a coordination headache, not a smaller change. Moving files into separate folders does nothing if the same pipeline, credentials, and approval path still govern everything. Decomposition should follow actual coupling, not just line count. So the next question is not "how many states should we create?" It is "which boundaries are real enough for teams to own, deploy, and recover independently?" If your Terraform monolith hurts, do not start by counting files or resources. Look at what is actually being coupled. Slow plans and unclear ownership are both signs that your Terraform boundaries no longer match your real ownership boundaries.

By Naveen Kalapala
Edge AI: Why Inference Is Moving Away From the Cloud
Edge AI: Why Inference Is Moving Away From the Cloud

Modern enterprise applications are increasingly running AI inference on-device rather than sending data to a central cloud. Improvements in hardware and model optimization have shifted the balance of compute. As one analysis notes, advances in 5G and edge hardware have made edge AI “a crucial technology for enabling intelligent applications.” Gartner predicts that by 2025 roughly 75% of enterprise data will originate at the edge rather than in traditional data centers. This data gravity, combined with emerging requirements for real-time response, privacy, and resilience, is driving inference tasks out of the cloud. Edge AI Reduces Latency, Bandwidth Costs, and Privacy Risks Running inference at the edge avoids the latency and network costs of cloud round-trips. For latency-critical use cases such as self-driving cars or augmented reality, even a few hundred milliseconds of delay is unacceptable. By processing sensor data locally, an edge device can make sub-10ms decisions for safety and interactivity. Similarly, streaming raw video or IoT sensor feeds to the cloud would incur massive bandwidth use and egress charges. Performing these analytics on-device eliminates that overhead. Local inference also enhances privacy and compliance as sensitive data (for example, images from a security camera or health readings from a wearable) can be analyzed on-premises without ever leaving the device. This helps meet data-sovereignty regulations and avoids exposing private information in transit. Finally, edge models continue to function even during network outages. When connectivity is lost, a device can still operate autonomously, maintaining “offline functionality” and zero downtime for critical tasks. Cloud Training and Edge Inference Enable Real-Time AI These factors have created a clear two-stage AI strategy for many organizations: train large models in the cloud (where vast compute and data are available) but run inference on the edge for real-time and private workloads. For example, one industry advisor observes that companies are already using clouds to develop and train models while then “optimizing, compressing, and deploying [them] to the edge for real-world application. This ensures sub-second decision-making, minimal data transfer, and continuous operation right where the value is delivered.” In practice, this can look like periodically syncing updated models from the cloud to a fleet of edge servers or devices, while daily operation happens locally. In fields like manufacturing, retail, or finance, this hybrid approach yields measurable ROI by applying AI where it matters most. AI Accelerators and Model Compression Make Edge Inference Practical Key technological advances have enabled this shift. On the hardware side, specialized AI accelerators have become commonplace in edge platforms. Mobile SoCs now include NPUs and DSPs for neural networks, while devices like NVIDIA Jetson or Google’s Coral Edge TPU provide GPU-like acceleration for embedded systems. A Qualcomm white paper notes that its chips combine CPUs, GPUs, and “neural processing units” specifically for edge AI, along with optimized frameworks and SDKs to deploy models on devices. These accelerators can execute inference algorithms much faster and more energy-efficiently than a general-purpose CPU. At the same time, neural network architectures have become more compact and efficient. Techniques such as model distillation, quantization, and pruning let developers shrink large models dramatically with little loss in accuracy. In practice, this means today’s “state-of-the-art smaller AI models” can outperform larger models from the cloud era while fitting on a phone or embedded board. For example, Qualcomm reports that many recent large generative models have been distilled down to versions under 100 billion parameters, yet still achieve performance comparable to much bigger models. Quantization (converting weights to 8-bit or 16-bit) and sparse pruning are now routine tools to reduce model size and latency. One survey explains that quantization “lowers power consumption and speeds up operations without significantly sacrificing accuracy, while pruning eliminates unnecessary parameters.” At the software level, lightweight inference frameworks and runtime libraries make deployment easier. TensorFlow Lite, PyTorch Mobile, ONNX Runtime, Intel OpenVINO, and similar toolkits offer optimized kernels for ARM processors, GPUs, and AI accelerators. These frameworks often include mobile-friendly model converters and delegate support for hardware acceleration. For example, TensorFlow Lite lets developers convert a trained TensorFlow model to a flatbuffer, then load it into a mobile app. On-device inference might look like: Java Interpreter interpreter = new Interpreter(modelBuffer); float[][] output = new float[1][NUM_CLASSES]; interpreter.run(inputData, output); This code snippet instantiates a TensorFlow Lite Interpreter with a pre-optimized model buffer and runs it on inputData, producing classification scores in output. Similarly, PyTorch Mobile can serialize a TorchScript model for Android/iOS, and ONNX Runtime can execute models across many hardware targets with reported speedups (Microsoft cites up to 17× faster inference). These mobile runtimes also leverage hardware delegates (GPU or NPU) under the hood when available. Edge AI Requires Careful Resource, Deployment, and Security Management Deploying and managing inference at the edge does introduce new engineering challenges. Resource constraints mean that models must be smaller and less complex than cloud counterparts, as even with quantization, a model that fits on a GPU server might need further compression for a microcontroller. Edge devices have limited memory and power budgets, so operators must balance accuracy against size and speed. The network of devices also requires orchestration, as software like Kubernetes (via lightweight distributions) or IoT platforms (AWS IoT Greengrass, Azure IoT Edge) can roll out updates and monitor health across fleets. For example, an edge deployment might containerize a TensorRT-based inference service and schedule it on Jetson nodes with GPU support, while sending telemetry to a central dashboard. Observability is critical as enterprises often integrate Prometheus/Grafana or cloud IoT logging to capture inference metrics and detect when models drift or hardware issues arise. Security must also be considered, as physical devices at the edge can be vulnerable, so measures like secure boot and authenticated OTA updates are important. Despite these complexities, many enterprises have already benefited. In retail, shops are using in-store edge cameras to detect incidents in real time without sending video to the cloud, saving bandwidth and complying with privacy rules. Industrial plants run anomaly-detection models on local PLCs to spot equipment faults instantly, ensuring operations can continue even if connectivity fails. Financial firms can do fraud checks on transaction terminals with millisecond latency. In all these cases, doing inference on-site is far cheaper and faster than piping every input to a data center. Conclusion In summary, ongoing trends in hardware, model design, and infrastructure are moving inference out of centralized clouds. Edge AI brings compute to the data, cutting latency and cost while meeting privacy requirements. That is not to say cloud AI is obsolete, as it remains essential for training, heavy analytics, and coordination, but the future of inference is local. By 2026, enterprises will likely adopt a hybrid model with cloud resources for development and big data tasks, with optimized models deployed to edge devices for live prediction. This shift requires new patterns of system design and monitoring, but it unlocks real-time intelligence and efficiency that cloud-only architectures can no longer match.

By Uthej Mopathi DZone Core CORE
Understand the Sidecar Pattern by Deploying n8n to AWS Fargate
Understand the Sidecar Pattern by Deploying n8n to AWS Fargate

A sidecar is a container that runs alongside another container as part of the same deployment unit. Just because two containers are in the same cluster or deployed around the same time doesn't make one a sidecar. There are two things that make a sidecar. First is that they share a network namespace, so they can reach each other over localhost rather than a network address. Second, they share a lifecycle. This means that they are created together, scaled together, and by default torn down together. Neither container has an existence independent of the other. The problem it solves is giving a specific concern its own boundary. For example, it can have its own filesystem, its own memory space, and often its own permissions or dependency set, without giving up the simplicity of deploying and operating one unit. You get isolation without paying for the operational overhead of running and coordinating a fully separate service. The test that defines the pattern across all of these is this: does it live and die with its partner container as one unit of deployment? If yes, it's a sidecar. If you have to reach it by hostname, through service discovery, or via a queue, it isn't one anymore. That is a separate service that happens to sit next to the first. That test matters because two adjacent patterns get called "sidecar" when they aren't: Decoupled worker/microservice. A separately deployed container, reached over the network, scaled on its own. A web application offloading work to Celery workers via Redis is a common instance of this: the app enqueues a job (send this signup email), a pool of workers pulls jobs off the queue independently, and neither side shares a network namespace or a lifecycle with the other. The workers scale on queue depth, not on how many web replicas are running, and a web app restart doesn't take queued or in-flight jobs down with it. n8n has its own version of the same shape: "queue mode," where a main node accepts webhooks and separate worker nodes pull jobs off a Redis queue. It's tempting to call either of these a sidecar relationship since the worker and the web app do feel paired, but neither qualifies: they don't share a deployment unit, and killing one doesn't touch the other.Ambassador/adapter. A container that proxies or translates traffic on its parent's behalf, like the Envoy example above, is actually this, more precisely. Structurally it's still a sidecar; it just gets a more specific name for what it does. Using n8n to Understand It What n8n Is n8n is a workflow automation platform like Zapier, but self-hostable and node-based rather than form-based. A handful of components make up a running instance: The editor/UI, where workflows are built visually as a graph of nodes.The main process, which serves that UI, listens for webhooks, and orchestrates workflow execution. The workflow execution decides what runs next, passing data between nodes and recording results.Nodes, the individual units of a workflow: trigger nodes (a webhook arrives, a schedule fires), action nodes (call an API, write to a database, send an email), and the Code node. The code node lets you drop in arbitrary JavaScript or Python to transform data however the built-in nodes can't. The code node is relevant in this article. The database, where workflow definitions, credentials, and execution history persist. In this article, Postgres is used. For most of what n8n does, the main process is the only thing doing work: routing a webhook, calling an API, writing a database row. The exception is the Code node, and that exception is the whole reason task runners exist. The Task Runner Feature and Its Use Case By default, a Code node's JavaScript or Python executes inside n8n's main process. This main process holds the database connection, the encryption key, and every credential stored in every workflow you've built. That's fine for trusted, well-understood scripts. It becomes a real problem the moment the code in that node is untrusted, third-party, or arbitrary enough that you can't fully audit it before it runs. By the way, that is how most Code nodes are used in practice. Task runners exist to solve exactly that use case: run Code node logic somewhere the main process's credentials and connections aren't reachable from it, without turning "write some JavaScript to reshape this JSON" into a separately deployed microservice every time. Going Deep on the Task Runner Feature n8n ships two modes for this: Internal mode (the default) runs Code nodes inline, in-process. No isolation. This is the fastest to set up, but the weakest boundary.External mode moves execution into a separate runner process entirely. That process connects back to the main n8n instance over a broker (an authenticated connection the main process listens on) and receives individual tasks to execute rather than having any standing access to n8n's internals. The runner never touches the database connection, the encryption key, or stored credentials directly; it only ever sees the specific input data for the task it's been handed. External mode goes further than just "a different process," too. The runner's own configuration (the n8n-task-runners.json file built in Phase 4) sets explicit allowlists — which environment variables the runner process can see at all, and which JavaScript built-ins or Python modules it's permitted to import, standard library and third-party tracked separately. So the boundary isn't just "different memory space," it's "different memory space, plus a declared, auditable list of exactly what this process is allowed to touch." That's a specific concern (arbitrary code execution) given its own boundary, without turning it into a fully independent service you have to deploy, discover, and monitor separately. It's the sidecar problem, stated exactly: external mode gives you the isolation; running the external runner as its own container in the same task definition is what makes that isolation a sidecar rather than just a separate process sharing a machine. Why This Needs to Scale Independently and Why "In the Same Container" Isn't Enough Most n8n deployment guides run n8n with task runners in internal mode, or with the external runner living inside the same container as the main process. For example, you will see guides about deploying n8n on a single EC2 instance, Render, DigitalOcean, or any platform's basic tier. That gets you the process isolation, which solves the security half of the problem. It doesn't solve the other half, which is that a runner sharing a container with the app can't be scaled, resourced, or restarted independently of it. That stops mattering the moment Code-node execution becomes the actual bottleneck rather than webhook handling or UI traffic. Imagine workflows doing heavy data transformation in Python, running numpy/pandas operations across large payloads, or executing many Code nodes concurrently. If the runner is bundled into the main container, giving it more CPU means giving the entire n8n instance more CPU, whether the UI and webhook layer need it or not. There's no way to say "the runner needs 2 more vCPUs, n8n itself is fine". Why AWS Fargate's Task Definition Is the Right Fit A Fargate task definition lets each container in the task carry its own CPU and memory reservation, its own health check, and its own essential flag governing what happens if it fails while still keeping every container in the task on one shared network interface. That's the sidecar promise made literal: isolation and independent resourcing for the runner, without losing the operational simplicity of one task, one deploy, one thing to scale as a unit when you do want to scale both together. The rest of this guide deploys exactly that: one Fargate task, two containers, wired together the way the definition above requires. Each infrastructure decision below gets tied back to a specific part of what's laid out here, so that by the end, the concept isn't something read once at the top, but it's something built. Prerequisites AWS account with billing enabledA domain you control, with DNS accessDocker installed locally, with docker buildx availableAWS CLI configured (aws configure) with permissions for ECR, ECS, RDS, ACM, and IAMThe runner image source (Dockerfile + n8n-task-runners.json) — built in Phase 4 Architecture Markdown User's Browser (HTTPS) | [Application Load Balancer] <- Certificate Manager (SSL Cert) | (Port 5678, HTTP internal) [ECS Fargate Task] |-- Container: n8n (main) <-- shared network namespace --> Container: n8n-runner (sidecar) | (Port 5432, PostgreSQL) [RDS PostgreSQL Database] The load balancer and RDS layers are ordinary AWS plumbing. The box in the middle is where the sidecar relationship actually lives. There is one task and two containers, each with its own resourcing. Phase 1: RDS PostgreSQL RDS Console → Create database → Standard create → Engine: PostgreSQLDB instance identifier: n8n-db. Master username: postgres. Generate and save a strong master password.Instance size: db.t4g.microStorage: 20 GB gp3, autoscaling on, max 100 GBConnectivity: the VPC you'll use throughout. Public access: No. New security group: n8n-db-sg, left empty for now.Additional configuration → Initial database name: n8n. Skip this and n8n fails on first connect with "database does not exist" — the DB instance identifier names the server, this field names the database inside it.Create, wait for "Available," copy the endpoint from Connectivity & security. Phase 2: ACM Certificate n8n requires HTTPS for webhooks to function Certificate Manager, in the same region you'll deploy the Load Balancer in → Request a public certificateDomain name: n8n.yourdomain.comValidation method: DNS validationCreate the CNAME record ACM provides at your registrar. If your registrar auto-appends your domain to the Host field, paste only the portion before your domain — the full string duplicates it and validation never completes.Wait for status: Issued Phase 3: Security Groups Two connections need rules: Security groupInbound rulePurposen8n-alb-sg443 from 0.0.0.0/0Public HTTPSn8n-ecs-sg5678 from n8n-alb-sgALB → n8n containern8n-db-sg (edit existing)5432 from n8n-ecs-sgn8n container → RDS Phase 4: Build and Push the Runner Image Dockerfile: Dockerfile FROM n8nio/runners:1.121.0 USER root RUN cd /opt/runners/task-runner-javascript && pnpm add moment uuid adm-zip RUN cd /opt/runners/task-runner-python && uv pip install numpy pandas pydantic requests boto3 certifi COPY n8n-task-runners.json /etc/n8n-task-runners.json ENV N8N_RUNNERS_CONFIG_FILE=/etc/n8n-task-runners.json USER runner It starts from n8n's own n8nio/runners base (containing the launcher and both runner processes), adds only the dependencies workflows actually need, and drops back to a non-root user once the root-only install steps finish. n8n-task-runners.json is where the isolation described above stops being architectural and becomes enforced: JSON { "task-runners": [ { "runner-type": "javascript", "health-check-server-port": "5681", "allowed-env": ["PATH", "GENERIC_TIMEZONE", "NODE_OPTIONS"], "env-overrides": { "NODE_FUNCTION_ALLOW_BUILTIN": "crypto,zlib", "NODE_FUNCTION_ALLOW_EXTERNAL": "moment,uuid,adm-zip" } }, { "runner-type": "python", "health-check-server-port": "5682", "env-overrides": { "N8N_RUNNERS_STDLIB_ALLOW": "json,zipfile,io,base64,datetime,re,math,random,statistics", "N8N_RUNNERS_EXTERNAL_ALLOW": "numpy,pandas,pydantic,requests,boto3,certifi" } } ] } allowed-env restricts which environment variables the runner process can see; N8N_RUNNERS_STDLIB_ALLOW / EXTERNAL_ALLOW restrict which Python modules it can import, stdlib and third-party separately. One container, two runner processes — the launcher inside n8nio/runners spawns both. Build and push: Shell docker buildx build -t n8nio/runners:custom . aws ecr create-repository --repository-name n8n-runners --region us-east-1 aws ecr get-login-password --region us-east-1 \ | docker login --username AWS --password-stdin <account-id>.dkr.ecr.us-east-1.amazonaws.com docker tag n8nio/runners:custom <account-id>.dkr.ecr.us-east-1.amazonaws.com/n8n-runners:custom docker push <account-id>.dkr.ecr.us-east-1.amazonaws.com/n8n-runners:custom --username AWS is a fixed literal, not your actual username — ECR auth always uses it. The password piped via --password-stdin is a short-lived token generated by the CLI, not your account password. Phase 5: The Task Definition This is where the two containers become an actual sidecar pair, and where the independent-resourcing argument from the introduction becomes a real field rather than a claim. JSON { "family": "n8n-task", "networkMode": "awsvpc", "requiresCompatibilities": ["FARGATE"], "cpu": "1024", "memory": "2048", "executionRoleArn": "arn:aws:iam::<account-id>:role/n8n-task-execution-role", "containerDefinitions": [ { "name": "n8n", "image": "n8nio/n8n:1.121.0", "essential": true, "entryPoint": ["sh", "-c"], "command": [ "mkdir -p /home/node/certs && wget https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem -O /home/node/certs/rds-ca.pem && /docker-entrypoint.sh" ], "portMappings": [{ "containerPort": 5678, "protocol": "tcp" }], "environment": [ { "name": "DB_TYPE", "value": "postgresdb" }, { "name": "DB_POSTGRESDB_HOST", "value": "<rds-endpoint>" }, { "name": "DB_POSTGRESDB_PORT", "value": "5432" }, { "name": "DB_POSTGRESDB_DATABASE", "value": "n8n" }, { "name": "DB_POSTGRESDB_USER", "value": "postgres" }, { "name": "DB_POSTGRESDB_SSL_CA", "value": "/home/node/certs/rds-ca.pem" }, { "name": "DB_POSTGRESDB_SSL_REJECT_UNAUTHORIZED", "value": "false" }, { "name": "WEBHOOK_URL", "value": "https://n8n.yourdomain.com/" }, { "name": "GENERIC_TIMEZONE", "value": "Africa/Lagos" }, { "name": "N8N_RUNNERS_ENABLED", "value": "true" }, { "name": "N8N_RUNNERS_MODE", "value": "external" }, { "name": "N8N_RUNNERS_BROKER_LISTEN_ADDRESS", "value": "0.0.0.0" }, { "name": "N8N_RUNNERS_BROKER_PORT", "value": "5679" } ], "secrets": [ { "name": "DB_POSTGRESDB_PASSWORD", "valueFrom": "arn:aws:secretsmanager:<region>:<account-id>:secret:n8n/db-password" }, { "name": "N8N_ENCRYPTION_KEY", "valueFrom": "arn:aws:secretsmanager:<region>:<account-id>:secret:n8n/encryption-key" }, { "name": "N8N_RUNNERS_AUTH_TOKEN", "valueFrom": "arn:aws:secretsmanager:<region>:<account-id>:secret:n8n/runners-auth-token" } ], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/n8n-task", "awslogs-region": "<region>", "awslogs-stream-prefix": "n8n" } } }, { "name": "n8n-runner", "image": "<account-id>.dkr.ecr.<region>.amazonaws.com/n8n-runners:custom", "cpu": 512, "memory": 1024, "essential": false, "dependsOn": [{ "containerName": "n8n", "condition": "START" }], "environment": [ { "name": "N8N_RUNNERS_TASK_BROKER_URI", "value": "http://localhost:5679" } ], "secrets": [ { "name": "N8N_RUNNERS_AUTH_TOKEN", "valueFrom": "arn:aws:secretsmanager:<region>:<account-id>:secret:n8n/runners-auth-token" } ], "healthCheck": { "command": ["CMD-SHELL", "curl -f http://localhost:5680/healthz || exit 1"], "interval": 30, "timeout": 5, "retries": 3, "startPeriod": 20 }, "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/n8n-task", "awslogs-region": "<region>", "awslogs-stream-prefix": "n8n-runner" } } } ] } Five fields here map directly back to the introduction: Per-container cpu/memory on n8n-runner. This is the independent-resourcing argument made literal. The runner gets its own 512 CPU units and 1024 MB, carved out of the task total, separate from whatever n8n is allotted. If Code-node execution turns out to be the bottleneck, this is the number you raise without touching the main container's allocation at all. That's the exact thing a same-container runner can't offer you. networkMode: awsvpc is the mechanical basis of "shared network namespace." Every container in the task gets one elastic network interface between them. This is the setting that makes Phase 3's missing security group rule make sense. There's one network surface, not two. N8N_RUNNERS_TASK_BROKER_URI: http://localhost:5679 only works because of the line above. The runner reaches n8n over localhost because they are the same task. If this pointed anywhere else, you would have built the decoupled-worker pattern from the introduction instead, no matter what you called the container. A shared N8N_RUNNERS_AUTH_TOKEN, pulled from Secrets Manager by both containers. Sharing a network namespace means the runner is reachable by anything else in the task. The isolation the whole pattern exists for still needs a trust boundary at the process level, not just the network level. A plaintext token here would defeat that, since task definitions are readable by anyone with ecs:DescribeTaskDefinition. essential: false on the runner. This governs how tightly the two containers' lifecycles are actually coupled. essential: true would mean a runner crash tears down the whole task, main container included. false means the runner can crash and recover independently: Code-node executions fail until it's back, but the UI and webhooks keep serving. The pattern doesn't mandate one answer; it just means this has to be a decision, not a default you inherited. The health check on port 5680 hits the launcher's own endpoint, separate from the per-runner-type ports (5681 JS, 5682 Python) set in Phase 4's config file. ECS is checking the supervisor, not each runner process individually. Register it: aws ecs register-task-definition --cli-input-json file://n8n-task-def.json Phase 6: Cluster, Service, and Load Balancer ECS → Create cluster → n8n-cluster → Infrastructure: AWS FargateCreate a service inside it: Task definition: n8n-task, latest revisionDesired tasks: 1Networking: your VPC, at least two subnets across AZs, security group n8n-ecs-sg, public IP onLoad balancing: Application Load Balancer, listener on 443 using the Phase 2 certificateTarget group: HTTP, port 5678, health check path /healthzCreate, wait for steady state. Notice the target group and health check only ever reference the n8n container. It did not mention n8n-runner at all. The n8n-runner container doesn't get a port that maps to the load balancer, doesn't get its own listener, doesn't get its own DNS entry. Everything that makes it reachable from outside the task goes through n8n . Phase 7: DNS At your registrar, add a CNAME: Host n8n, Value = your Load Balancer's DNS name. Confirm with nslookup n8n.yourdomain.com once it propagates. Verifying the Sidecar Relationship Visiting https://n8n.yourdomain.com and completing owner setup confirms the main container and database are working. To confirm the runner specifically: Create a workflow with a Code node (JavaScript or Python), and run it.Pull CloudWatch logs for both streams (/ecs/n8n-task, prefixes n8n and n8n-runner). The n8n-runner stream should show the launcher starting both runner processes and reporting a broker connection. The n8n stream should show the Code node's execution dispatched out rather than run inline. If the workflow completes but nothing appears in n8n-runner's logs, check N8N_RUNNERS_MODE=external on the main container first. That's the setting that actually hands execution off instead of running it in-process regardless of what else is configured.

By Iyanuoluwa Ajao
Architecting Production AI Across Clouds: Patterns That Decide System Survival
Architecting Production AI Across Clouds: Patterns That Decide System Survival

Most enterprise AI post-mortems do not blame the model. They blame the storage tier that starved the accelerators, the identity policy that over-granted access, the cost model that ignored egress, the forecast that leaked future data, or the region that failed and took a business process with it. The hard part of production AI was never intelligence. It was the engineering discipline around it. This article distills the architectural patterns that decide whether a cloud AI system is trustworthy at scale, spanning infrastructure, identity, cost, operations, the applied domains, low-code assembly, platform selection, and multi-cloud resilience. It is written for engineers who have to keep these systems running, not for a keynote. Infrastructure: The Interconnect Is the Bottleneck Distributed training is a systems problem before it is a machine learning problem. When a job spans many graphics processing units (GPUs), the fabric connecting them (e.g., NVLink within a node, InfiniBand, or a vendor fabric across nodes) frequently caps throughput more than raw compute does. Accelerators wired through an ordinary network idle while they wait to synchronize gradients. Storage is the symmetric constraint. If the file system cannot deliver data at the rate the accelerators consume it, utilization collapses. The pattern is a tiered design: Hot tier: parallel or block storage feeding active training at high input/output operations per second (IOPS).Warm tier: recent data staged for quick promotion.Durable lake: object storage providing petabyte-scale durability, partitioned and lifecycle-managed underneath. Two cost drivers hide from the pricing page: data egress (moving data across regions or out of a provider) and idle warm capacity. Optimizing only the advertised compute line item guarantees a surprise on the invoice. Identity Is the Perimeter In a service-to-service AI architecture, the network perimeter is gone; identity is the boundary. A zero-trust posture, where every request authenticates and receives least privilege, contains the blast radius when a component is compromised. Across providers, identity federation is the load-bearing pattern: a principal authenticates once and is recognized everywhere, so access is granted and revoked centrally instead of reconciled across three identity systems. Policy must travel with the workload; a rule enforced on one cloud and forgotten on another is not a policy. Model authorization is the emerging frontier. As models call tools and take actions, the question moves from who can query this model to what may this model do on a user's behalf. Least privilege applied to an autonomous agent is the boundary between useful and unbounded. Cost and Operations Are a Control Loop Cost management is not a spreadsheet; it is automation. Consistent resource tagging across every cloud is the prerequisite for attribution. On top sit budgets, alerts, and automated remediation that throttles runaway spend before it escalates. Site reliability engineering (SRE) supplies measurable targets. For AI workloads, the golden signals extend beyond latency and errors to accelerator utilization, queue depth, and prediction quality. A model can be fully available and quietly wrong, so define a service level objective (SLO) for output quality, not just uptime. Three techniques earn their complexity: Spot or preemptible capacity plus checkpointing cuts training cost sharply when jobs resume cleanly after reclamation.Predictive scaling anticipates load instead of reacting to it.LLM inference optimization becomes architectural: batch requests, cache frequent responses, route easy queries to smaller models, reserve the expensive model for queries that need it. The Applied Domains Share a Spine, Differ in Physics Vision is byte-heavy. High-resolution images and video streams make the data and network layers dominant. For real-time video, decouple frame capture from analysis and sample frames rather than processing every one. Critically, a business-rule layer, never the model alone, owns consequential decisions. Every extraction should carry a confidence score used as a routing gate: Python def route_extraction(field, threshold=0.90): if field["confidence"] >= threshold: return "auto_process" return "human_review" Language is byte-light but semantically treacherous, and because it replies directly to users, errors are visible. The defining risk of generative systems is hallucination. The strongest architectural defense is retrieval grounding, forcing answers from verified sources with citations: Python def answer(question, knowledge_base): passages = knowledge_base.search(question, top_k=3) context = "\n".join(p.text for p in passages) prompt = f"Answer using ONLY this context.\n{context}\n\nQ: {question}" return model.generate(prompt), [p.source for p in passages] Forecasting is defined by time order. You cannot shuffle a time series into random splits, and the most common failure is data leakage, using information unavailable at prediction time. Test on a fair, time-ordered holdout, and always emit a prediction interval; a point forecast that hides its uncertainty invites overconfident decisions. No-Code and Low-Code: Governed or Ungoverned No-code and low-code platforms collapse build cost from a scoped project to an afternoon, which is why adoption is exploding. The symmetric risk is sprawl: hundreds of ungoverned flows handling sensitive data, owned by no one. Govern with guardrails, not gates. Restrict which connectors and data sources are permitted, assign an owner and an SLO to every production flow, then let builders move freely inside the boundary. The goal is to make the safe path the easy path. Platform Selection Without Self-Deception Vendors all claim to be fastest, cheapest, and most reliable. Benchmark to replace claims with evidence: Latency: report percentiles (p95, p99), never averages that hide the slow tail.Quality: measure on your own representative data, not a public leaderboard.Cost: model total cost of ownership, including transfer, storage, idle capacity, operations, and migration, not the headline compute rate.Reliability: verify the platform meets your recovery time objective (RTO) and recovery point objective (RPO). Combine dimensions in a weighted scorecard whose weights are fixed before scores are seen. Adjusting weights afterward to crown a favorite converts analysis into rationalization. Multi-Cloud Resilience: Design for the Day a Cloud Fails For systems a business cannot lose, a single provider is a gamble. Multi-cloud resilience deliberately places critical workloads so no single provider failure takes the business down, applied only where the cost of failure exceeds the cost of prevention. Predict rather than react. Combine leading signals into a health score and fail over proactively: Python def health_score(latency_ms, error_rate, saturation): latency_factor = max(0, 1 - (latency_ms / 1000)) error_factor = max(0, 1 - (error_rate / 0.05)) saturation_factor = max(0, 1 - saturation) return round(0.4*latency_factor + 0.4*error_factor + 0.2*saturation_factor, 3) Kubernetes makes workloads portable; data replication (with the consistency-versus-availability trade-off decided per workload) keeps data ready on the other side; and a portable foundation of federated identity, uniform policy, and centralized monitoring makes failover routine rather than heroic. The discipline that separates real resilience from a slide deck is rehearsing failure on purpose. An untested failover path is a promise, not a capability. The Judgment Layer Across every layer, value came not from the most powerful component but from the judgment applied to it: matching effort to problem difficulty, keeping humans on consequential decisions, measuring before deciding, building governance in early, and designing for change. Tools will churn; foundation models will make today's designs look quaint. That is precisely why principles outlast product knowledge. The scarce resource in enterprise AI was never intelligence. It was judgment, and judgment does not ship from the cloud.

By VenkataSrinivas Kantamneni
Porting GPU Drivers to Rust on ARM64: The Hardest Trial for Kernel-Level Computing
Porting GPU Drivers to Rust on ARM64: The Hardest Trial for Kernel-Level Computing

Rust Has Entered the Kernel. Now Comes the Dangerous Part. A kernel driver does not fail politely. It does not throw a friendly exception, generate a neat stack trace, and ask whether you would like to restart. It corrupts memory, wedges hardware, leaks secrets, freezes the compositor, and leaves engineers spelunking through logs at 2 a.m. with the emotional range of a haunted printer. Nowhere is this more obvious than in GPU drivers. GPU drivers are some of the most complex pieces of kernel-level software in modern systems. They sit between userspace graphics APIs, memory managers, firmware, hardware queues, display engines, DMA buffers, synchronization fences, interrupts, power states, and error recovery paths. They are not “just drivers.” They are operating systems within the operating system. That is why Rust in GPU drivers matters. Rust support exists in the Linux kernel documentation today, but the kernel documentation is careful about its scope: Rust support is still primarily aimed at kernel developers and maintainers building abstractions, drivers, infrastructure, and tooling. It is not a blanket promise that every Rust kernel module is production-ready everywhere. That caution is healthy. Kernel engineering is allergic to magic, and rightly so. Rust does not sprinkle safety dust on MMIO registers. It does not fix firmware bugs. It does not turn a bad architecture into a good one. But Rust does attack one of the oldest and most expensive problems in systems software: memory unsafety. And if Rust can survive in GPU drivers on ARM64, it can survive almost anywhere. Why GPU Drivers Are the Real Test Most Rust-in-kernel discussions start too gently. They talk about simple drivers, toy modules, or safe wrappers around existing C APIs. Useful, yes. Convincing, not enough. The real test is graphics. A modern GPU driver must handle: - Device probing and initialization - Firmware loading and communication - Command submission queues - Shared memory between userspace, kernel, and device - DMA buffer ownership - Synchronization fences - Interrupt handling - Runtime power management - GPU reset and recovery - Userspace ABI compatibility - Performance under real graphical workloads This is where C’s sharp edges show up in full costume. A buffer may outlive the object that owns it, a command queue may still point to freed memory, or firmware may enter a state it should never reach. Reset and teardown paths add more chances for things to go wrong, especially if userspace still holds handles, an interrupt arrives at the wrong moment, or a fence is never signaled. These are not rare problems. They are normal driver engineering problems. The security motivation is also real. Google reported that memory-safety vulnerabilities accounted for 76% of Android vulnerabilities in 2019 and 24% in 2024 after shifting new development toward memory-safe languages. Google later reported major reductions in memory-safety vulnerability density for Rust code compared with Android’s C and C++ code, along with lower rollback rates and less code-review time for Rust changes. Do not overread that. Android is not the Linux DRM subsystem. A phone platform is not a GPU kernel driver. But the broader lesson is hard to ignore: when memory-safety bugs dominate the risk profile, changing the language of new code can change the shape of future vulnerability data. GPU drivers are exactly the kind of high-risk subsystem where that bet deserves serious attention. Why ARM64 Makes the Story More Important ARM64 is no longer just “the phone architecture.” It is in laptops, cloud servers, edge systems, automotive platforms, developer boards, AI devices, and embedded systems. On many ARM64 systems, the GPU is not a discrete PCIe card sitting at a safe distance. It is part of a tightly integrated SoC, sharing memory, power constraints, thermal limits, and firmware relationships with the rest of the system. That changes the stakes. A GPU driver bug on an ARM64 SoC can affect: - System memory safety - Display stability - Battery life - Thermal behavior - Compositor responsiveness - Application latency - AI and graphics workloads sharing the same memory fabric Rust support for AArch64 entered the Linux kernel development story as part of the broader Rust-for-Linux effort, and Linux kernel documentation now includes Rust materials for kernel developers working on Rust abstractions and drivers. That makes ARM64 GPU work more than a curiosity. It is a practical proving ground for the next decade of heterogeneous computing. The future machine is not CPU-only. It is CPU plus GPU plus NPU plus DSP plus video accelerator plus firmware-controlled subsystems. The kernel increasingly becomes an orchestration layer for compute fabrics. That means more shared memory, more queues, more firmware protocols, and more places for C lifetime bugs to hide like raccoons in ductwork. The Serious Case Study: Tyr for Arm Mali The strongest ARM64 GPU example today is Tyr, a Rust-based DRM driver for CSF-based Arm Mali GPUs. Tyr is especially interesting because it is not a random greenfield fantasy. It is a Rust port of Panthor, the C driver for the same class of hardware. Tyr is being developed as a joint effort involving Collabora, Arm, and Google engineers, and it aims to implement the same userspace API as Panthor for compatibility, so it can eventually be used as a drop-in replacement by PanVK, the Vulkan driver. That one design choice is the difference between serious engineering and conference glitter. Tyr is not trying to rewrite the entire graphics stack at once. It is trying to preserve the userspace contract while changing the kernel implementation language. That is exactly how infrastructure migration should be done. Change one major variable. Measure the result. Then decide. The target is also meaningful. CSF-based Arm Mali GPUs use a command-stream frontend where the driver must coordinate with firmware and hardware scheduling mechanisms. That naturally creates state-machine-heavy code, shared buffers, queues, and lifetime-sensitive resource handling. In other words: the kind of code where Rust’s ownership model is not academic. It is directly relevant. The Other Serious Case Study: Nova for NVIDIA GSP GPUs The second important example is Nova, a Rust-based driver for NVIDIA GPUs that use the GPU System Processor, or GSP. Nova is intended to become the successor to Nouveau for GSP-based NVIDIA GPUs in Linux and targets NVIDIA GPUs beginning with the GeForce RTX 20-series Turing family and newer. Nova is not primarily an ARM64 story, but it matters because it shows Rust entering serious DRM and GPU-driver territory, not just small demo modules. Together, Tyr and Nova point toward the same pattern: Rust is being explored where GPU drivers interact with firmware protocols, memory objects, queues, and kernel graphics APIs. This is the important architectural shift. As GPU firmware takes on more low-level responsibilities, host drivers often become protocol coordinators. They manage firmware boot, message queues, device objects, error states, recovery paths, memory handles, and userspace interfaces. That is a very Rust-shaped problem. Not because Rust is trendy. Trendy is how JavaScript frameworks reproduce. Rust is relevant because protocol state, resource ownership, and invalid transitions can often be modeled explicitly in the type system. The Core Engineering Idea: Make Illegal States Hard to Represent Here is the difference between a shallow Rust port and a serious one. A shallow port translates C into Rust line by line and celebrates because the file extension changed. A serious port rethinks dangerous state transitions. GPU drivers are full of implicit states: Buffer: Allocated -> Mapped -> Submitted -> Retired -> Freed Queue: Created -> Active -> Hung -> Recovering -> Destroyed Firmware: Absent -> Loaded -> Booting -> Running -> Failed Device: Probed -> Initialized -> Suspended -> Resuming -> Resetting -> Removed In C, these states are often spread across flags, pointers, locks, comments, and prayers. In Rust, they can be modeled more directly: enum BufferState { Allocated, Mapped, Submitted, Retired, } struct GpuBuffer<S> { handle: BufferHandle, size: usize, state: S, } struct Allocated; struct Mapped; struct Submitted; struct Retired; impl GpuBuffer<Allocated> { fn map(self) -> Result<GpuBuffer<Mapped>, DriverError> { // Map buffer into GPU-visible address space. Ok(GpuBuffer { handle: self.handle, size: self.size, state: Mapped, }) } } impl GpuBuffer<Mapped> { fn submit(self, queue: &mut CommandQueue) -> Result<GpuBuffer<Submitted>, DriverError> { queue.push(self.handle)?; Ok(GpuBuffer { handle: self.handle, size: self.size, state: Submitted, }) } } This is simplified, but the principle is powerful: make the dangerous lifecycle visible in the type system. In C, the rule might live in a comment: /* Do not free this buffer after submission until the fence signals. */ That comment is useful until someone edits a cleanup path six months later and accidentally turns it into historical fiction. Rust lets engineers encode more of that rule into APIs. It does not remove the need for review. It makes review more focused. Unsafe Rust Is Not a Loophole. It Is the Blast Radius. Kernel Rust still needs unsafe. Anyone claiming otherwise should be escorted away from the whiteboard. Drivers touch hardware. They read and write MMIO registers. They interact with C APIs. They manage DMA. They cross boundaries where the compiler cannot verify everything. The right goal is not “no unsafe code.” The right goal is small, explicit, audited unsafe code. Example: struct RegisterBlock { base: *mut u32, } impl RegisterBlock { unsafe fn read_raw(&self, offset: usize) -> u32 { core::ptr::read_volatile(self.base.add(offset)) } fn read_status(&self) -> DeviceStatus { let raw = unsafe { self.read_raw(STATUS_REGISTER_OFFSET) }; DeviceStatus::from_bits(raw) } } The outer driver should not scatter volatile pointer arithmetic everywhere. It should interact with typed operations such as: read_status() submit_queue() reset_engine() map_buffer() signal_fence() This is the real win: concentrate unsafety behind abstractions whose invariants can be documented, reviewed, and tested. Diffuse unsafety is archaeology. Concentrated unsafety is engineering. Research around Rust safety continues to focus on the fact that unsafe Rust and linked unsafe libraries can still compromise memory safety if not isolated or analyzed properly. That is directly relevant to kernel work. Rust is not a force field. It is a tool for shrinking the zone where humans must be perfect. Humans are bad at being perfect. That is why we invented compilers. What a Real Porting Plan Looks Like A credible GPU subsystem port should not start with “rewrite the driver.” That sentence is how you summon budget demons. A better plan looks like this: Phase 1: Choose a narrow subsystem Start where Rust gives clear leverage: - Buffer lifetime tracking - Command submission validation - Firmware message queues - Fence ownership - Reset and recovery state machines Do not begin with the entire DRM subsystem. That is not bravery. That is poor impulse control. Phase 2: Preserve the userspace API Tyr’s compatibility goal with Panthor’s userspace API is exactly the right instinct. If the userspace API remains stable, the migration can focus on kernel-internal safety and maintainability rather than forcing the whole graphics stack to change at once. Stable outside. Safer inside. That is the migration pattern. Phase 3: Wrap unsafe boundaries Every unsafe block should answer three questions: 1. What invariant must be true before this code runs? 2. Who guarantees that invariant? 3. How do we test that the invariant remains true? If the answer is “trust me,” the code is not ready. Trust is not a test strategy. Phase 4: Measure performance with real workloads A Rust GPU driver that is safer but introduces unacceptable frame-time spikes will not survive. Kernel developers care about safety, but they also care about latency, throughput, and not humiliating themselves in front of perf. A real benchmark plan should include: - Frame-time mean, p95, and p99 - Command submission latency - CPU cycles during graphics workloads - Context switches - Interrupt rate - GPU reset recovery time - Firmware boot time - Memory bandwidth - Power draw under sustained load - Thermal throttling behavior A basic harness might look like this: #!/usr/bin/env bash set -euo pipefail DRIVER="${1:?usage: ./bench.sh <driver-name>}" FRAMES="${2:-600}" OUT="results-${DRIVER}-$(date +%Y%m%d-%H%M%S)" mkdir -p "${OUT}" echo "Driver: ${DRIVER}" | tee "${OUT}/metadata.txt" uname -a | tee -a "${OUT}/metadata.txt" lscpu | tee "${OUT}/cpu.txt" sudo dmesg -C perf stat -d \ -o "${OUT}/perf.txt" \ -- ./gpu_workload_runner \ --driver "${DRIVER}" \ --frames "${FRAMES}" \ --json "${OUT}/frames.json" dmesg > "${OUT}/dmesg.log" echo "Benchmark complete: ${OUT}" That is not enough for a final paper-quality result, but it is the start of an honest engineering conversation. One run is a screenshot. Ten controlled runs are evidence. A bar chart with no methodology is decorative nonsense. The ARM64 Benchmark Matrix For an ARM64-focused evaluation, use a matrix like this: Hardware: - Rockchip RK3588 board, such as Rock 5B - Stable power supply - Active cooling - Fixed CPU governor - Fixed GPU governor, where available Software: - Same kernel baseline for C and Rust driver tests - Same Mesa version - Same compositor setting - Same Vulkan or OpenGL workload - Same thermal constraints Workloads: - Synthetic command submission stress test - Vulkan sample workload - Mesa or IGT graphics tests - Real application trace - Forced GPU reset and recovery test Metrics: - Mean frame time - p95 frame time - p99 frame time - CPU cycles - Context switches - Interrupts - GPU resets - Kernel warnings - Power draw This is where many articles fail. They show a chart but hide the setup. That is amateur hour. For DZone, the article can be powerful even without original benchmark results if it clearly presents the benchmark plan. But to become exceptional, it needs real numbers from a reproducible setup. No fake numbers. No “up to 5x faster” nonsense unless measured. The internet already has enough performance astrology. What Rust Will Not Fix Rust can reduce some classes of bugs, but it does not solve the hard parts of driver development by itself. It cannot compensate for poor hardware documentation, opaque firmware behavior, bad scheduling decisions, or flawed abstractions, and it does not make DRM any less complex. Low-level drivers will still require unsafe code, and deadlocks, design mistakes, and logic errors remain very much on the table. The Linux kernel documentation itself remains cautious: Rust support is still aimed at developers and maintainers working on abstractions, drivers, infrastructure, and tools, and it notes that Rust support is still under development, especially for certain configurations. That caution should be repeated, not buried. The correct argument is not: Rust makes kernel drivers safe. The correct argument is: Rust can reduce specific classes of memory and lifetime bugs in new kernel driver code, especially when unsafe hardware access is isolated behind reviewed abstractions. That is less flashy. It is also true. True wins. Why This Matters Beyond GPUs GPU drivers are a proxy for where kernel-level computing is headed. Modern systems are becoming accelerator orchestras. The CPU no longer owns the whole performance story. Work moves across GPUs, NPUs, DSPs, video encoders, SmartNICs, security processors, and firmware-managed islands. That means kernel software must manage: - Shared memory across devices - Complex queue lifetimes - Cross-device synchronization - Firmware protocols - Userspace-visible handles - Device reset semantics - Security boundaries around accelerators If Rust helps in GPU drivers, it can help in other accelerator drivers too. That includes: - AI accelerator drivers - Media encode and decode engines - Camera pipelines - SmartNIC offload paths - Storage acceleration - Embedded display controllers The real innovation is not “Rust replaces C.” That is a bumper sticker. The real innovation is selective memory-safe kernel development for high-risk hardware boundaries. That is a more mature thesis, and it is the one engineering leaders should care about. The Takeaway For application developers, this matters because kernel reliability eventually becomes application reliability. A browser tab, a video call, a game engine, a dashboard, a vision model, or an edge AI pipeline can all be ruined by a GPU stack that mishandles memory or fails recovery. For systems developers, this matters because Rust offers a practical way to encode ownership and state transitions that C leaves to discipline and code review. For engineering leaders, this matters because memory-safety work is not just a security initiative. It is a maintenance initiative. Safer new code can reduce future review burden, rollback risk, and vulnerability exposure, as Google’s Android reporting suggests. For kernel maintainers, this matters because the only acceptable Rust adoption path is incremental, measurable, and compatible with existing kernel development culture. The path is not revolution. It is disciplined infiltration. A Practical Checklist for Teams Before starting a Rust GPU driver effort, answer these questions: 1. What exact bug class are we trying to reduce? 2. Which subsystem has the worst lifetime complexity? 3. Can we preserve the userspace API? 4. Where must unsafe code exist? 5. Can unsafe code be isolated behind reviewed abstractions? 6. What real workload will prove performance? 7. What metric would make us abandon or redesign the port? 8. Who will maintain the Rust abstractions after the prototype hype fades? That last question is brutal and necessary. A prototype is easy. Maintenance is the boss fight. Conclusion: Rust Must Earn Its Place Where the Bugs Are Worst Rust in the Linux kernel should not be judged by toy modules. It should be judged where kernel bugs are expensive: drivers, firmware interfaces, DMA, shared memory, synchronization, and recovery paths. That is why GPU drivers on ARM64 are such an important proving ground. They combine modern hardware complexity with exactly the memory and lifecycle hazards Rust was designed to reduce. Tyr shows the most direct ARM Mali path: a Rust DRM driver for CSF-based Arm Mali GPUs, designed as a port of the C Panthor driver while aiming for userspace API compatibility. Nova shows that Rust GPU driver work is also moving into NVIDIA GSP territory, with ambitions to succeed Nouveau for supported modern NVIDIA GPUs. The lesson is not that Rust is perfect. The lesson is that the next generation of kernel-level computing is becoming too heterogeneous, too concurrent, and too security-sensitive to keep writing every new dangerous subsystem in C by default. Rust will not replace engineering discipline. It will punish the lack of it earlier. And in kernel development, earlier is everything. The Buffer That Came Back From the Dead The board had been running for nine hours. Same workload. Same scene. Same cursed GPU path that used to crash whenever the moon was wrong and the scheduler sneezed. In the old driver, the bug never arrived on command. It preferred drama. Sometimes frame 417. Sometimes frame 12,004. Sometimes only when the engineer walked away, because bugs respect neither science nor lunch. The Rust port failed too, at first. But it failed differently. Not with a corrupted pointer three layers below the crime scene. Not with a dead queue holding a ghost reference to a buffer that should have been buried. It failed at compile time, loudly, rudely, and before anyone had to read 900 lines of logs with the dead-eyed stare of a person reconsidering their career. The compiler pointed at the illegal transition. The engineer stared at it. The GPU kept rendering. For once, the monster had left footprints.

By Aakash Chaudhary

Monthly Top Cloud Architecture Experts

expert thumbnail

Raghava Dittakavi

Manager , Release Engineering & DevOps,
TraceLink

expert thumbnail

Srinivas Chippagiri

Sr. Member of Technical Staff

Srinivas Chippagiri is a highly skilled software engineering leader with over a decade of experience in cloud computing, distributed systems, virtualization, and AI/ML-applications across multiple industries, including telecommunications, healthcare, energy, and CRM software. He is currently involved in the development of core features for analytics products, at a Fortune 500 CRM company, where he collaborates with cross-functional teams to deliver innovative, scalable solutions. Srinivas has a proven track record of success, demonstrated by multiple awards recognizing his commitment to excellence and innovation. With a strong background in systems and cloud engineering at GE Healthcare, Siemens, and RackWare Inc, Srinivas also possesses expertise in designing and developing complex software systems in regulated environments. He holds an Master's degree from the University of Utah, where he was honored for his academic achievements and leadership contributions. The views expressed are his own and do not represent those of any employer, government agency or affiliated organization.
expert thumbnail

Vidyasagar (Sarath Chandra) Machupalli FBCS

Software Developer Operations Manager | Executive IT Architect,
IBM

Executive IT Architect, IBM Cloud | BCS Fellow, Distinguished Architect (The Open Group Certified)
expert thumbnail

Pruthvi Raj Seknametla

Site Reliability Engineer,
National Institute of Health (contractor)

The Latest Cloud Architecture Topics

article thumbnail
Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A
Multi-agent systems are common. Here we build a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol.
October 9, 2026
by Purnanga Borah
· 289 Views · 1 Like
article thumbnail
Supercharging AI Agents with Azure Context: A Hands-On Guide to Azure MCP
Here’s a step-by-step guide for cloud engineers on how to bridge local AI assistants with live Azure infrastructure using the Model Context Protocol.
October 8, 2026
by Ammar Ekbote
· 494 Views · 1 Like
article thumbnail
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
This guide walks through custom Azure ML model training and deployment to a Managed Online Endpoint and connecting it to a Microsoft Foundry agent as a function tool.
October 6, 2026
by Jubin Soni, FBCS DZone Core CORE
· 4,929 Views · 3 Likes
article thumbnail
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
In this article, we will discuss how to run your coding agents in the cloud using sbx. Cloud compute is usage-billed, so keep track of your sandboxes accordingly.
October 2, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 1,244 Views · 1 Like
article thumbnail
Your Cloud Diagram Is Already Out of Date: An Operating Model for Continuous Security Architecture
A five-stage operating model — Define, Prevent, Observe, Validate, and Improve — for keeping multi-account cloud environments aligned with architectural intent.
October 1, 2026
by Avik Mukherjee
· 1,249 Views · 1 Like
article thumbnail
The Silent Container Death: A TCP Dial That Never Times Out
A pod goes into CrashLoopBackOff. You pull the logs expecting a stack trace, a panic, an error string — and then nothing. No error. No exit message. Magic.
September 30, 2026
by Alexander Fo
· 1,617 Views · 1 Like
article thumbnail
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
Learn how to classify workloads, choose the right migration path, and avoid the traps that turn 6-week projects into 6-month ones.
September 30, 2026
by Jerzy Kopaczewski
· 952 Views · 2 Likes
article thumbnail
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A hands-on guide to Microsoft Foundry Document Intelligence SDK for extracting text, structured fields, and document data from PDFs for RAG and AI pipelines.
September 29, 2026
by Jubin Soni, FBCS DZone Core CORE
· 6,288 Views · 3 Likes
article thumbnail
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults
Golden paths standardize software delivery with self-service workflows, deployment guardrails, and observability while preserving team autonomy.
September 25, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 1,618 Views · 1 Like
article thumbnail
Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction
Cloud-native platforms need more than mature infrastructure. Learn how shared standards, platform engineering, and self-service can reduce delivery friction at scale.
September 24, 2026
by Igboanugo David Ugochukwu DZone Core CORE
· 1,499 Views · 2 Likes
article thumbnail
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes operations can drift as teams scale. Use this checklist to standardize clusters, releases, observability, access, reliability, and cost.
September 23, 2026
by Abhishek Gupta DZone Core CORE
· 2,559 Views · 1 Like
article thumbnail
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Terraform monoliths hurt when one state couples too many resources and owners. Split around ownership boundaries, not size.
September 22, 2026
by Naveen Kalapala
· 2,171 Views · 1 Like
article thumbnail
Edge AI: Why Inference Is Moving Away From the Cloud
Edge inference thrives on-device for real-time, private AI. Advances in hardware and compression cut latency and costs, pushing AI away from the cloud.
September 21, 2026
by Uthej Mopathi DZone Core CORE
· 2,172 Views · 3 Likes
article thumbnail
Understand the Sidecar Pattern by Deploying n8n to AWS Fargate
Learn how to deploy n8n Task Runners as AWS Fargate sidecars for isolated code execution, independent resources, and scalable workflow automation.
September 17, 2026
by Iyanuoluwa Ajao
· 2,948 Views · 2 Likes
article thumbnail
Architecting Production AI Across Clouds: Patterns That Decide System Survival
In production, enterprise AI rarely fails at the model. It fails in the architecture around it. Here are the cross-cutting patterns that work.
September 16, 2026
by VenkataSrinivas Kantamneni
· 2,964 Views · 1 Like
article thumbnail
Porting GPU Drivers to Rust on ARM64: The Hardest Trial for Kernel-Level Computing
Rust in the Linux kernel is gaining traction, but GPU drivers are the real test, where complex memory, synchronization, and recovery paths make C bugs costly.
September 14, 2026
by Aakash Chaudhary
· 1,810 Views · 1 Like
article thumbnail
Member Spotlight: Abhishek Sharma
Meet DZone community member Abhishek Sharma as he shares his tech journey, continuous learning, enterprise architecture insights, and life beyond work.
September 11, 2026
by Dominique Pugh
· 3,125 Views · 2 Likes
article thumbnail
Kubernetes Says Ready. Your LLM Still Isn’t.
Kubernetes can say Ready before an LLM can infer. Measure the gap, then make the readiness check a real inference in production.
September 9, 2026
by Shamsher Khan DZone Core CORE
· 3,190 Views · 2 Likes
article thumbnail
The Startup Time Trick Hiding Inside Your Docker Build
Spring Boot pods reload the same classes on every start. A CDS training run inside your Dockerfile caches that work once and cuts startup time roughly in half.
September 3, 2026
by Garima Agarwal
· 3,531 Views · 4 Likes
article thumbnail
Making Running Optional: Scaling AI Agents on Kubernetes With Agent Substrate
Learn how an early-stage open-source project separates workload lifecycle from compute allocation for bursty, stateful, and massively concurrent AI workloads.
September 3, 2026
by Mayowa Fajobi
· 2,776 Views · 1 Like
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×