In the SDLC, deployment is the final lever that must be pulled to make an application or system ready for use. Whether it's a bug fix or new release, the deployment phase is the culminating event to see how something works in production. This Zone covers resources on all developers’ deployment necessities, including configuration management, pull requests, version control, package managers, and more.
AI on Top of a Dysfunctional System
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
AI-generated software needs provenance that survives beyond the chat window. A code review can show what changed, but it rarely shows which model produced a fragment, what prompt and repository context influenced it, which agent or tool executed the change, which intent was being implemented, or whether the recorded history was altered later. Provenance fills that gap by treating generation as a supply-chain event rather than an ephemeral interaction. The core idea is already established in adjacent standards where W3C PROV models provenance through entities, activities, and agents, while SLSA records how software artifacts were produced so downstream consumers can verify expected processes and inputs. Generation Provenance as Engineering Metadata The first design rule is to separate authorship from provenance. Provenance answers where code came from and how it was produced, and it does not, by itself, determine legal ownership. The U.S. Copyright Office states that generative-AI output is copyrightable only where sufficient human-authored expressive elements exist, and that prompting alone is not enough. Employment agreements, contributor agreements, licenses, and jurisdictional law still govern ownership questions. Provenance instead supplies evidence for attribution, review, audit, and accountability. A minimal record should bind the generated artifact to the model provider, model identifier and revision, agent identity and version, prompt digest, context digests, execution trace, repository commit, intent digest, timestamp, and approving human or service identity. Hosted model aliases can change over time, so a provider-returned model identifier or immutable deployment revision is preferable to a friendly model name alone. The record should also contain cryptographic digests for generated files so later edits cannot silently inherit stale provenance. JSON { "artifact": "src/billing/CancelService.java", "sha256": "7e91...c42a", "commit": "9f3c1ad", "model": {"provider": "acme-ai", "id": "code-model", "revision": "2026-08-14"}, "agent": {"id": "repo-agent", "version": "3.7.2"}, "prompt": "sha256:18ab...90ef", "context": ["git:9f3c1ad^", "sha256:44c2...bb10"], "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1...", "approvedBy": "team:payments-reviewers" } Git commit trailers provide a low-friction place to attach pointers because Git supports structured token-value trailers at the end of commit messages. The commit should store references and digests rather than sensitive prompts themselves. Plain Text AI-Provenance: sha256:5df0...a992 AI-Model: acme-ai/code-model@2026-08-14 AI-Agent: [email protected] Prompt-Digest: sha256:18ab...90ef Context-Digest: sha256:44c2...bb10 AI-Trace: urn:uuid:2cf1... Intent-Digest: sha256:a771...0d61 File-level attribution can use a compact pointer rather than duplicating the full record. A generated region can carry a comment such as // ai-provenance: urn:gen:2cf1..., while the referenced sidecar record maps that generation to file hashes and, when needed, line ranges. This keeps source readable and prevents model metadata from becoming scattered, inconsistent comments. From SBOM to Generation BOM AI-generated code needs an additional description of the generation process. CycloneDX already supports source code, machine-learning models, component provenance, formulation describing how objects were created, and citations that attribute supplied information to entities or processes. Its ML-BOM capability records models, datasets, configurations, and provenance, while the earlier model-card framework established the broader practice of documenting model identity, intended use, evaluation, and limitations. A generation BOM can therefore be implemented as a small signed sidecar linked to the repository and release artifact, rather than inventing a second source-control system. JSON { "bomFormat": "GenerationBOM", "specVersion": "0.1", "subject": {"path": "src/billing/CancelService.java", "sha256": "7e91...c42a"}, "generator": {"agent": "[email protected]", "model": "acme-ai/code-model@2026-08-14"}, "inputs": {"prompt": "sha256:18ab...90ef", "context": ["sha256:44c2...bb10"]}, "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1..." } The intent digest is especially important. A prompt records instructions presented to a model, but an intent contract records the behavior that must remain true after generation. Such a contract can contain permitted change scope, protected behaviors, security constraints, and acceptance criteria. Provenance then connects the produced code not merely to an AI request, but to a reviewable engineering objective. This mirrors data-lineage systems such as OpenLineage, which associate runs, jobs, datasets, and extensible facets so downstream analysis can reconstruct how an output was produced. Artifact hashes alone cannot establish reproducibility when generation depends on mutable infrastructure. Provenance should therefore bind execution parameters such as model configuration, decoding settings, tool versions, retrieval indexes, and policy revisions. Capturing these values converts provenance from a historical label into a verifiable reconstruction boundary for later audits and incident analysis. Tamper Evidence and CI Enforcement Metadata becomes trustworthy only when alteration is detectable. SLSA explicitly treats provenance authenticity and digital-signature verification as mechanisms for detecting tampering, and recommends approaches that improve compromise detection, including transparency logs. Sigstore provides signing with short-lived identity-bound certificates and records signing events in Rekor, an append-only transparency log. A provenance file can be signed as a blob during CI: Shell cosign sign-blob \ --bundle generation-provenance.sigstore.json \ generation-provenance.json Verification should occur before merge or release, not after an incident. SLSA similarly emphasizes that provenance has little value unless a consumer verifies it against expected properties. Shell provctl verify \ --commit "$GIT_COMMIT" \ --require-model \ --require-context-digest \ --require-intent \ --require-signature \ --max-unattributed-lines 0 A practical gate should reject AI-marked changes when the artifact digest no longer matches, required model or agent fields are absent, the provenance signature fails, the intent contract is missing, or the trace cannot be resolved. Human-edited code should not be forced into artificial AI attribution; instead, the policy should distinguish generated, transformed, and manually authored regions. Execution traces can preserve tool calls, retrievals, test runs, and agent steps and are designed to make software-supply-chain steps transparent by recording what happened, by whom, and in what order, and their runtime-trace predicate can describe system events associated with a supply-chain step. Runtime verification closes another gap. Provenance can prove which generation path produced a deployment, but not that the resulting behavior remains correct under production conditions. Release telemetry should therefore link runtime incidents back to commit, provenance record, model revision, and intent contract. That correlation turns an AI-related defect from an unstructured forensic exercise into a query over lineage. Accountability Without Capturing Everything Capturing every prompt verbatim is usually the wrong default. Prompts and retrieved context may contain credentials, personal data, proprietary code, customer information, or licensed material. A safer design stores encrypted source material in an access-controlled evidence store and places digests, object references, retention class, and classification labels in Git-visible provenance. High-sensitivity environments can retain only keyed digests and approved summaries where reproduction is less important than proof of correspondence. Storage and performance costs also require boundaries. Full agent traces can be large, while line-level metadata can become noisy after refactoring. The durable unit should normally be a generation event bound to artifact digests and commits, with finer-grained ranges reserved for high-risk code. Developer ergonomics matter equally, as provenance capture should be automatic in IDE agents, repository bots, and CI runners rather than dependent on manual form filling. Regulation strengthens the case for disciplined records without creating a universal rule that every AI-generated source line must carry a label. The EU AI Act requires general-purpose AI model providers to maintain technical documentation and copyright-compliance policies, while NIST SP 800-218A extends secure development practices specifically for generative AI across the software lifecycle. These frameworks reinforce documentation, traceability, and governance, but a code-provenance system should be treated as engineering evidence rather than a substitute for legal analysis. AI-generated code should enter a repository with the same expectations applied to any other supply-chain artifact: origin, inputs, process, identity, integrity, and approval must be recoverable later. The strongest implementation is not a comment saying that AI was used, but a signed provenance chain linking model and agent identity, prompt and context digests, execution trace, intent contract, commit, generated artifact, review decision, and runtime evidence. Teams adopting AI-assisted development should make that chain automatic, verify it in CI, protect sensitive evidence separately, and fail closed for unattributed high-risk changes. That converts provenance from documentation into an enforceable engineering control and makes accountability possible long after the generation session has disappeared.
Agent systems are increasingly decomposed into specialized services: a planning agent delegates research, a research agent invokes retrieval and synthesis, and a compliance agent validates the result before an action is allowed. Open standards such as A2A formalize communication between independent agents, but transport interoperability is only part of the production problem. Requests can be lost after acceptance, retries can duplicate expensive work, callers can disappear while remote tasks continue, and multi-agent chains can become difficult to reconstruct. Temporal Nexus addresses a different layer. It connects Temporal applications through durable service contracts, allowing an agent capability to behave less like a fragile HTTP handoff and more like a reliable distributed operation. When an Agent Call Becomes a Distributed Transaction Boundary A conventional agent handoff often looks like ordinary RPC: serialize a task, send it to another service, wait for a response, and retry failures. That model works for short, stateless interactions. It becomes fragile when delegated work lasts minutes or hours, crosses ownership boundaries, depends on databases or model providers, or triggers side effects. Retry behavior, deduplication, cancellation, timeout ownership, and result delivery then become part of the application protocol. Nexus moves those concerns into Temporal’s execution model. A Nexus Endpoint routes a request to a target Namespace and Task Queue while hiding those implementation details from the caller. The caller depends on a named service contract rather than the handler’s Workflow type, queue, or deployment topology. Temporal describes the relationship as peer-to-peer: caller and handler Workflows remain independent executions while Nexus provides the durable boundary between them. This distinction matters for agent platforms because a durable agent service should expose a business capability, not internal orchestration mechanics. A “research” operation can remain stable even if its implementation changes from one Workflow to a graph of Activities, model calls, human approval, or additional Nexus calls. Temporal supports multi-level Nexus composition, with each hop represented as a separate durable operation. A Nexus Contract Fits the Agent Capability Boundary In the Java SDK, a Nexus Service defines the callable contract and Nexus Operations define individual capabilities. A compact contract for a research agent can remain deliberately narrow: Java @Service interface ResearchAgentService { @Operation AgentResult research(AgentTask task); } The contract carries only the input and output required by the collaboration boundary. Temporal’s Java guidance recommends service APIs containing operation names plus serializable input and output types; JSON and Protobuf are practical choices when multiple SDK languages are involved. The caller therefore does not need access to the handler’s prompts, memory representation, tool graph, or Workflow implementation. Long-running agent work should normally use an asynchronous Nexus Operation backed by a Workflow. Temporal reserves synchronous operations for reliable, predictably low-latency paths that finish within the ten-second handler deadline. Asynchronous operations can represent long-running work and return an operation token for tracking completion without holding a conventional request open. In Temporal Cloud, the maximum Nexus Schedule-to-Close timeout is 60 days. A handler can expose an agent Workflow without introducing an HTTP controller or polling API: Java @OperationImpl public OperationHandler<AgentTask, AgentResult> research() { return WorkflowRunOperation.fromWorkflowMethod( (ctx, details, task) -> Nexus.getOperationContext() .getWorkflowClient() .newWorkflowStub( ResearchAgentWorkflow.class, WorkflowOptions.newBuilder() .setWorkflowId("research-" + details.getRequestId()) .build()) ::run); } WorkflowRunOperation.fromWorkflowMethod maps the Nexus Operation to a Workflow execution. Temporal documents the Nexus request ID as stable across retries, making it suitable for constructing a deduplication-oriented Workflow ID when no stronger business identifier exists. A business identifier is generally preferable because it aligns duplicate suppression and operational correlation with domain semantics. Durability Changes the Meaning of Retry and Cancellation The most important Nexus behavior is execution semantics. Once a caller Workflow schedules an operation, Temporal atomically hands the command to Nexus machinery. Delivery uses at-least-once execution, automatic retry, rate limiting, concurrency limiting, load balancing, and circuit breaking. A handler can therefore be invoked more than once for the same operation, which makes idempotency essential for external side effects. Temporal also documents stronger duplicate protection by backing the operation with a Workflow whose ID reuse policy rejects duplicates. That behavior is especially relevant to agents because model inference may be nondeterministic and repeated tool calls can duplicate actions such as ticket creation, notifications, or database mutations. Durable execution does not remove the need for idempotency keys at external boundaries. Instead, it provides a stable execution context in which prior decisions and completed steps can be recorded and replayed consistently. Cancellation also becomes a protocol-level feature rather than a best-effort HTTP convention. Canceling a caller Workflow propagates cancellation to pending Nexus Operations and their backing handler Workflows. Termination is different: terminating the caller abandons pending operations and does not send cancellation to the handler Namespace, so remote work can continue. Graceful cancellation is therefore the safer control path for agent chains that may need cleanup or compensation. Timeouts express separate service-level expectations. Schedule-to-Start bounds how long an operation may wait to begin, Start-to-Close bounds asynchronous execution after start, and Schedule-to-Close bounds the complete lifecycle. Nexus retries retryable failures within those limits, while non-retryable failures resolve the operation and become visible to the caller. The Caller Stays Simple While the Runtime Carries State From a caller Workflow, the remote agent looks like a typed service stub. Retry loops, callback plumbing, and result polling do not need to appear in business code: Java private final ResearchAgentService research = Workflow.newNexusServiceStub( ResearchAgentService.class, NexusServiceOptions.newBuilder() .setOperationOptions( NexusOperationOptions.newBuilder() .setScheduleToCloseTimeout(Duration.ofHours(2)) .build()) .build()); public AgentResult delegate(AgentTask task) { return research.research(task); } The endpoint mapping is configured when the caller Workflow implementation is registered, allowing deployment configuration to bind a service name to a Nexus Endpoint without changing business logic. Temporal recommends a collocated pattern by default, where Nexus handlers live beside the Workflows they expose, and a router-queue pattern when routing requires separate scaling, permissions, or deployment ownership. Operationally, Nexus creates a stronger debugging surface than an opaque chain of HTTP requests. Caller histories record Nexus scheduling, start, completion, failure, timeout, and cancellation events. Bidirectional links connect caller events to handler Workflow histories, pending operations expose retry state, and OpenTelemetry integration can visualize call graphs across Nexus Operations, Activities, and Child Workflows. Security follows the service boundary. In Temporal Cloud, Nexus Endpoints use caller-Namespace allowlists, Workers authenticate with mTLS or API keys, and cross-Namespace Nexus traffic is secured by the platform. Nexus payloads use the same Data Converter model as Workflows and Activities, including codec-based encryption. Cloud Nexus Endpoints are not general public HTTP endpoints; they are reached through Temporal SDK execution inside the account. Durable Agent Services, Not Just Durable Requests Temporal Nexus does not replace agent interoperability protocols such as A2A. A2A standardizes how independent agentic applications discover capabilities, delegate tasks, and exchange results, while Nexus connects Temporal applications through durable operations. The boundary is complementary: an external A2A-facing layer can provide ecosystem interoperability, while Temporal Workflows and Nexus provide durable execution for agent services running inside Temporal. The deeper shift is from treating agent delegation as message delivery to treating it as durable work. HTTP can move a request across a network, but production agent collaboration also needs remembered state, bounded retries, deduplication, cancellation propagation, durable completion, access control, and traceable execution across service boundaries. Nexus puts those semantics directly into the invocation model. For long-running or side-effecting agent systems, that changes the handoff from an unreliable gap between services into a first-class execution boundary that can survive process crashes, worker outages, and delayed completion without losing the thread of the operation.
Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq
A few months ago, one of our teams celebrated a milestone in their quarterly review. AI adoption was up. The productivity dashboard showed that developers were generating, on average, 40% more code per sprint. The tech lead showed this as a major win. Three weeks later, I was on a call for a production issue. The challenge was that errors were coming in three different formats depending on which endpoint you hit. The alerting was blind to a category of failure it had always caught before. We had a centralized exception handler. It logged context, mapped it to the right HTTP status, and pushed these alerts to our observability stack. When we investigated, we found that recent AI-assisted PRs had started introducing their own try-catch blocks inline. Each one caught exceptions locally, logged in a slightly different format, and returned a slightly different error shape. Some swallowed the exception instead of letting it propagate up to the handler that would have alerted us. Each one of those PRs was correct. Every one passed review, including the reviews I did myself. We were all checking for correctness, and consistency isn't the kind of thing that shows up in a git diff. Cleaning it up took most of a sprint. And that productivity dashboard? It counted the original generation and the cleanup as output. Twice the code, twice the "productivity," for a net loss of engineering time. The dashboard was going up while the system got worse underneath it. That's the mistake I keep seeing where teams confuse code generation with engineering progress. The LOC Trap, Reloaded Fred Brooks called this out in The Mythical Man-Month decades ago. He explicitly mentioned that measuring programming productivity by lines of code is nonsensical. Everyone agreed and then somehow forgot. So why did we rebuild the exact same dashboard the moment AI arrived? We are using the same flawed metric, with AI branding, and presenting it to boards. When a team lead reports that AI tools helped produce 40% more code, the follow-up I want to hear is “Did we actually need 40% more code?” Usually the answer is no. What we needed was the same outcomes with less effort, and effort in software lives overwhelmingly outside the act of typing. I’ll come back to that. There is a difference this time, and it is worth naming. Teams aren’t defending line counts out loud anymore. The dashboards have moved on to merged pull requests, agent tasks completed, and suggestions accepted. It is the same instinct in a unit that sounds more respectable in front of a board, counting the artifacts of work and reporting the count as progress. Why Senior Engineers Delete Code Here’s a pattern you’ll recognize if you have led engineering teams for any length of time. Your best engineers often produce fewer lines of code than anyone else on the team. Some of their most impactful weeks come out to negative line counts. That’s expertise rather than laziness. I haven’t fully worked out why the instinct for deletion over addition takes years to develop, but it does. GitClear has been measuring this rather than speculating about it. Their 2026 analysis covers 623 million changes from 2023 to 2026, and the figure that stopped me had nothing to do with volume. Moved code, their proxy for refactoring, dropped from 21% of all changes in 2022 to 3.8% by the middle of this year. Duplicated blocks are up 81% across the same window. Cross-file function calls, which is roughly what reuse looks like in a diff, are down 35%. Those three together describe a codebase that has stopped being rearranged when work goes in, and very little gets moved or deleted. When someone takes 2,000 lines of tangled logic and replaces it with 200 clean ones, that looks like a loss on any volume metric. It is an enormous win for the system, and it is exactly the activity that has gone quiet. One correction I owe, since I quoted the earlier version of this research at people for the better part of a year. GitClear’s 2024 report predicted two-week code churn would double in the AI era. It didn’t double. It went up 15%. The headline projection was too aggressive, and the part almost nobody quoted — refactoring falling off a cliff — turned out worse than predicted. Senior engineers get this intuitively. Every line of code is a liability, because every line has to be read, understood, tested, and maintained. So the best solution often makes code disappear, and that takes different forms like a well-chosen abstraction that kills duplication or a config change that removes a custom implementation. Sometimes it’s just a conversation with the PO that drops the requirement entirely. Now think about what AI coding metrics would say about this. An engineer spends a day understanding a system, realizes three services can collapse into one, and deletes 4,000 lines. By every AI productivity metric in use today, that engineer had a terrible day. In reality, they may have saved the organization months of future pain. Gergely Orosz tells a revealing story about what happens when you optimize for the wrong signal. When Uber introduced diff-count metrics, engineers started creating more, smaller changes to look productive. They flooded CI systems, driving up costs. The metric improved and engineering got worse. We are setting ourselves up for the same trap with AI-generated LOC. AI’s Tendency Toward Verbose Implementations This gets worse when you look at what AI coding tools actually excel at, which is producing plausible code quickly. That skill carries a built-in bias toward verbosity. Sometimes more code is genuinely the right call. Explicit beats implicit. A verbose but readable implementation can be better than a clever one-liner that nobody understands at 3 am when production is on fire. I’m not arguing for code golf. But AI-generated verbosity is a specific kind of bad, because it is default verbosity from ignorance of context rather than chosen verbosity for clarity. This distinction matters a lot more than I initially thought. Ask an AI assistant to implement a feature, and you’ll get a complete, working solution, longer than what an experienced developer would write, because it optimizes for correctness and completeness in isolation. It may not pick up the utility you wrote last month to do exactly this. It doesn’t realize the framework provides a one-liner if you structure the problem slightly differently. It can’t tell the difference between “I should be explicit here for readability” and “I am reinventing something that already exists three directories over.” The handler drift I opened with is the cleanest example I have of it. The AI did exactly what it was asked, every single time. Each PR added code, each one passed review on its own terms, and nothing was wrong inside any of them. What we lost lived across them. One handler gave us one error shape, and one error shape gave our alerting something to fire on. No single diff broke that, and all of them together did. GitClear tracks error-masking constructs, which is the failure mode buried in there, and they are up 47% since 2023. Our inline handlers were exactly that. I’d like to think we were an unlucky outlier, and the data says we were ordinary. I’m still not sure how you review for a property that isn’t visible in the file in front of you. Complexity as the Hidden Cost Code volume isn’t a perfect proxy for system complexity, and I acknowledged that above. But it is a directional one, and in aggregate it holds. When your codebase grows by 30-40% in a quarter without a corresponding growth in functionality, complexity is almost certainly growing with it. And system complexity is the single biggest thing determining how fast your team can move over time. I haven’t found a way around that in eighteen years. Every line of code carries ongoing costs that nobody puts on a dashboard: Cognitive load for anyone working nearbyTest coverage, without which it becomes a ticking riskReview time on every future change to itDependencies it drags in that need constant updatingMigration effort during every platform change When AI tools grow your code volume by 30-40%, each of these costs grows. The productivity gain at the moment of writing is real, and I’m not denying that. But it can be entirely eaten up by the downstream cost of maintaining a bigger, more complex system. Then there is the study I keep coming back to, and the update to it that I nearly missed. In mid-2025, METR found that experienced developers using AI coding tools took 19% longer on real-world tasks. The setup was specific. It had 16 open-source contributors working in their own repositories, code they’d lived in for years, 246 real issues averaging about two hours each, with Cursor Pro and Claude. In February 2026, METR published an update that takes a fair amount of it back. They are redesigning the experiment, and the reasons aren’t flattering to the original. Developers who would no longer work without AI declined to take part at all. Somewhere between 30% and 50% of participants avoided submitting exactly the tasks where they expected to want AI. The pay rate for the follow-up work dropped from $150 an hour to $50, which made recruitment worse again. Their own summary is that the newer data amounts to very weak evidence in either direction, and that developers are probably more sped up now, in early 2026, than the early-2025 estimate suggested. So the 19% was never a fact about AI-assisted development. It was a measurement of sixteen people in one setting, and the people who ran it now think it read low. What survives is the part that makes me think harder. Those developers believed they were 20% faster. Whatever the true effect was, it wasn’t the effect they perceived, and not one of them could feel the gap while it was happening. Reviewing and integrating suggestions consumed time none of them accounted for. Selection bias moved the headline number around, but it does not explain away a room full of experienced engineers being wrong about their own week. I could never square that finding with my own experience, because I do feel faster on certain tasks. The update moves me off the fence, slightly. Maybe I wasn’t fooling myself. I would still put no weight at all on my own estimate of how much faster I am, and that’s close to the only thing here I am confident about. None of that is an argument against the tools, only against measuring them by how much code they produce. The Strongest Number Against Me If you want to argue the other side, the best evidence available today is Microsoft’s. Early in 2026, they rolled Claude Code and GitHub Copilot CLI out across the organization and studied what happened. Tens of thousands of engineers, four months, and the ones who adopted merged roughly 24% more pull requests than the counterfactual said they would have. That is not a lab, and it is not a small group of volunteers. It’s the largest measurement of agentic coding tools anyone has published, the effect is large, and it points the right way. I take it seriously. I also notice what the unit is. The authors get there ahead of any critic. Their paper says a merged PR isn’t the same as the value it delivers, which is the argument of this entire post, conceded inside the study that is supposed to answer it. Twenty-four percent more merged pull requests is consistent with 24% more delivered value. It is equally consistent with the same work arriving in smaller slices, which is what happened at Uber the moment diff count landed on a dashboard. There are narrower caveats, and I won’t pretend I have chased all of them. The comparison is against engineers who already had AI in their IDE, so what it measures is the increment from adding an agent rather than the effect of AI from zero. Four months is not long enough for maintenance cost to turn up. And engineers chose for themselves whether to adopt. None of that makes the study wrong. It is a good measurement of pull request volume, and pull request volume behaves the way lines of code always did. It is easy to report and showcase, if somebody decides that moving it matters. Blind to what the system underneath is doing. What This Actually Means If You Are Leading a Team If you’re six months into your AI investment and your main evidence of ROI is that the output counter went up, whether that counter says lines or pull requests or tasks completed, that should worry you more than it reassures you. You might be measuring the accumulation of future cost and calling it present-day value. Here is what I would look at instead. Cycle time as value reaching production sooner, not just code getting written faster? Those two come apart more often than anyone expects. Rework rate counts if you are fixing more bugs in AI-assisted code? If generated code carries a higher defect rate, your productivity gain is a mirage. Cognitive complexity trend is when your application is getting harder to understand. Tools like SonarQube measure this. If complexity is climbing faster than it did before AI adoption, you’ve got a compounding problem. Developer effort distribution is where the time is actually going. If writing dropped from 20% to 10% of developer time while code review grew from 15% to 30%, you have moved the burden rather than reduced it. None of that is exotic, and it is not just my read. DORA’s 2026 work on the ROI of AI-assisted development arrives at something similar from a different direction. The return on these tools tracks the strength of the engineering system around them rather than the tools themselves. The things that decide it are unglamorous. Whether code review has any slack left in it. Whether people trust the test suite enough to act on a red build, which is the one I have never seen anybody audit. How much sits between a merge and production. Where those are weak, faster generation fills a queue and waits there, and the measurement science is still catching up to the tooling. The Uncomfortable Question Here is what I would ask any engineering leader who reports AI productivity gains based on code output: “If your best engineer spent last week deleting 3,000 lines of AI-generated code and replacing them with 300 lines that do the same thing better, would your dashboard show that as a win or a loss?” On a line count, that week is a catastrophe. On a PR count, it is one merged pull request, which is what a typo fix is worth too. Neither number has any way of seeing what actually happened. If your dashboard shows a loss, you are measuring the wrong thing, and you are rewarding your team for building a larger, slower, more fragile system in exchange for a chart that goes up and to the right. The goal was never more code. It was better systems that deliver business value and are maintained cheaply. Next in this series: Optimizing Benchmark Tasks Instead of Real Delivery Work — why the fact that coding is not the bottleneck makes most AI productivity claims irrelevant to actual delivery speed.
Mobile networks are notoriously unreliable. A common scenario is when a user taps "Pay Now" in an app: the payment request reaches the server and is processed, but the network response never reaches the phone. The client assumes the request failed and retries, leading to the charge running twice. This is precisely the kind of bug that idempotency solves. Idempotency means that repeating the same operation has no additional effect, and the second attempt should recognize it's a duplicate and do nothing new. In practice, mobile engineers must treat payment or order APIs as idempotent by attaching unique operation identifiers to requests and deduplicating them on the backend. With this approach, even if the network drops a response or the user double-taps a button, the user is charged only once. Achieving this involves coordination between the app and the server. On the client side, every payment or mutation request is given a persistent unique ID. For example, the app might generate a new UUID when the user submits a payment, save that operation in a local "outbox" or queue, and include the ID in the HTTP request: Swift let opID = UUID().uuidString var request = URLRequest(url: URL(string: "/api/payments")!) request.httpMethod = "POST" request.setValue(opID, forHTTPHeaderField: "Idempotency-Key") request.httpBody = /* JSON payload of the payment */ This custom header (Idempotency-Key) carries the operation's identity to the server. (Modern APIs often expect this header to ensure that POST is treated safely.) The client's logic must be persistent; before sending, it writes the operation (ID and payload) into local storage (e.g., SQLite or SharedPreferences) so that it can recover and retry if the app closes or the network is down. This pattern, sometimes called the outbox pattern, means the user's intent is recorded immediately. A background dispatcher can then drain the queue on network availability or app restart; it tries each pending request. If a send fails (timeout, no internet, 5xx error), the entry remains in the queue for a later retry. This ensures at-least-once delivery and the server will eventually see the request, even if the app crashes or the network is flaky. But at-least-once alone would cause duplicates, so the server must be ready. When the backend receives the request with its Idempotency-Key, it first checks a deduplication store (for example, a database table keyed by this ID). Java String key = request.getHeader("Idempotency-Key"); PaymentResponse prev = idempotencyStore.lookup(key); if (prev != null) { // We have processed this request before - return the original response return prev; } // No record of this key; proceed with processing PaymentResponse result = processPayment(request.getBody()); // Store the result before returning it idempotencyStore.insert(key, result); return result; If the key already exists, the server simply returns the stored result without charging again. This ensures that the second (or third) time the client re-sends, the user doesn't get double-charged. Stripe's API, for instance, works exactly this way: it saves the outcome of the first request for a given idempotency key, and any retry with the same key returns the same result. In effect, the combination of at-least-once delivery (the client keeps retrying) plus idempotent handling on the server yields an effectively-once outcome. The server's idempotency store can be implemented with a simple database table that records each key and the operation's result. For instance, a processed_payments table might use the idempotency key as a primary key or unique constraint. The service then does an atomic INSERT ... ON CONFLICT DO NOTHING (PostgreSQL syntax) or equivalent. If the insert succeeds, the code proceeds with the payment and stores the result; if it fails because the key already exists, it knows this is a duplicate and can fetch the prior result. Wrapping the insert and the business operation in one database transaction avoids a race condition, as either both the key and payment record are written, or neither is. In SQL terms: SQL BEGIN; INSERT INTO payments(idempotency_key, user_id, amount) VALUES (:key, :userId, :amount) ON CONFLICT (idempotency_key) DO NOTHING; -- Check how many rows were inserted: IF (INSERT was successful) THEN -- This is the first time seeing this key; perform the payment CALL process_payment(...); -- The payment service may record a transaction ID, etc. COMMIT; ELSE -- Key already existed: rollback any partial work ROLLBACK; -- Retrieve and return the original payment result END IF; Even if two identical requests arrive concurrently, the unique constraint ensures only one succeeds in its insert. The other can detect the conflict and simply return the saved response. The system design sandbox guide describes this approach as "The database enforces uniqueness and no separate check needed. This works well when the idempotency record belongs in the same database as the business data, since you can wrap both in a single transaction." On the mobile side, it's also wise to guard against duplicates before the request is even sent. A simple in-memory or on-disk set of "seen" IDs can help reject retry loops after a crash or double tap. In Swift: Swift final class OperationDeduplicator { private var seen: Set<String> = [] func shouldProcess(_ id: String) -> Bool { return seen.insert(id).inserted } } This OperationDeduplicator returns true only the first time an ID appears. Persisting this set across app launches (for example in Core Data or a file) makes the app resilient to a crash after the payment is sent but before the response arrives. On relaunch, the app knows it already handled that operation and won't enqueue it again. It's important to integrate these pieces smoothly. A typical mobile flow might look something like this: the user submits a payment form, the app immediately generates a new opID (a UUID) and creates an operation record { id: opID, payload: {amount, items, ...} }. This record is saved locally. Then a background task picks it up, attaches opID as the Idempotency-Key header, and sends it. If the network call times out, the record stays queued. When the app regains connectivity or restarts, the dispatcher tries again. Because the same opID is used each time, the server knows to treat all retries as one. Only after the server successfully processes the payment does the app remove the operation from its queue. This pattern ensures retries and crashes do not cause duplicate side effects. Some systems even use more granular controls. For example, if the backend involves multiple microservices, one service might call others, and each service should propagate the same idempotency key or a related correlation ID so that the entire transaction remains idempotent. Distributed tracing can help debug how a request flowed through the system. Ultimately, the goal is to capture the entire user action from UI tap through backend processing and make sure it's only applied once globally. This often means also having the backend return the same HTTP status and response body on every retry, so the client never gets an unexpected error. Developers should simulate network failures and verify that retries do not cause double effects. Most important is to observe real production behavior, as logs or traces with the operation ID can tie multiple client attempts to a single transaction. If everything is correct, the system achieves effectively-once behavior where the payment occurs exactly once no matter how many times the client tries. As systemdesignsandbox summarizes, "at-least-once delivery + idempotent consumer = effective exactly-once". In practice, this means mobile apps can assume failures are not fatal and they can safely retry with the same key, knowing the server will protect against duplicates. Conclusion In summary, retry-safe mobile operations require treating each user action as an idempotent transaction. The client must persist a unique operation key and reuse it across retries, while the server must detect previously processed keys and prevent duplicate side effects. Combining durable client operations with server-side idempotency allows payments and other critical transactions to survive timeouts, crashes, and unreliable networks without being executed twice.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
One of the most important parts of API test automation is validating the response body to ensure data integrity. This step plays a key role in functional API testing, as it helps confirm that the API is returning the right data in the expected format. Response body validation isn’t limited to a specific request type; it applies equally to POST, GET, PUT, and PATCH APIs. The same validation approach can be used for any API response to verify the data returned by the service. Playwright offers multiple ways to validate response bodies. In this tutorial, I’ll walk you through these approaches to help you efficiently perform assertions on the response data using best practices. Checkout the previous tutorial blog to learn about Installation, the demo application, and how to send GET API requests with Playwright. How to Verify the Response Structure Response structure checks ensure that an API consistently returns data in the expected format, protecting the contract between backend services and their consumers. They help catch breaking changes early, such as missing or renamed fields, even when the API still returns a successful status code. TypeScript test("GET Order details and perform structure check", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toHaveProperty("message"); expect(responseBody).toHaveProperty("orders"); expect(responseBody.orders[0]).toHaveProperty("id"); expect(responseBody.orders[0]).toHaveProperty("product_name"); }); This test focuses on validating the structure of the API response. It validates that the response body contains the expected top-level keys and that each order object includes the required fields. Basic Assertions The basic assertions validate API success and data presence, making them a good first layer of verification before deeper structure or data-level checks. TypeScript test("Get order details and perform basic level verification", async ({ request, }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(Array.isArray(responseBody.orders)).toBeTruthy(); expect(responseBody.orders.length).toBeGreaterThan(0); }); This test performs a basic level check to confirm that the endpoint works as expected and returns the expected data in the response. After parsing the response body, the assertions focus on the following essential basic-level checks: TypeScript expect(responseBody.message).toBe("Order found!!"); The above line of code verifies that the API returns the expected message text in the response body. TypeScript expect(Array.isArray(responseBody.orders)).toBeTruthy(); This line of code ensures that the orders field in the response is an array, validating the basic response format. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This part of the test confirms that at least one order is returned in the orders array, ensuring the response contains required data. How to Verify Response Data With Details Validating the actual data returned in the response is essential to ensure that the API response contains the correct values. TypeScript test("Get order and verify order details", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); }); The following code ensures that the response has a valid identifier and it is not missing or empty. TypeScript expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); This check is required because the API generates the order ID when a new order is created in the system. It ensures that the “id” field has a valid value generated and assigned to it, since this “id” is used to retrieve, update, or delete order data. TypeScript expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); These statements assert that the order details are retrieved correctly for the respective request. The “user_id” - “1” was sent in the request, and verifying it in the response, along with the other order details such as “product_id” and “product_name,” ensures that the correct data is returned. How to Verify Response Data by Matching Objects and Arrays Playwright allows response data verification by matching objects and arrays partially within the API response. This approach is useful because it makes tests more flexible and confirms that the API returns the correct data structure and values. TypeScript test("Get order and verify matching object and array", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toMatchObject({ message: "Order found!!", orders: expect.arrayContaining([ expect.objectContaining({ product_id: "79", product_name: "5 star 10gm Chocobar", product_amount: 5, qty: 1, tax_amt: 0.5, total_amt: 5.5, }), ]), }); }); In this test, the toMatchObject assertion verifies that the response contains a “message” with the expected value “Order found!!” and an orders array. Within the array, "expect.arrayContaining" ensures that at least one order matches the expected data, while "expect.objectContaining" verifies only the values in the specified fields of that order. Using Best Practices to Perform Assertions Best practices create stable, maintainable API automation tests by combining basic checks with flexible data matching. TypeScript test("Get Order details API test with best practice", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); }); The test sends a GET request to fetch order details for “user_id”-“1". The use of failOnStatusCode: true ensures the test fails immediately if the API does not return a 2xx status code. The response is then parsed into a JSON object for validation. The assertions are structured in layers: TypeScript expect(responseBody.message).toBe("Order found!!"); This assertion verifies the message text, confirming that the API returns the correct message when an order is found. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This statement ensures meaningful data is returned and avoids false positives when the array is empty. TypeScript expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); The final part of the code performs the final assertion using arrayContaining and objectContaining to verify that at least one order has the expected “id” and “product_name”, without asserting every field. These layered validations improve clarity by verifying structure, data presence, and key data values in sequence. Extracting Data From the Response Extracting data from the API response is a common and widely used pattern in API test automation. It is important in multiple ways, such as reusing the data in further tests for dynamic testing and end-to-end validation. TypeScript test('Get order details and extract the order id', async({request}) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); const product_name = order.product_name console.log(product_name) }); This test sends a GET API request and performs basic validations to ensure the API response is reliable. TypeScript const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); The code above extracts the “order_id” from the order object in the response. Before accessing it, an assertion is made to verify that the value is not null. Finally, the value of the order_id is printed in the console. TypeScript const product_name = order.product_name console.log(product_name) Similarly, other values, such as product_name, can also be extracted. Attaching the Response Body to the Playwright Report The Playwright report, by default, shows the steps executed, the number of tests run, pass/fail status, and time taken to run the tests. However, it does not attach the response body to the test report. Attaching the response body to the report improves visibility and makes the test report more informative and transparent. The following code shows how to extract the required metadata and attach it to the Playwright report. TypeScript test("Get order details API and attach the response details to the report", async ({ request, }, testInfo) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, }); expect(response.status()).toBe(200); const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); const fullResponse = { status, statusText, headers, body, }; await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); }); The testInfo is a built-in Playwright fixture and provides utilities to manage and inspect test execution, such as attaching files to reports, updating test timeouts, and identifying the currently running test. The following lines of code extract the response metadata, such as the status code, status text, headers, and response body. TypeScript const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); Next, let’s combine all response details and create a single object containing: Status codeStatus textHeadersResponse body TypeScript const fullResponse = { status, statusText, headers, body, }; Finally, let’s attach these details to the report using the testInfo.attach() method as shown below: TypeScript await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); The testInfo.attach() adds an attachment to the Playwright report. The attach() method has 3 parameters: Name of the attachment: The first parameter is the name, “Full API Response”, that will be shown for the attachment.Body of the attachment: The second parameter is for the body of the attachment. The JSON.stringify(fullResponse, null, 2) has 3 arguments. The first argument converts the fullResponse object into a readable, pretty-formatted JSON. The second argument is the replacer, which is null. It ensures that all properties from the fullResponse object are included as they are, without modifying anything. The third argument controls pretty-printing. Here, “2” means indent nested JSON by 2 spaces.Content type: This parameter ensures that the report treats the attachment as JSON. The following screenshot is generated after the tests are run: Test Execution Running the tests in Playwright is simple and easy. We can run the following command from the terminal: Plain Text npx playwright test To generate the report, the following command can be used: Plain Text npx playwright show-report Summary Playwright provides multiple approaches, including structure checks and matching objects and arrays for verifying response data. The right strategy should be chosen based on your project’s requirements. Based on my experience, combining response structure checks with response data validation, including the matching object and array strategy, can be used as an effective approach for validating API responses. Happy testing!
A green Ansible run can hide an operation with no clear owner. The tasks completed. Every target reported success. The requested change happened. Yet nobody can say with confidence which system now owns the resource state, watches the service, controls the credential, or decides whether the next action is safe. This is how useful configuration management becomes an operational liability. The problem is rarely that Ansible cannot run the command. It is that successful execution gets mistaken for durable control. Ansible can create cloud resources, build images, launch migrations, rotate passwords, and promote databases. Its flexibility encourages teams to keep adding tasks until the playbook becomes the resource ledger, runtime controller, artifact system, credential authority, and approval workflow. Being able to express an operation does not make the playbook its correct owner. The Question Beneath the Playbook Ansible's own playbook documentation describes playbooks as a repeatable configuration-management and multi-machine deployment system. It also makes a narrower point about idempotency: most modules check whether the desired state already exists, but not every module or playbook behaves that way. Where modules support it, check mode can report proposed changes before execution. That is a strong execution model. A playbook receives inventory and variables, connects to targets, executes ordered tasks, reports a result, and exits. Automation controllers add scheduling, role-based access, managed credentials, workflows, and event triggers. Those capabilities improve how playbooks run, but they do not automatically give the playbook the state model of every domain it touches. A simple review question exposes the boundary: After this automation exits, what must remain true, and which system keeps it true? This is the exit test. Plain Text Required behavior Natural owner Host configuration convergence Configuration management Resource graph and replacement plan Stateful provisioning engine Continuous observation and correction Runtime controller Versioned machine or container output Artifact build pipeline Schema history and transactional order Domain migration system Credential issuance and rotation Secret or identity authority Approval and decision policy Governance workflow Ansible can participate in every row without becoming the authority for every row. In A Tool Is Not a Platform, I argued that a platform is defined by its contract rather than its technology. The exit test applies the same reasoning to operations: the execution contract can complete while the wider operational contract remains open. A Recovery Drill That Required Several Authorities A recorded HybridOps PostgreSQL HA recovery cycle on March 31, 2026 rebuilt a three-node recovery cluster in Google Cloud from pgBackRest, took a fresh backup from the recovered primary, and returned service on premises. The restore completed in 26 minutes 58 seconds, the fresh backup in 27 seconds, and failback in 9 minutes 38 seconds. Configuration management prepared the nodes and executed bounded steps. It did not own every operational truth. The provisioning layer retained resource state. pgBackRest retained recovery lineage. Patroni retained cluster leadership. DNS retained the active service endpoint. The cutover procedure required the original primary to be fenced before traffic moved. That final boundary was critical. Every configuration task could succeed while the original primary remained writable. The playbook would be green, but the database estate would carry split-brain risk. The example is not an argument for less automation. It is an argument for explicit authority. The executor should not silently inherit responsibilities that belong to the systems around it. The blueprint ordered provisioning, restore, validation, backup, cutover, and failback, while structured run records captured the outcome across those handoffs. Configuration management remained one bounded implementation path. It did not become the resource ledger, database controller, backup authority, or DNS state model. Resource State Should Survive the Executor Ansible cloud modules can create networks, virtual machines, identity bindings, and managed services. That can be appropriate for a bounded or ephemeral operation. It becomes harder to defend when the workload needs a durable resource graph, replacement planning, state locking, imports, and a predictable destroy path. DZone's IaC platform example using Terraform, Ansible, and GitLab shows this division in practice: the provisioning layer retains infrastructure state while Ansible roles handle software provisioning and configuration. A stateful provisioning engine retains the relationship between declared resources and provider objects. HashiCorp describes this state mapping as the binding between configured resource instances and remote objects, together with supporting metadata. That memory allows the engine to calculate a plan and reason about the next change. Without that memory, a partial run can leave the next operator reconstructing ownership from cloud inventory, task output, and assumptions about which steps completed. The automation worked until recovery required information it did not retain. Ansible remains useful after provisioning. It can configure the operating system, install packages, place files, manage services, and verify readiness. Resource lifecycle and host convergence are clearer as separate responsibilities. Runtime Control Must Outlive the Run A playbook can inspect a service, restart it, and confirm that it is healthy. The ordinary run stops observing after it exits. Kubernetes documents a controller as a non-terminating control loop that watches current state and moves it toward desired state. The persistent loop, observed state, and domain model are the important parts of that definition. Leader election, database failover, autoscaling, and cluster reconciliation require an active control loop with domain knowledge. A database cluster manager understands membership, replication health, promotion safety, and split-brain risk. Remote tasks do not acquire those semantics because they can call the same commands. Configuration management can install and validate the controller. The controller should retain authority over live decisions. DZone's introduction to event-driven Ansible automation shows the model clearly: event sources feed rulebooks, and matched rules trigger actions. That is useful for bounded remediation and evidence collection. It still depends on the quality of the event source and the safety of the rule. A faster trigger cannot make an unsafe promotion condition safe. Artifacts and Transactions Need Their Own Histories Building an image is not the same operation as configuring a running host. The output is a versioned artifact that needs known inputs, build metadata, tests, checksums, and a publication path. Ansible can provision the filesystem during the build. The image pipeline should retain artifact identity and release history. Otherwise, a successful build can produce an image that nobody can reproduce or confidently roll back to later. Database migrations expose a similar boundary. A playbook can copy a migration and invoke a command. The difficult work is knowing which migrations ran, enforcing order, acquiring locks, coordinating concurrent releases, and recovering from a partial failure. A domain migration system is designed around that history. Ansible may install or invoke it, but reproducing its state model in task conditions creates a weaker version of the same mechanism. Encryption Is Not a Credential Lifecycle Ansible's Vault documentation defines Vault around encrypting and managing sensitive variables and files. That solves an important storage problem. It does not provide issuance, scoped access, expiry, rotation, revocation, or an audit trail by itself. An encrypted variable file should not quietly become the organization's credential authority. A secret manager, certificate authority, or identity provider should manage the lifecycle. Ansible can configure clients, deliver references, and consume short-lived credentials during execution. When encrypted variables become the credential system, expiry and revocation tend to become manual cleanup. The playbook protects stored content, but the wider credential lifecycle remains unowned. Execution Is Not Authorization Some operations are easy to automate and unsafe to trigger from one signal. Disaster-recovery failover, destructive teardown, data promotion, and wide-blast-radius changes fall into this category. A playbook can execute a prepared sequence consistently. It does not decide whether an outage signal is trustworthy, whether a recovery target is current enough to promote, or whether the business impact justifies the action. A confirmation prompt records consent at one moment; it does not establish that the decision was sound. The decision belongs in a policy or workflow layer that evaluates the required signals, records the decision class, applies the approval boundary, and then authorizes execution. Ansible may remain the executor. A reliable sequence can still execute the wrong decision perfectly. Keep Ansible in Its Strongest Position Ansible is a strong choice for repeatable configuration across reachable systems: packages, users, files, services, operating-system settings, application prerequisites, and post-provision checks. It also works well as a bounded orchestrator when each underlying system retains its own state. It can coordinate provisioning, image, cluster, migration, and secret operations without replacing the authorities behind them. The exit test belongs in design review: After the playbook exits, what must remain true, and which system keeps it true? If the answer depends on continuous observation, durable state, transaction history, artifact identity, credential lifecycle, or a policy decision, another mechanism probably needs to remain responsible. Ansible can configure it, invoke it, or verify it. Configuration management becomes an operational liability when successful runs hide missing ownership. Knowing where the playbook should stop is part of using it well.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Kubernetes environments can drift, accumulate one-off fixes, and diverge across teams until a routine deploy breaks or a cost spike forces a review. This checklist gives platform, SRE, and engineering teams a way to keep clusters, deployments, and automation manageable as Kubernetes operations scale and teams grow. It covers standards, observability, releases, access, drift, and cost. Review it before promoting a service to production and revisit it as your environments shift. Cluster Standards and Environment Discipline Most teams run more than one Kubernetes cluster, and those clusters diverge over time as they are upgraded and modified independently. At that point, a fix or runbook that works on one cluster can’t be trusted to work on another. Keeping the fleet operable requires every cluster to run a supported Kubernetes version and follow the same approved platform settings and policies. Document centrally controlled settings (e.g., Kubernetes versions, networking, admission policies) separately from service-team settings (e.g., pod resource requests, autoscaling, ConfigMaps)Standardize namespace, labeling, and resource-quota conventions so workloads are identified and bounded consistently across clustersMaintain each cluster’s baseline configuration in version control; use reconciliation to apply it and correct untracked changesMaintain an approved Kubernetes version range across environments; track each cluster against this range and upgrade it before its current version reaches end of supportRecord any cluster setting that differs from the standard baseline, including the justification, approver, expiry date, and whether it must be restored or reapproved control areawhat to standardizeminimum evidence Kubernetes version Supported version range and upgrade cadence Version inventory showing every cluster within the supported range Cluster baseline Networking, ingress, and baseline policies Declarative config in version control, reconciled to live state Namespaces and quotas Naming, labels, and resource quotas Quota and label audit across clusters Exceptions Approved deviations from baseline Record of approved deviations with justification and expiry date Deployment Consistency and Release Safety Kubernetes makes it easy to ship a change to production several ways: a CI/CD pipeline, a Helm upgrade run by hand, or a kubectl apply straight from a laptop. Each runs different checks, but the manual ones skip the tests and approvals that a pipeline would enforce. A repeatable release path applies the same gates every time and provides a reliable way to recover when a deployment fails. Require every service to follow the same approved deployment path from commit to production, with consistent release steps and controls across teams and environmentsPromote the same versioned, immutable artifact through every environment without rebuilding it at each stageRequire every change to clear the same automated gates (e.g., tests, policy checks, health checks) before reaching productionRoll out production changes in stages (e.g., canary release, percentage-based traffic shift); automatically stop or roll back when predefined health criteria are not metFor every production change, require a rollback, feature disablement, or recovery path that has been tested before releaseFor each deployment, assign an owner accountable for monitoring it through release and triggering rollback on failureRecord every production deployment with its artifact version, approver, and timestamp so the active release stays auditable Observability and Operational Readiness A Kubernetes cluster keeps workloads running by restarting and rescheduling them, so a service can keep failing without the failure ever becoming obvious. A pod stuck in CrashLoopBackOff or failing its readiness probe can remain unhealthy for hours, and if it emits no metrics or logs of its own, there’s nothing to tell you what went wrong. Catching that early depends on each service surfacing its own signals rather than waiting for the cluster to show something is wrong. Require every new service to ship with a minimum observability baseline before production: metrics, structured logs, traces, and liveness and readiness probesDefine service health signals (e.g., latency, traffic, errors, saturation), each with a threshold and assigned team that responds when it is breachedStandardize structured logging and trace context so a request can be followed end to endRoute every alert to an on-call rotation or runbook; retire alerts no one acts onMaintain a quarterly reviewed runbook for each service, including known failure modes, escalation contacts, and recovery stepsSet minimum retention periods for metrics, logs, and traces, with documented justification and explicit approval for shorter retention periodsRun a post-incident review after every major outage; apply findings to update runbooks, alerts, and service baselines Access Controls and Automation Guardrails A Kubernetes cluster usually serves many teams and workloads through a single shared control plane. A role with too much access, for example, can affect them all at once. And when the credential is shared, there’s no way to tell later who actually made the change. Access that stays narrow and tied to a single identity keeps a mistake or a compromised account from impacting the whole cluster. Use namespace-scoped RBAC roles with only the required permissions; grant cluster-wide administrator access only through logged, justified, time-limited exceptionsGive each automation its own scoped service account so automated and privileged actions trace to a distinct identity instead of shared credentialsReserve break-glass access for emergency production changes, with time limits and post-use reviewUse admission policies to reject workloads with unsigned images, privileged containers, or settings barred by platform standardsRecord the actor, target, and timestamp for every privileged or automated action in the Kubernetes audit log; regularly review for activity that does not match an approved change or access requestUse short-lived, automatically rotated ServiceAccount tokens for workloads; revoke credentials and RBAC bindings when a person, workload, or automated process is decommissioned Drift and Failure Management Over time, a Kubernetes cluster’s live state can drift from the configuration stored in version control. This could be due to a hotfix applied directly to a live resource during an incident or an incomplete rollout that leaves the cluster partially updated. If those differences are not fixed, a subsequent deployment may conflict with the live state or overwrite a manual change, and version control may no longer accurately reflect what is running in the cluster. Use automated checks to compare live cluster state with the version-controlled baseline at defined intervals; record each mismatch and notify the team responsible for the affected resourceSet risk-based remediation deadlines for detected drift, requiring teams to restore the baseline or approve a time-limited exception for the changed configuration before the deadlineLog every manual production change and resolve it within a defined period by updating the baseline or reverting the live resource to its declared stateSet an SLO and error budget for each service, identify the team tracking budget use, and pause feature work to prioritize reliability fixes when the budget is exhaustedRun root-cause reviews for recurring failures and apply findings to update baselines, policies, and admission checks instead of patching each instanceTest failure scenarios (e.g., pod disruption, node loss, dependency outages) on a defined schedule, confirm services recover as expected, and track remediation for any gaps example drift patternwhat usually reveals it Manual live-resource change Reconciliation diff against declared state Version or baseline skew Scheduled cluster inventory audit Expired break-glass fix Exception register entry past its window Repeated failure patched one service at a time Same root cause across incident reviews Cost Awareness and Resource Discipline In Kubernetes, resource requests for CPU and memory determine how much cluster capacity is reserved for a workload. Teams may size these requests for peak demand and leave them unchanged even when normal usage is much lower. Across many workloads, this unused capacity adds up and can cause the cluster to run more nodes than actual demand requires, increasing infrastructure costs. Set CPU and memory requests based on representative usage data; set limits where appropriate based on workload behavior and reliability requirementsReview workloads whose requests exceed observed use by a defined threshold, accounting for traffic patterns and reliability needsRequire cost-allocation labels for each workload by team and namespace; correct unallocated spend and missing or inaccurate labelsReclaim idle and orphaned resources (e.g., unused volumes, stale namespaces, oversized nodes) on a monthly cadenceSet autoscaling thresholds based on demand and reliability requirements; periodically review settings that fall outside the approved rangeRegularly review sustained overprovisioning or low utilization; reduce excess capacity or record why it must be retained when avoidable cost exceeds a set threshold Closing Run this checklist before a service enters production and at regular intervals afterward. Repeat it when clusters are upgraded, team responsibilities change, or services are added or retired. Resolve failed checks and revisit approved exceptions before they expire. Unresolved configuration drift can accumulate across environments until teams begin to treat it as the intended baseline. This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
The first warning sign wasn't an outage. It was a boring pull request. We changed one App Service setting. It was the sort of change that should have resulted in a small plan and a quick review. Instead, Terraform refreshed networking, private endpoints, DNS, Key Vaults, storage accounts, app services, and monitoring before showing what would actually change. Nothing was broken; that was the point. Terraform did exactly what it was designed to do: account for everything represented in state before calculating change. The problem was that our Terraform state had become a single, platform-sized boundary that every small change had to pass through, and one no team could fully own. If you have run a landing zone as a single Terraform configuration, you have probably had a version of that pull request. The instinct afterward is to blame size: the configuration has grown too large, so break it up. That instinct is wrong, or at least incomplete. Size is uncomfortable, but coupling is what actually hurts. Nothing in the change touched networking, DNS, or those key vaults. They were dragged into the plan because everything was bound together through one state. At first, that coupling just means slow plans and noisy reviews. Later, it raises a harder question: who actually owns this? Where the Coupling Shows Up Start with the plan. In a monolith, Terraform has to account for everything represented in the state before it can tell you what changed. You can target a single resource, but that is an escape hatch, not a way to run a platform. So the wait scales with the size of the estate, not your change. Both a one-line edit and a fifty-resource migration get stuck behind the same refresh before the diff appears. Provider upgrades show the same problem. A single root configuration pins one set of provider versions, so you cannot move networking to a newer azurerm version and leave everything else behind. Every upgrade becomes all-or-nothing, which means it keeps losing to smaller, safer priorities. Ours sat on azurerm 2.97 and only moved to the 4.x line once the upgrade could no longer be put off. The monolith had made the jump too big to schedule any sooner. The bigger concern is blast radius. One state file, one lock, one plan. A bad apply, a corrupted state, a destroy that catches more than you aimed at: whatever goes wrong can reach more of the platform than the change was ever meant to touch, because nothing in the layout is there to contain it. The dependency graph suffers too. Unrelated resources get sequenced together just because they share a graph. A network change might wait on unrelated compute, DNS on policy. The graph ends up reflecting accidental grouping rather than real dependencies. The result is clear. There is no small change. You cannot ship a DNS record or a new Key Vault without running the entire configuration through plan and apply. Every change is a platform change, carrying platform risk and requiring review, no matter how minor. These look like separate problems, but all come from the same design choice: too many unrelated concerns tied into one Terraform boundary. Where Coupling Becomes Ownership It is easy to call these operational annoyances: slow plans, awkward upgrades, risky applies, the tax you pay for a big configuration. But the same coupling appears in review and approval, where it stops being just an operational problem. Once too many concerns share the same state, pipeline, and approval path, the question is no longer only "how long did the plan take?" It becomes "who is accountable for the boundary this change is crossing?" Take private connectivity. A single private endpoint on Azure isn't handled by just one team. The application team owns the service behind it. The platform team manages the landing zone, subnet, and endpoint placement. Private DNS zones might be managed centrally or by another team. Security or governance may require the service to be private. How these map to teams varies, but in a monolith, everything ends up in the same state, pipeline, and plan. So "who owns this?" rarely has a clear answer. However you split teams, they are coupled through a single configuration that none can truly own. When the application team changes its service, the same config still carries platform connectivity and governance controls. You cannot draw ownership along your real organizational boundaries, because the code does not have them. Both slow plans and unclear ownership trace back to the same issue: shared concerns treated as if they belong to just one team. Figure 1: When Terraform boundaries stop matching ownership boundaries. The monolith gives Terraform one boundary. Organizations have several. The pain comes when small changes have to cross boundaries that no team fully owns. Reach for the Coupling, Not the Size The reflex now is to split the state and move on. But splitting a landing zone poorly can be worse than leaving it alone. If you split along the wrong lines, you trade one blast radius for tangled cross-state dependencies. You also lose the single plan that at least showed the whole graph in one place. For example, splitting private endpoints into one state and private DNS zones into another may look clean on paper. But if different teams deploy them without a clear agreement, every new endpoint becomes a coordination headache, not a smaller change. Moving files into separate folders does nothing if the same pipeline, credentials, and approval path still govern everything. Decomposition should follow actual coupling, not just line count. So the next question is not "how many states should we create?" It is "which boundaries are real enough for teams to own, deploy, and recover independently?" If your Terraform monolith hurts, do not start by counting files or resources. Look at what is actually being coupled. Slow plans and unclear ownership are both signs that your Terraform boundaries no longer match your real ownership boundaries.
Senior Staff Engineer,
Marqeta
Manager , Release Engineering & DevOps,
TraceLink