DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Software Design and Architecture

Software design and architecture focus on the development decisions made to improve a system's overall structure and behavior in order to achieve essential qualities such as modifiability, availability, and security. The Zones in this category are available to help developers stay up to date on the latest software design and architecture trends and techniques.

Functions of Software Design and Architecture

Cloud Architecture

Cloud Architecture

Cloud architecture refers to how technologies and components are built in a cloud environment. A cloud environment comprises a network of servers that are located in various places globally, and each serves a specific purpose. With the growth of cloud computing and cloud-native development, modern development practices are constantly changing to adapt to this rapid evolution. This Zone offers the latest information on cloud architecture, covering topics such as builds and deployments to cloud-native environments, Kubernetes practices, cloud databases, hybrid and multi-cloud environments, cloud computing, and more!

Containers

Containers

Containers allow applications to run quicker across many different development environments, and a single container encapsulates everything needed to run an application. Container technologies have exploded in popularity in recent years, leading to diverse use cases as well as new and unexpected challenges. This Zone offers insights into how teams can solve these challenges through its coverage of container performance, Kubernetes, testing, container orchestration, microservices usage to build and deploy containers, and more.

Integration

Integration

Integration refers to the process of combining software parts (or subsystems) into one system. An integration framework is a lightweight utility that provides libraries and standardized methods to coordinate messaging among different technologies. As software connects the world in increasingly more complex ways, integration makes it all possible facilitating app-to-app communication. Learn more about this necessity for modern software development by keeping a pulse on the industry topics such as integrated development environments, API best practices, service-oriented architecture, enterprise service buses, communication architectures, integration testing, and more.

Microservices

Microservices

A microservices architecture is a development method for designing applications as modular services that seamlessly adapt to a highly scalable and dynamic environment. Microservices help solve complex issues such as speed and scalability, while also supporting continuous testing and delivery. This Zone will take you through breaking down the monolith step by step and designing a microservices architecture from scratch. Stay up to date on the industry's changes with topics such as container deployment, architectural design patterns, event-driven architecture, service meshes, and more.

Performance

Performance

Performance refers to how well an application conducts itself compared to an expected level of service. Today's environments are increasingly complex and typically involve loosely coupled architectures, making it difficult to pinpoint bottlenecks in your system. Whatever your performance troubles, this Zone has you covered with everything from root cause analysis, application monitoring, and log management to anomaly detection, observability, and performance testing.

Security

Security

The topic of security covers many different facets within the SDLC. From focusing on secure application design to designing systems to protect computers, data, and networks against potential attacks, it is clear that security should be top of mind for all developers. This Zone provides the latest information on application vulnerabilities, how to incorporate security earlier in your SDLC practices, data governance, and more.

Latest Premium Content
Trend Report
Cloud-Native Foundations
Cloud-Native Foundations
Trend Report
Security by Design
Security by Design
Refcard #291
Code Review Core Practices
Code Review Core Practices
Refcard #392
Software Supply Chain Security
Software Supply Chain Security

DZone's Featured Software Design and Architecture Resources

Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction

Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction

By Igboanugo David Ugochukwu DZone Core CORE
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. After a few years of operating a shared Kubernetes environment, the shift in the center of gravity becomes clear. Cluster provisioning, container scheduling, and upgrades become routine, yet releases still stall over ownership, access, telemetry, and cost allocation. Consider a hypothetical product team adding a stateful order-processing service to a shared platform. The service has an API, a worker, database migrations, and a data store backed by cluster-managed persistent storage, and it must run in staging and production. We will follow that service through its delivery path to examine where mature infrastructure stops helping, how local workflow differences compound, and which operating model decisions restore consistency without stripping teams of useful autonomy. When Infrastructure Maturity Stops Solving the Hard Part At first glance, onboarding the service should be routine. The cluster exists, the CI system can build an image, and infrastructure as code can create the namespace. Then the service reaches production and encounters a different StorageClass, quota profile, network policy, or service account configuration from staging. Each difference may be valid, but the delivery workflow didn’t surface the environment contract early enough. This is the practical limit of infrastructure maturity. Reliable clusters provide capable building blocks, while reliable delivery also requires a shared agreement about how teams use those blocks, what evidence a release produces, where exceptions go, and who owns the outcome. How Cloud Complexity Starts to Compound Follow the service through one release and the pattern becomes clearer: The same components use inconsistent service and environment identifiers across logs and traces.Ownership labels exist in one cluster but not the other.The team copies a pipeline because the shared template cannot sequence migrations.Production access and policy exceptions move through separate ticket queues. The operational cost shows up in the manual coordination required before each deployment. An engineer has to reconstruct which rules apply every time. During an incident, responders can’t move cleanly from an alert to the owning team, deployment record, runbook, and cost center. Finance sees shared-cluster spend that cannot be attributed reliably, while the security team receives evidence in different formats. OpenTelemetry semantic conventions and FinOps allocation practices rely on consistent service, environment, and allocation metadata, so local naming schemes undercut the value of the underlying tools. As the same pattern spreads across clusters and cloud accounts, small differences become a persistent operating burden. Operating Models Set the Terms of Scale The operating model decides who turns those building blocks into a usable delivery system. For our example, the product team owns the order domain, data model, migration safety, scaling behavior, SLOs, and on-call response. The platform team owns the interface through which the service receives a namespace, workload identity, baseline policy, deployment workflow, and telemetry defaults. Security, SRE, and FinOps teams contribute requirements and review the evidence that the workflow produces. This split keeps service-specific decisions close to the people who understand them while centralizing cross-cutting capabilities that every team would otherwise rebuild. CNCF’s platform guidance makes an important distinction here: A platform team is responsible for the interfaces and experience around shared capabilities, even when another team or provider operates the backing service. In practice, a mature platform offers versioned workflows, clear support boundaries, self-service for common requests, and feedback loops based on real usage. In this way, the platform team is an enabler of consistency rather than the operator of every component. Standardization, Autonomy, and Shared Operating Logic The defensible baseline is narrower than a universal application architecture. For this service, shared standards should cover: Service and environment identityWorkload identity and minimum network policyResource requests, quota expectations, and cost-allocation metadataRelease evidence, rollback behavior, and minimum telemetry These rules belong in the shared workflow because inconsistency affects other teams and complicates incident response, security, and cost allocation. The product team still chooses its schema, partitioning strategy, cache design, scaling thresholds, and release timing, and defines SLOs around the behavior users experience. The stateful workload then tests that boundary. A default pipeline built for stateless HTTP services may need a supported hook for migrations and worker rollout. An overly broad standard becomes an approval layer or bottleneck, while an overly narrow one leaves every team maintaining its own release and recovery logic. A bounded extension with an owner, tests, constraints, and review date preserves autonomy without creating an unsupported parallel system. Why Team Friction Turns Into a Scaling Tax Weaknesses in the operating model become most visible in the friction between teams. If a developer must request a namespace, ask another team for credentials, copy a pipeline, and find a production approver, the architecture may be automated while delivery remains ticket-driven. Each handoff adds queue time and loses context. Adding a portal without changing that path gives the developer one more place to check. Effective self-service completes the request, applies policy, records the change, and returns a clear support path. To see whether self-service is reducing friction, track metrics like request-to-environment time, time to first production deployment, exception rate, support demand, and failed-deployment recovery time. CNCF recommends tracking fulfillment and new-service delivery latency; DORA advises applying delivery metrics in the context of a specific service. Together, these measures show whether the workflow reduced coordination overhead or moved it to another queue. When Control Models Backfire The order-processing service example exposes two ways the control model can fail: Overly rigid standardization. A workflow designed only for stateless services forces the team to create a separate migration path, fragmenting release evidence.Unbounded local variation. Unrestricted cluster access allows identity, policy, and resource controls to drift between teams. The scalable approach pairs a narrow baseline, enforced through mechanisms like admission policies, with a documented extension path for legitimate workload-specific behavior, keeping the standard credible without turning each exception into a permanent fork. Operating Assumptions That Fail at Scale Old Assumption Why It Breaks Operating Model Replacement Healthy clusters make a workload portable Storage, identity, policy, and quota profiles differ by environment Versioned environment contract with a shared baseline One shared pipeline can serve every workload Stateful rollout and migration steps don’t fit the default sequence Core workflow with bounded, tested hooks A portal provides self-service Tickets and manual approvals remain behind the interface Workflow that provisions, enforces policy, and records evidence Local conventions remain harmless when teams own their services Metadata and controls drift across services Small enforced baseline with governed exceptions Making Cloud Complexity More Manageable To begin, you don’t need to redesign your entire platform. You can trace one representative delivery workflow and find where coordination breaks. For the order-processing service example, map the path from repository creation to production, including owners, queues, controls, evidence, and exceptions. Improvements should then be tested through adoption and outcomes such as lead time, failed-deployment recovery time, support demand, exception volume, and cost-attribution coverage. This sequence shows whether the platform is reducing operational variation for real workloads before the model expands to more teams and environments. References: Platforms for Cloud-Native Computing, CNCFResource Quotas, KubernetesStorage Classes, KubernetesAdmission Control in Kubernetes, KubernetesResource Semantic Conventions, OpenTelemetryAllocation FinOps Framework Capability, FinOps FoundationService Level Objectives, Google SRESoftware Delivery Performance Metrics, DORA This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report More
6 Techniques To Reduce LLM API Costs With the Python Library

6 Techniques To Reduce LLM API Costs With the Python Library

By Somnath Banerjee
Before diving into solutions, it helps to understand the scale of the problem. Take a common production pattern: a customer support bot that processes 10,000 messages per day, each with a 2,000-token system prompt and a 200-token user message. Cost breakdown by the numbers: scenariomodeldaily costNo optimizationClaude Opus ($5/1M)$11.00Right model for taskClaude Haiku ($1/1M)$2.20Add prompt cachingHaiku + caching ($0.10/1M cached)$0.42Combined savings96% reduction That's not a benchmark; it's arithmetic. The techniques don't require magic; they require applying what providers already offer. Technique 1: Prompt Caching, Up to 90% Off Repeated Content The Problem Most LLM applications send the same system prompt on every request. If your system prompt is 2,000 tokens and you make 10,000 requests per day, you're paying for 20 million input tokens daily, even though the content never changes. How It Works Anthropic's prompt caching lets you mark content blocks with cache_control. The first request pays full price and writes to cache. Every subsequent request that hits the same cached content pays 10% of the normal input price. Cache entries last 5 minutes and reset on each hit. The key insight: cache the most stable content first. Your base instructions change rarely. Your few-shot examples change occasionally. Your per-request context changes every time. Structure your prompt from most stable to least stable. Python from llm_optimizer import OptimizedClient, build_cached_system_prompt import anthropic client = OptimizedClient(anthropic_client=anthropic.Anthropic()) # Build an optimally structured cached system prompt system = client.build_cached_system( base_instructions=""" You are an expert customer support agent for a SaaS company. You have deep knowledge of our product, billing, and technical issues. Always be empathetic, clear, and solution-focused. [... 1,500 more tokens of stable instructions ...] """, # ← cached after first request — 10% cost on all subsequent calls few_shot_examples=""" Example 1: Billing question → here's how to handle it Example 2: Technical issue → here's the escalation path [... 500 tokens of examples ...] """, # ← also cached separately ) # First call: pays full price, writes to cache response1 = client.complete(messages=[{"role": "user", "content": "How do I cancel?"}], system=system) # Second+ calls: system prompt served from cache at 10% cost response2 = client.complete(messages=[{"role": "user", "content": "Where's my invoice?"}], system=system) What the Library Does llm-optimizer automatically injects cache_control breakpoints at optimal positions, system prompt, few-shot examples, and long conversation history, respecting Anthropic's 4-breakpoint limit. You don't touch the API directly. Savings Calculation Shell 2,000 token system prompt × 10,000 requests/day = 20M tokens/day Without caching: 20M × $3.00/1M (Sonnet) = $60.00/day With caching: 2M × $3.00 + 18M × $0.30 = $11.40/day Savings: $48.60/day = $17,739/year Technique 2: Model Routing, 60% to 80% Off by Using the Right Model The Problem Routing every request to your best model is the most common and most expensive mistake. Claude Opus costs 5x more than Claude Haiku. For tasks that Haiku handles perfectly, such as classification, extraction, translation, and simple Q&A, you're paying a 500% premium for no benefit. The Naive Approach and Why It Fails The obvious solution is to route by keyword: if the prompt contains "classify," use Haiku; if it contains "analyze," use Sonnet. This works until it doesn't. A prompt like "Explain the constitutional implications of this clause" is 8 words. Short, simple-looking. A keyword router sees no complexity signals and routes it to Haiku. But the task requires expert-level legal reasoning. This class of error intent-heavy short prompts is the primary failure mode of heuristic routing. The Better Approach: Use a Classifier llm-optimizer solves this by optionally using Haiku itself to classify task complexity before routing. The cost is approximately 15 tokens, about $0.000015. If that classification prevents one wrong Opus call (2,000 tokens × $5/1M = $0.01), it pays for itself 666 times over. Python from llm_optimizer import OptimizedClient, Provider client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_llm_classifier=True, # uses Haiku to assess complexity — ~$0.000015/call preferred_provider=Provider.ANTHROPIC, ) # Short prompt, complex intent → correctly routed to Opus response = client.complete( messages=[{"role": "user", "content": "Explain the constitutional implications of this clause"}] ) # You can audit the routing decision from llm_optimizer import ModelRouter router = ModelRouter(enable_llm_classifier=False) print(router.explain("classify this email as spam or not")) # { # "detected_complexity": "simple", # "routed_model": "claude-haiku-4-5", # "keyword_signals_fired": {"simple": ["classify"]}, # "token_count": 8 # } Complexity Tiers A closer look at complexity: Tierexamplesdefault modelSimpleClassification, extraction, yes/no, translationClaude HaikuMediumSummarization, paraphrasing, short Q&AClaude HaikuComplexCode generation, analysis, evaluationClaude SonnetExpertLegal reasoning, research, system design, math proofsClaude Opus Shell 1,000 requests/day — mixed complexity Without routing: all → Opus ($5/1M input) 1,000 × 500 tokens = 500K tokens × $5 = $2.50/day With routing: 70% Haiku, 20% Sonnet, 10% Opus 700 × 500 × $1 + 200 × 500 × $3 + 100 × 500 × $5 = $1.05/day Savings: 58% reduction Technique 3: Prompt Optimization, 5% to 20% Off Token Count The Problem Prompts written by humans, especially in collaborative or enterprise settings, accumulate filler. Phrases like "please note that", "it is important to note that", "in order to", and "due to the fact that" add tokens without adding meaning. At scale, this is a measurable cost. What the Library Strips Python from llm_optimizer import PromptOptimizer opt = PromptOptimizer() result = opt.optimize(""" In order to complete this task, please note that you should carefully analyze the following text. It is important to note that accuracy matters. Please be aware that your response should be concise. """) print(result.optimized_text) # "To complete this task, carefully analyze the following text. # Accuracy matters. Your response should be concise." print(f"Saved {result.tokens_saved} tokens ({result.savings_pct}%)") # Saved 18 tokens (31%) What is never touched: Code blocks, factual content, user-specified phrasing. The optimizer is conservative by default. It only removes patterns with no semantic value. Conversation history trimming: In long conversations, the library keeps the last N turns and drops older context, preventing unbounded token grow. Technique 4: Batch Processing, 50% Off Non-Urgent Requests The Problem Not every LLM call needs an immediate response. Nightly report generation, document indexing, data enrichment pipelines, and offline classification jobs all of these run fine with a delay. But most teams send them as real-time requests anyway, paying full price. How Anthropic's Batch API Works Anthropic's Message Batch API processes up to 10,000 requests per batch at 50% of the normal price. Results are available within minutes to hours. The trade-off is explicit: cost for latency. Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_batching=True, ) # Queue 1,000 document summaries throughout the day for doc in documents: client.queue( custom_id=doc["id"], messages=[{"role": "user", "content": f"Summarize: {doc['text']}"}], max_tokens=200, ) # Submit as one batch — 50% cheaper than 1,000 individual calls batch_id = client.submit_batch() # Poll when ready — minutes to hours depending on load results = client.poll_batch(batch_id, wait=True) for r in results: print(f"{r.custom_id}: {r.content}") When to Use It Nightly data processing pipelinesDocument indexing and enrichment Offline classification and tagging Report generation Technique 5: Document Compression to Reduce Context Before Sending The Honest Tradeoff This technique requires a direct warning: document compression is lossy. Removing content from a document to reduce token count means the model works with less information. For some tasks this is fine; for others it produces wrong answers. Use it only when: You've verified empirically that compression doesn't hurt your answer quality.You're doing rough extraction where completeness isn't required.You have re-ranking downstream (e.g., RAG pipelines). Do not use it for legal documents, compliance reviews, or any task where every sentence may be relevant. TF-IDF Extractive Compression When you do use compression, naive truncation (cutting from the end) is the worst strategy. llm-optimizer implements TF-IDF paragraph scoring where each paragraph is scored by its term overlap with your query, weighted by how unique those terms are across the document. The most relevant paragraphs fill the token budget; the rest are dropped. Python from llm_optimizer import DocumentCompressor # ⚠️ Read the accuracy warning before using in production comp = DocumentCompressor( max_tokens=4000, strategy="extractive", # TF-IDF scoring — best accuracy ) compressed, tokens_saved = comp.compress( document=long_contract, # 50,000 tokens query="payment terms and termination clauses" # focus compression here ) print(f"Compressed to {4000} tokens, saved {tokens_saved} tokens") # All compressed output includes a visible [⚠️ COMPRESSION WARNING] marker There are three strategies available: Strategy options: StrategyHow it worksbest forExtractiveTF-IDF scoring against queryWhen you have a specific querySmartKeeps first 60% + last 20%Structured documents with summariesTruncateHard cutoffWhen you need predictable behavior Technique 6: Cost Tracking That Observes Before You Optimize Why This Matters You can't optimize what you don't measure. Before applying any of the above techniques, you need to know: Which models you're actually usingWhere your token spend is goingWhether your optimizations are working Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), persist_tracking="usage.jsonl", # survives restarts ) # ... run your application ... client.print_summary() # ═══════════════════════════════════════════════════════ # LLM Cost Optimizer — Usage Summary # ═══════════════════════════════════════════════════════ # Total Requests : 1,247 # Total Cost : $0.8432 # Total Saved : $7.2180 (89.5% savings) # Cached Tokens : 8,432,000 # By Model : haiku: 891 reqs ($0.12) | sonnet: 312 ($0.58) # Optimizations : prompt_caching: 1247x | model_routing: 1247x # ═══════════════════════════════════════════════════════ The tracker records every request's tokens, cost, cached tokens, savings, latency, and which optimizations fired. Data persists to JSONL so you can analyze it across sessions or pipe it to your observability stack. Architecture: Why Not Just Use LiteLLM? The obvious question. LiteLLM is excellent and covers a lot of ground, including unified provider API, routing, cost tracking, batch processing. If you're not already using it, you should evaluate it. llm-optimizer does three things LiteLLM doesn't: Automatic cache_control injection: LiteLLM passes caching headers through but doesn't inject breakpoints at optimal positions automatically.Prompt filler stripping: LiteLLM has no token-level prompt optimization.TF-IDF document compression: LiteLLM has no query-aware document compression. The intended use is actually as a complement: llm-optimizer can sit on top of a LiteLLM setup, handling the prompt-level optimizations that LiteLLM doesn't touch. Streaming Support For user-facing applications, the library supports streaming: Python with client.stream( messages=[{"role": "user", "content": "Explain quantum entanglement"}], system="You are a physics tutor.", max_tokens=512, ) as stream: for chunk in stream: print(chunk, end="", flush=True) # Access token usage after stream completes usage = stream.usage() All optimizations, including caching, routing, prompt, and optimization, apply identically to streaming requests. Error Handling Production LLM applications need to handle rate limits and model overloads gracefully. The library handles this automatically: Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), max_retries=3, # retry on rate limit with exponential backoff retry_base_delay=1.0, # 1s, 2s, 4s ) # Rate limit (429) → retried with backoff # Model overloaded (529) → falls back to next capable model automatically # Non-retriable error → raises immediately Limitations Honest about what this doesn't do yet: Not org-scale validated: v0.4.0 is tested against ai-core Bedrock and Anthropic direct (14 live tests, all passing). Not yet run against production workloads at team scale. The pilot measures this.No async support: complete() and stream() are synchronous. Async support planned for a future release.OpenAI and Google partially tested: Anthropic and AWS Bedrock are the validated providers. OpenAI is implemented but not end-to-end tested in CI. Google Gemini streaming is not yet implemented.Token counting is approximate: The default estimator is within ~20% of the actual count. Install tiktoken for exact counts: pip install llm-optimizer[tiktoken].Pricing data can go stale: Stored in pricing.json with a version stamp. The library warns automatically if data is older than 30 days.Model allowlist: Only models listed in pricing.json can be routed to. Mythos and Fable 5 are not in the registry and cannot be called. Adding a new model requires a deliberate update to pricing.json. Python # Basic install pip install llm-optimizer # With exact token counting pip install llm-optimizer[tiktoken] # All providers pip install llm-optimizer[all] Python import anthropic from llm_optimizer import OptimizedClient client = OptimizedClient( anthropic_client=anthropic.Anthropic(), # All optimizations on by default except compression (lossy — opt-in) ) response = client.complete( messages=[{"role": "user", "content": "Your prompt here"}], system="Your system prompt here", ) client.print_summary() Links: PyPI: https://pypi.org/project/llm-optimizeGitHub: https://github.com/banerjeeso/llm-optimiz What's Next Async support (acomplete(), astream())LiteLLM adapterBudget guard: raise before a request exceeds a cost thresholdReal production benchmarks once I've run this against a live workload. Feedback welcome, especially from anyone who works with LLM APIs in production and can stress-test the routing logic or compression accuracy. Published on PyPI as llm-optimizer. MIT license. Contributions welcome. More
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
By Abhishek Gupta DZone Core CORE
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
By Naveen Kalapala
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
By Peter Ndumia
How to Perform Response Verification in REST-Assured Java for API Testing: Part 2
How to Perform Response Verification in REST-Assured Java for API Testing: Part 2

API testing is an essential part of modern software development. While sending requests and receiving responses is straightforward, the real value of API automation comes from response verification. A test is meaningful only when it validates that the API returns the correct data, structure, status codes, and business rules. In Java-based API automation, REST Assured combined with Hamcrest Matchers provides a clean and expressive way to verify API responses. These matchers help testers write readable assertions that validate numbers, strings, arrays, JSON objects, and collections with minimal code. This tutorial explains how to perform response verification in REST Assured using the following Hamcrest Matchers: NumericStringCollectionsJSON Object validationsNegative validation By the end of this article, you will be able to write powerful and maintainable API assertions in your automation tests. If you have not checked, click here to read Part 1 of this blog post. What Is Response Verification in API Testing? Response verification is the process of validating the API response returned from the server. This includes checking status codes, response body values, JSON structure, headers, data types, arrays, objects, and business validations. The verification includes checking: Does the API response return a 200 OK status code?Does the response contain the expected value for the fields?Is the list size greater than zero?Does every object contain a specific key? Without assertions, an API test is just sending requests and receiving responses without actually checking whether the API behaves correctly. How to Use Hamcrest Matchers With Rest-Assured for Response Verification in REST-Assured Java Hamcrest Matchers improve readability and make assertions more expressive. To use Hamcrest, the following dependency should be added to the pom.xml in the Maven project: XML <dependency> <groupId>org.hamcrest</groupId> <artifactId>hamcrest</artifactId> <version>3.0</version> <scope>test</scope> </dependency> Numeric Matchers In this section, we’ll learn to use numeric matchers in Rest-Assured tests, including greaterThan (), greaterThanOrEqualTo(), lessThan(), and lessThanOrEqualTo(). These assertions help in validating numerical values returned in API responses. Using greaterThan() and greaterThanOrEqualTo() The greaterThan() matcher verifies that a numeric value is greater than the expected value. Similarly, the greaterThanOrEqualTo() matcher validates that the value is either greater than or equal to the expected number. Java @Test public void testGreaterThanAssertions () { given ().when () .get ("https://api.restful-api.dev/objects") .then () .statusCode (200) .and () .assertThat () .body ("[2].data['capacity GB']", greaterThan (500)) .body ("[5].data['price']", greaterThanOrEqualTo (120)); } In this test, the greaterThan () method from the Hamcrest library verifies that the capacity GB value in the third JSON object is greater than 500. The greaterThanOrEqualTo matcher checks whether the price value in the sixth object is 120 or more. These assertions help validate numerical values returned by the API without relying on exact matches. Numeric matchers are useful for testing values such as prices, counts, capacities, and response times. Using lessThan() and lessThanOrEqualTo() The lessThan() matcher validates that the value is below the expected number. Likewise, the lessThanOrEqualTo() matcher validates that the number is less than or equal to the expected value. Java @Test public void testLessThanAssertions () { given ().when () .log () .all () .get ("https://api.restful-api.dev/objects") .then () .log () .all () .statusCode (200) .and () .assertThat () .body ("[4].data['price']", lessThan (700f)) .body ("[6].data['year']", lessThanOrEqualTo (2019)); } In this test, the lessThan() method from the Hamcrest library verifies that the price value in the fifth JSON object is less than 700, while lessThanOrEqualTo() checks whether the year value in the seventh object is 2019 or lower. The value 700f is written with the “f” suffix because the API returns the price as a float, and using “f” ensures the expected value is also treated as a float during comparison. These assertions help ensure that the numerical values returned by the API remain within the expected limits. String Matchers In this section, we’ll learn to use String matchers in Rest-Assured tests, including equalToIgnoringCase(), containsString(), startsWith(), endsWith(), and equalToCompressingWhiteSpace(). These assertions are useful for validating text-based values returned in API responses. Java @Test public void testStringAssertion() { given ().when () .log () .all () .queryParam ("id", 3) .get ("https://api.restful-api.dev/objects") .then () .log () .all () .statusCode (200) .and () .assertThat () .body ("[0].name", equalTo ("Apple iPhone 12 Pro Max")) .body ("[0].name", equalToIgnoringCase ("ApPLE IPhone 12 pro MAX")) .body ("[0].data.color", containsString ("White")) .body ("[0].name", startsWith ("A")) .body ("[0].name", endsWith ("x")) .body ("[0].name", equalToCompressingWhiteSpace (" Apple iPhone 12 Pro Max ")); } The testStringAssertion() method demonstrates different ways to validate string values in an API response using REST Assured and Hamcrest matchers: body(“[0].name”, equalTo (“Apple iPhone 12 Pro Max”)): Verifies that the name field exactly matches the expected string, including the letter casing and spaces.body(“[0].name”, equalToIgnoringCase(“ApPLE IPhone 12 pro MAX”)): Validates the string value while ignoring differences in uppercase and lowercase characters.body(“[0].data.color”, containsString(“White”)): Verifies whether the color field contains the text White anywhere within the string.body(“[0].name”, startsWith(“A”)): Verifies that the name field begins with the letter “A”.body(“[0].name”, endsWith(“x”)): Validates that the name field ends with the letter “x”.body(“[0].name”, equalToCompressingWhiteSpace(“ Apple iPhone 12 Pro Max ”)): Compares the string values after removing extra spaces and compressing multiple whitespaces into a single space, making the assertion more flexible for formatting differences. These matchers help verify exact text, partial text, prefixes, suffixes, case sensitivity, and whitespace formatting. Collection Matchers In this section, we’ll learn to use collection matchers in Rest-Assured tests, including hasSize(), hasItem(), hasKey(), and everyItem(hasKey()). These assertions help in validating arrays and collections returned in API responses, such as verifying the number of items, checking for specific values, and ensuring required keys are present. Using hasSize() and hasItem() matchers Java @Test public void testHasSizeAndHasItem () { given ().when () .queryParam ("id", 3) .queryParam ("id", 5) .get ("https://api.restful-api.dev/objects") .then () .statusCode (200) .and () .assertThat () .body ("$", hasSize (2)) .body ("name", hasItem ("Apple iPhone 12 Pro Max")); } The testHasSizeAndHasItem() method demonstrates how to validate collections and arrays returned in the API response using Hamcrest matchers in REST Assured. It uses the hasSize() and hasItem() methods from the Hamcrest matchers for verifying the size of the response collection and whether specific items exist within it. body(“$”, hasSize(2)): The hasSize() matcher verifies that the response array contains exactly “2” objects. As the request is sent with two query params (id=3 and id=5), the API is expected to return two matching records.body(“name”, hasItem(“Apple iPhone 12 Pro Max”)): The hasItem() matcher checks whether the name collection in the response contains the value “Apple iPhone 12 Pro Max”. This assertion helps in validating that a specific item exists within the returned response. Using hasKey(), and everyItem(hasKey()) matchers Java @Test public void testHasKeyAssertions () { given ().when () .log () .all () .queryParam ("id", 3) .get ("https://api.restful-api.dev/objects") .then () .log () .all () .statusCode (200) .and () .assertThat () .body ("$", everyItem (hasKey ("id"))) .body ("[0].data", hasKey ("capacity GB")) .body ("$", everyItem (hasKey ("name"))); } The testHasKeyAssertions() method shows how to validate the presence of keys in the JSON objects returned by an API response. The hasKey() matcher is commonly used to ensure that the required fields are present in the response. body(“$”, everyItem(hasKey(“id”))): The everyItem(hasKey()) assertion verifies that every object in the response array contains the “id” key. This helps ensure consistency across all returned objects.body(“[0].data”, hasKey(“capacity GB”)): The hasKey() matcher checks whether the data object of the first response item contains the key “capacity GB”. This assertion validates the presence of a specific field inside a nested JSON object.body(“$”, everyItem(hasKey(“name”))): This assertion verifies that all objects in the response array contain the name key. It ensures that the expected field is available in every returned record in the API response. Negative Validations Negative validation in Rest-Assured is commonly performed using the not() negation matcher from Hamcrest to verify that an API response does not contain certain values or conditions. Using the not() matcher, the condition can be inverted, and accordingly, the assertion validates that the specified value or condition is not present in the API response. Java @Test public void testNotAssertions () { given ().when () .log () .all () .queryParam ("id", 3) .get ("https://api.restful-api.dev/objects") .then () .log () .all () .statusCode (200) .and () .assertThat () .body ("$", not (emptyArray ())) .body ("[0].id", notNullValue ()) .body ("[0].name", not (equalTo ("Samsung"))) .body ("[0].data['capacity GB']", not (greaterThan (550))); } The testNotAssertions() method demonstrates how to perform negative validations in Rest-Assured using the not() matcher and related assertions. body(“$”, not(emptyArray())): This assertion verifies that the response array is not empty and contains at least one object.body(“[0].id”, notNullValue()): The notNullValue() matcher verifies that the “id” field in the first response object is not null.body(“[0].name”, not (equalTO (“Samsung”))): This assertion validates that the name field is not equal to “Samsung”.body(“[0].data[‘capacity GB’]”, not(greaterThan(550))): The not(greaterThan()) assertion verifies that the “capacity GB” value is not greater than 550. This means that the value should be less than or equal to 550. Summary Response verification is what transforms an API test from simply sending requests into actually validating application behavior. In this tutorial, we explored how REST Assured and Hamcrest Matchers make assertions more readable and powerful by validating numbers, strings, arrays, JSON keys, and response structures. In my experience, learning these matchers significantly improves the quality and maintainability of API automation frameworks. Numeric, String, Collection, and Negative Matchers are especially useful in real-world testing because they help create validations that are both flexible and easy to understand, making debugging and test maintenance much simpler over time. Happy testing!!

By Faisal Khatri DZone Core CORE
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures

A few months ago, I watched a senior engineer spend forty-five minutes reviewing a single pull request — a PR that an AI assistant had generated in under two minutes. The code looked clean. The tests passed. But she kept cross-referencing an incident postmortem from eight months earlier, muttering something about retry amplification. She caught a real production risk. The AI reviewer had flagged nothing. That moment stuck with me. We've spent years optimizing how fast we can write code. But we haven't seriously reckoned with what happens when review can't keep up. The Bottleneck Has Shifted A single engineer with AI assistance can now produce hundreds of lines of code, large refactors, infrastructure changes, and test suites — all within minutes. Review complexity, however, grows exponentially with change size and system interdependency. The core problem is no longer "Can AI write code?" It's "Can humans reliably validate what AI wrote?" Code generation speed increases. Human cognitive review capacity stays flat. That imbalance is quietly accumulating risk in engineering organizations everywhere. Why Current AI Reviewers Fall Short Most AI PR review systems today operate on static diffs, syntax-level reasoning, and shallow best-practice detection. They produce comments like: "Potential null pointer.""Consider renaming this variable.""Possible optimization opportunity." Occasionally useful. Rarely sufficient for production-critical systems. The structural problem is that these tools treat PR review as a language problem instead of a systems reasoning problem. They assume software correctness is inferable from local code semantics alone. In reality, production safety emerges from interactions between architecture, runtime behavior, operational history, and organizational context. The Shallow Review Problem in Practice Here's a concrete example. An AI assistant generates this database query optimization: Python # AI-optimized version def get_user_orders(user_id): return db.query(""" SELECT o.*, p.*, i.* FROM orders o JOIN payments p ON o.id = p.order_id JOIN items i ON o.id = i.order_id WHERE o.user_id = ? """, user_id) Typical AI reviewer comment: "Query optimized with JOIN to reduce round trips." What a senior engineer sees: "This will cause a Cartesian explosion. The orders table has 50M rows, items averages 8 per order. This returns 400M+ rows for power users. We had a nearly identical incident (INC-287) that took down the read replica. Needs pagination and selective columns." The difference isn't token count or model size. It's operational memory and causal reasoning. The Real Challenge Is Not Context Windows Many people assume the fix is larger context windows. Feed the model the whole repo, and it'll review like a senior engineer. But experienced engineers don't review code by loading entire systems into working memory. They use abstraction, selective attention, and compressed mental models. A senior engineer reviewing a Kafka retry change doesn't reread the entire messaging subsystem — they remember prior incidents, retry amplification risks, and historical outages. That's cognitive compression, not token recall. Modern LLMs are exceptional at syntax fluency, pattern completion, and probabilistic association — what you might call token intelligence. But effective PR review requires something deeper: causal reasoning, architectural abstraction, operational memory, risk forecasting. Call it cognitive intelligence — persistent contextual reasoning grounded in operational history and causality. The distinction matters because it changes what we need to build. What a Cognitive Review Architecture Looks Like Instead of: Plain Text Large Prompt + Large LLM → Review We need: Plain Text Structured Memory + Semantic Retrieval + Runtime Context + Specialized Review Agents + Reasoning Layer + LLM → Review The LLM should not be the memory. It should be the reasoning interface over structured engineering knowledge. Intent Reconstruction Before reviewing code, the system needs to understand why the change exists. Business intent, bug root cause, architectural motivation. Inputs include Jira tickets, PR descriptions, ADRs, incident reports, and commit timelines. Without intent, review quality stays shallow regardless of model size. Engineering Knowledge Graphs Human reviewers carry organizational memory: fragile services, latency-sensitive paths, scaling bottlenecks, previous outages, dangerous dependencies. AI reviewers need persistent semantic memory systems encoding the same — service relationships, API contracts, operational metadata, incident history, ownership boundaries. This creates an engineering cognition layer far richer than raw repository context. Multi-Agent Review Systems A single reviewer model is insufficient. Future systems will consist of specialized agents working together: Architecture Reviewer – dependency boundaries, coupling risk, architectural driftReliability Reviewer – retries, backpressure, idempotency, failover behaviorSecurity Reviewer – injection risks, auth issues, secret exposurePerformance Reviewer – memory growth, query amplification, scaling regressionsHistorical Regression Reviewer – correlation with past outages, postmortems, incident fingerprints This begins to approximate how experienced engineering organizations actually review software. Runtime-Aware Review Static analysis alone misses emergent runtime behavior. Future cognitive review systems will integrate observability telemetry, tracing data, production metrics, and traffic patterns. Compare these two responses to a retry configuration change: Traditional AI reviewer: "Code follows retry best practices." Cognitive AI reviewer with operational memory: "HIGH RISK: Similar retry configuration caused incident on 2023-09-15. This service processes 2M messages/hour at peak. 10 retries with exponential backoff = up to 17 minutes per message. Previous incident resulted in 8M message consumer lag and cascading downstream failures. Recommend: max 3 retries, circuit breaker, dead letter queue, idempotency check before db.save(). See ADR-089." That is a fundamentally different class of intelligence — and a fundamentally different class of safety. Engineering Memory Is the Missing Piece One of the biggest gaps in current AI systems is durable operational memory. Experienced engineers develop intuition through outages, failed deployments, debugging sessions, and production emergencies. These experiences become compressed heuristics: "This retry increase feels dangerous" — not because of syntax, but because of remembered causal relationships. Replicating this requires episodic memory systems, incident-aware reasoning, and causal knowledge graphs. Much of this mirrors practices long established in Site Reliability Engineering, where institutional learning from incidents is treated as critical infrastructure. Incident postmortems aren't just documentation — they're organizational immune system responses. Getting AI systems to genuinely learn from incidents rather than just pattern-match against them remains one of the harder open problems in this space. What Teams Can Do Today Fully cognitive review systems don't exist yet. But organizations can meaningfully improve AI-assisted review quality right now: Capture architectural knowledge in machine-readable form. Service boundaries, retry policies, timeout configurations, scaling assumptions — not just in wikis, but in structured formats AI systems can query.Link PRs explicitly to incident history. Build connections between code changes and the incidents they caused or prevented. This is organizational memory that AI systems can leverage today.Tag services with operational metadata. Criticality tier, traffic patterns, known failure modes, blast radius. Treat repositories as systems, not just files.Integrate observability into review pipelines. Connect production metrics and tracing data to code review. Runtime context dramatically improves review quality.Prioritize high-signal AI feedback. Review fatigue from noisy, low-signal comments is a real trust problem. Focus AI comments on incident-correlated patterns, architectural violations, and operational risks. The Trust Calibration Problem One concern I keep coming back to: bad AI reviewers are dangerous not because they miss things, but because they sound confident while missing things. They reduce human vigilance through automation bias. They generate fatigue through noise. They normalize shallow approval. Future cognitive review systems need to be not just more accurate, but properly calibrated — knowing when they lack sufficient context and escalating accordingly. An AI reviewer should be able to say: "I may not have enough confidence to validate this safely." That self-awareness may matter more than raw capability. The Road Ahead The next era of AI software engineering will not be defined by who generates the most code. It will be defined by trust, reasoning quality, and operational awareness. The future belongs to systems capable of understanding not just what changed — but why it changed, what it affects, and whether it's safe. That's the difference between code generation and engineering intelligence. And honestly, solving it seems harder and more interesting than anything we've built so far. Key Takeaways The bottleneck has shifted from code generation to code review and validation.Larger context windows alone won't bridge token intelligence and cognitive intelligence.Human-like review requires structured memory, causal reasoning, and operational awareness.Multi-agent architectures with specialized reviewers mirror how engineering teams actually work.Runtime-aware systems integrating production telemetry represent the next frontier.Engineering memory — learning from incidents — is critical for trust and safety.Teams can start today by capturing architectural knowledge and linking incidents to code changes. References Vaswani, A., et al. (2017). "Attention Is All You Need." NeurIPS.Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.Lewis, P., et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS.Shinn, N., et al. (2023). "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv.Beyer, B., et al. (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media.Allspaw, J. (2012). "Blameless PostMortems and a Just Culture." Etsy Engineering.

By Sayan Chatterjee
How to Build a Production-Ready iOS App With AI-Generated Code
How to Build a Production-Ready iOS App With AI-Generated Code

Vibe coding has compressed the distance between an idea and a runnable application. Natural-language instructions can now produce SwiftUI screens, networking code, persistence, authentication flows, and deployment configuration with very little manual typing. That acceleration changes the bottleneck rather than removing it. A build that launches successfully is not evidence that the application handles hostile inputs, unreliable networks, concurrency boundaries, credential storage, production failures, or future changes safely. Recent research makes the distinction concrete. A June 2026 preprint studying 200 deployed applications sampled from 10,517 open-source vibe-coded projects reported 1,471 manually validated vulnerabilities, including broken access control, cryptographic failures, injection, and secret exposure. A separate 2025 benchmark found a large gap between functional correctness and security in agent-generated solutions. These studies are not specific to iOS, but they reinforce a useful engineering principle that generated code still requires independent verification. The Real Production Boundary The most dangerous property of weak generated code is often that it looks ordinary. Consider a networking fragment that compiles, returns data on a healthy connection, and decodes the expected payload: Swift let (data, _) = try await URLSession.shared.data(from: endpoint) return try JSONDecoder().decode(Profile.self, from: data) The missing response handling is easy to overlook. URLSession exposes HTTP metadata through HTTPURLResponse, and server-side failures must be interpreted from that response rather than treated as transport failures. Apple explicitly advises inspecting the response for server-side errors. A production boundary should therefore make success semantics explicit: Swift let (data, response) = try await session.data(for: request) guard let http = response as? HTTPURLResponse, 200..<300 ~= http.statusCode else { throw APIError.unexpectedResponse } return try decoder.decode(Profile.self, from: data) That correction is small, but production hardening goes further. Request timeouts require deliberate policy, connectivity can change while a request is active, and retries must respect HTTP semantics. Apple provides waitsForConnectivity so a session can wait for viable connectivity instead of failing immediately, while RFC 9110 defines idempotency as the property that makes automatic repetition safe in the intended server effect. Blindly retrying a purchase or account-creation POST can therefore be materially different from retrying an idempotent operation. Security Cannot Be Inferred From a Successful Login Authentication is another area where generated implementations can satisfy the visible requirement while violating the operational one. A token persisted like this remains syntactically valid: Swift UserDefaults.standard.set(accessToken, forKey: "access_token") For sensitive credentials, that storage choice should fail a production-readiness check. Apple describes Keychain Services as storage for passwords, keys, certificates, identities, and authentication tokens, while OWASP MASVS requires sensitive local data to be stored securely and protected from leakage. The safer implementation places credential persistence behind a dedicated abstraction backed by Keychain: Swift try credentialStore.save( accessToken, account: "session.access-token" ) The abstraction matters because static checking can then enforce a project invariant as credential-bearing values must not flow directly into UserDefaults, logs, or ad hoc files. Network configuration needs similar scrutiny. App Transport Security is enabled for URLSession connections and is designed to improve privacy and data integrity through secure transport requirements, as broad ATS exceptions should therefore be treated as review findings, not convenient defaults. Logging deserves the same treatment. Production diagnostics need context, but authentication tokens, personal data, and identifiers should not become unrestricted log payloads. Apple’s logging APIs provide privacy controls specifically because generated log messages can be accessible beyond the immediate code path. A readiness gate can reject obvious secret logging patterns while still allowing structured, privacy-aware operational telemetry. Turning Review Knowledge Into Executable Rules A useful production-readiness gate should convert engineering expectations into checks that run on every change. SwiftSyntax is well suited to syntax-level rules because SwiftLint itself uses SwiftSyntax for most of its rules, while type-sensitive analyzer rules can rely on deeper compiler information. That distinction is important: syntax analysis can identify suspicious patterns, but it cannot prove application behavior. A concise SwiftSyntax rule can flag direct token persistence: Swift override func visit( _ node: FunctionCallExprSyntax ) -> SyntaxVisitorContinueKind { let call = node.calledExpression.description if call.contains("UserDefaults.standard.set"), node.arguments.description .localizedCaseInsensitiveContains("token") { findings.append(.insecureCredentialStorage) } return .visitChildren } The same mechanism can flag URLSession.shared calls inside SwiftUI view declarations, force operations in production targets, oversized view bodies, or unstructured Task creation in lifecycle-sensitive code. These should be findings with severity and location, not claims of certainty. SwiftLint’s own design illustrates why most rules operate from syntax, while analyzer rules exist separately when type information is required. Compiler checks should complement those heuristics. Swift 6 strengthens data-race safety through actor isolation and Sendable checking, and Apple’s migration guidance explicitly calls out MainActor and Sendable audits. A production CI job should therefore compile with the intended Swift language mode and strict concurrency settings rather than attempting to reproduce concurrency correctness with custom pattern matching. A Gate That Measures Evidence, Not Polish Static analysis catches source-level risks, but production readiness also depends on evidence from tests and runtime diagnostics. Xcode can collect code coverage through test plans, and command-line test execution produces .xcresult bundles containing results and coverage data. Coverage alone should not become a release score; critical behaviors matter more than a single percentage. Authentication refresh, corrupted responses, offline startup, cancellation, persistence migration, and destructive operations should have explicit automated tests. Operational readiness begins after those tests pass. MetricKit provides real-user metrics and diagnostics, including launch behavior, responsiveness, crashes, hangs, disk writes, and memory-related termination information. A vibe-coded application that catches errors with print(error) has almost no diagnostic value once failures occur outside a development machine. Structured logging, crash diagnostics, release identifiers, request correlation, and privacy-safe error classification turn an unknown failure into an actionable production signal. The final gate can combine these sources without pretending that a single score proves safety. A failed secret-storage rule can block release outright. Strict-concurrency compiler failures can block release. Missing tests around critical flows can block release. Lower-severity architecture findings can remain warnings requiring review. This approach resembles a policy engine more than a linter, as the purpose is not stylistic consistency, but repeatable evidence that generated code satisfies agreed production invariants. Apple’s App Review guidance reinforces the practical value of that discipline as Apple reports that more than 40% of unresolved review issues are associated with App Completeness, including crashes, placeholder content, and incomplete information. Production Readiness Is a Verification Problem Vibe coding can make software generation dramatically faster, but production software still has to survive conditions that a successful demo does not exercise. The reliable response is not to reject generated code, nor to trust it because it compiles. The stronger model is to place an executable verification boundary between generation and release. On iOS, that boundary can combine SwiftSyntax rules, compiler-enforced concurrency checks, secure-storage and transport policies, automated tests, and production diagnostics. The result is a development process in which AI-generated code is treated like any other untrusted change that's useful immediately, releasable only after independent evidence demonstrates that the required engineering invariants hold.

By Uthej Mopathi DZone Core CORE
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team

Enterprise platforms don't fail gracefully. A stalled integration, a malformed payload, a service that silently drops a field — these aren't abstract bugs; they're business processes that stop moving and stakeholders who start calling. Leading a multinational engineering team across two regions that keeps a large platform's web service integrations running, I've spent the last several months evaluating where an agentic coding tool like Claude Code actually earns its place in that world — not as a novelty, but as infrastructure. Here's what that looks like in practice. Development: From Description to Diff The obvious use case is still the most valuable one. Claude Code reads an entire codebase — including the mix of legacy and modern service layers that most enterprise platforms actually run on — and can trace a bug from a symptom description down to the root cause across files it wasn't explicitly pointed at. Describe the integration you want built, and it plans an approach, writes the code, and verifies its own work before handing it back. For a legacy-adjacent stack paired with modern services, this matters more than it would in a greenfield project. Institutional knowledge about why an integration was built a certain way often lives in people's heads, not in comments. A CLAUDE.md file at the project root — read at the start of every session — becomes a place to encode that: architecture decisions, naming conventions, which endpoints are safe to retry and which aren't. Claude also builds its own memory as it works, picking up build commands and debugging patterns without being told to. Support: Turning Log Noise Into Signal Enterprise integration incidents rarely announce themselves cleanly. They show up as a batch job that partially completed, or a spike in errors from a downstream system. Claude Code's headless mode (claude -p) lets you pipe raw log output straight into a triage step: Shell tail -200 integration.log | claude -p "flag anything that looks like a failed payload" That same pattern scripts into runbooks, scheduled checks, or CI jobs — useful when your support rotation spans time zones, and nobody wants to be the one manually grepping logs at 2 a.m. local time. The Distributed Team Problem Leading engineers across multiple regions means standards drift is a constant risk — not from lack of skill, but from lack of shared context. Two teams solving the same class of integration problem several time zones apart will converge on different patterns unless something forces alignment. This is where Claude Code's team-facing features do real work: Skills package a repeatable workflow — a /review-integration command that checks a new service against your platform's conventions — so the process isn't re-explained every time it's needed.GitHub Actions or GitLab CI/CD integration lets Claude respond to @claude mentions on a pull request or run automated code review with severity-tagged findings, which matters when reviewers in one region are asleep while another region is shipping.Managed settings push CLAUDE.md files, permission rules, and MCP server allowlists centrally, so a Tech Lead sets the guardrails once instead of hoping every developer's local config agrees. MCP: The Part That Matters Most for Integration Work If your job is service integrations, the Model Context Protocol (MCP) is arguably more relevant than the coding assistance itself. MCP lets Claude Code connect directly to the systems around your codebase — ticketing tools, internal databases, Slack — so it can pull incident context or check a downstream system's schema as part of diagnosing a fix, instead of an engineer manually assembling that context first. For a platform with a web of service dependencies, that's the difference between "here's a plausible fix" and "here's a fix that accounts for what the downstream service actually expects today." Governance Isn't Optional Here Enterprise data often brings compliance weight that a typical greenfield codebase doesn't carry. Before rolling this out past a pilot, the governance layer needs real attention: Permission modes and sandboxing control what Claude can execute autonomously versus what needs a human in the loop — worth tuning tighter than the defaults for anything touching production data.Data handling and Zero Data Retention options are worth understanding fully if your platform has regulatory constraints on where data can transit.Analytics (on Team/Enterprise plans) gives adoption and PR-attribution data, which is useful less for surveillance and more for showing leadership that a pilot is actually paying off before expanding it. Where to Start Don't roll this out platform-wide on day one. Pick a single integration workstream, write the CLAUDE.md that captures your legacy-plus-modern conventions, and run a two-week pilot with one automated review pipeline (GitHub Actions or GitLab CI/CD, whichever matches where your repos already live). Measure whether incident triage time drops and whether code review turnaround improves across the time-zone gap. Expand from there. The tooling is capable enough now that the harder problem isn't whether it can help — it's building the guardrails and shared conventions that let a distributed enterprise team trust it with production-critical integrations.

By Balaji Venkatasubramaniyar DZone Core CORE
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest

A mobile request can fail without the server-side work failing. An iOS app may time out, lose the response after a POST has reached the service, or retry after connectivity changes while the original execution is still progressing. Apple explicitly distinguishes safe retry behavior by HTTP method and notes that URLSession can retry requests in some connection-loss cases, waitsForConnectivity can also cause the system to continue a request when connectivity returns. The dangerous state is therefore not “request failed,” but “completion is unknown.” If that request starts an agent that charges an account, reserves inventory, sends a message, or invokes an MCP tool, a second submission can become a second side effect. The Retry Boundary Is the Real Transaction Boundary “Exactly once” is too strong for a workflow crossing an iPhone, HTTP, an agent runtime, an MCP server, Kafka, a database, and an external API. Kafka can provide exactly-once guarantees within defined Kafka processing boundaries, but those guarantees do not atomically include arbitrary remote tool effects. The practical target is effectively-once behavior, and retries are expected, but every effect is guarded by a stable operation identity and converges on one committed outcome. Kafka’s idempotent producer suppresses duplicate records caused by producer retries, while transactional producers can atomically publish across Kafka partitions; the producer documentation also limits idempotence guarantees to a producer session and requires read_committed consumers for end-to-end transactional visibility. The operation identity must exist before the first network attempt. An iOS client can create an operationId when an action becomes durable local intent, persist it, and reuse it across transport retries. Transport material such as a server challenge may change, but the business ID must not. The server treats (subjectId, operationId) as a uniqueness boundary and stores a canonical payload hash with it. PostgreSQL unique constraints enforce row uniqueness, while INSERT ... ON CONFLICT provides an atomic conflict path under concurrency. SQL INSERT INTO agent_operation(subject_id, operation_id, payload_hash, status) VALUES (:subject, :operationId, :payloadHash, 'ACCEPTED') ON CONFLICT (subject_id, operation_id) DO NOTHING; A conflict with the same payload hash returns the existing operation; a different hash rejects key reuse. The record should exist before agent execution starts, and the accepted response should expose the durable operation identity. Let LangGraph Resume Without Repeating Effects LangGraph persistence is useful precisely because durable execution can replay code. With a checkpointer, LangGraph saves state at super-step boundaries; if execution resumes after a failure, an affected node can run again from the beginning. Official guidance consequently requires idempotent node logic, and task results can be checkpointed so completed task work can be reused during resume instead of recomputed. Replaying from an earlier checkpoint can also re-trigger later LLM calls and API requests. A stable business operation should therefore map to a stable LangGraph thread, while every effectful tool boundary receives the same operation ID. Python config = {"configurable": {"thread_id": operation_id} result = graph.invoke( {"operation_id": operation_id, "command": command}, config ) Checkpointing reduces recomputation but does not replace downstream idempotency. A reservation can succeed before its task result is durably checkpointed. LangGraph’s functional API therefore recommends idempotent tasks because an incomplete task can execute again during resume. Python @task def reserve_inventory(operation_id, sku, quantity): return mcp.call_tool("reserve_inventory", { "operationId": operation_id, "sku": sku, "quantity": quantity }) The significant property in this snippet is not the decorator. The important part is that the business identity crosses the graph boundary and reaches the tool implementation. A downstream inventory service can then use that identity to return a previously committed reservation rather than creating another one. MCP Tasks Are Durable Handles, Not Deduplication Keys The current MCP Tasks design is especially relevant to long-running agent tools. In the July 28, 2026 protocol revision, Tasks moved into the io.modelcontextprotocol/tasks extension. A server can return a durable task handle, and the client can poll with tasks/get, provide input with tasks/update, or request cancellation with tasks/cancel. The task is durably created before its handle is returned, which allows polling after a disconnect. That durability solves result retrieval after task creation, but it does not by itself deduplicate the request that creates the task. The task ID is server-generated. If the server creates task A, the response disappears, and the original tools/call is sent again, a naïve implementation can create task B. Therefore, the business operationId must be part of the tool arguments or equivalent application metadata, and task creation must first look up an existing operation. This follows directly from MCP’s server-generated task-ID model combined with retry ambiguity at the HTTP boundary. The MCP server can return an existing task handle for the same authenticated subject, operation ID, and payload hash, and later return the stored terminal result. Cancellation should also be idempotent because MCP defines it as cooperative rather than a guarantee that underlying work stops immediately. Keep Kafka Guarantees Inside Kafka Kafka is most valuable after the operation has been claimed. A database transaction can persist operation state with an outbox row carrying the same ID. Kafka producer idempotence protects against duplicates caused by producer retries, while consumers can still use the operation ID for application-level deduplication. Kafka transactions can atomically cover Kafka writes, but they do not extend over an MCP server or payment API. The event contract should preserve causality rather than inventing a new identity at each hop. JSON { "operationId": "8E7B6D9E-...", "type": "AgentToolCompleted", "tool": "reserve_inventory", "status": "SUCCEEDED" } A consumer can enforce uniqueness on (consumerName, operationId, eventType) or make the state transition conditional. Kafka delivery guarantees and application idempotency then reinforce each other instead of being treated as interchangeable. Bind Retry Identity to App Attest Without Blocking Legitimate Retries App Attest addresses a different failure mode: whether a request comes from a legitimate app instance and whether signed request material has been replayed or altered. Apple’s current guidance uses a server-provided challenge for assertions and requires the server to validate a strictly increasing assertion counter; that counter is specifically an anti-replay signal. Assertions are generated locally on the device after key attestation. The App Attest assertion must not become the business idempotency token. A legitimate retry should obtain fresh challenge material and generate a fresh assertion while retaining the original operation ID. The data hashed for the assertion can bind the server challenge, operation ID, and canonical payload hash together. Swift let payloadHash = SHA256.hash(data: body) let clientData = challenge + operationID.data + Data(payloadHash) let clientDataHash = Data(SHA256.hash(data: clientData)) let assertion = try await service.generateAssertion( keyID, clientDataHash: clientDataHash ) Apple recommends server-controlled challenges, server-side validation, and assertion-counter tracking as assertions are generated on demand without a round trip to Apple’s servers. The server verifies App Attest, checks that the challenge binds the operation ID and payload, then performs the idempotency lookup. A fresh assertion can retry the same operation; a replayed assertion fails anti-replay validation; an altered payload fails the hash check. Effectively-Once Behavior Is a Composition Property Reliable agent execution does not come from asking iOS to retry less often or from labeling a Kafka pipeline “exactly once.” It comes from carrying one durable business identity across every retry and every boundary, claiming that identity atomically before execution, making LangGraph effects idempotent under resume, using MCP Tasks as durable result handles rather than creation-time deduplication keys, restricting Kafka’s exactly-once guarantees to Kafka’s transactional domain, and using App Attest to prove request integrity without confusing anti-replay state with business deduplication. When those boundaries align, a lost mobile response can cause another HTTP attempt, another graph invocation, or another poll, but it does not cause another business effect. That is the operational meaning of effectively once.

By Uthej Mopathi DZone Core CORE
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?

Security teams usually describe an application through the assets they know about. This includes the production domain, documented APIs, the services currently in use, and the repositories connected to the latest release. But applications leave things behind as they change. A staging environment created for an old release may still be online months later, alongside an API version that was supposed to be retired. Other forgotten parts of the application can surface through DNS records, certificate data, or information left in client-side code. Most of these assets were created for perfectly legitimate reasons. Trouble starts when the work moves on, but the infrastructure doesn't. Ownership becomes unclear, security controls fall behind, and eventually an internet-facing component can remain active without appearing in the inventory used by the team responsible for it. I use unindexed attack surface to describe the gap between the application a team actively manages and the parts of it that are still reachable. Unindexed Does Not Mean Inaccessible It is easy to assume that an application resource is relatively safe when it doesn't appear in search results or isn't linked from the main application. That assumption falls apart once someone discovers the address. The same idea comes up when explaining the deep web. A large amount of online content sits outside conventional search indexes while remaining accessible through a direct URL, login, or other route. Application infrastructure can end up in a similar position. A staging host or old endpoint may be absent from the normal user journey and still be exposed to the internet. Finding these assets doesn't always require sophisticated techniques. During reconnaissance, an attacker can piece together clues from certificate records, DNS data, JavaScript, public repositories, and documentation. One discovery can lead to another until parts of the application that developers rarely think about become visible. robots.txt is a simple example. It tells compliant crawlers what they should avoid crawling, but it doesn't prevent someone from requesting those paths directly. OWASP's web security testing guidance includes reviewing web server metadata, identifying application entry points, and mapping execution paths during reconnaissance. Development and security teams can use similar techniques to see what their application exposes from the outside. Modern Delivery Creates Assets Faster Than Inventories Can Follow Modern development makes it easy to create infrastructure quickly. A pull request may generate a temporary preview environment, while a migration can leave /api/v1/ running as clients move to /api/v2/. During an incident, a debugging endpoint might be created and never removed afterward. Even a short-lived cloud experiment can leave behind a hostname outside the infrastructure account monitored by security. Keeping track of all of this gets harder as the application changes. Cloud platforms, API gateways, deployment configurations, and DNS may each hold a different piece of the inventory. Something created for one sprint can still be reachable several releases later. OWASP addresses this problem in API9:2023 Improper Inventory Management. Its guidance covers outdated API versions, exposed hosts, and missing documentation that can leave older parts of an application running without the security attention given to current services. Multiple development teams make that inventory harder to maintain. An environment may remain active after the developer who created it has moved to another project. Unless deployment and retirement update the inventory along with the infrastructure, these assets can stay online far longer than anyone intended. Forgotten Assets Often Retain Real Trust An old endpoint can outlive its original purpose without losing the access it was given. An earlier API version, for example, may still connect to the production database even though it hasn't received the authorization checks, rate limits, or input validation added to the current version. Staging environments can have the same problem when they use production-like data or continue running with an identity configuration that hasn't been reviewed in some time. Once an environment falls outside normal development work, security updates and monitoring are easier to miss. OWASP gives a useful example in its guidance on improper inventory management. A beta API host exposes the same password-reset capability as the production API, but the beta version lacks the rate limiting applied to production. Anyone who discovers the older host gets another route to the same function with fewer protections. The same situation can appear elsewhere in an application. A preview deployment might contain credentials left in an old build, while a forgotten administrative interface could still be reachable through a load balancer. These assets become even harder to manage when logging is no longer checked, or alerts still point to a team that has stopped owning the service. The longer an asset sits outside normal development and security workflows, the easier it is for its permissions, dependencies, and controls to fall behind the rest of the application. The Frontend Can Reveal the Backend’s Missing Map The browser often reveals more about an application than teams realize. It needs enough information to communicate with backend services, so production JavaScript can include API URLs, route names, environment identifiers, feature flags, and references to functionality users no longer see. This becomes interesting when the frontend has moved on, but the backend hasn't. A feature may disappear from the interface while its endpoint continues to respond. Commenting out a button or removing a route from the visible application doesn't remove the server-side functionality behind it. Source maps can make this easier to investigate. They help browsers reconstruct optimized JavaScript into something closer to the original source, making debugging easier. MDN's documentation explains how the SourceMap header and sourceMappingURL annotation point developer tools to these files. When source maps are publicly available in production, they can reveal original filenames and make the application's client-side structure easier to follow. That exposure doesn't automatically mean the application is vulnerable. Problems arise when an old endpoint is still reachable, authorization depends too heavily on what the frontend displays, or functionality exposed through the client was never included in the team's current security review. One useful check is to compare what appears in production JavaScript, browser traffic, and available source maps with the routes the team expects to have deployed. Unexpected endpoints deserve a closer look, especially when nobody can immediately explain why they are still there. Make Inventory Part of the Delivery Process Finding forgotten assets starts with looking past the list provided to the vulnerability scanning tool. If that list is incomplete, even a successful scan may leave parts of the application untouched. It is important to continuously verify the accuracy of the inventory against the discoverable assets outside the organization. This means collecting information from cloud accounts, deployment configurations, API gateways, DNS and certificate records, and exploring all hosts/endpoints that are reachable but cannot be placed anywhere in those records. The inventory needs to be closely tied to the deployment process as well. Whenever a service gets deployed, information should be collected about who owns the service, where it runs, what environment it belongs to, and whether it is public or private. This can be accomplished through a CI/CD pipeline as infrastructure is created or changed rather than having someone manage the spreadsheet manually. The same applies to API documentation. OWASP recommends generating API documentation automatically and including it in the CI/CD process. Teams can also compare deployed routes to the approved specification of the API so that an unexpected endpoint becomes visible while the application is still being worked on. From there, a few controls can catch problems early: Require ownership of public services before deploying them to production.Put expiration dates on preview and temporary environments.Flag unexpected new DNS or certificate records outside of the expected deployment process.Track deprecated API versions until they are removed.Keep production data in non-production environments only if it is absolutely necessary. Whenever something unexpected gets detected, assign it to someone who can determine why it is still running. Active services must receive the same level of security attention as the rest of the application. The ones that have served their purpose should be removed. Make Sure Retired Assets Are Actually Gone Removing a service from a repository or architecture diagram doesn't mean the service has disappeared from the internet. Old DNS records can remain, gateway routes may still forward traffic, and credentials created for the service can continue working after the team considers the project finished. Before shutting anything down, check whether it is still receiving traffic. An old API version may have clients nobody remembered, and immediately removing it could break an integration that is still in use. If the service needs to stay online, it should remain under the same monitoring, patching, and access controls as other active systems until those dependencies are dealt with. Once the service is retired, verify the result from outside the environment. Confirm that its hostname no longer resolves where it shouldn't, old routes no longer respond, credentials have been revoked, and any associated storage or third-party integrations have been removed. This step is easy to overlook during migrations or team changes, when responsibility for older infrastructure can become unclear. A service nobody considers active can still be reachable months later if no one checks that the shutdown actually happened. Conclusion Applications change constantly, and some of the infrastructure created along the way will eventually be forgotten. Problems begin when those old hosts, endpoints, and environments remain reachable without anyone checking whether they still need to exist. The inventory needs to change with the application. Build discovery into the delivery process, keep ownership clear, and verify that retired assets are actually gone. If something connected to your application is still reachable from the internet, your team should know why it is there and who is responsible for it.

By Igboanugo David Ugochukwu DZone Core CORE
MCP Is the USB-C of AI — Here's What That Actually Means for Your Architecture
MCP Is the USB-C of AI — Here's What That Actually Means for Your Architecture

Three weeks. That's how long it took my team to wire Claude into our internal ticketing system last year. Not because the API was hard. Because every layer of the stack was speaking a different dialect — custom function schemas on one side, brittle REST wrappers on the other, and a Python shim in the middle that I was too embarrassed to commit without a comment that said: "don't look at this." We shipped it. It worked. For about four days, until the vendor updated their response payload and our parser silently swallowed the change. Tickets started routing to the wrong queue at 2 AM on a Tuesday. I learned about it from Slack, not monitoring. That experience is why Model Context Protocol (MCP) landed so differently for me than it did for the people writing blog posts about it from a fresh MacBook. This wasn't "interesting new protocol." It was a direct answer to a specific, grinding pain. Stop Calling It a Framework The USB-C analogy gets repeated so often it's starting to lose meaning. Let me make it concrete. USB-C solved a problem the tech industry had been ignoring for a decade: every device spoke a slightly different power/data dialect, and the combinatorial explosion of adapters was genuinely slowing things down. USB-C collapsed that N×M adapter problem into a single connector. One port. Any cable. Any device. You still need to negotiate speeds and capabilities over the wire — but the physical contract is shared, which means you can stop thinking about connectors and start thinking about what you're actually moving. MCP does exactly that for AI tool integration. Before it, connecting an LLM to a tool meant writing a custom schema for that LLM's function-call format, a custom parsing layer for that tool's response shape, and — if you wanted to switch providers — starting over. Six integrations across three LLM providers meant eighteen combinations to maintain. The N×M problem. MCP's answer is a shared protocol layer: one JSON-RPC 2.0 contract, negotiated at initialization, that any compliant client can speak to any compliant server. Tools become server capabilities, not one-off function schemas. The LLM doesn't care whether it's talking to a Salesforce connector or a PostgreSQL server — both speak MCP, both expose the same tool-call lifecycle, both fail in predictable ways. That last part matters more than people give it credit for. The Protocol Stack, Actually Explained Most articles stop at "MCP uses JSON-RPC 2.0." That's true, but it's like saying "HTTP uses TCP." Correct. Not sufficient. Layer 1: JSON-RPC 2.0 Messaging JSON-RPC 2.0 is a stateless, lightweight remote procedure call protocol. It predates AI by over a decade — Ethereum uses it, Ethereum Classic uses it, VS Code's Language Server Protocol is built on it. Anthropic's team made a smart choice borrowing from LSP specifically, because LSP proved that you could build richly typed, bidirectional tooling protocols on top of a dead-simple message format. Every MCP message is one of three shapes: JSON // Request (client → server) { "jsonrpc": "2.0", "id": 42, "method": "tools/call", "params": { "name": "search_tickets", "arguments": { "query": "priority:high assignee:me" } } } // Response (server → client) { "jsonrpc": "2.0", "id": 42, "result": { "content": [{ "type": "text", "text": "Found 3 tickets..." }], "isError": false } } // Notification (no id — fire and forget, no response expected) { "jsonrpc": "2.0", "method": "notifications/tools/list_changed" } The id field is doing important work there. Requests have IDs; notifications don't. The client matches responses to requests by ID — which means you can pipeline multiple concurrent requests without ordering guarantees. That's relevant once you start running parallel tool calls, which is exactly what modern agent orchestrators do. Layer 2: Transport Options MCP supports two transports, and picking the wrong one is one of the most common production mistakes I see. stdio is for developer tooling. Cursor uses it. Claude Desktop uses it for local servers. The server runs as a child process, stdin/stdout are the pipe. Zero network overhead, instant startup, trivially secure. Wrong choice for anything multi-tenant or horizontally scaled. Streamable HTTP is what you deploy to production. Single HTTPS endpoint, HTTP POST for client-to-server, optional SSE (Server-Sent Events) stream for server-to-client pushes. The March 2025 spec update replaced the earlier dedicated SSE-only transport — importantly, Streamable HTTP added support for stateless operation, which is the feature that makes real horizontal scaling possible. More on that in a minute. HTTP # Client → Server: negotiate capabilities POST /mcp HTTP/1.1 Content-Type: application/json Authorization: Bearer eyJhbGc... { "jsonrpc": "2.0", "id": 1, "method": "initialize", "params": { "protocolVersion": "2025-11-25", "capabilities": { "tools": {} }, "clientInfo": { "name": "my-agent", "version": "1.4.0" } } } # Server → Client: confirm supported capabilities HTTP/1.1 200 OK Content-Type: application/json Mcp-Session-Id: a3f9-c2d1-8b04 # only in stateful mode { "jsonrpc": "2.0", "id": 1, "result": { "protocolVersion": "2025-11-25", "capabilities": { "tools": { "listChanged": true } }, "serverInfo": { "name": "ticketing-mcp", "version": "2.1.0" } } } Notice Mcp-Session-Id. That header only appears in stateful mode. In stateless mode — which you want for any horizontally scaled deployment — there's no session header. Every request is self-contained. Critically, that means load balancers can route requests to any instance without sticky sessions. That's the architectural unlock. Layer 3: The Three Primitives MCP servers expose exactly three types of capabilities. This is deliberate. The constraint is the feature. The distinction between Tools and Resources isn't cosmetic. Tools can have side effects. Resources can't. The November 2025 spec update formalized tool annotations — you now declare whether a tool is read-only, destructive, or idempotent in the schema itself. That annotation is what lets your gateway apply different rate limits and audit policies per tool class without building bespoke middleware. OAuth 2.1: Why It's Here and What It Costs You Auth was technically optional in early MCP. The community paid for that decision: trojanized packages, unauthenticated community servers running wide open in local dev environments, and at least one incident report I've seen from an enterprise pilot that I won't name where an MCP server was reachable from a public IP with no credentials required. The November 2025 spec update made OAuth 2.1 the recommended standard for remote servers. In practice, if you're deploying Streamable HTTP in a production environment, treat it as mandatory. A few things worth knowing before you implement this: OAuth 2.1 drops implicit flow entirely. If you have legacy client code that used implicit — and plenty of older enterprise apps do — you're rewriting that before you go live. Plan a sprint.PKCE is mandatory for public clients even with authorization code flow. The spec doesn't give you a waiver for this.Server discovery at /.well-known/oauth-authorization-server is how clients find your token endpoint without hardcoding. Don't skip implementing this. Dynamic client registration makes onboarding new agent clients 10x less painful.Tokens are per-user context, not per-MCP-server. Your gateway needs to thread the right token to the right downstream server. That routing logic is where I've seen the most production bugs — specifically, token scope mismatches that silently returned empty results instead of 403s. Production Architecture: The Full Stack Here's what the actual architecture looks like once you move past single-developer demos. The Stateless Scaling Model This is the piece most tutorials gloss over, and it's the piece that will bite you at 3 AM. The original MCP spec used session IDs. Every client got pinned to a server instance via Mcp-Session-Id. That's great for local development. For Kubernetes? It means sticky sessions, broken pod rollouts, and a load balancer that has to track which client is where. The November 2025 spec update added stateless operation as a first-class option — no session IDs; every request carries all context it needs. Stateless vs. Stateful: The Decision Tree Choose stateless(no session ID) if your tools are side-effect-free queries or short-lived mutations. Your load balancer routes freely, Kubernetes rolling deployments work cleanly, horizontal scale is trivial. Choose stateful(session pinned) only when you genuinely need server-side context across calls — browser automation, long-running file operations, or multi-step transactions where partial state lives on the server. For stateful deployments, you need Redis-backed session storage and sticky session config at the ingress level. Python from fastmcp import FastMCP from fastmcp.server.auth import BearerAuthProvider import httpx, os # All state lives in downstream systems. Zero server-side session state. mcp = FastMCP( "ticketing-mcp", auth=BearerAuthProvider( jwks_uri="https://auth.corp.example/.well-known/jwks.json", required_scopes=["mcp:ticketing:read"], ), ) @mcp.tool( description="Search tickets by JQL query. Read-only.", annotations={"readOnlyHint": True, "idempotentHint": True}, ) async def search_tickets(query: str, max_results: int = 20) -> list[dict]: # Auth context injected per-request by the BearerAuthProvider. # No session object. No global state. Safe for any pod to handle. async with httpx.AsyncClient() as client: resp = await client.get( f"https://jira.corp.example/rest/api/3/search", params={"jql": query, "maxResults": max_results}, headers={"Authorization": f"Bearer {os.environ['JIRA_API_TOKEN']}"}, ) resp.raise_for_status() return resp.json()["issues"] Where MCP Actually Breaks I've been building on this protocol for over a year. Here's the honest list of failure modes nobody talks about until they've hit them. The SSE timeout one catches almost everyone. You configure Streamable HTTP, everything works in dev, you push to prod, and suddenly long-running tool calls are silently dying. The load balancer's idle connection timeout — usually 60 seconds — kills the SSE stream before your database export finishes. The fix is simple once you know it: push heartbeat notifications every 30 seconds, and bump your ingress idle timeout to at least 5 minutes. The discovery process is not simple. MCP vs. the Alternatives: An Honest Comparison The column that matters most in that table is the one everyone argues about least: LLM portability. Right now you might be locked into Claude or GPT-4o. Six months from now, there'll be a model from a lab you've never heard of that outperforms both on your specific task. If your tool integrations are written against a provider's function-call schema, you're rewriting them. If they're MCP servers, you're updating a client config file. The Migration Playbook: 3 Days to 11 Minutes Here's exactly how we did our migration. Not the sanitized version. The version that includes the detour through a broken approach we had to back out. Audit your existing tool integrations – catalog every function schema, every parsing layer, every auth mechanism. We found 14 distinct integrations in our codebase, 6 of which were duplicates with slightly different error handling. Don't migrate duplicates; kill them first.Pick FastMCP, not the raw SDK – We initially tried building directly against the TypeScript SDK for more control. That cost us a week. FastMCP's Python decorator model handles 90% of the scaffolding — schema generation, transport setup, error wrapping. Use it. You can always drop to the raw SDK for edge cases.Deploy your gateway first, before any servers – the gateway is where your auth, rate limiting, and audit logging live. Getting it right before servers come online means you're not retrofitting security. We used a simple FastAPI proxy with httpx for upstream calls. Took two days. Worth every hour.Migrate one server per sprint, not all at once — We tried a big-bang migration on our second attempt. It failed. One server per sprint gives you a working fallback and real production data on how each integration behaves under MCP before you cut over.Instrument tool calls from day one – every tools/call should emit a structured log with: tool name, calling agent, token scope used, response time, and whether it succeeded. That data will save you during the first production incident, which will happen. What's Coming — And What to Plan For MCP governance moved to the Linux Foundation's Agentic AI Foundation in December 2025. That matters because it de-risks the protocol against any single vendor's agenda. OpenAI adopted it in April 2025. Google DeepMind's Vertex AI came on board in March 2026. AWS Bedrock in November 2025. This is no longer Anthropic's protocol. It's infrastructure. The 2026 roadmap has four working-group priorities worth knowing: MCP Server Cards – Machine-readable server manifests at a /.well-known/mcp.json endpoint. Think package.json for your MCP server: capabilities, auth requirements, tool list, rate limits. Enables automatic gateway discovery and policy enforcement without configuration drift.Stateless transport formalization – The current stateless mode is in spec but not yet standardized in behavior across SDKs. The Q2 2026 working group is closing those gaps. Wait for this before going all-in on multi-cloud stateless deployments.A2A protocol integration – Google's Agent-to-Agent protocol handles horizontal agent coordination. MCP handles vertical tool connection. The integration point is where agents hand off tasks to subagents that themselves use MCP. Plan for this architecture now, even if you don't need it yet.Audit extensions – Structured compliance fields for tool calls: user context, data classification, retention tags. If you're building in a regulated industry, this will make your compliance team significantly less anxious. Targeting Q3 2026.

By Dinesh Elumalai DZone Core CORE
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration

Every major AI vendor now supports the Model Context Protocol. The framing is almost always the same: MCP is the universal connector for AI agents in the enterprise. That framing sets up a false choice. MCP, REST/HTTP APIs, and Apache Kafka are not alternatives. They solve different problems at different layers of the architecture. Treating them as competing options produces systems that are fragile exactly where they need to be reliable. These three technologies can and do coexist in the same architecture. The question is not which one to pick. It is which one belongs where, and what the tradeoffs are when more than one could technically do the job. This article maps that decision: what each technology is built for, where the boundaries are, and where the genuine gray areas lie. 1. What Is MCP and What Is It Built For? Anthropic introduced the Model Context Protocol in November 2024 as an open standard for connecting AI assistants to external tools and data sources. Before MCP, every AI model required a custom connector to each external system. Three models, ten systems: thirty custom integrations to build and maintain. MCP collapses that to one standard interface. Any compliant client talks to any compliant server without prior coordination. OpenAI adopted MCP in March 2025. Google DeepMind confirmed support in April 2025. By December 2025, MCP had reached over 97 million monthly SDK downloads across Python, TypeScript, Java, Kotlin, C#, and Swift. Anthropic donated the protocol to the Agentic AI Foundation under the Linux Foundation, with AWS, Google, Microsoft, Bloomberg, and OpenAI as platinum members. MCP is no longer a developer experiment. Signals of enterprise maturity are arriving quickly: AI agents paying for API access autonomously, cross-SDK interoperability between Anthropic and OpenAI converging on MCP Resources, composable enterprise workflows where agents read tool signatures and compose cross-system flows without predefined paths, and an official MCP Registry launched in late 2025 as the community-driven server directory. The 2026 roadmap focuses on scalable transport, agent-to-agent communication, governance maturation, and enterprise readiness covering audit trails and SSO-integrated authentication. MCP handles tool access: how an agent calls an external capability. It does not handle agent-to-agent coordination, which is the domain of protocols like Google's Agent-to-Agent (A2A). MCP and A2A are complementary and address different layers of agentic architecture. The moment MCP is asked to do more than tool access, the architecture starts to break. Security Maturity Is Still Catching Up With Adoption Most incidents disclosed in 2025 and early 2026 are implementation failures, not protocol flaws. An Endor Labs analysis of 2,614 MCP implementations found 82% use file system operations prone to path traversal and 67% use APIs related to code injection. Enterprise-grade authentication with OAuth 2.1 and SAML/OIDC is on the 2026 roadmap but still in progress. The practical controls for today: apply least privilege, limit MCP server access to only the systems and data each tool requires, and monitor tool definitions for unexpected changes. 2. MCP vs. REST/HTTP API MCP and REST/HTTP APIs serve different consumers and should not be treated as interchangeable. REST is an architectural style built on HTTP, widely adopted but with no fixed conventions for discovery, error formats, or method naming. Well-designed REST APIs backed by OpenAPI specifications work well for direct, programmatic data access when a native SDK or versioned API already exists and teams know how to operate it. MCP enforces consistency at the interface level because the consumer is an AI model that cannot tolerate creative API interpretation. MCP standardizes how a tool is called. It does not standardize what the tool returns, how fresh that data is, or whether two agents calling the same tool simultaneously see the same state. For direct data access to vector stores, databases, or business application APIs, a well-governed REST API, native SDK, or Kafka Connect integration is almost always the better choice: lower latency, no protocol overhead, mature tooling. For giving AI agents standardized, discoverable access to a broader set of tools across vendors and frameworks, MCP is the right layer. The two are complementary, not competing. Tool Design Matters as Much as the Protocol Choice One important nuance on tool design: mapping one-to-one from existing APIs to MCP tools rarely works well. What matters is tool granularity, smart metadata, and thoughtful assembly of the MCP layer. An MCP server that exposes well-structured, semantically rich tools lets an AI agent reason about capabilities and compose workflows. This is reminiscent of the composability questions from the enterprise SOA (Service-oriented Architecture) era. SOA promised flexible service composition but delivered integration chaos when governance, metadata quality, and service granularity were treated as afterthoughts. MCP faces the same risk. The protocol is sound; what determines success is the discipline applied to how tools are defined, documented, and assembled. What MCP Does Not Do What MCP does not do matters as much as what it does. It does not manage data, guarantee message delivery, enforce governance, or guarantee consistency across systems. It is an interface layer, not a data pipeline. That boundary becomes even clearer when looking at what Kafka does, which is structurally different from both MCP and REST. 3. Apache Kafka: Event Broker, Decoupling, and the Backbone Role Operational data is the live data that runs business processes: order states, inventory levels, transaction records, customer accounts, risk scores. It originates in systems like SAP, Salesforce, Oracle, and mainframes, and it changes continuously. Kafka is architecturally different from both HTTP and MCP in one way that matters most: it decouples producers and consumers through a persistent, ordered, append-only log. With HTTP or MCP, the caller and the callee are coupled at request time. Every integration is point-to-point. If the target system is slow or unavailable, the caller is directly affected. Kafka breaks that coupling entirely. A producer writes an event once. Any number of consumers read it independently, at their own pace, using their own communication paradigm. One consumer processes records in real time. Another runs nightly batch analytics over the same events. A third powers a stream processing pipeline. A fourth writes results to a data lake via Apache Iceberg. All of them consume the same underlying data product. None of them affects the others. Kafka supports three consumption patterns from a single event stream: streaming, request-response, and batch. The event exists once; each consumer is independent. This is the pub/sub event broker model, and it is what makes Kafka the integration backbone between operational and analytical systems. The diagram below shows this decoupling: a single Kafka topic serving real-time applications, HTTP-based consumers, batch analytics, and MCP agent interfaces simultaneously. Stream Processing With Kafka Streams and Apache Flink Stream processing is a core complement to Apache Kafka, extending the platform from event transport into real-time data processing and decisioning. Kafka Streams is a lightweight Java library embedded in applications. It is well-suited for streaming ETL and simple to medium stateful stream processing without requiring a separate cluster. It integrates closely with existing JVM-based services. Apache Flink is a distributed stream processing engine designed for more complex workloads. It supports Java, Python, and SQL APIs, making it accessible to both application developers and data engineers. Flink runs as a dedicated cluster or in managed environments and is built for high-scale scenarios such as multi-stream joins, event-time processing, large state management, exactly-once semantics, Complex Event Processing (CEP), real-time analytics, and AI model inference. Both approaches extend Kafka with processing capabilities. The choice depends on workload complexity, required deployment model, and preferred programming language, not on replacing Kafka’s role as the event streaming backbone. A detailed comparison is available in the post Apache Kafka and Apache Flink: A Match Made in Heaven. Operational and Analytical Integration, Including the Data Lakehouse Kafka is not only for operational data integration. It serves as the ingestion layer into data lakes, feeds real-time analytical pipelines, enables stream processing with embedded AI models, and connects business applications bidirectionally. A governed data streaming platform provides schema registry, lineage tracking, role-based access control, and exactly-once delivery semantics across all of that. It serves both operational and analytical use cases and acts as the bridge between those two worlds. For how streaming and the lakehouse converge via Apache Iceberg, see Data Streaming Meets Lakehouse. Kafka's append-only commit log is the foundation of data consistency across the enterprise. Every downstream consumer sees the same data in the same order. That is not just a performance feature. It is what prevents the architecture where every system has its own version of the truth. 4. The Tradeoffs: It Is Not Black and White The choice between MCP, REST/HTTP APIs, and Kafka is rarely clean. All three can play a role in the same architecture. REST/HTTP APIs work well for operational data access when volume is moderate and a well-governed API already exists. A REST API backed by a Kafka-derived serving layer can return consistent, current data. The API is the interface; the streaming platform is what makes the data trustworthy behind it. A financial services firm exposing account balances via REST is not doing it wrong, as long as those balances are derived from a governed, consistent data source rather than pulled directly from a source system on every request. Kafka becomes the clear choice when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when the same events need to feed operational applications, analytical pipelines, and AI agents simultaneously. MCP fits best when access is supplementary, loosely coupled, and low-frequency. A support agent looking up a ServiceNow ticket before drafting a response, or a sales assistant pulling the latest slide deck from Google Drive before a call, are good fits. The key test is simple: does it matter if the data the agent receives is a few seconds or minutes old? If yes, MCP should not own that responsibility. If no, MCP is the right interface. SAP: Clean Separation Between ERP Integration and Developer Tooling The boundary between MCP and REST is not a choice between two equivalent options for the same integration. SAP is the clearest example of a clean separation. SAP exposes extensive REST and OData APIs for ERP integration: order management, finance, supply chain, procurement, and HR data flowing bidirectionally between SAP and other enterprise systems. SAP's MCP servers serve an entirely different purpose: developer tooling for ABAP code generation, CAP application development, UI5 and Fiori assistance, and operational tasks like transport validation and incident management. An architect connecting SAP order events to downstream systems uses OData and Kafka Connect. A developer asking an AI coding assistant to generate ABAP code uses the SAP MCP server. Different consumers, different use cases, different data. No overlap. Salesforce and ServiceNow: Same Data, Different Consumer Salesforce and ServiceNow follow a different pattern. Their MCP servers wrap the same underlying REST APIs and expose the same underlying data, but for a different consumer. A developer-written integration calls the Salesforce REST API directly with known endpoints and hardcoded logic. An AI agent calls the Salesforce MCP server, which wraps that same API to make it discoverable and stateful for an agent that cannot read documentation or manage its own session state. The data is identical. The access path differs based on who is consuming it. This is not a free choice between equivalent options. It is the same system serving two different client types through two different interface layers. REST vs. Kafka for Operational Data: The Harder Call The harder boundary is between REST and Kafka for operational data. Both can technically serve it, and that is where the real architectural decision lies. REST is simpler to start with but introduces point-to-point coupling, integration spaghetti at scale, and consistency risks when the same data needs to reach multiple consumers. Kafka is more complex to operate but provides the decoupling, consistency, and governance that enterprise architectures require when the same data needs to reach many consumers reliably. The two are not mutually exclusive. A common and well-proven pattern combines both: Kafka handles the event backbone, decoupling, and consistency, while a REST layer sits on top for synchronous request-response access, API management integration, or compatibility with systems that cannot speak the native Kafka protocol. This is particularly common in mobile applications, legacy system integration, and API gateway architectures. For a detailed look at how REST and Kafka complement each other in practice, see Request-Response with REST/HTTP vs. Data Streaming with Apache Kafka. 5. Decision Framework: MCP, REST/HTTP, or Kafka? Choosing between MCP, REST/HTTP, and Kafka is not a single decision but a set of tradeoffs that depend on data volume, consumer type, consistency requirements, and what is already in production. The comparison table below makes those tradeoffs concrete across eight dimensions. When to Use Which: A Guide to the Decision Tree The decision tree below walks through the same logic as a series of questions, routing to the right choice based on the integration's actual requirements. Use MCP when the integration is supplementary and tool-like: Slack, Google Drive, ServiceNow tickets, internal knowledge bases. The agent needs context to act, not a stream of events to react to. Eventual consistency is acceptable. Apply least privilege, monitor tool definitions for changes, and isolate MCP servers from production systems. Use a REST/HTTP API or native SDK when a well-documented API or SDK already exists and the engineering team knows how to operate it. The access pattern is direct, moderate-volume, and latency-sensitive. REST is also a reasonable choice for operational data when the backend is a governed Kafka-derived serving layer and consistency properties are inherited, not assumed. Use Apache Kafka when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when governance, lineage, and auditability are non-negotiable. Kafka is also the right choice when the same data needs to feed operational applications, real-time analytics, data lakes, and AI agents simultaneously. Use the real-time context engine when an AI agent needs current, consistent operational context for autonomous decisions. Kafka and Flink govern the data. MCP provides the agent interface. The consistency guarantee comes from the streaming layer, not from MCP. The practical question is not which protocol to choose. It is whether the data architecture underneath the agents can be trusted. Agents making autonomous decisions about inventory, risk, or customer service are only as reliable as the data they act on. 6. Where MCP and Kafka Work Together: The Real-Time Context Engine There is one pattern where MCP and data streaming complement each other directly: the real-time context engine. Kafka and Flink process and govern the data: ingesting from operational systems, applying transformations and filters, producing real-time materialized views. Those views are then exposed to AI agents through a standardized MCP interface. The streaming platform owns the data, its freshness, and its consistency guarantees. MCP owns the interface to the agent. Neither layer bleeds into the other's responsibility. Data consistency is not delegated to MCP. The streaming platform enforces it upstream before the MCP interface comes into play. The agent calls a tool and receives context that is current, governed, and consistent, not because MCP guarantees it, but because the streaming platform does. Any compliant AI agent, whether Claude, ChatGPT, Amazon Bedrock, LlamaIndex, or CrewAI, can call the context engine and receive current context from operational systems without needing to understand Kafka topics, Flink jobs, or schema evolution. An agent routing shipments from yesterday's inventory, approving transactions against a risk score from three hours ago, or reading an account balance that has not propagated: none of these is reliable. A real-time context engine eliminates this class of error at the source, reduces hallucinations, lowers inference cost, and anchors decisions to current operational reality. From Data Freshness to Agent Governance Enterprise readiness for this pattern also depends on how agents are governed once deployed. Trust, control, and accountability become central once agents start chaining decisions across domains. The context engine is the data layer of that answer. Governance of the agents themselves, covering what they are permitted to do, under what conditions, and with what audit trail, is the other half. This is the dimension enterprise buyers are actively evaluating when selecting agent orchestration platforms. The diagram below shows how the three layers fit together: the streaming platform as the data backbone, the context engine as the governed serving layer, and MCP as the clean interface to agents. 7. Conclusion: One Protocol, One Job MCP has earned its place in the enterprise architecture stack. What it has not yet earned is the role of universal integration layer, and understanding that distinction is what this article has been about. The broader architecture this sits inside connects three interdependent pillars. Event-driven data integration, with Kafka as the backbone, moves data reliably between operational and analytical systems and delivers governed data products to every consumer. Process intelligence is the orchestration layer that determines which decisions to automate, in what sequence, and under what conditions, giving agentic workflows the structure and governance they need to be trustworthy. Trusted agentic AI is where MCP plays its role: the standardized, governed interface through which agents access external tools and context, anchored to real data by the streaming layer beneath it. For a vendor-by-vendor analysis of trust and lock-in across the major AI platforms, see the Enterprise Agentic AI Landscape 2026. For a deeper look at how the three pillars fit together as an enterprise architecture framework, see The Trinity of Modern Data Architecture: Process Intelligence, Event-Driven Integration, and Trusted Agentic AI. One protocol, one job. That is the right way to use MCP.

By Kai Wähner DZone Core CORE
Multi-Agent Systems: Architecture Patterns for Developers
Multi-Agent Systems: Architecture Patterns for Developers

Most production agent projects do not fail because the model is weak. They fail because one agent was asked to hold too much at once: routing, planning, tool use, memory, and error recovery all inside a single growing prompt. By 2026, this failure mode shows up in nearly every engineering retro, and the fix is usually the same. Split the work across several coordinated agents. The numbers back this up. Gartner reports that roughly 80% of enterprise applications shipped or updated in early 2026 embed at least one AI agent, up from about a third in 2024. Yet a figure cited across IDC and Forrester research puts pilot-to-production failure near 88%, and the root causes cluster on orchestration, data access, and evaluation gaps, not model quality. Architecture, not model choice, is where most of these systems are won or lost. This piece walks through the multi-agent patterns worth knowing, with notes on when each one fits and where it tends to break. What Is a Multi-Agent System? A multi-agent system is a set of specialized agents that split a task, coordinate through shared state or messages, and combine their outputs into one result. Each agent owns a narrow job: a planner decides steps, a researcher gathers context, a writer drafts, a critic reviews. This keeps prompts short, makes behavior easier to test, and lets you retry or swap one part without rerunning the whole chain. Why Single-Agent Designs Hit a Ceiling A single agent works well until the task branches. Add several tools, conditional logic, and long context, and the model starts to lose the thread. Instructions compete, the context window fills with irrelevant history, and one bad tool call derails everything downstream. Splitting responsibilities gives each agent a smaller decision space, which is easier to reason about and cheaper to debug. Core Architecture Patterns for Multi-Agent Systems 1. Orchestrator (Supervisor) Pattern A central agent receives the request, decides which worker should handle it, and routes accordingly. Workers do not talk to each other; they report back to the supervisor, which picks the next move. Python def supervisor(task, state): route = router_model(task, state) # pick the next worker if route == "research": return research_agent(task) if route == "code": return code_agent(task) if route == "done": return finalize(state) This is the most common starting point. Centralized control makes logging and human review straightforward. The tradeoff: the supervisor becomes a bottleneck and a single point of failure. 2. Sequential (Pipeline) Pattern Agents run in a fixed order, each consuming the previous output: extraction, then validation, then summary. Use it when steps are stable and order matters. It is simple to trace, but rigid. A change in requirements often means rewriting the chain. 3. Hierarchical Agent Teams Supervisors manage sub-supervisors, which manage workers. A top planner splits a goal into subgoals, hands each to a team lead, and each lead coordinates its own workers. This scales to larger problems and mirrors how organizations already divide labor, at the cost of more coordination overhead and latency. Anthropic's Claude Agent SDK added hierarchical subagent spawning in 2026 for exactly this shape of problem. 4. Network (Peer-to-Peer) Pattern Agents hand control directly to one another based on the task, with no fixed hub. The handoff model in the OpenAI Agents SDK works this way: a triage agent passes a conversation to a billing or support agent, which can pass it on again. It fits open-ended, conversational AI agents where the next step is not known in advance. The risk is loops and unclear ownership, so you need turn limits and explicit exit conditions. 5. Blackboard (Shared State) Pattern Agents read from and write to one shared store instead of messaging each other directly. Each agent watches the board, contributes when it can help, and stops when the goal is met. This decouples agents cleanly but makes state management the hard part. Concurrent writes and stale reads cause most of the bugs. State and Communication: The Real Design Decision Patterns are the visible layer. Beneath them sits the question that decides how hard your system is to operate: how do agents share information? Two options dominate. Shared state keeps one structured object that every agent updates, which is easy to inspect and checkpoint; LangGraph builds on this with checkpointing and time-travel debugging. Message passing sends discrete messages between agents, which maps well to conversational and event-driven designs such as AutoGen and its successor AG2. Shared state is easier to audit. Message passing is easier to distribute. Pick based on which one your team can debug at 2 a.m. Choosing the Right Pattern If you need... Reach for Central control and easy logging Orchestrator Fixed, ordered steps Sequential pipeline Large tasks split across teams Hierarchical Open-ended, conversational flow Network/handoffs Loose coupling, many contributors Blackboard A few rules hold across all of them. Start with the simplest pattern that could work, usually an orchestrator, and add structure only when a real limit appears. Give every agent a narrow role and a clear stop condition. And treat evaluation as part of the architecture, not an afterthought. Why This Matters in 2026 Teams that cross from pilot to production share one habit: they instrument everything. Failure analyses in 2026 point to observability and evaluation coverage as the largest single blocker, ahead of tool access and data quality. In practice, that means logging every agent decision, running automated evals on each step, and putting human review gates where a wrong action is expensive. Generative AI agents are only as trustworthy as the traces they leave behind. Multi-agent architecture is moving from research demos to standard practice, and the frameworks now converge on the same primitives: state, handoffs, checkpoints, subagents. That convergence means the durable skill is not in any single library. It is knowing which pattern fits the problem in front of you and being able to explain why.

By Matthew Truong
Stop Blaming Executor Memory: The Real Reasons Your Spark Jobs Are Slow
Stop Blaming Executor Memory: The Real Reasons Your Spark Jobs Are Slow

After a decade of building and debugging large-scale data pipelines across financial services, payments processing, and analytics platforms, I can tell you that almost every slow Spark job I've investigated had the same root cause — and it wasn't the one the team thought it was. The default response when a Spark job is slow is to add more executor memory, increase the number of executors, or bump spark.sql.shuffle.partitions. Sometimes that helps. Usually it doesn't. What I've found, consistently, is that the real problems are structural — a join strategy mismatch that silently multiplies your intermediate dataset by ten times, a single slow task on a degraded node that holds an entire stage hostage, or a decrypt chain that re-reads source data six times when it only needed to read it once. This article is organized around five patterns I keep seeing across teams. Each one looks different on the surface but traces back to a misunderstanding of how Spark actually executes your code. For each pattern, I'll describe what it looks like, when it bites you, the failure mode, and how to fix it. Pattern 1: The OR Join That Quietly Multiplies Your Data What It Looks Like A join condition with an OR clause. Usually introduced when a business requirement adds a secondary matching rule — match on primary card number, or if the transaction is a virtual card transaction, match on the underlying physical PAN. The SQL looks reasonable. The engineer tests it on a sample, and it returns the right rows. When It Bites You At scale. With 100 million transaction rows and 50 million account rows, this query starts running for hours. The output size is also wrong — much larger than expected before DISTINCT trims it down. The Failure Mode Spark cannot use a hash join or sort-merge join when the join condition contains OR. It falls back to BroadcastNestedLoopJoin — for every row in the left table, scan every row in the right table. That's O(n x m). On real datasets, this produces an intermediate result in the hundreds of GB before any downstream filter runs. I've watched a pipeline that should produce 8 GB of output generate 400 GB of intermediate data because of exactly this pattern, taking a 20-minute job to 4 hours. You can verify this in 30 seconds: run df.explain(formatted) and look for BroadcastNestedLoopJoin in the physical plan. If you see it on a join involving any table over a few million rows, it's almost certainly unintentional. The Fix Split the join into two equi-join legs and UNION ALL the results: SQL -- Leg 1: primary match (equi-join — uses SortMergeJoin or BroadcastHashJoin) SELECT txn.*, acct.* FROM transactions txn JOIN accounts acct ON txn.card_number = acct.card_number UNION ALL -- Leg 2: fallback match, filtered scope only SELECT txn.*, acct.* FROM transactions txn JOIN accounts acct ON txn.fpan = acct.physical_pan WHERE txn.transaction_type = 'VIRTUAL' Each leg is a proper equi-join. Apply DISTINCT at the end to deduplicate rows that matched both. The performance difference is routinely an order of magnitude. Pattern 2: The Straggler Task That Nobody Notices Until It's Too Late What It Looks Like A stage that should take 10 minutes takes 3 hours. The Spark UI shows nearly all tasks completed quickly. One or two tasks are still running with a disproportionately long duration. When It Bites You Jobs running on shared YARN or cloud infrastructure where any node can have a bad disk, a noisy neighbor, or degraded network throughput. Also common in stages that call external services per partition — one slow API response can cause a single partition's tasks to take 100x longer than the others. The Failure Mode A stage doesn't complete until the last task completes. Not the median. Not p95. The absolute last one. If 2,200 tasks finish in under 2 minutes and one takes 3 hours and 7 minutes, the stage takes 3 hours and 7 minutes. The other 2,199 executors sit idle. This is the straggler problem, and it's distinct from data skew. The diagnostic: in the Stage detail view, check the task duration distribution. If MAX is dramatically higher than p99, that's a straggler (hardware or external service issue). If p75 is already much higher than p50, that's skew (data distribution issue). They require different fixes, and many teams treat them identically. The Fix For stragglers caused by degraded infrastructure, enable Spark speculation: Properties files spark.speculation=true spark.speculation.multiplier=3 # task must be 3x slower than median spark.speculation.quantile=0.9 # wait for 90% completion before speculating Speculation re-launches slow tasks on a different executor and uses whichever copy finishes first. The caveat: don't use this on stages that write to non-idempotent sinks. For read-heavy or compute-heavy stages — including external decryption calls — it's often the single most impactful config change you can make. Pattern 3: The df.rdd Decrypt Chain That Recomputes Everything Six Times What It Looks Like A pipeline that calls an external encryption or decryption service per record, implemented as a series of df.rdd.mapPartitions() calls, one per column that needs to be processed. When It Bites You When you have multiple columns to decrypt. Each .rdd call creates a new computation starting from the original DataFrame — Spark re-reads from source, re-executes all upstream joins and filters, and then runs the decryption for that column. With six columns to decrypt, you're doing that six times. The Failure Mode Two distinct sub-problems compound each other. First, going to RDD bypasses Catalyst entirely — no predicate pushdown, no column pruning, no Tungsten execution. Second, without a persist checkpoint before the chain, every decrypt call lineages all the way back to the source. I've seen this double the runtime of a job compared to the same pipeline with a single persist() before the decrypt chain. On top of that, the external call latency per partition is dominated by the number of HTTP round trips, not the payload size. Cutting your batch size in half doubles your request count and roughly doubles your wall-clock time for that stage. Most teams set an initial batch size and never revisit it. The Fix Two changes, applied together: Persist the input DataFrame before starting the decrypt chain. This means the join and filter logic runs once, and each decrypt call reads from the cached result.Increase the batch size for external calls. Test at several sizes — going from 20,000 to 40,000 records per batch often cuts stage time by 30-50% with no change to correctness. Scala val base = rawDf.filter(...).join(key1, ...).persist(StorageLevel.MEMORY_AND_DISK) val step1 = decryptColumn(base, secret1) // reads from cache val step2 = decryptColumn(step1, secret2) // reads from cache val step3 = decryptColumn(step2, secret3) // reads from cache Without persist, step2 re-executes everything step1 did from source. With persist, each step reads from the in-memory result of the previous. Pattern 4: The shuffle.partitions Setting That Nobody Updates What It Looks Like A job that works fine in staging — where data volumes are 10% of production — but runs slowly, spills to disk, or produces thousands of tiny output files in production. When It Bites You When the default spark.sql.shuffle.partitions=200 is left unchanged. 200 partitions made sense as a default for medium datasets but is almost always wrong at production scale — either too few (huge partitions, memory pressure) or too many (tiny partitions, scheduling overhead, small files problem). The Failure Mode Too few partitions means each executor handles a disproportionately large chunk of data. With 200 partitions on a 1 TB shuffle, each partition is 5 GB. That will spill to disk. Too many partitions means thousands of 1 MB tasks — the scheduling overhead becomes significant, and your output has thousands of tiny files that hurt downstream readers. With Adaptive Query Execution (AQE) enabled in Spark 3.2+, this problem largely manages itself. AQE merges small post-shuffle partitions automatically and can handle modest skew. But AQE can't help if it's disabled, and it can't fix the upstream causes of extreme skew. The Fix Enable AQE if you're on Spark 3.2+: Properties files spark.sql.adaptive.enabled=true spark.sql.adaptive.coalescePartitions.enabled=true spark.sql.adaptive.skewJoin.enabled=true If you need to set shuffle.partitions manually, target roughly 128-256 MB per partition post-shuffle. For a 500 GB shuffle, that means 2,000-4,000 partitions. Set it high and let AQE coalesce down — that's cheaper than setting it low and getting OOM errors. Pattern 5: The Incremental Job That Degrades Silently Over Time What It Looks Like A job that runs in 15 minutes when first deployed and runs in 4 hours six months later. No code changes. No obvious data quality issues. The team attributes it to data growth. When It Bites You When the job fails a few times in a row, and the recovery accumulates multiple windows' worth of data. Or when the watermark logic was designed for small windows but nobody anticipated that the underlying join tables would grow significantly. The Failure Mode Two separate causes, often confused. First, if the watermark is a single timestamp and the job has been failing, recovery runs can accumulate large backlogs. A job that normally processes 2 hours of data may need to process 48 hours on first successful recovery, with no change to the resource configuration. Second, growth in reference data (like an accounts table or lookup table used in a join) increases the size of every run regardless of whether the incremental input grew. I've seen a 30-minute job become a 3-hour job purely because the accounts table grew from 10 million rows to 80 million rows over 18 months, while the OR join condition (see Pattern 1) meant that growth was amplified into the intermediate result. The Fix Two design principles that pay off over the lifetime of the pipeline: Track processed partitions explicitly rather than using a single timestamp watermark. This makes recovery granular — you can replay specific missing partitions without re-processing everything after them.Add a fast-path no-op check before initializing the full Spark session. Check whether any new partitions exist first. A 5-second check that exits early is much better than a 2-minute executor startup that discovers there's nothing to process. For the reference table growth problem: if your lookup table grows significantly, revisit whether it can be broadcast (small enough to fit in executor memory) or whether the join itself needs to be redesigned. Quick Diagnostic Reference Use this table to map what you observe in the Spark UI to the likely pattern and first action to take: WHat you observeLikely patternconfirm withfirst action MAX task duration >> p99 Straggler (Pattern 2) Task timeline in Stage UI Enable spark.speculation p75 >> p50 task duration Data skew Input bytes per task Repartition on join key; AQE skewJoin BroadcastNestedLoopJoin in explain() OR join (Pattern 1) df.explain( formatted) Rewrite as UNION of equi-joins Stage runtime grows week on week; no code change Incremental accumulation or reference table growth (Pattern 5) Input bytes trend in History Server Audit watermark logic; check reference table size OOM errors or heavy disk spill Too few shuffle partitions (Pattern 4) Spill metrics in Stage UI Enable AQE or increase shuffle.partitions The Common Thread Every pattern here traces back to the same underlying issue: Spark is executing something different from what the engineer intended. The OR join was intended as a flexible matching rule; Spark turned it into a nested loop. The decrypt chain was intended as six independent transformations; Spark turned it into six full re-reads of source data. The incremental job was intended to process one window of data; without proper watermark design, it occasionally processes twelve. The Spark UI has everything you need to see this — task distribution, input and output sizes, physical plans, spill metrics. Most teams open it when something breaks and close it once they find the obvious error. Opening it proactively, forming a hypothesis, and then confirming or refuting it in the metrics is the practice that separates engineers who consistently improve pipeline performance from those who add executor memory and hope for the best. The mistake isn't choosing the wrong config. It's not understanding what Spark is actually doing with your code.

By Swaminathan Sethuraman

The Latest Software Design and Architecture Topics

article thumbnail
Stop Paying a Model to Make Decisions You Already Made
A skill that spells out a fixed procedure in prose makes Claude re-decide it every run. Here's how to measure that cost using data Claude Code already emits.
September 28, 2026
by Amith Reddy Ravuru
· 107 Views
article thumbnail
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning
Batch processing works well for many workloads, but real-time decisioning requires event-driven architecture designed for resilience, observability, and failure handling.
September 25, 2026
by Prem Kumar Gadhanki
· 817 Views
article thumbnail
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults
Golden paths standardize software delivery with self-service workflows, deployment guardrails, and observability while preserving team autonomy.
September 25, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 826 Views
article thumbnail
How to Verify Response Data in API Testing With Playwright TypeScript
Learn how to verify the response data, including structure checks, basic validations, and more, in API Testing with Playwright TypeScript
September 25, 2026
by Faisal Khatri DZone Core CORE
· 786 Views
article thumbnail
Locking Down the Enterprise: Data Security Patterns for AI Integrations
Learn how to secure enterprise data in AI systems using data classification, access controls, encryption, vendor controls, prompt security, and output validation.
September 25, 2026
by Balaji Venkatasubramaniyar DZone Core CORE
· 835 Views
article thumbnail
Cloud Complexity Is an Operating Model Problem: Why Infrastructure Maturity Alone Can’t Solve Scale, Reliability, and Team Friction
Cloud-native platforms need more than mature infrastructure. Learn how shared standards, platform engineering, and self-service can reduce delivery friction at scale.
September 24, 2026
by Igboanugo David Ugochukwu DZone Core CORE
· 1,044 Views · 1 Like
article thumbnail
6 Techniques To Reduce LLM API Costs With the Python Library
Six techniques to cut LLM API costs by up to 90%: prompt caching, model routing, batch processing, and more. (Includes a pip-installable Python library.)
September 23, 2026
by Somnath Banerjee
· 1,909 Views · 1 Like
article thumbnail
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes operations can drift as teams scale. Use this checklist to standardize clusters, releases, observability, access, reliability, and cost.
September 23, 2026
by Abhishek Gupta DZone Core CORE
· 1,604 Views
article thumbnail
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Terraform monoliths hurt when one state couples too many resources and owners. Split around ownership boundaries, not size.
September 22, 2026
by Naveen Kalapala
· 1,929 Views
article thumbnail
The Hidden Production Risks of Third-Party SDKs
Third-party SDKs speed up development, but they also introduce performance, security, reliability, and maintenance risks that teams must actively manage.
September 22, 2026
by Satyam Nikhra
· 2,787 Views
article thumbnail
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
Building an MCP server that processes files introduces problems a typical read-only API doesn't have. Here's what mattered.
September 22, 2026
by Peter Ndumia
· 1,761 Views · 1 Like
article thumbnail
Beyond Token Intelligence: Why AI Code Review Needs Cognitive Architectures
AI is generating code faster than humans can review it. The fix is cognitive architectures that understand not just "what changed" but "why" and whether it's safe.
September 22, 2026
by Sayan Chatterjee
· 1,892 Views · 2 Likes
article thumbnail
Architecting for <1s Latency: Managing Eventual Consistency in Distributed Search Platforms
To maintain sub-second search freshness, logistics systems must actively manage eventual consistency across Kafka ordering, search indexing, and cache invalidation.
September 22, 2026
by Dhruv Goel
· 1,797 Views · 1 Like
article thumbnail
How to Build a Production-Ready iOS App With AI-Generated Code
AI-generated iOS apps need rigorous engineering across security, architecture, testing, observability, and reliability before production deployment.
September 21, 2026
by Uthej Mopathi DZone Core CORE
· 2,221 Views · 1 Like
article thumbnail
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team
Learn how Claude Code helps enterprise teams build, troubleshoot, and manage service integrations with MCP, CI/CD, automated reviews, and stronger governance.
September 21, 2026
by Balaji Venkatasubramaniyar DZone Core CORE
· 2,455 Views · 1 Like
article thumbnail
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest
Stable operation IDs prevent iOS retries from duplicating agent tools across LangGraph, MCP Tasks, Kafka, and App Attest.
September 21, 2026
by Uthej Mopathi DZone Core CORE
· 1,808 Views · 2 Likes
article thumbnail
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
Learn how forgotten internet-facing assets expand your attack surface and how continuous asset discovery, inventory, and ownership can reduce security risks.
September 21, 2026
by Igboanugo David Ugochukwu DZone Core CORE
· 1,402 Views · 1 Like
article thumbnail
MCP Is the USB-C of AI — Here's What That Actually Means for Your Architecture
A senior engineer's guide to production MCP: JSON-RPC 2.0 transport, OAuth 2.1 auth, stateless horizontal scaling, and where the protocol genuinely breaks.
September 21, 2026
by Dinesh Elumalai DZone Core CORE
· 1,657 Views · 1 Like
article thumbnail
Edge AI: Why Inference Is Moving Away From the Cloud
Edge inference thrives on-device for real-time, private AI. Advances in hardware and compression cut latency and costs, pushing AI away from the cloud.
September 21, 2026
by Uthej Mopathi DZone Core CORE
· 1,895 Views · 1 Like
article thumbnail
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration
MCP, Kafka, and REST APIs are not the same: this comparison maps each to the right layer of your agentic AI architecture.
September 18, 2026
by Kai Wähner DZone Core CORE
· 3,581 Views
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×