DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Integration

Integration refers to the process of combining software parts (or subsystems) into one system. An integration framework is a lightweight utility that provides libraries and standardized methods to coordinate messaging among different technologies. As software connects the world in increasingly more complex ways, integration makes it all possible facilitating app-to-app communication. Learn more about this necessity for modern software development by keeping a pulse on the industry topics such as integrated development environments, API best practices, service-oriented architecture, enterprise service buses, communication architectures, integration testing, and more.

icon
Latest Premium Content
Trend Report
Modern API Management
Modern API Management
Refcard #303
API Integration Patterns
API Integration Patterns
Refcard #249
GraphQL Essentials
GraphQL Essentials

DZone's Featured Integration Resources

Gossips on Cryptography: Part 4

Gossips on Cryptography: Part 4

By Sahil Aggarwal
In this blog, we will continue our discussion from the previous parts. If you have not read them, please read them first. Parts 1 & 2 – Caesar Cipher, Vigenere Cipher, Symmetric Encryption, AES, Convergent Encryption, IVPart 3 – Hashing, Salting, Rainbow Table Attacks, Asymmetric Encryption, RSA In Part 3, we teased a few topics for Part 4 — Envelope Encryption, PKI, and more. Today we gossip about exactly those! Let's go. First, A Quick Revisit: Convergent Encryption We discussed Convergent Encryption back in Parts 1 & 2, but let's revisit it here because it connects beautifully to Envelope Encryption. Remember? Convergent Encryption means — if you encrypt the same plaintext with the same key and the same IV, you will always get the same ciphertext. So Why Is That Useful? Imagine you work at a big company. 500 employees all upload the same file — let's say the company's HR policy PDF. If you use regular encryption (different ciphertext every time), your storage system stores 500 different encrypted copies. That's 500x storage wasted! With Convergent Encryption, since the same file + same key = same ciphertext, the storage system realizes — "Hey, I already have this encrypted file!" — and stores only ONE copy. All 500 employees point to the same encrypted blob. This is called deduplication. This is exactly how Dropbox, Google Drive, and AWS S3 save enormous amounts of storage at their scale. Another Use Case — Searching Over Encrypted Data Here is another very powerful use case of Convergent Encryption that most people don't think about — searching. Imagine you have a database where all the data is encrypted. A user wants to search for records where the email is "[email protected]." With regular encryption, every time "[email protected]" is encrypted, it produces a **different** ciphertext (because of a random IV). So to search, you would have to: Decrypt every single record in the databaseCompare the plaintextReturn the matches That is insanely expensive! Imagine doing this on a database with 100 million records. Your server will cry. Now with Convergent Encryption — "[email protected]" always produces the **same** ciphertext. So to search, you just: Encrypt the search term "[email protected]" onceLook for that ciphertext in the database — just like a normal indexed search!Return the matches No decryption needed at all! The data stays encrypted at rest, and you can still do fast, exact-match searches on it. This is called searchable encryption, and it is used in scenarios like: Encrypted databases where you still need to support queriesHealthcare systems — searching patient records without ever exposing raw dataEmail systems — searching your encrypted inbox without the server ever seeing your emails in plain text Pretty powerful, right? Same property (deterministic output) — two completely different superpowers (deduplication + searchable encryption). But wait — there's a catch. If two people can produce the same ciphertext, can someone guess your file? Yes, this is called a confirmation attack. Someone could hash a known file, compare it with stored hashes, and confirm whether you uploaded that file. So Convergent Encryption is great for performance and deduplication but is used carefully in highly sensitive scenarios. Now Let's Talk About Envelope Encryption Okay, so now we know — encryption needs keys. And those keys need to be stored somewhere safely. But here's the problem — who encrypts the key itself? If your key is lying around in plain text, a hacker who gets access to your server gets everything. So the answer is — we encrypt the key too! This is the core idea of Envelope Encryption. The Two Keys in Envelope Encryption DEK — Data Encryption Key. This is the key that directly encrypts your actual data. Think of it as the key to your diary.KEK — Key Encryption Key. This is the master key that encrypts the DEK. Think of it as the key to your locker — inside which you keep your diary key. So the flow looks like this: Your Data → encrypted with DEK → Encrypted Data DEK → encrypted with KEK → Encrypted DEK You store both — the Encrypted Data and the Encrypted DEK — together. The KEK lives safely inside a highly secure system (like AWS KMS or Vault). Real-Life Example — The Bank Locker Imagine you have an important document (your data). You put it in a box and lock it with a small key (DEK). Now you don't want to carry this small key everywhere — so you put the small key inside your bank locker (encrypt DEK with KEK). The bank locker key (KEK) stays with the bank in a highly secure vault. To read your document: Go to the bank → get your small key out (decrypt DEK using KEK)Use the small key to open the box (decrypt data using DEK) Simple! And very secure. Why Not Just Encrypt Data Directly With KEK? Two very practical reasons: 1. Performance The KEK usually lives inside a secure hardware vault or cloud service (like AWS KMS). If you send your entire 10GB file to KMS every time you want to encrypt or decrypt — that's painfully slow and expensive. Instead, you only send the tiny DEK (a few bytes) to KMS. The heavy lifting of encrypting actual data is done locally with the DEK. 2. Key Rotation Say after 6 months you want to change your encryption key (key rotation is a security best practice). Without envelope encryption — you'd have to decrypt ALL your data and re-encrypt it with a new key. Imagine doing that for terabytes of data! With Envelope Encryption — you only re-encrypt the DEK with the new KEK. Your actual data stays untouched. Much faster, much cheaper. Where Is Envelope Encryption Used? Literally everywhere in the cloud world: AWS S3 – When you enable server-side encryption on a bucketAWS RDS – When you enable encryption on a databaseGCP Cloud Storage – Envelope encryption is the defaultAzure Key Vault – Same pattern Every time you see that little "encryption enabled" checkbox on a cloud service — envelope encryption is what's happening under the hood. Now Let's Talk About PKI PKI stands for Public Key Infrastructure. In Part 3, we discussed Asymmetric Encryption — where you have a Public Key and a Private Key. Sounds great in theory. But here's a real problem. The Trust Problem Imagine Rahul wants to send an encrypted message to Priya. Priya shares her public key with Rahul. Rahul encrypts the message with Priya's public key and sends it. But wait — how does Rahul know that the public key he received is actually Priya's? What if a hacker intercepted the communication and swapped Priya's public key with their own? Rahul encrypts with the hacker's public key → hacker decrypts → reads the message. This is called a Man-in-the-Middle (MITM) Attack. We need someone that both Rahul and Priya trust, who can say — "Yes, this public key truly belongs to Priya." That trusted someone is called a certificate authority (CA). PKI — The Complete Picture PKI is a system made up of several components that together solve the trust problem. Let's go one by one. 1. Certificate Authority (CA) A CA is a trusted organization whose job is to verify identities and issue digital certificates. Think of them like the government passport office — they verify who you are and give you an official identity document (passport). Well-known CAs in the real world: DigiCert, Let's Encrypt, GlobalSign, Comodo. Your browser/OS comes pre-loaded with a list of trusted CAs. That's how your browser automatically trusts websites — because their certificates were signed by a CA your browser already trusts. 2. Digital Certificate A Digital Certificate is like a government-issued ID card for websites (or people or servers). It contains: The owner's name (e.g., google.com)The owner's Public KeyThe CA's name (who issued it)Expiry dateA digital signature from the CA When you open https://google.com — your browser checks Google's certificate. It sees the CA that signed it. It checks if that CA is in its trusted list. If yes — green light, connection is secure! That's the lock you see. 3. Digital Signature We talked about hashing in Part 3 — that you cannot reverse a hash. Digital Signatures use this + asymmetric encryption together in a clever way. When a CA wants to sign a certificate, it: Takes the certificate content and hashes itEncrypts that hash with its own Private Key That encrypted hash = Digital Signature. Anyone can verify the signature using the CA's Public Key (which is publicly available). If the decrypted hash matches the actual certificate content → the certificate is genuine and untampered! Think of it like a wax seal on an envelope. Anyone can see the seal, but only the king's ring (private key) could have made it. 4. Certificate Chain (Chain of Trust) In the real world, CAs have a hierarchy: Root CA → Intermediate CA → Your Website Certificate The Root CA is the ultimate trusted authority. It signs Intermediate CAs. Intermediate CAs sign individual website certificates. This chain is called the Chain of Trust. Why this hierarchy? Security! Root CA private keys are kept in ultra-secure, air-gapped hardware. They are almost never used directly. Intermediate CAs do the day-to-day certificate signing. If an Intermediate CA is ever compromised, it can be revoked without affecting the Root CA. Where Is PKI Used in Real Life? PKI is literally everywhere. You just don't see it because it works silently in the background. 1. HTTPS Websites (SSL/TLS) Every https:// website uses PKI. When you open your bank's website — a PKI handshake happens in milliseconds: Your browser asks the bank's server for its certificateThe bank's server sends its Digital CertificateBrowser verifies the certificate using the CA's public keyIf valid → browser and server agree on a secret key (using asymmetric encryption)All further communication uses that secret key with AES (fast symmetric encryption) That's TLS in a nutshell. And PKI is the backbone of all of it. 2. Email Signing (S/MIME) When your company sends you a digitally signed email — PKI is involved. The sender signs the email with their private key. You verify it with their public key from their certificate. You can be sure the email is genuinely from them and was not tampered with in transit. 3. Code Signing When you download an app or a software update — how does your phone/OS know it's legit and not malware? The developer signs the app with their private key. Your phone verifies it using the developer's certificate. This is why on Android you see the "Install from unknown sources" warning — no valid certificate found! 4. Government and Banking Your Aadhaar card, digital signatures on GST filings, net banking OTPs — all of these use PKI infrastructure behind the scenes. India's government runs its own CA called CCA (Controller of Certifying Authorities) under the IT Act. 5. VPNs and Internal Company Networks When you connect to your company's VPN, PKI certificates are used to verify that you are connecting to the genuine company server and not an imposter. Terms We Have Learned So Far (All 4 Parts) Cryptography, Algorithm, Plain Text, Key, Cipher TextSymmetric Encryption, Convergent Encryption, Initialization Vector (IV), Searchable EncryptionHashing, Hash/Digest, Avalanche EffectSalt, Rainbow Table Attack, Confirmation AttackAsymmetric Encryption, Public Key, Private Key, RSADEK (Data Encryption Key)KEK (Key Encryption Key)Envelope EncryptionKey Rotation, DeduplicationPKI (Public Key Infrastructure)Certificate Authority (CA)Digital CertificateDigital SignatureChain of TrustTLS/SSL That's a solid vocabulary now! Coming in Part 5... (Part 5 is in progress — stay tuned!) In the next part, we will gossip about: SSL/TLS Deep Dive – The full step-by-step TLS handshake explained simplymTLS (Mutual TLS) – How microservices talk to each other securelyHSM (Hardware Security Module) – The physical vault where the most sensitive keys liveZero-Knowledge Proofs – Proving you know something without revealing what it isAnd more... Stay tuned for Part 5! If you liked this blog, do give it a like and share it with someone who you think should learn this. Let's spread the knowledge! Read the previous parts here: Part 1 & 2 Part 3 More
Build Software Faster With Three Simple Principles

Build Software Faster With Three Simple Principles

By Ilia Ivankin
Development in small companies and startups often slows down at the boundaries between people and teams. A developer waits for a product decision. Frontend work stalls because the API response is unclear. QA discovers that two teams interpreted the same requirement differently. Three practical habits can reduce these delays: document the feature, discuss its technical risks, and agree on interface contracts before dependent implementations diverge. Keep each step proportional to the size and uncertainty of the task. Start With the Biggest Uncertainty Before distributing work, identify what the team needs to learn first. Unclear user experience? Start with a UI prototype. Frontend developers can use mocks to explore the flow while the team clarifies requirements.Uncertain business logic, performance, or integration? Start with a focused backend investigation or technical prototype.Unclear user problem or business value? Start with product discovery, involving users, product stakeholders, and technical specialists as needed. Having a UI does not automatically make frontend work the first priority. Choose the starting point that resolves the most important uncertainty, then agree on enough shared detail to let work proceed in parallel. 1. Maintain Feature Documentation and Define Responsibilities A concise Product Requirements Document (PRD) gives the team a shared explanation of what to build, why it matters, and how to evaluate it. For a small feature, a short page may be enough. QuestionWhat to recordWhat are we building?The intended behavior, MVP scope, and explicit exclusions.Who is it for?The users and the problem they face.Why is it needed?The expected benefit to users and the business.How will we evaluate it?Acceptance criteria, success metrics, a baseline or comparison group, and a measurement period. Name the person responsible for product decisions and identify the people needed to implement and validate the feature. These are responsibilities; a small team may combine several of them in one role. ResponsibilityTypical ownerSet goals, decide scope, and resolve product questionsProduct ownerAssess feasibility, design interfaces, and implement the featureDevelopers and technical leadDefine test scenarios and verify behaviorQA and developersDesign the user experienceDesigner, where neededDefine instrumentation and evaluate impactAnalyst or another explicitly assigned team member Example: A Post Recommendation System The following is an illustrative PRD. Its numerical targets are examples to agree on for a particular product, rather than measured results or universal benchmarks. SectionExampleProblem and hypothesisUsers may struggle to find relevant posts in a chronological feed. We expect recommendations based on their interests to improve content discovery.UsersReaders discovering posts; creators whose eligible posts can be recommended.MVP behaviorReturn up to 20 eligible posts, ranked by popularity within topics inferred from likes and subscriptions. Recompute lists every 24 hours. Exclude deleted posts and posts the requesting user cannot access when serving the response. Use a general popularity list when there is insufficient interest history.Experiment supportAssign eligible users to stable control and treatment groups and record recommendation impressions and clicks. Include this in the MVP so the first version can be evaluated.Out of scopeML models, similarity between users, and advertising recommendations. Decide whether to add these after evaluating the MVP.PerformanceIllustrative target: server-side p95 response time below 200 ms for serving precomputed lists at 1,000 requests per second on a representative test dataset. Measure the batch recomputation job separately and require it to finish within the 24-hour refresh window.Acceptance criteriaTests verify ranking on a fixed dataset, the popularity fallback, access filtering, and feedback recording. Likes and subscription changes affect the next scheduled recomputation. Load tests meet the stated latency target.Success measurementIllustrative primary target: a 10% relative increase in recommendation CTR versus the control group. Define CTR as clicks divided by recorded impressions. Plan an initial two-week experiment; estimate sample size before launch and report uncertainty if the result is inconclusive. Monitor seven-day retention and API error rate as guardrails, with acceptable thresholds agreed before launch.RisksWeak relevance, overexposure of already popular posts, and load spikes. Inspect recommendation diversity, provide a fallback, and test capacity before rollout. Technical notes can accompany the PRD, with implementation decisions reviewed during technical refinement. For example, a Go service could use PostgreSQL for source data and Redis for precomputed lists. A draft API might expose GET /recommendations for the authenticated user and POST /feedback for interactions. The team still needs to define schemas, error responses, authorization behavior, and pagination before implementation. Keep release acceptance separate from business success. A correctly implemented feature can fail to improve the chosen metric. That result should inform the next product decision. What if the Product Has No Documentation? Start with the workflow you are changing and the business rules most likely to be misunderstood. Link the relevant code, record open questions, and expand the documentation as the team learns. Documenting the entire system does not need to become a prerequisite for the next useful change. After release, compare outcomes with the original hypothesis. A weak result is a reason to investigate the feature, measurement, and assumptions. Some work provides value through reliability, lower operating costs, or reduced risk, and its evaluation should reflect that purpose. 2. Conduct Technical Refinement Technical refinement, sometimes called technical grooming, connects product expectations with implementation decisions. Its output should be a workable approach and a clear record of remaining questions. Ask a developer or technical lead to review the PRD and identify: Existing constraints and integration dependencies.Performance, security, and operational risks relevant to the change.Decisions that require input from product or another team.Unknowns that need a short investigation or prototype. Bring the relevant people together to resolve those questions. If the implementation cost changes the original assumptions, revisit the scope with the product owner. Keep the Process Proportional Agree on a timebox based on scope and uncertainty. For example, a modest feature might receive two days of initial technical analysis, with a named owner and deadline for each product question. A complex migration may require several investigations. These are planning choices to review as new information appears. Record decisions, tradeoffs, unresolved questions, and owners. There is usually no need to transcribe every discussion. Finish with a small implementation plan: what can run in parallel, what must happen first, and what evidence will show that a risky assumption holds. This process can reduce avoidable rework. It does not guarantee that every issue will be discovered in advance, so leave room to revise the plan during development. 3. Develop Using the Specification-First Principle For work that crosses an API boundary, agree on the contract early. An OpenAPI description can capture HTTP operations, request and response schemas, and other interface details. The contract should also be supported by examples and documented behavior where a schema alone is insufficient. Backend and frontend engineers review the contract together, including empty states, errors, and compatibility expectations.Frontend developers build against mocks that reflect the agreed contract.Backend developers implement the API and verify that its behavior matches the contract.QA and developers prepare API contract checks and derive end-to-end scenarios from the PRD, user flows, and business rules. API contract checks and end-to-end tests serve different purposes. Matching a response schema does not establish that a complete user journey works correctly. How This Reduces Waiting Consider an illustrative recommendation feature. If frontend developers invent a response shape while waiting for the backend, they may need to rewrite rendering and error handling during integration. Agreeing on the response schema, empty-list behavior, and refresh semantics first lets both sides work against the same assumptions. Mocks still need to match the implementation. Run compatibility checks in CI and update the specification, mocks, and tests together when the contract changes. Integrate early enough to expose mismatches before release. Specification-first development also leaves room for exploratory code. A short prototype may be necessary to discover whether a proposed interface is feasible. The aim is to agree on the contract before teams invest heavily in dependent implementations. For more on the approach, see Boost Efficiency With the Specification-First Principle. Keep Documentation Useful as the Project Evolves Documentation helps future team members understand behavior and the reasons behind earlier decisions. DORA's research on documentation quality links high-quality internal documentation with organizational performance and finds that it strengthens the impact of technical practices. Keep a small set of maintained resources close to the work: The PRD and the results of the feature's evaluation.Architecture decisions, API contracts, and relevant data models.Deployment and recovery instructions.Links to changes and the decisions behind them, in GitLab, a README, or another shared system.Test scenarios and acceptance checklists. Assign ownership and update these resources when behavior changes. Outdated documentation can mislead the next person just as missing documentation can leave them guessing. Conclusion These three habits address common sources of delay: Concise feature documentation gives the team a shared goal, scope, and definition of success.Technical refinement surfaces constraints and assigns owners to unresolved questions.Agreed API contracts support parallel development and reduce integration rework. Try them on one feature. Track time spent waiting for decisions, integration rework, and time from an agreed scope to release, alongside quality and product outcomes. Use what you learn to adjust the process. The practical goal is a team that can make decisions, build, and validate changes with less avoidable waiting. More
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
By Kai Wähner DZone Core CORE
Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus
Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus
By Akhil Madineni DZone Core CORE
How to Test POST API Requests With Playwright TypeScript
How to Test POST API Requests With Playwright TypeScript
By Faisal Khatri DZone Core CORE
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads

Batch processing remains vital because many business operations aren't suited to interactive requests. Tasks such as recalculating prices, reconciling transactions, migrating records, generating reports, processing invoices, reclassifying customers, or applying rules across millions of records may require considerable time. Handling these as standard requests leads to fragile systems, increased user wait times, frequent timeouts, challenging retries, and possible data inconsistencies. A batch model handles large workloads predictably, incrementally, and with control over progress and recovery. Rather than processing a massive operation as a single loop, batch processing uses jobs, steps, chunks, checkpoints, filtering, and restartability. This approach separates long-running data tasks from the user experience while delivering a structured execution model. In this article, we will focus on Jakarta Batch and examine its sustained relevance for modern enterprise applications. Why Batch Processing Still Matters in Enterprise Systems Modern applications offer various methods for background processing, such as message queues, event-driven architectures, schedulers, reactive pipelines, and distributed stream-processing platforms. While each addresses specific needs, batch processing is most effective when operations have a defined start and end, involve a known or discoverable dataset, and require controlled execution, progress tracking, restartability, or periodic processing. Batch processing remains essential in enterprise systems. Workloads such as financial reconciliation, billing, payroll, reporting, data migration, regulatory processing, catalog updates, and large-scale reclassification are still prevalent. In these scenarios, the priority is to process large volumes of work safely and predictably, rather than responding to individual events quickly. Batch provides a model specifically designed for these requirements. How Jakarta Batch Works Jakarta Batch organizes background processing into jobs and steps. A job defines the overall batch operation, while each step represents a specific stage. In chunk-oriented processing, a step follows a simple pipeline: read, process, write, and repeat until it processes all input. The Jakarta Batch runtime manages this lifecycle so application code can focus on reading, transforming, and persisting data. A job is the top-level unit of execution and represents a complete business operation, such as importing records, recalculating customer classifications, processing invoices, or reconciling transactions. Jobs can accept parameters at startup, allowing the same batch definition to run with different inputs or business rules. A job consists of one or more steps, each representing a distinct phase of the workload. Simple jobs may have a single step, while complex processes can use multiple steps in sequence, such as importing data, validating it, and generating a final report. Within a chunk-oriented step, the ItemReader supplies data to the runtime one item at a time, from sources such as a database or file. The reader only retrieves the next item and does not need to know how it will be processed or persisted.The ItemProcessor receives each item and applies business rules, such as validation, transformation, classification, enrichment, or filtering. It may return a modified item or null if the item should be excluded from writing.The ItemWriter receives processed items and persists or exports them. Unlike the reader and processor, which handle items individually, the writer typically receives a group of items from the current chunk. This enables more efficient database or bulk operations. Jakarta Batch adds features around this pipeline to support enterprise workloads. The runtime manages chunk boundaries, transactions, checkpoints, execution status, failures, and restart behavior. Chunk size determines how much work is grouped before a write and checkpoint, making it a key parameter for balancing throughput, memory usage, database cost, and recovery. The core model is straightforward: Job → Step → Read → Process → Write → Repeat Jakarta Batch keeps the business pipeline simple while the runtime manages the execution mechanics needed for reliable, long-running data processing. The Sample: Customer Segmentation with Jakarta Batch This example demonstrates the Jakarta Batch model using an e-commerce customer segmentation scenario. Customers are assigned to tiers such as Bronze, Silver, Gold, and Platinum based on configurable spending thresholds. When thresholds change, the application reevaluates the customer base and updates only customers whose classification has changed. The full application includes MongoDB integration, a Jakarta Faces UI, a preview workflow, validation, and supporting services. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. This section focuses on the classes directly involved in Jakarta Batch execution. Starting the Batch Job The application initiates the batch process through CustomerSegmentationService. Unlike the reader, processor, and writer, this class is not a batch artifact. Instead, it is an application service that retrieves Jakarta Batch’s JobOperator from BatchRuntime to start and monitor job executions. Java @ApplicationScoped public class CustomerSegmentationService { public static final String JOB_NAME = "customer-segmentation"; private volatile CustomerSegmentationPolicy currentPolicy; // initialization and status methods omitted public long start(CustomerSegmentationPolicy policy) { if (isRunning()) { throw new IllegalStateException( "A customer segmentation batch is already running"); } Properties parameters = new Properties(); parameters.setProperty( CustomerSegmentationPolicy.JOB_PARAMETER, policy.toJson()); long executionId = BatchRuntime.getJobOperator() .start(JOB_NAME, parameters); currentPolicy = policy; return executionId; } public boolean isRunning() { // implementation omitted } } The key API here is JobOperator, which Jakarta Batch provides as the interface for starting, stopping, restarting, and inspecting jobs. In this example, the segmentation policy is serialized into the job parameters to ensure each execution gets the correct business rules. Reading the Input The first batch artifact, CustomerItemReader, extends Jakarta Batch’s AbstractItemReader to implement a chunk-oriented reader. Java @Named("customerItemReader") @Dependent public class CustomerItemReader extends AbstractItemReader { @Inject private CustomerRepository customerRepository; private List<Customer> customers = List.of(); private int nextIndex; @Override public void open(Serializable checkpoint) { try (Stream<Customer> customerStream = customerRepository.findAll()) { customers = customerStream .sorted(Comparator.comparing(Customer::getId)) .toList(); } nextIndex = checkpoint instanceof Integer index ? index : 0; } @Override public Customer readItem() { if (nextIndex >= customers.size()) { return null; } return customers.get(nextIndex++); } @Override public Serializable checkpointInfo() { return nextIndex; } } These methods are part of the Jakarta Batch reader lifecycle defined by AbstractItemReader. open() prepares the reader and accepts a previous checkpoint if available. readItem() provides the next item to the runtime; returning null indicates there is no more input. checkpointInfo() reports the reader’s current position for checkpointing. For simplicity, this sample loads customers into memory. For larger workloads, the implementation might use pagination or a MongoDB cursor without changing the Jakarta Batch model. Processing Each Customer The next artifact implements Jakarta Batch’s ItemProcessor interface. Java @Named("customerTierProcessor") @Dependent public class CustomerTierProcessor implements ItemProcessor { @Inject @BatchProperty( name = CustomerSegmentationPolicy.JOB_PARAMETER) private String thresholdsJson; private CustomerSegmentationPolicy policy; @PostConstruct void initialize() { policy = CustomerSegmentationPolicy.fromJson( thresholdsJson); } @Override public Customer processItem(Object item) { if (!(item instanceof Customer customer)) { throw new IllegalArgumentException( "Expected a Customer item"); } CustomerTier calculatedTier = policy.tierFor(customer.getTotalSpent()); if (calculatedTier == customer.getTier()) { return null; } return Customer.builder() .id(customer.getId()) .name(customer.getName()) .totalSpent(customer.getTotalSpent()) .tier(calculatedTier) .build(); } } Here the Jakarta Batch contract is explicit: ItemProcessor defines processItem(). The runtime calls that method for every item produced by the reader. The processor applies the segmentation rule and either returns the transformed customer or null. Returning null has a specific meaning in Jakarta Batch: the item is filtered and does not continue to the writer. The @BatchProperty is also part of the Batch integration. It receives the thresholds property defined for this job execution, allowing the processor to reconstruct the CustomerSegmentationPolicy before processing begins. Writing the Results The final artifact extends AbstractItemWriter, Jakarta Batch’s base implementation for writing a chunk. Java @Named("customerItemWriter") @Dependent public class CustomerItemWriter extends AbstractItemWriter { @Inject private CustomerRepository customerRepository; @Override public void writeItems(List<Object> items) { List<Customer> customers = items.stream() .map(this::toCustomer) .toList(); customerRepository.saveAll(customers); } private Customer toCustomer(Object item) { if (item instanceof Customer customer) { return customer; } throw new IllegalArgumentException( "Expected a Customer item"); } } writeItems() is defined by the Jakarta Batch writer contract inherited from AbstractItemWriter. Unlike the processor, which receives one item at a time, the writer receives a collection of processed items. In this case, the collection contains only customers whose classification changed, as the processor has already filtered the others. At this point, the Java components of the pipeline are as follows: Plain Text CustomerItemReader extends AbstractItemReader ↓ CustomerTierProcessor implements ItemProcessor ↓ CustomerItemWriter extends AbstractItemWriter These types are what connect the application code to the Jakarta Batch runtime. Connecting the Artifacts With JSL The Java classes define the behavior, but Jakarta Batch requires explicit mapping of the reader, processor, and writer to each job. This orchestration is described in JSL: XML <?xml version="1.0" encoding="UTF-8"?> <job id="customer-segmentation" xmlns="https://jakarta.ee/xml/ns/jakartaee" version="2.0"> <step id="recalculate-customer-tiers"> <chunk item-count="20"> <reader ref="customerItemReader"/> <processor ref="customerTierProcessor"> <properties> <property name="thresholds" value="#{jobParameters['thresholds']}"/> </properties> </processor> <writer ref="customerItemWriter"/> </chunk> </step> </job> The ref values correspond directly to the names declared with @Named in the Java classes: Java @Named("customerItemReader") @Named("customerTierProcessor") @Named("customerItemWriter") The XML therefore tells the Jakarta Batch runtime: for this step, use this reader, then this processor, and finally this writer. It also maps the thresholds job parameter into the processor property. The item-count="20" sets the chunk size for this sample. Jakarta Batch coordinates reading and processing, periodically invoking the writer according to the chunk lifecycle and establishing transaction and checkpoint boundaries. The value 20 is for demonstration; real applications should tune chunk size based on processing cost, database behavior, transaction size, throughput, and recovery requirements. This structure is recommended for the article: present the class declaration first, then describe the lifecycle methods inherited from or required by Jakarta Batch. This approach helps the sample teach the API rather than simply presenting isolated methods. Conclusion Jakarta Batch is valuable because it transforms large-scale data processing into a structured execution model, eliminating the need for custom loops and ad hoc background logic. By separating reading, processing, and writing, and introducing runtime concepts such as jobs, steps, checkpoints, restartability, and chunk-oriented execution, it provides enterprise applications with a predictable approach to handling workloads involving thousands or millions of records. This allows implementations to focus on business logic, while the Batch runtime manages repetitive execution concerns, making the model easier to understand, optimize, and scale as workloads increase.

By Otavio Santana DZone Core CORE
The Request Timed Out, But the Payment Succeeded: Building Retry-Safe Mobile APIs
The Request Timed Out, But the Payment Succeeded: Building Retry-Safe Mobile APIs

Mobile networks are notoriously unreliable. A common scenario is when a user taps "Pay Now" in an app: the payment request reaches the server and is processed, but the network response never reaches the phone. The client assumes the request failed and retries, leading to the charge running twice. This is precisely the kind of bug that idempotency solves. Idempotency means that repeating the same operation has no additional effect, and the second attempt should recognize it's a duplicate and do nothing new. In practice, mobile engineers must treat payment or order APIs as idempotent by attaching unique operation identifiers to requests and deduplicating them on the backend. With this approach, even if the network drops a response or the user double-taps a button, the user is charged only once. Achieving this involves coordination between the app and the server. On the client side, every payment or mutation request is given a persistent unique ID. For example, the app might generate a new UUID when the user submits a payment, save that operation in a local "outbox" or queue, and include the ID in the HTTP request: Swift let opID = UUID().uuidString var request = URLRequest(url: URL(string: "/api/payments")!) request.httpMethod = "POST" request.setValue(opID, forHTTPHeaderField: "Idempotency-Key") request.httpBody = /* JSON payload of the payment */ This custom header (Idempotency-Key) carries the operation's identity to the server. (Modern APIs often expect this header to ensure that POST is treated safely.) The client's logic must be persistent; before sending, it writes the operation (ID and payload) into local storage (e.g., SQLite or SharedPreferences) so that it can recover and retry if the app closes or the network is down. This pattern, sometimes called the outbox pattern, means the user's intent is recorded immediately. A background dispatcher can then drain the queue on network availability or app restart; it tries each pending request. If a send fails (timeout, no internet, 5xx error), the entry remains in the queue for a later retry. This ensures at-least-once delivery and the server will eventually see the request, even if the app crashes or the network is flaky. But at-least-once alone would cause duplicates, so the server must be ready. When the backend receives the request with its Idempotency-Key, it first checks a deduplication store (for example, a database table keyed by this ID). Java String key = request.getHeader("Idempotency-Key"); PaymentResponse prev = idempotencyStore.lookup(key); if (prev != null) { // We have processed this request before - return the original response return prev; } // No record of this key; proceed with processing PaymentResponse result = processPayment(request.getBody()); // Store the result before returning it idempotencyStore.insert(key, result); return result; If the key already exists, the server simply returns the stored result without charging again. This ensures that the second (or third) time the client re-sends, the user doesn't get double-charged. Stripe's API, for instance, works exactly this way: it saves the outcome of the first request for a given idempotency key, and any retry with the same key returns the same result. In effect, the combination of at-least-once delivery (the client keeps retrying) plus idempotent handling on the server yields an effectively-once outcome. The server's idempotency store can be implemented with a simple database table that records each key and the operation's result. For instance, a processed_payments table might use the idempotency key as a primary key or unique constraint. The service then does an atomic INSERT ... ON CONFLICT DO NOTHING (PostgreSQL syntax) or equivalent. If the insert succeeds, the code proceeds with the payment and stores the result; if it fails because the key already exists, it knows this is a duplicate and can fetch the prior result. Wrapping the insert and the business operation in one database transaction avoids a race condition, as either both the key and payment record are written, or neither is. In SQL terms: SQL BEGIN; INSERT INTO payments(idempotency_key, user_id, amount) VALUES (:key, :userId, :amount) ON CONFLICT (idempotency_key) DO NOTHING; -- Check how many rows were inserted: IF (INSERT was successful) THEN -- This is the first time seeing this key; perform the payment CALL process_payment(...); -- The payment service may record a transaction ID, etc. COMMIT; ELSE -- Key already existed: rollback any partial work ROLLBACK; -- Retrieve and return the original payment result END IF; Even if two identical requests arrive concurrently, the unique constraint ensures only one succeeds in its insert. The other can detect the conflict and simply return the saved response. The system design sandbox guide describes this approach as "The database enforces uniqueness and no separate check needed. This works well when the idempotency record belongs in the same database as the business data, since you can wrap both in a single transaction." On the mobile side, it's also wise to guard against duplicates before the request is even sent. A simple in-memory or on-disk set of "seen" IDs can help reject retry loops after a crash or double tap. In Swift: Swift final class OperationDeduplicator { private var seen: Set<String> = [] func shouldProcess(_ id: String) -> Bool { return seen.insert(id).inserted } } This OperationDeduplicator returns true only the first time an ID appears. Persisting this set across app launches (for example in Core Data or a file) makes the app resilient to a crash after the payment is sent but before the response arrives. On relaunch, the app knows it already handled that operation and won't enqueue it again. It's important to integrate these pieces smoothly. A typical mobile flow might look something like this: the user submits a payment form, the app immediately generates a new opID (a UUID) and creates an operation record { id: opID, payload: {amount, items, ...} }. This record is saved locally. Then a background task picks it up, attaches opID as the Idempotency-Key header, and sends it. If the network call times out, the record stays queued. When the app regains connectivity or restarts, the dispatcher tries again. Because the same opID is used each time, the server knows to treat all retries as one. Only after the server successfully processes the payment does the app remove the operation from its queue. This pattern ensures retries and crashes do not cause duplicate side effects. Some systems even use more granular controls. For example, if the backend involves multiple microservices, one service might call others, and each service should propagate the same idempotency key or a related correlation ID so that the entire transaction remains idempotent. Distributed tracing can help debug how a request flowed through the system. Ultimately, the goal is to capture the entire user action from UI tap through backend processing and make sure it's only applied once globally. This often means also having the backend return the same HTTP status and response body on every retry, so the client never gets an unexpected error. Developers should simulate network failures and verify that retries do not cause double effects. Most important is to observe real production behavior, as logs or traces with the operation ID can tie multiple client attempts to a single transaction. If everything is correct, the system achieves effectively-once behavior where the payment occurs exactly once no matter how many times the client tries. As systemdesignsandbox summarizes, "at-least-once delivery + idempotent consumer = effective exactly-once". In practice, this means mobile apps can assume failures are not fatal and they can safely retry with the same key, knowing the server will protect against duplicates. Conclusion In summary, retry-safe mobile operations require treating each user action as an idempotent transaction. The client must persist a unique operation key and reuse it across retries, while the server must detect previously processed keys and prevent duplicate side effects. Combining durable client operations with server-side idempotency allows payments and other critical transactions to survive timeouts, crashes, and unreliable networks without being executed twice.

By Uthej Mopathi DZone Core CORE
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning

The Batch Processing Problem Batch processing isn't inherently a disadvantage. It becomes a problem when the business needs a decision now, but the architecture was designed to make that information available later. Picture an enterprise system processing millions of customer interactions. Transactions land across multiple systems throughout the day. Every few hours, a scheduled job extracts the data, transforms it, updates another system, and eventually makes it available downstream. This works fine — until the business asks: "Why can't we react to this the moment it happens?" A suspicious transaction. A changed preference. A completed payment. A signed document. A submitted service request. An account state change. In large enterprise environments, I've watched a fairly consistent pattern play out. Teams initially focus on throughput and infrastructure capacity — can the pipeline handle the volume, can it finish the batch window in time? As the systems mature, the harder questions shift elsewhere entirely: who owns a given event, how failures get recovered, how schemas evolve without breaking consumers nobody remembers exist, and — most importantly — what the actual business impact is when a consumer falls behind. Running the batch more frequently doesn't answer any of those questions. Eventually, the architecture itself has to change. Event-Driven Doesn't Mean "Install Kafka" This is where most transformations quietly stall. A common pattern looks like: batch system → add Kafka → the same tightly coupled design underneath. The organization now calls itself event-driven, but nothing structural has actually changed. Real event-driven architecture requires rethinking state ownership, service boundaries, data contracts, failure handling, consistency assumptions, observability, and operational responsibility — not just swapping the transport layer. Plain Text Business Action | v Producer Service | v Event Backbone / | \ v v v Risk Customer Analytics Svc Svc Svc The producer shouldn't need to know or care who's downstream. That's the real architectural benefit — a producer that stays ignorant of its consumers is what genuine decoupling looks like. If the producer still has to know which five systems need to be updated and in what order, you haven't built an event-driven system — you've built a batch job that happens to run on Kafka. Business Events Are Not Technical Messages There's an important distinction between commands and events. A command — UpdateCustomerProfile, SendNotification — says do something. An event — PaymentAuthorized, DocumentSigned — says something happened. Well-designed events represent durable business facts, not implementation instructions. Publish PaymentAuthorized, and Fraud Detection, Notifications, Analytics, Accounting, and Audit can all react independently, without the producer orchestrating any of them. That's the difference between an event-driven system and a batch system wearing a streaming costume. The Hard Problems Start After the First Event Duplicates Most messaging systems guarantee at-least-once delivery, so PaymentAuthorized may legitimately arrive twice. The customer shouldn't be charged twice. Idempotency — via event IDs, business transaction IDs, or a processed-event store — isn't optional polish. Duplicate delivery is a normal condition in a distributed system, not an edge case you occasionally trip over. Ordering If AccountClosed is processed before AccountCreated ever arrives, the consumer ends up holding a state that shouldn't be able to exist — an account that's closed but was never opened. The instinct is to enforce global ordering everywhere, but that kills scalability. The better question is narrower: what actually needs to be ordered? Usually it's events for the same business entity — the same customer, the same account — not the entire enterprise-wide stream. Schema Evolution An event schema gains a new field six months after a consumer was deployed against the old one. Does it break? Backward compatibility, schema registries, and contract testing matter here because events tend to outlive the applications that created them. Treat event contracts like governed APIs, not like internal implementation details nobody needs to track. Failure Don't retry forever. A sane strategy escalates in stages: an initial attempt, then a short retry, then backoff, then a dedicated retry queue, then a dead-letter queue for anything that still hasn't succeeded, then manual investigation or replay. A malformed or logically invalid "poison" event shouldn't be allowed to block the pipeline indefinitely just because it keeps failing the same way. Worth watching closely: retry count, dead-letter volume, consumer failure rate, and the age of the oldest unprocessed event. Eventual Consistency Changes How Teams Think In a synchronous system, an update and its visibility happen together — you write, you read back the new value, done. In an asynchronous architecture, that guarantee disappears. One consumer might reflect a change in twenty milliseconds; another might take two seconds; a third might be temporarily unavailable and catch up later. Different systems can legitimately hold different states for a period of time, and that isn't automatically a defect. The real architectural question is: how stale can this information safely become? Fraud decisioning tolerates almost none — a few hundred milliseconds of staleness can be the difference between catching and missing something. Marketing analytics can tolerate a great deal more. Audit cares more about completeness than about speed. This needs to be decided per business function, not applied as one blanket policy across the platform. It's also worth being honest about what "real-time" actually means in practice. A system that processes an event in milliseconds isn't meaningfully real-time if the downstream systems that act on that event take minutes to reflect the result. I've seen teams celebrate a fast event pipeline while the actual customer-facing decision — the offer shown, the risk flag raised — still lagged well behind because a downstream dependency hadn't caught up. Real-time decisioning has to be measured end-to-end, at the point where the business decision is made, not just at the point where the event was published. Migrating Off Legacy Without a Big-Bang Cutover Ripping out a legacy system in one motion rarely goes well. A more workable path is incremental: capture changes from the legacy system as events, route them through the event backbone, and let new and existing systems consume from the same stream during the transition. Plain Text Legacy System | v Change / Event Capture | v Event Backbone / | \ v v v New New Existing Svc Svc Systems The transactional outbox pattern is useful here: write the event to an outbox table in the same database transaction as the business update, then publish from the outbox separately. That avoids the classic dual-write problem, where the database commit succeeds but the event publish fails, silently leaving downstream systems out of sync. Change Data Capture can also help expose changes from a legacy system as a migration bridge. But it's worth being deliberate about this: a raw database row change is not automatically a well-designed business event. CDC tells you a row changed; it doesn't tell you why, or whether that change represents something a downstream consumer should actually care about. Treating every CDC record as a business event is one of the more common ways these migrations end up producing noisy, low-value streams. Observability Has to Be Designed In, Not Added Later A customer says: "My transaction disappeared." Where do you look, across a chain of services and events? You need correlation IDs, trace IDs, event IDs, business transaction IDs, timestamps, producer identity, and schema versions threaded through everything — plus the standard infrastructure metrics: consumer lag, event age, processing latency, retry rates, dead-letter volume, error rate. But infrastructure telemetry on its own isn't enough. Knowing "consumer lag is 12,000" is far less useful than knowing "12,000 customer transactions are currently delayed." That translation — from technical signal to business impact — is what tends to separate a platform that's merely instrumented from one that's genuinely observable. It's also usually the gap that shows up first when something goes wrong in production: the engineering team sees a metric, and it takes real effort to connect that metric to what a customer or a business stakeholder is actually experiencing. Security and Governance Have to Follow the Data Event-driven architecture multiplies how much data moves around a system, so security has to travel with the data rather than sit only at the application perimeter. That means clear authentication and authorization for who can publish and consume which topics, encryption both in transit and at rest, discipline about not routinely copying sensitive or personal information into every event just because it's convenient, defined retention policies, and clear auditability of who produced what and when. When Not to Use Event-Driven Architecture Don't adopt EDA because it's fashionable. A synchronous API is often the better choice when immediate request-response is required, the workflow is simple, only one system needs the result, or strong immediate consistency is essential. Batch remains entirely appropriate for monthly statements, historical reporting, bulk reconciliation, archival, and much regulatory reporting. The mature position isn't "everything must become event-driven." It's choosing synchronous, asynchronous, and batch patterns based on what the business actually requires — and having a clear answer for why. A Practical Decision Framework Before converting a workload, it's worth asking a short set of questions: Does the business genuinely require lower latency — or would nobody notice the difference between seconds and hours?Do multiple independent consumers need the same business change? If so, event-driven design becomes attractive.Can the business tolerate eventual consistency? If not, the workflow needs closer examination before proceeding.Can the organization actually operate distributed, asynchronous systems — with the observability, on-call practices, and schema governance that requires?What happens when one component fails? If the design can't answer that clearly before production, it isn't ready for production. A Reference Architecture Plain Text ┌──────────────┐ │ Channels │ └───────┬──────┘ │ v ┌──────────────┐ │ API / Domain │ │ Services │ └───────┬──────┘ │ Business Events │ v ┌────────────────────────┐ │ Event Backbone │ └────────────────────────┘ │ │ │ ┌────┘ │ └────┐ v v v Decisioning Notifications Analytics │ │ │ v v v Data Store Data Store Data Store ──── Observability ──── ────── Security ─────── ───── Governance ────── Observability, security, and governance aren't downstream services bolted onto the diagram — they span the whole architecture, or they don't really work. Conclusion The real transformation isn't batch → Kafka. It's delayed processing → continuous business awareness, and central orchestration → autonomous consumers responding to business facts. That shift comes at a cost: more distribution, more asynchronous behavior, more operational complexity, more governance overhead. So the goal was never to produce more events. The goal is systems capable of making timely, reliable decisions — while staying understandable and operable when, inevitably, something fails.

By Prem Kumar Gadhanki
How to Verify Response Data in API Testing With Playwright TypeScript
How to Verify Response Data in API Testing With Playwright TypeScript

One of the most important parts of API test automation is validating the response body to ensure data integrity. This step plays a key role in functional API testing, as it helps confirm that the API is returning the right data in the expected format. Response body validation isn’t limited to a specific request type; it applies equally to POST, GET, PUT, and PATCH APIs. The same validation approach can be used for any API response to verify the data returned by the service. Playwright offers multiple ways to validate response bodies. In this tutorial, I’ll walk you through these approaches to help you efficiently perform assertions on the response data using best practices. Checkout the previous tutorial blog to learn about Installation, the demo application, and how to send GET API requests with Playwright. How to Verify the Response Structure Response structure checks ensure that an API consistently returns data in the expected format, protecting the contract between backend services and their consumers. They help catch breaking changes early, such as missing or renamed fields, even when the API still returns a successful status code. TypeScript test("GET Order details and perform structure check", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toHaveProperty("message"); expect(responseBody).toHaveProperty("orders"); expect(responseBody.orders[0]).toHaveProperty("id"); expect(responseBody.orders[0]).toHaveProperty("product_name"); }); This test focuses on validating the structure of the API response. It validates that the response body contains the expected top-level keys and that each order object includes the required fields. Basic Assertions The basic assertions validate API success and data presence, making them a good first layer of verification before deeper structure or data-level checks. TypeScript test("Get order details and perform basic level verification", async ({ request, }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(Array.isArray(responseBody.orders)).toBeTruthy(); expect(responseBody.orders.length).toBeGreaterThan(0); }); This test performs a basic level check to confirm that the endpoint works as expected and returns the expected data in the response. After parsing the response body, the assertions focus on the following essential basic-level checks: TypeScript expect(responseBody.message).toBe("Order found!!"); The above line of code verifies that the API returns the expected message text in the response body. TypeScript expect(Array.isArray(responseBody.orders)).toBeTruthy(); This line of code ensures that the orders field in the response is an array, validating the basic response format. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This part of the test confirms that at least one order is returned in the orders array, ensuring the response contains required data. How to Verify Response Data With Details Validating the actual data returned in the response is essential to ensure that the API response contains the correct values. TypeScript test("Get order and verify order details", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); }); The following code ensures that the response has a valid identifier and it is not missing or empty. TypeScript expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); This check is required because the API generates the order ID when a new order is created in the system. It ensures that the “id” field has a valid value generated and assigned to it, since this “id” is used to retrieve, update, or delete order data. TypeScript expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); These statements assert that the order details are retrieved correctly for the respective request. The “user_id” - “1” was sent in the request, and verifying it in the response, along with the other order details such as “product_id” and “product_name,” ensures that the correct data is returned. How to Verify Response Data by Matching Objects and Arrays Playwright allows response data verification by matching objects and arrays partially within the API response. This approach is useful because it makes tests more flexible and confirms that the API returns the correct data structure and values. TypeScript test("Get order and verify matching object and array", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toMatchObject({ message: "Order found!!", orders: expect.arrayContaining([ expect.objectContaining({ product_id: "79", product_name: "5 star 10gm Chocobar", product_amount: 5, qty: 1, tax_amt: 0.5, total_amt: 5.5, }), ]), }); }); In this test, the toMatchObject assertion verifies that the response contains a “message” with the expected value “Order found!!” and an orders array. Within the array, "expect.arrayContaining" ensures that at least one order matches the expected data, while "expect.objectContaining" verifies only the values in the specified fields of that order. Using Best Practices to Perform Assertions Best practices create stable, maintainable API automation tests by combining basic checks with flexible data matching. TypeScript test("Get Order details API test with best practice", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); }); The test sends a GET request to fetch order details for “user_id”-“1". The use of failOnStatusCode: true ensures the test fails immediately if the API does not return a 2xx status code. The response is then parsed into a JSON object for validation. The assertions are structured in layers: TypeScript expect(responseBody.message).toBe("Order found!!"); This assertion verifies the message text, confirming that the API returns the correct message when an order is found. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This statement ensures meaningful data is returned and avoids false positives when the array is empty. TypeScript expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); The final part of the code performs the final assertion using arrayContaining and objectContaining to verify that at least one order has the expected “id” and “product_name”, without asserting every field. These layered validations improve clarity by verifying structure, data presence, and key data values in sequence. Extracting Data From the Response Extracting data from the API response is a common and widely used pattern in API test automation. It is important in multiple ways, such as reusing the data in further tests for dynamic testing and end-to-end validation. TypeScript test('Get order details and extract the order id', async({request}) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); const product_name = order.product_name console.log(product_name) }); This test sends a GET API request and performs basic validations to ensure the API response is reliable. TypeScript const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); The code above extracts the “order_id” from the order object in the response. Before accessing it, an assertion is made to verify that the value is not null. Finally, the value of the order_id is printed in the console. TypeScript const product_name = order.product_name console.log(product_name) Similarly, other values, such as product_name, can also be extracted. Attaching the Response Body to the Playwright Report The Playwright report, by default, shows the steps executed, the number of tests run, pass/fail status, and time taken to run the tests. However, it does not attach the response body to the test report. Attaching the response body to the report improves visibility and makes the test report more informative and transparent. The following code shows how to extract the required metadata and attach it to the Playwright report. TypeScript test("Get order details API and attach the response details to the report", async ({ request, }, testInfo) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, }); expect(response.status()).toBe(200); const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); const fullResponse = { status, statusText, headers, body, }; await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); }); The testInfo is a built-in Playwright fixture and provides utilities to manage and inspect test execution, such as attaching files to reports, updating test timeouts, and identifying the currently running test. The following lines of code extract the response metadata, such as the status code, status text, headers, and response body. TypeScript const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); Next, let’s combine all response details and create a single object containing: Status codeStatus textHeadersResponse body TypeScript const fullResponse = { status, statusText, headers, body, }; Finally, let’s attach these details to the report using the testInfo.attach() method as shown below: TypeScript await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); The testInfo.attach() adds an attachment to the Playwright report. The attach() method has 3 parameters: Name of the attachment: The first parameter is the name, “Full API Response”, that will be shown for the attachment.Body of the attachment: The second parameter is for the body of the attachment. The JSON.stringify(fullResponse, null, 2) has 3 arguments. The first argument converts the fullResponse object into a readable, pretty-formatted JSON. The second argument is the replacer, which is null. It ensures that all properties from the fullResponse object are included as they are, without modifying anything. The third argument controls pretty-printing. Here, “2” means indent nested JSON by 2 spaces.Content type: This parameter ensures that the report treats the attachment as JSON. The following screenshot is generated after the tests are run: Test Execution Running the tests in Playwright is simple and easy. We can run the following command from the terminal: Plain Text npx playwright test To generate the report, the following command can be used: Plain Text npx playwright show-report Summary Playwright provides multiple approaches, including structure checks and matching objects and arrays for verifying response data. The right strategy should be chosen based on your project’s requirements. Based on my experience, combining response structure checks with response data validation, including the matching object and array strategy, can be used as an effective approach for validating API responses. Happy testing!

By Faisal Khatri DZone Core CORE
6 Techniques To Reduce LLM API Costs With the Python Library
6 Techniques To Reduce LLM API Costs With the Python Library

Before diving into solutions, it helps to understand the scale of the problem. Take a common production pattern: a customer support bot that processes 10,000 messages per day, each with a 2,000-token system prompt and a 200-token user message. Cost breakdown by the numbers: scenariomodeldaily costNo optimizationClaude Opus ($5/1M)$11.00Right model for taskClaude Haiku ($1/1M)$2.20Add prompt cachingHaiku + caching ($0.10/1M cached)$0.42Combined savings96% reduction That's not a benchmark; it's arithmetic. The techniques don't require magic; they require applying what providers already offer. Technique 1: Prompt Caching, Up to 90% Off Repeated Content The Problem Most LLM applications send the same system prompt on every request. If your system prompt is 2,000 tokens and you make 10,000 requests per day, you're paying for 20 million input tokens daily, even though the content never changes. How It Works Anthropic's prompt caching lets you mark content blocks with cache_control. The first request pays full price and writes to cache. Every subsequent request that hits the same cached content pays 10% of the normal input price. Cache entries last 5 minutes and reset on each hit. The key insight: cache the most stable content first. Your base instructions change rarely. Your few-shot examples change occasionally. Your per-request context changes every time. Structure your prompt from most stable to least stable. Python from llm_optimizer import OptimizedClient, build_cached_system_prompt import anthropic client = OptimizedClient(anthropic_client=anthropic.Anthropic()) # Build an optimally structured cached system prompt system = client.build_cached_system( base_instructions=""" You are an expert customer support agent for a SaaS company. You have deep knowledge of our product, billing, and technical issues. Always be empathetic, clear, and solution-focused. [... 1,500 more tokens of stable instructions ...] """, # ← cached after first request — 10% cost on all subsequent calls few_shot_examples=""" Example 1: Billing question → here's how to handle it Example 2: Technical issue → here's the escalation path [... 500 tokens of examples ...] """, # ← also cached separately ) # First call: pays full price, writes to cache response1 = client.complete(messages=[{"role": "user", "content": "How do I cancel?"}], system=system) # Second+ calls: system prompt served from cache at 10% cost response2 = client.complete(messages=[{"role": "user", "content": "Where's my invoice?"}], system=system) What the Library Does llm-optimizer automatically injects cache_control breakpoints at optimal positions, system prompt, few-shot examples, and long conversation history, respecting Anthropic's 4-breakpoint limit. You don't touch the API directly. Savings Calculation Shell 2,000 token system prompt × 10,000 requests/day = 20M tokens/day Without caching: 20M × $3.00/1M (Sonnet) = $60.00/day With caching: 2M × $3.00 + 18M × $0.30 = $11.40/day Savings: $48.60/day = $17,739/year Technique 2: Model Routing, 60% to 80% Off by Using the Right Model The Problem Routing every request to your best model is the most common and most expensive mistake. Claude Opus costs 5x more than Claude Haiku. For tasks that Haiku handles perfectly, such as classification, extraction, translation, and simple Q&A, you're paying a 500% premium for no benefit. The Naive Approach and Why It Fails The obvious solution is to route by keyword: if the prompt contains "classify," use Haiku; if it contains "analyze," use Sonnet. This works until it doesn't. A prompt like "Explain the constitutional implications of this clause" is 8 words. Short, simple-looking. A keyword router sees no complexity signals and routes it to Haiku. But the task requires expert-level legal reasoning. This class of error intent-heavy short prompts is the primary failure mode of heuristic routing. The Better Approach: Use a Classifier llm-optimizer solves this by optionally using Haiku itself to classify task complexity before routing. The cost is approximately 15 tokens, about $0.000015. If that classification prevents one wrong Opus call (2,000 tokens × $5/1M = $0.01), it pays for itself 666 times over. Python from llm_optimizer import OptimizedClient, Provider client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_llm_classifier=True, # uses Haiku to assess complexity — ~$0.000015/call preferred_provider=Provider.ANTHROPIC, ) # Short prompt, complex intent → correctly routed to Opus response = client.complete( messages=[{"role": "user", "content": "Explain the constitutional implications of this clause"}] ) # You can audit the routing decision from llm_optimizer import ModelRouter router = ModelRouter(enable_llm_classifier=False) print(router.explain("classify this email as spam or not")) # { # "detected_complexity": "simple", # "routed_model": "claude-haiku-4-5", # "keyword_signals_fired": {"simple": ["classify"]}, # "token_count": 8 # } Complexity Tiers A closer look at complexity: Tierexamplesdefault modelSimpleClassification, extraction, yes/no, translationClaude HaikuMediumSummarization, paraphrasing, short Q&AClaude HaikuComplexCode generation, analysis, evaluationClaude SonnetExpertLegal reasoning, research, system design, math proofsClaude Opus Shell 1,000 requests/day — mixed complexity Without routing: all → Opus ($5/1M input) 1,000 × 500 tokens = 500K tokens × $5 = $2.50/day With routing: 70% Haiku, 20% Sonnet, 10% Opus 700 × 500 × $1 + 200 × 500 × $3 + 100 × 500 × $5 = $1.05/day Savings: 58% reduction Technique 3: Prompt Optimization, 5% to 20% Off Token Count The Problem Prompts written by humans, especially in collaborative or enterprise settings, accumulate filler. Phrases like "please note that", "it is important to note that", "in order to", and "due to the fact that" add tokens without adding meaning. At scale, this is a measurable cost. What the Library Strips Python from llm_optimizer import PromptOptimizer opt = PromptOptimizer() result = opt.optimize(""" In order to complete this task, please note that you should carefully analyze the following text. It is important to note that accuracy matters. Please be aware that your response should be concise. """) print(result.optimized_text) # "To complete this task, carefully analyze the following text. # Accuracy matters. Your response should be concise." print(f"Saved {result.tokens_saved} tokens ({result.savings_pct}%)") # Saved 18 tokens (31%) What is never touched: Code blocks, factual content, user-specified phrasing. The optimizer is conservative by default. It only removes patterns with no semantic value. Conversation history trimming: In long conversations, the library keeps the last N turns and drops older context, preventing unbounded token grow. Technique 4: Batch Processing, 50% Off Non-Urgent Requests The Problem Not every LLM call needs an immediate response. Nightly report generation, document indexing, data enrichment pipelines, and offline classification jobs all of these run fine with a delay. But most teams send them as real-time requests anyway, paying full price. How Anthropic's Batch API Works Anthropic's Message Batch API processes up to 10,000 requests per batch at 50% of the normal price. Results are available within minutes to hours. The trade-off is explicit: cost for latency. Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_batching=True, ) # Queue 1,000 document summaries throughout the day for doc in documents: client.queue( custom_id=doc["id"], messages=[{"role": "user", "content": f"Summarize: {doc['text']}"}], max_tokens=200, ) # Submit as one batch — 50% cheaper than 1,000 individual calls batch_id = client.submit_batch() # Poll when ready — minutes to hours depending on load results = client.poll_batch(batch_id, wait=True) for r in results: print(f"{r.custom_id}: {r.content}") When to Use It Nightly data processing pipelinesDocument indexing and enrichment Offline classification and tagging Report generation Technique 5: Document Compression to Reduce Context Before Sending The Honest Tradeoff This technique requires a direct warning: document compression is lossy. Removing content from a document to reduce token count means the model works with less information. For some tasks this is fine; for others it produces wrong answers. Use it only when: You've verified empirically that compression doesn't hurt your answer quality.You're doing rough extraction where completeness isn't required.You have re-ranking downstream (e.g., RAG pipelines). Do not use it for legal documents, compliance reviews, or any task where every sentence may be relevant. TF-IDF Extractive Compression When you do use compression, naive truncation (cutting from the end) is the worst strategy. llm-optimizer implements TF-IDF paragraph scoring where each paragraph is scored by its term overlap with your query, weighted by how unique those terms are across the document. The most relevant paragraphs fill the token budget; the rest are dropped. Python from llm_optimizer import DocumentCompressor # ⚠️ Read the accuracy warning before using in production comp = DocumentCompressor( max_tokens=4000, strategy="extractive", # TF-IDF scoring — best accuracy ) compressed, tokens_saved = comp.compress( document=long_contract, # 50,000 tokens query="payment terms and termination clauses" # focus compression here ) print(f"Compressed to {4000} tokens, saved {tokens_saved} tokens") # All compressed output includes a visible [⚠️ COMPRESSION WARNING] marker There are three strategies available: Strategy options: StrategyHow it worksbest forExtractiveTF-IDF scoring against queryWhen you have a specific querySmartKeeps first 60% + last 20%Structured documents with summariesTruncateHard cutoffWhen you need predictable behavior Technique 6: Cost Tracking That Observes Before You Optimize Why This Matters You can't optimize what you don't measure. Before applying any of the above techniques, you need to know: Which models you're actually usingWhere your token spend is goingWhether your optimizations are working Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), persist_tracking="usage.jsonl", # survives restarts ) # ... run your application ... client.print_summary() # ═══════════════════════════════════════════════════════ # LLM Cost Optimizer — Usage Summary # ═══════════════════════════════════════════════════════ # Total Requests : 1,247 # Total Cost : $0.8432 # Total Saved : $7.2180 (89.5% savings) # Cached Tokens : 8,432,000 # By Model : haiku: 891 reqs ($0.12) | sonnet: 312 ($0.58) # Optimizations : prompt_caching: 1247x | model_routing: 1247x # ═══════════════════════════════════════════════════════ The tracker records every request's tokens, cost, cached tokens, savings, latency, and which optimizations fired. Data persists to JSONL so you can analyze it across sessions or pipe it to your observability stack. Architecture: Why Not Just Use LiteLLM? The obvious question. LiteLLM is excellent and covers a lot of ground, including unified provider API, routing, cost tracking, batch processing. If you're not already using it, you should evaluate it. llm-optimizer does three things LiteLLM doesn't: Automatic cache_control injection: LiteLLM passes caching headers through but doesn't inject breakpoints at optimal positions automatically.Prompt filler stripping: LiteLLM has no token-level prompt optimization.TF-IDF document compression: LiteLLM has no query-aware document compression. The intended use is actually as a complement: llm-optimizer can sit on top of a LiteLLM setup, handling the prompt-level optimizations that LiteLLM doesn't touch. Streaming Support For user-facing applications, the library supports streaming: Python with client.stream( messages=[{"role": "user", "content": "Explain quantum entanglement"}], system="You are a physics tutor.", max_tokens=512, ) as stream: for chunk in stream: print(chunk, end="", flush=True) # Access token usage after stream completes usage = stream.usage() All optimizations, including caching, routing, prompt, and optimization, apply identically to streaming requests. Error Handling Production LLM applications need to handle rate limits and model overloads gracefully. The library handles this automatically: Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), max_retries=3, # retry on rate limit with exponential backoff retry_base_delay=1.0, # 1s, 2s, 4s ) # Rate limit (429) → retried with backoff # Model overloaded (529) → falls back to next capable model automatically # Non-retriable error → raises immediately Limitations Honest about what this doesn't do yet: Not org-scale validated: v0.4.0 is tested against ai-core Bedrock and Anthropic direct (14 live tests, all passing). Not yet run against production workloads at team scale. The pilot measures this.No async support: complete() and stream() are synchronous. Async support planned for a future release.OpenAI and Google partially tested: Anthropic and AWS Bedrock are the validated providers. OpenAI is implemented but not end-to-end tested in CI. Google Gemini streaming is not yet implemented.Token counting is approximate: The default estimator is within ~20% of the actual count. Install tiktoken for exact counts: pip install llm-optimizer[tiktoken].Pricing data can go stale: Stored in pricing.json with a version stamp. The library warns automatically if data is older than 30 days.Model allowlist: Only models listed in pricing.json can be routed to. Mythos and Fable 5 are not in the registry and cannot be called. Adding a new model requires a deliberate update to pricing.json. Python # Basic install pip install llm-optimizer # With exact token counting pip install llm-optimizer[tiktoken] # All providers pip install llm-optimizer[all] Python import anthropic from llm_optimizer import OptimizedClient client = OptimizedClient( anthropic_client=anthropic.Anthropic(), # All optimizations on by default except compression (lossy — opt-in) ) response = client.complete( messages=[{"role": "user", "content": "Your prompt here"}], system="Your system prompt here", ) client.print_summary() Links: PyPI: https://pypi.org/project/llm-optimizeGitHub: https://github.com/banerjeeso/llm-optimiz What's Next Async support (acomplete(), astream())LiteLLM adapterBudget guard: raise before a request exceeds a cost thresholdReal production benchmarks once I've run this against a live workload. Feedback welcome, especially from anyone who works with LLM APIs in production and can stress-test the routing logic or compression accuracy. Published on PyPI as llm-optimizer. MIT license. Contributions welcome.

By Somnath Banerjee
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency

Most write-ups on building an MCP server focus on the protocol itself: defining tools, handling requests, wiring up a client. That part is genuinely straightforward. What gets skipped over far more often is what changes when the tool you are exposing operates on files rather than returning data. File processing introduces a specific set of security and reliability problems that a typical read-only API does not have to think about, and getting them wrong is easy to miss until something goes badly. This is a rundown of the decisions that mattered most while building an MCP server that exposes document processing tools, merge, convert, OCR, and similar operations, and why a few of the obvious approaches turned out to be the wrong ones. Why File-Processing Tools Are a Different Security Case A typical MCP tool that queries a database or calls a read-only API has a bounded, predictable attack surface. A tool that accepts a file, or worse, a URL pointing to a file, and processes it does not. Two problems show up immediately that a simpler API rarely has to deal with. First, any tool parameter that accepts a URL is a potential SSRF vector. An MCP client could be tricked, directly or through a compromised upstream model response, into passing a URL pointing at an internal service, a cloud metadata endpoint, or an otherwise unreachable internal address. If the server naively fetches whatever URL it is given, that request happens from inside your infrastructure with whatever network access your server has. Treating every incoming URL as untrusted input, resolving it before fetching, and explicitly blocking private IP ranges and metadata endpoints is not optional for a tool like this, it is baseline. Second, file processing is expensive relative to a typical API call. Merging PDFs, running OCR, converting between formats, these all consume real CPU and memory per request in a way that a database lookup does not. That changes how rate limiting needs to work, which is worth its own section below. Auth: Why API Keys Plus JWT, Not Just One or the Other A single long-lived API key is simple to implement and simple to leak. Once issued, it is valid until manually revoked, and if a key ends up in a log file, a committed config, or a client-side integration by accident, there is no time-boxing to limit the damage. The approach that held up better in practice: bcrypt-hashed API keys for the initial authentication step, then a short-lived JWT issued from that exchange for the actual session. The API key never gets passed around on every request, only at the start, and it is never stored in plaintext server-side, so a database compromise does not directly expose usable credentials. The JWT that follows has a real expiry, which bounds how long a leaked token stays useful and gives you a natural mechanism for revocation without needing to invalidate the underlying key. This is not a novel pattern. It is standard practice in plenty of API design. The point worth making is that it is easy to skip for an MCP server specifically, because the tooling and examples in most MCP documentation default to a single static key for simplicity, and that default quietly becomes the shipped implementation if nobody revisits it. Idempotency: The Requirement Everyone Forgets Until It Bites MCP clients retry. Network hiccups, timeouts, a model deciding to re-invoke a tool call, all of these mean the same logical request can arrive at your server more than once. For a read-only tool, that is harmless, you just return the same data twice. For a tool that processes and charges against a file, a duplicate request means duplicate processing, potentially duplicate output files, and depending on your billing model, duplicate charges for a single user action. The fix is an idempotency key attached to each request, generated client-side and checked server-side before any processing begins. If a request with a given idempotency key has already been handled, the server returns the cached result rather than reprocessing. This sounds obvious once stated, but it is very easy to build a working MCP server that passes every test in development, where retries are rare, and only discover the gap once it is handling real, occasionally flaky client connections in production. Rate Limiting That Doesn't Punish Legitimate Use Because file processing is CPU and memory intensive per request, generic per-minute rate limits borrowed from a typical REST API tend to either allow abuse or block legitimate batch workflows, and it is hard to tune a single number that avoids both. Someone processing twenty files in a genuine batch workflow looks identical, from a naive rate limiter's perspective, to a script hammering the endpoint. What worked better was tracking limits per API key with enough granularity to distinguish sustained high-frequency abuse from a legitimate burst of activity, rather than a single flat request-per-minute ceiling applied uniformly. This is a harder problem to get exactly right than it sounds, and it is one area worth revisiting periodically as real usage patterns become clearer, rather than treating the initial configuration as final. Audit Logging as a Design Decision, Not an Afterthought It is tempting to treat logging as something you bolt on once a security question actually comes up. For a tool that processes user files, that is backwards. Knowing which API key touched which file, when, and what operation was performed needs to exist from the first deployment, not added retroactively after an incident makes it obvious it should have been there. This matters for debugging as much as for security, since a surprising number of support questions end up being answerable directly from audit logs rather than requiring back-and-forth with the user. What Would Have Saved Time in Hindsight Two things, if starting over. The first is deciding on the auth pattern, API key exchange plus short-lived JWT versus a single static key, before writing a single tool handler, rather than starting with the simpler static key for speed and migrating later. The migration is not hard technically, but it touches every existing integration and every piece of client documentation, so the cost of delaying the decision is mostly organizational rather than technical. The second is building the idempotency check in from the first tool, rather than adding it once a duplicate-processing report surfaces. It is a small amount of code, a lookup and a cache write around the start of request handling, but retrofitting it means auditing every existing tool for where duplicate execution would actually cause a visible problem versus where it is harmless, which takes longer than just building it in from the start would have. Putting It Together None of these individually are exotic ideas. Short-lived tokens over static keys, treating URL inputs as untrusted, idempotency keys for retryable operations, audit logging from day one, all of these are well-understood patterns in API design generally. What is specific to building an MCP server for file processing is that the combination matters more here than it does for a typical read-only integration, because the failure modes are more expensive: a duplicated file, a leaked key with no expiry, an SSRF hole reachable through a tool parameter, or an untracked operation on a user's document. If you are building or evaluating an MCP server that touches files rather than just data, these are the questions worth asking early, before the first real client connects to it, rather than after.

By Peter Ndumia
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team

Enterprise platforms don't fail gracefully. A stalled integration, a malformed payload, a service that silently drops a field — these aren't abstract bugs; they're business processes that stop moving and stakeholders who start calling. Leading a multinational engineering team across two regions that keeps a large platform's web service integrations running, I've spent the last several months evaluating where an agentic coding tool like Claude Code actually earns its place in that world — not as a novelty, but as infrastructure. Here's what that looks like in practice. Development: From Description to Diff The obvious use case is still the most valuable one. Claude Code reads an entire codebase — including the mix of legacy and modern service layers that most enterprise platforms actually run on — and can trace a bug from a symptom description down to the root cause across files it wasn't explicitly pointed at. Describe the integration you want built, and it plans an approach, writes the code, and verifies its own work before handing it back. For a legacy-adjacent stack paired with modern services, this matters more than it would in a greenfield project. Institutional knowledge about why an integration was built a certain way often lives in people's heads, not in comments. A CLAUDE.md file at the project root — read at the start of every session — becomes a place to encode that: architecture decisions, naming conventions, which endpoints are safe to retry and which aren't. Claude also builds its own memory as it works, picking up build commands and debugging patterns without being told to. Support: Turning Log Noise Into Signal Enterprise integration incidents rarely announce themselves cleanly. They show up as a batch job that partially completed, or a spike in errors from a downstream system. Claude Code's headless mode (claude -p) lets you pipe raw log output straight into a triage step: Shell tail -200 integration.log | claude -p "flag anything that looks like a failed payload" That same pattern scripts into runbooks, scheduled checks, or CI jobs — useful when your support rotation spans time zones, and nobody wants to be the one manually grepping logs at 2 a.m. local time. The Distributed Team Problem Leading engineers across multiple regions means standards drift is a constant risk — not from lack of skill, but from lack of shared context. Two teams solving the same class of integration problem several time zones apart will converge on different patterns unless something forces alignment. This is where Claude Code's team-facing features do real work: Skills package a repeatable workflow — a /review-integration command that checks a new service against your platform's conventions — so the process isn't re-explained every time it's needed.GitHub Actions or GitLab CI/CD integration lets Claude respond to @claude mentions on a pull request or run automated code review with severity-tagged findings, which matters when reviewers in one region are asleep while another region is shipping.Managed settings push CLAUDE.md files, permission rules, and MCP server allowlists centrally, so a Tech Lead sets the guardrails once instead of hoping every developer's local config agrees. MCP: The Part That Matters Most for Integration Work If your job is service integrations, the Model Context Protocol (MCP) is arguably more relevant than the coding assistance itself. MCP lets Claude Code connect directly to the systems around your codebase — ticketing tools, internal databases, Slack — so it can pull incident context or check a downstream system's schema as part of diagnosing a fix, instead of an engineer manually assembling that context first. For a platform with a web of service dependencies, that's the difference between "here's a plausible fix" and "here's a fix that accounts for what the downstream service actually expects today." Governance Isn't Optional Here Enterprise data often brings compliance weight that a typical greenfield codebase doesn't carry. Before rolling this out past a pilot, the governance layer needs real attention: Permission modes and sandboxing control what Claude can execute autonomously versus what needs a human in the loop — worth tuning tighter than the defaults for anything touching production data.Data handling and Zero Data Retention options are worth understanding fully if your platform has regulatory constraints on where data can transit.Analytics (on Team/Enterprise plans) gives adoption and PR-attribution data, which is useful less for surveillance and more for showing leadership that a pilot is actually paying off before expanding it. Where to Start Don't roll this out platform-wide on day one. Pick a single integration workstream, write the CLAUDE.md that captures your legacy-plus-modern conventions, and run a two-week pilot with one automated review pipeline (GitHub Actions or GitLab CI/CD, whichever matches where your repos already live). Measure whether incident triage time drops and whether code review turnaround improves across the time-zone gap. Expand from there. The tooling is capable enough now that the harder problem isn't whether it can help — it's building the guardrails and shared conventions that let a distributed enterprise team trust it with production-critical integrations.

By Balaji Venkatasubramaniyar DZone Core CORE
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest

A mobile request can fail without the server-side work failing. An iOS app may time out, lose the response after a POST has reached the service, or retry after connectivity changes while the original execution is still progressing. Apple explicitly distinguishes safe retry behavior by HTTP method and notes that URLSession can retry requests in some connection-loss cases, waitsForConnectivity can also cause the system to continue a request when connectivity returns. The dangerous state is therefore not “request failed,” but “completion is unknown.” If that request starts an agent that charges an account, reserves inventory, sends a message, or invokes an MCP tool, a second submission can become a second side effect. The Retry Boundary Is the Real Transaction Boundary “Exactly once” is too strong for a workflow crossing an iPhone, HTTP, an agent runtime, an MCP server, Kafka, a database, and an external API. Kafka can provide exactly-once guarantees within defined Kafka processing boundaries, but those guarantees do not atomically include arbitrary remote tool effects. The practical target is effectively-once behavior, and retries are expected, but every effect is guarded by a stable operation identity and converges on one committed outcome. Kafka’s idempotent producer suppresses duplicate records caused by producer retries, while transactional producers can atomically publish across Kafka partitions; the producer documentation also limits idempotence guarantees to a producer session and requires read_committed consumers for end-to-end transactional visibility. The operation identity must exist before the first network attempt. An iOS client can create an operationId when an action becomes durable local intent, persist it, and reuse it across transport retries. Transport material such as a server challenge may change, but the business ID must not. The server treats (subjectId, operationId) as a uniqueness boundary and stores a canonical payload hash with it. PostgreSQL unique constraints enforce row uniqueness, while INSERT ... ON CONFLICT provides an atomic conflict path under concurrency. SQL INSERT INTO agent_operation(subject_id, operation_id, payload_hash, status) VALUES (:subject, :operationId, :payloadHash, 'ACCEPTED') ON CONFLICT (subject_id, operation_id) DO NOTHING; A conflict with the same payload hash returns the existing operation; a different hash rejects key reuse. The record should exist before agent execution starts, and the accepted response should expose the durable operation identity. Let LangGraph Resume Without Repeating Effects LangGraph persistence is useful precisely because durable execution can replay code. With a checkpointer, LangGraph saves state at super-step boundaries; if execution resumes after a failure, an affected node can run again from the beginning. Official guidance consequently requires idempotent node logic, and task results can be checkpointed so completed task work can be reused during resume instead of recomputed. Replaying from an earlier checkpoint can also re-trigger later LLM calls and API requests. A stable business operation should therefore map to a stable LangGraph thread, while every effectful tool boundary receives the same operation ID. Python config = {"configurable": {"thread_id": operation_id} result = graph.invoke( {"operation_id": operation_id, "command": command}, config ) Checkpointing reduces recomputation but does not replace downstream idempotency. A reservation can succeed before its task result is durably checkpointed. LangGraph’s functional API therefore recommends idempotent tasks because an incomplete task can execute again during resume. Python @task def reserve_inventory(operation_id, sku, quantity): return mcp.call_tool("reserve_inventory", { "operationId": operation_id, "sku": sku, "quantity": quantity }) The significant property in this snippet is not the decorator. The important part is that the business identity crosses the graph boundary and reaches the tool implementation. A downstream inventory service can then use that identity to return a previously committed reservation rather than creating another one. MCP Tasks Are Durable Handles, Not Deduplication Keys The current MCP Tasks design is especially relevant to long-running agent tools. In the July 28, 2026 protocol revision, Tasks moved into the io.modelcontextprotocol/tasks extension. A server can return a durable task handle, and the client can poll with tasks/get, provide input with tasks/update, or request cancellation with tasks/cancel. The task is durably created before its handle is returned, which allows polling after a disconnect. That durability solves result retrieval after task creation, but it does not by itself deduplicate the request that creates the task. The task ID is server-generated. If the server creates task A, the response disappears, and the original tools/call is sent again, a naïve implementation can create task B. Therefore, the business operationId must be part of the tool arguments or equivalent application metadata, and task creation must first look up an existing operation. This follows directly from MCP’s server-generated task-ID model combined with retry ambiguity at the HTTP boundary. The MCP server can return an existing task handle for the same authenticated subject, operation ID, and payload hash, and later return the stored terminal result. Cancellation should also be idempotent because MCP defines it as cooperative rather than a guarantee that underlying work stops immediately. Keep Kafka Guarantees Inside Kafka Kafka is most valuable after the operation has been claimed. A database transaction can persist operation state with an outbox row carrying the same ID. Kafka producer idempotence protects against duplicates caused by producer retries, while consumers can still use the operation ID for application-level deduplication. Kafka transactions can atomically cover Kafka writes, but they do not extend over an MCP server or payment API. The event contract should preserve causality rather than inventing a new identity at each hop. JSON { "operationId": "8E7B6D9E-...", "type": "AgentToolCompleted", "tool": "reserve_inventory", "status": "SUCCEEDED" } A consumer can enforce uniqueness on (consumerName, operationId, eventType) or make the state transition conditional. Kafka delivery guarantees and application idempotency then reinforce each other instead of being treated as interchangeable. Bind Retry Identity to App Attest Without Blocking Legitimate Retries App Attest addresses a different failure mode: whether a request comes from a legitimate app instance and whether signed request material has been replayed or altered. Apple’s current guidance uses a server-provided challenge for assertions and requires the server to validate a strictly increasing assertion counter; that counter is specifically an anti-replay signal. Assertions are generated locally on the device after key attestation. The App Attest assertion must not become the business idempotency token. A legitimate retry should obtain fresh challenge material and generate a fresh assertion while retaining the original operation ID. The data hashed for the assertion can bind the server challenge, operation ID, and canonical payload hash together. Swift let payloadHash = SHA256.hash(data: body) let clientData = challenge + operationID.data + Data(payloadHash) let clientDataHash = Data(SHA256.hash(data: clientData)) let assertion = try await service.generateAssertion( keyID, clientDataHash: clientDataHash ) Apple recommends server-controlled challenges, server-side validation, and assertion-counter tracking as assertions are generated on demand without a round trip to Apple’s servers. The server verifies App Attest, checks that the challenge binds the operation ID and payload, then performs the idempotency lookup. A fresh assertion can retry the same operation; a replayed assertion fails anti-replay validation; an altered payload fails the hash check. Effectively-Once Behavior Is a Composition Property Reliable agent execution does not come from asking iOS to retry less often or from labeling a Kafka pipeline “exactly once.” It comes from carrying one durable business identity across every retry and every boundary, claiming that identity atomically before execution, making LangGraph effects idempotent under resume, using MCP Tasks as durable result handles rather than creation-time deduplication keys, restricting Kafka’s exactly-once guarantees to Kafka’s transactional domain, and using App Attest to prove request integrity without confusing anti-replay state with business deduplication. When those boundaries align, a lost mobile response can cause another HTTP attempt, another graph invocation, or another poll, but it does not cause another business effect. That is the operational meaning of effectively once.

By Uthej Mopathi DZone Core CORE
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?

Security teams usually describe an application through the assets they know about. This includes the production domain, documented APIs, the services currently in use, and the repositories connected to the latest release. But applications leave things behind as they change. A staging environment created for an old release may still be online months later, alongside an API version that was supposed to be retired. Other forgotten parts of the application can surface through DNS records, certificate data, or information left in client-side code. Most of these assets were created for perfectly legitimate reasons. Trouble starts when the work moves on, but the infrastructure doesn't. Ownership becomes unclear, security controls fall behind, and eventually an internet-facing component can remain active without appearing in the inventory used by the team responsible for it. I use unindexed attack surface to describe the gap between the application a team actively manages and the parts of it that are still reachable. Unindexed Does Not Mean Inaccessible It is easy to assume that an application resource is relatively safe when it doesn't appear in search results or isn't linked from the main application. That assumption falls apart once someone discovers the address. The same idea comes up when explaining the deep web. A large amount of online content sits outside conventional search indexes while remaining accessible through a direct URL, login, or other route. Application infrastructure can end up in a similar position. A staging host or old endpoint may be absent from the normal user journey and still be exposed to the internet. Finding these assets doesn't always require sophisticated techniques. During reconnaissance, an attacker can piece together clues from certificate records, DNS data, JavaScript, public repositories, and documentation. One discovery can lead to another until parts of the application that developers rarely think about become visible. robots.txt is a simple example. It tells compliant crawlers what they should avoid crawling, but it doesn't prevent someone from requesting those paths directly. OWASP's web security testing guidance includes reviewing web server metadata, identifying application entry points, and mapping execution paths during reconnaissance. Development and security teams can use similar techniques to see what their application exposes from the outside. Modern Delivery Creates Assets Faster Than Inventories Can Follow Modern development makes it easy to create infrastructure quickly. A pull request may generate a temporary preview environment, while a migration can leave /api/v1/ running as clients move to /api/v2/. During an incident, a debugging endpoint might be created and never removed afterward. Even a short-lived cloud experiment can leave behind a hostname outside the infrastructure account monitored by security. Keeping track of all of this gets harder as the application changes. Cloud platforms, API gateways, deployment configurations, and DNS may each hold a different piece of the inventory. Something created for one sprint can still be reachable several releases later. OWASP addresses this problem in API9:2023 Improper Inventory Management. Its guidance covers outdated API versions, exposed hosts, and missing documentation that can leave older parts of an application running without the security attention given to current services. Multiple development teams make that inventory harder to maintain. An environment may remain active after the developer who created it has moved to another project. Unless deployment and retirement update the inventory along with the infrastructure, these assets can stay online far longer than anyone intended. Forgotten Assets Often Retain Real Trust An old endpoint can outlive its original purpose without losing the access it was given. An earlier API version, for example, may still connect to the production database even though it hasn't received the authorization checks, rate limits, or input validation added to the current version. Staging environments can have the same problem when they use production-like data or continue running with an identity configuration that hasn't been reviewed in some time. Once an environment falls outside normal development work, security updates and monitoring are easier to miss. OWASP gives a useful example in its guidance on improper inventory management. A beta API host exposes the same password-reset capability as the production API, but the beta version lacks the rate limiting applied to production. Anyone who discovers the older host gets another route to the same function with fewer protections. The same situation can appear elsewhere in an application. A preview deployment might contain credentials left in an old build, while a forgotten administrative interface could still be reachable through a load balancer. These assets become even harder to manage when logging is no longer checked, or alerts still point to a team that has stopped owning the service. The longer an asset sits outside normal development and security workflows, the easier it is for its permissions, dependencies, and controls to fall behind the rest of the application. The Frontend Can Reveal the Backend’s Missing Map The browser often reveals more about an application than teams realize. It needs enough information to communicate with backend services, so production JavaScript can include API URLs, route names, environment identifiers, feature flags, and references to functionality users no longer see. This becomes interesting when the frontend has moved on, but the backend hasn't. A feature may disappear from the interface while its endpoint continues to respond. Commenting out a button or removing a route from the visible application doesn't remove the server-side functionality behind it. Source maps can make this easier to investigate. They help browsers reconstruct optimized JavaScript into something closer to the original source, making debugging easier. MDN's documentation explains how the SourceMap header and sourceMappingURL annotation point developer tools to these files. When source maps are publicly available in production, they can reveal original filenames and make the application's client-side structure easier to follow. That exposure doesn't automatically mean the application is vulnerable. Problems arise when an old endpoint is still reachable, authorization depends too heavily on what the frontend displays, or functionality exposed through the client was never included in the team's current security review. One useful check is to compare what appears in production JavaScript, browser traffic, and available source maps with the routes the team expects to have deployed. Unexpected endpoints deserve a closer look, especially when nobody can immediately explain why they are still there. Make Inventory Part of the Delivery Process Finding forgotten assets starts with looking past the list provided to the vulnerability scanning tool. If that list is incomplete, even a successful scan may leave parts of the application untouched. It is important to continuously verify the accuracy of the inventory against the discoverable assets outside the organization. This means collecting information from cloud accounts, deployment configurations, API gateways, DNS and certificate records, and exploring all hosts/endpoints that are reachable but cannot be placed anywhere in those records. The inventory needs to be closely tied to the deployment process as well. Whenever a service gets deployed, information should be collected about who owns the service, where it runs, what environment it belongs to, and whether it is public or private. This can be accomplished through a CI/CD pipeline as infrastructure is created or changed rather than having someone manage the spreadsheet manually. The same applies to API documentation. OWASP recommends generating API documentation automatically and including it in the CI/CD process. Teams can also compare deployed routes to the approved specification of the API so that an unexpected endpoint becomes visible while the application is still being worked on. From there, a few controls can catch problems early: Require ownership of public services before deploying them to production.Put expiration dates on preview and temporary environments.Flag unexpected new DNS or certificate records outside of the expected deployment process.Track deprecated API versions until they are removed.Keep production data in non-production environments only if it is absolutely necessary. Whenever something unexpected gets detected, assign it to someone who can determine why it is still running. Active services must receive the same level of security attention as the rest of the application. The ones that have served their purpose should be removed. Make Sure Retired Assets Are Actually Gone Removing a service from a repository or architecture diagram doesn't mean the service has disappeared from the internet. Old DNS records can remain, gateway routes may still forward traffic, and credentials created for the service can continue working after the team considers the project finished. Before shutting anything down, check whether it is still receiving traffic. An old API version may have clients nobody remembered, and immediately removing it could break an integration that is still in use. If the service needs to stay online, it should remain under the same monitoring, patching, and access controls as other active systems until those dependencies are dealt with. Once the service is retired, verify the result from outside the environment. Confirm that its hostname no longer resolves where it shouldn't, old routes no longer respond, credentials have been revoked, and any associated storage or third-party integrations have been removed. This step is easy to overlook during migrations or team changes, when responsibility for older infrastructure can become unclear. A service nobody considers active can still be reachable months later if no one checks that the shutdown actually happened. Conclusion Applications change constantly, and some of the infrastructure created along the way will eventually be forgotten. Problems begin when those old hosts, endpoints, and environments remain reachable without anyone checking whether they still need to exist. The inventory needs to change with the application. Build discovery into the delivery process, keep ownership clear, and verify that retired assets are actually gone. If something connected to your application is still reachable from the internet, your team should know why it is there and who is responsible for it.

By Igboanugo David Ugochukwu DZone Core CORE
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration

Every major AI vendor now supports the Model Context Protocol. The framing is almost always the same: MCP is the universal connector for AI agents in the enterprise. That framing sets up a false choice. MCP, REST/HTTP APIs, and Apache Kafka are not alternatives. They solve different problems at different layers of the architecture. Treating them as competing options produces systems that are fragile exactly where they need to be reliable. These three technologies can and do coexist in the same architecture. The question is not which one to pick. It is which one belongs where, and what the tradeoffs are when more than one could technically do the job. This article maps that decision: what each technology is built for, where the boundaries are, and where the genuine gray areas lie. 1. What Is MCP and What Is It Built For? Anthropic introduced the Model Context Protocol in November 2024 as an open standard for connecting AI assistants to external tools and data sources. Before MCP, every AI model required a custom connector to each external system. Three models, ten systems: thirty custom integrations to build and maintain. MCP collapses that to one standard interface. Any compliant client talks to any compliant server without prior coordination. OpenAI adopted MCP in March 2025. Google DeepMind confirmed support in April 2025. By December 2025, MCP had reached over 97 million monthly SDK downloads across Python, TypeScript, Java, Kotlin, C#, and Swift. Anthropic donated the protocol to the Agentic AI Foundation under the Linux Foundation, with AWS, Google, Microsoft, Bloomberg, and OpenAI as platinum members. MCP is no longer a developer experiment. Signals of enterprise maturity are arriving quickly: AI agents paying for API access autonomously, cross-SDK interoperability between Anthropic and OpenAI converging on MCP Resources, composable enterprise workflows where agents read tool signatures and compose cross-system flows without predefined paths, and an official MCP Registry launched in late 2025 as the community-driven server directory. The 2026 roadmap focuses on scalable transport, agent-to-agent communication, governance maturation, and enterprise readiness covering audit trails and SSO-integrated authentication. MCP handles tool access: how an agent calls an external capability. It does not handle agent-to-agent coordination, which is the domain of protocols like Google's Agent-to-Agent (A2A). MCP and A2A are complementary and address different layers of agentic architecture. The moment MCP is asked to do more than tool access, the architecture starts to break. Security Maturity Is Still Catching Up With Adoption Most incidents disclosed in 2025 and early 2026 are implementation failures, not protocol flaws. An Endor Labs analysis of 2,614 MCP implementations found 82% use file system operations prone to path traversal and 67% use APIs related to code injection. Enterprise-grade authentication with OAuth 2.1 and SAML/OIDC is on the 2026 roadmap but still in progress. The practical controls for today: apply least privilege, limit MCP server access to only the systems and data each tool requires, and monitor tool definitions for unexpected changes. 2. MCP vs. REST/HTTP API MCP and REST/HTTP APIs serve different consumers and should not be treated as interchangeable. REST is an architectural style built on HTTP, widely adopted but with no fixed conventions for discovery, error formats, or method naming. Well-designed REST APIs backed by OpenAPI specifications work well for direct, programmatic data access when a native SDK or versioned API already exists and teams know how to operate it. MCP enforces consistency at the interface level because the consumer is an AI model that cannot tolerate creative API interpretation. MCP standardizes how a tool is called. It does not standardize what the tool returns, how fresh that data is, or whether two agents calling the same tool simultaneously see the same state. For direct data access to vector stores, databases, or business application APIs, a well-governed REST API, native SDK, or Kafka Connect integration is almost always the better choice: lower latency, no protocol overhead, mature tooling. For giving AI agents standardized, discoverable access to a broader set of tools across vendors and frameworks, MCP is the right layer. The two are complementary, not competing. Tool Design Matters as Much as the Protocol Choice One important nuance on tool design: mapping one-to-one from existing APIs to MCP tools rarely works well. What matters is tool granularity, smart metadata, and thoughtful assembly of the MCP layer. An MCP server that exposes well-structured, semantically rich tools lets an AI agent reason about capabilities and compose workflows. This is reminiscent of the composability questions from the enterprise SOA (Service-oriented Architecture) era. SOA promised flexible service composition but delivered integration chaos when governance, metadata quality, and service granularity were treated as afterthoughts. MCP faces the same risk. The protocol is sound; what determines success is the discipline applied to how tools are defined, documented, and assembled. What MCP Does Not Do What MCP does not do matters as much as what it does. It does not manage data, guarantee message delivery, enforce governance, or guarantee consistency across systems. It is an interface layer, not a data pipeline. That boundary becomes even clearer when looking at what Kafka does, which is structurally different from both MCP and REST. 3. Apache Kafka: Event Broker, Decoupling, and the Backbone Role Operational data is the live data that runs business processes: order states, inventory levels, transaction records, customer accounts, risk scores. It originates in systems like SAP, Salesforce, Oracle, and mainframes, and it changes continuously. Kafka is architecturally different from both HTTP and MCP in one way that matters most: it decouples producers and consumers through a persistent, ordered, append-only log. With HTTP or MCP, the caller and the callee are coupled at request time. Every integration is point-to-point. If the target system is slow or unavailable, the caller is directly affected. Kafka breaks that coupling entirely. A producer writes an event once. Any number of consumers read it independently, at their own pace, using their own communication paradigm. One consumer processes records in real time. Another runs nightly batch analytics over the same events. A third powers a stream processing pipeline. A fourth writes results to a data lake via Apache Iceberg. All of them consume the same underlying data product. None of them affects the others. Kafka supports three consumption patterns from a single event stream: streaming, request-response, and batch. The event exists once; each consumer is independent. This is the pub/sub event broker model, and it is what makes Kafka the integration backbone between operational and analytical systems. The diagram below shows this decoupling: a single Kafka topic serving real-time applications, HTTP-based consumers, batch analytics, and MCP agent interfaces simultaneously. Stream Processing With Kafka Streams and Apache Flink Stream processing is a core complement to Apache Kafka, extending the platform from event transport into real-time data processing and decisioning. Kafka Streams is a lightweight Java library embedded in applications. It is well-suited for streaming ETL and simple to medium stateful stream processing without requiring a separate cluster. It integrates closely with existing JVM-based services. Apache Flink is a distributed stream processing engine designed for more complex workloads. It supports Java, Python, and SQL APIs, making it accessible to both application developers and data engineers. Flink runs as a dedicated cluster or in managed environments and is built for high-scale scenarios such as multi-stream joins, event-time processing, large state management, exactly-once semantics, Complex Event Processing (CEP), real-time analytics, and AI model inference. Both approaches extend Kafka with processing capabilities. The choice depends on workload complexity, required deployment model, and preferred programming language, not on replacing Kafka’s role as the event streaming backbone. A detailed comparison is available in the post Apache Kafka and Apache Flink: A Match Made in Heaven. Operational and Analytical Integration, Including the Data Lakehouse Kafka is not only for operational data integration. It serves as the ingestion layer into data lakes, feeds real-time analytical pipelines, enables stream processing with embedded AI models, and connects business applications bidirectionally. A governed data streaming platform provides schema registry, lineage tracking, role-based access control, and exactly-once delivery semantics across all of that. It serves both operational and analytical use cases and acts as the bridge between those two worlds. For how streaming and the lakehouse converge via Apache Iceberg, see Data Streaming Meets Lakehouse. Kafka's append-only commit log is the foundation of data consistency across the enterprise. Every downstream consumer sees the same data in the same order. That is not just a performance feature. It is what prevents the architecture where every system has its own version of the truth. 4. The Tradeoffs: It Is Not Black and White The choice between MCP, REST/HTTP APIs, and Kafka is rarely clean. All three can play a role in the same architecture. REST/HTTP APIs work well for operational data access when volume is moderate and a well-governed API already exists. A REST API backed by a Kafka-derived serving layer can return consistent, current data. The API is the interface; the streaming platform is what makes the data trustworthy behind it. A financial services firm exposing account balances via REST is not doing it wrong, as long as those balances are derived from a governed, consistent data source rather than pulled directly from a source system on every request. Kafka becomes the clear choice when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when the same events need to feed operational applications, analytical pipelines, and AI agents simultaneously. MCP fits best when access is supplementary, loosely coupled, and low-frequency. A support agent looking up a ServiceNow ticket before drafting a response, or a sales assistant pulling the latest slide deck from Google Drive before a call, are good fits. The key test is simple: does it matter if the data the agent receives is a few seconds or minutes old? If yes, MCP should not own that responsibility. If no, MCP is the right interface. SAP: Clean Separation Between ERP Integration and Developer Tooling The boundary between MCP and REST is not a choice between two equivalent options for the same integration. SAP is the clearest example of a clean separation. SAP exposes extensive REST and OData APIs for ERP integration: order management, finance, supply chain, procurement, and HR data flowing bidirectionally between SAP and other enterprise systems. SAP's MCP servers serve an entirely different purpose: developer tooling for ABAP code generation, CAP application development, UI5 and Fiori assistance, and operational tasks like transport validation and incident management. An architect connecting SAP order events to downstream systems uses OData and Kafka Connect. A developer asking an AI coding assistant to generate ABAP code uses the SAP MCP server. Different consumers, different use cases, different data. No overlap. Salesforce and ServiceNow: Same Data, Different Consumer Salesforce and ServiceNow follow a different pattern. Their MCP servers wrap the same underlying REST APIs and expose the same underlying data, but for a different consumer. A developer-written integration calls the Salesforce REST API directly with known endpoints and hardcoded logic. An AI agent calls the Salesforce MCP server, which wraps that same API to make it discoverable and stateful for an agent that cannot read documentation or manage its own session state. The data is identical. The access path differs based on who is consuming it. This is not a free choice between equivalent options. It is the same system serving two different client types through two different interface layers. REST vs. Kafka for Operational Data: The Harder Call The harder boundary is between REST and Kafka for operational data. Both can technically serve it, and that is where the real architectural decision lies. REST is simpler to start with but introduces point-to-point coupling, integration spaghetti at scale, and consistency risks when the same data needs to reach multiple consumers. Kafka is more complex to operate but provides the decoupling, consistency, and governance that enterprise architectures require when the same data needs to reach many consumers reliably. The two are not mutually exclusive. A common and well-proven pattern combines both: Kafka handles the event backbone, decoupling, and consistency, while a REST layer sits on top for synchronous request-response access, API management integration, or compatibility with systems that cannot speak the native Kafka protocol. This is particularly common in mobile applications, legacy system integration, and API gateway architectures. For a detailed look at how REST and Kafka complement each other in practice, see Request-Response with REST/HTTP vs. Data Streaming with Apache Kafka. 5. Decision Framework: MCP, REST/HTTP, or Kafka? Choosing between MCP, REST/HTTP, and Kafka is not a single decision but a set of tradeoffs that depend on data volume, consumer type, consistency requirements, and what is already in production. The comparison table below makes those tradeoffs concrete across eight dimensions. When to Use Which: A Guide to the Decision Tree The decision tree below walks through the same logic as a series of questions, routing to the right choice based on the integration's actual requirements. Use MCP when the integration is supplementary and tool-like: Slack, Google Drive, ServiceNow tickets, internal knowledge bases. The agent needs context to act, not a stream of events to react to. Eventual consistency is acceptable. Apply least privilege, monitor tool definitions for changes, and isolate MCP servers from production systems. Use a REST/HTTP API or native SDK when a well-documented API or SDK already exists and the engineering team knows how to operate it. The access pattern is direct, moderate-volume, and latency-sensitive. REST is also a reasonable choice for operational data when the backend is a governed Kafka-derived serving layer and consistency properties are inherited, not assumed. Use Apache Kafka when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when governance, lineage, and auditability are non-negotiable. Kafka is also the right choice when the same data needs to feed operational applications, real-time analytics, data lakes, and AI agents simultaneously. Use the real-time context engine when an AI agent needs current, consistent operational context for autonomous decisions. Kafka and Flink govern the data. MCP provides the agent interface. The consistency guarantee comes from the streaming layer, not from MCP. The practical question is not which protocol to choose. It is whether the data architecture underneath the agents can be trusted. Agents making autonomous decisions about inventory, risk, or customer service are only as reliable as the data they act on. 6. Where MCP and Kafka Work Together: The Real-Time Context Engine There is one pattern where MCP and data streaming complement each other directly: the real-time context engine. Kafka and Flink process and govern the data: ingesting from operational systems, applying transformations and filters, producing real-time materialized views. Those views are then exposed to AI agents through a standardized MCP interface. The streaming platform owns the data, its freshness, and its consistency guarantees. MCP owns the interface to the agent. Neither layer bleeds into the other's responsibility. Data consistency is not delegated to MCP. The streaming platform enforces it upstream before the MCP interface comes into play. The agent calls a tool and receives context that is current, governed, and consistent, not because MCP guarantees it, but because the streaming platform does. Any compliant AI agent, whether Claude, ChatGPT, Amazon Bedrock, LlamaIndex, or CrewAI, can call the context engine and receive current context from operational systems without needing to understand Kafka topics, Flink jobs, or schema evolution. An agent routing shipments from yesterday's inventory, approving transactions against a risk score from three hours ago, or reading an account balance that has not propagated: none of these is reliable. A real-time context engine eliminates this class of error at the source, reduces hallucinations, lowers inference cost, and anchors decisions to current operational reality. From Data Freshness to Agent Governance Enterprise readiness for this pattern also depends on how agents are governed once deployed. Trust, control, and accountability become central once agents start chaining decisions across domains. The context engine is the data layer of that answer. Governance of the agents themselves, covering what they are permitted to do, under what conditions, and with what audit trail, is the other half. This is the dimension enterprise buyers are actively evaluating when selecting agent orchestration platforms. The diagram below shows how the three layers fit together: the streaming platform as the data backbone, the context engine as the governed serving layer, and MCP as the clean interface to agents. 7. Conclusion: One Protocol, One Job MCP has earned its place in the enterprise architecture stack. What it has not yet earned is the role of universal integration layer, and understanding that distinction is what this article has been about. The broader architecture this sits inside connects three interdependent pillars. Event-driven data integration, with Kafka as the backbone, moves data reliably between operational and analytical systems and delivers governed data products to every consumer. Process intelligence is the orchestration layer that determines which decisions to automate, in what sequence, and under what conditions, giving agentic workflows the structure and governance they need to be trustworthy. Trusted agentic AI is where MCP plays its role: the standardized, governed interface through which agents access external tools and context, anchored to real data by the streaming layer beneath it. For a vendor-by-vendor analysis of trust and lock-in across the major AI platforms, see the Enterprise Agentic AI Landscape 2026. For a deeper look at how the three pillars fit together as an enterprise architecture framework, see The Trinity of Modern Data Architecture: Process Intelligence, Event-Driven Integration, and Trusted Agentic AI. One protocol, one job. That is the right way to use MCP.

By Kai Wähner DZone Core CORE

Monthly Top Integration Experts

expert thumbnail

John Vester

Senior Staff Engineer,
Marqeta

IT professional with 30+ years expertise in app design and architecture, feature development, and project and team management. Currently focusing on establishing resilient cloud-based services running across multiple regions and zones. Additional expertise architecting (Spring Boot) Java and .NET APIs against leading client frameworks, CRM design, and Salesforce integration.
expert thumbnail

Thomas Jardinet

IT Architect,
Rhapsodies Conseil

As an IT Architect with strong experience in Integration topics (with multiple contributions for Dzone Tech and Ref Cards), I accompany business projects in defining their architectures, whether functional, application or technical, by studying with them the best path. I also have more than I also accompany them in the organizational side, and above all I seek intellectual and human exchange. I am also a supporter of flattened organizations, as I think it greatly improves productivity, robustness, and resilience of companies

The Latest Integration Topics

article thumbnail
Gossips on Cryptography: Part 4
In this blog, we will continue our discussion from the previous parts. If you have not read them, please read them first.
October 1, 2026
by Sahil Aggarwal
· 1,083 Views · 1 Like
article thumbnail
Build Software Faster With Three Simple Principles
Reduce rework with three practical habits: concise feature documentation, technical refinement, and API contracts that let teams develop in parallel.
October 1, 2026
by Ilia Ivankin
· 1,053 Views · 1 Like
article thumbnail
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
Databricks and Snowflake speak the Kafka protocol, but Kafka for lakehouse ingestion is not Kafka as an event-driven architecture.
September 30, 2026
by Kai Wähner DZone Core CORE
· 2,889 Views · 1 Like
article thumbnail
Beyond HTTP Handoffs: Build Durable Agent-to-Agent Services With Temporal Nexus
Temporal Nexus enables durable agent-to-agent handoffs, managing long-running operations, retries, timeouts, and cancellation across service boundaries.
September 30, 2026
by Akhil Madineni DZone Core CORE
· 866 Views · 3 Likes
article thumbnail
How to Test POST API Requests With Playwright TypeScript
Learn how to test POST API requests in Playwright with TypeScript using static JSON objects and arrays, JSON.stringify(), JSON files, and the Faker library.
September 29, 2026
by Faisal Khatri DZone Core CORE
· 821 Views · 1 Like
article thumbnail
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch gives enterprise apps a standard model for long-running data processing with jobs, steps, readers, processors, writers, checkpoints, and tunable execution.
September 29, 2026
by Otavio Santana DZone Core CORE
· 1,203 Views · 3 Likes
article thumbnail
The Request Timed Out, But the Payment Succeeded: Building Retry-Safe Mobile APIs
Build retry-safe mobile APIs that prevent duplicate transactions during network failures, timeouts, and automatic retries.
September 28, 2026
by Uthej Mopathi DZone Core CORE
· 795 Views · 3 Likes
article thumbnail
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning
Batch processing works well for many workloads, but real-time decisioning requires event-driven architecture designed for resilience, observability, and failure handling.
September 25, 2026
by Prem Kumar Gadhanki
· 1,379 Views
article thumbnail
How to Verify Response Data in API Testing With Playwright TypeScript
Learn how to verify the response data, including structure checks, basic validations, and more, in API Testing with Playwright TypeScript
September 25, 2026
by Faisal Khatri DZone Core CORE
· 1,361 Views · 1 Like
article thumbnail
6 Techniques To Reduce LLM API Costs With the Python Library
Six techniques to cut LLM API costs by up to 90%: prompt caching, model routing, batch processing, and more. (Includes a pip-installable Python library.)
September 23, 2026
by Somnath Banerjee
· 2,451 Views · 2 Likes
article thumbnail
Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
Building an MCP server that processes files introduces problems a typical read-only API doesn't have. Here's what mattered.
September 22, 2026
by Peter Ndumia
· 1,946 Views · 2 Likes
article thumbnail
Architecting for <1s Latency: Managing Eventual Consistency in Distributed Search Platforms
To maintain sub-second search freshness, logistics systems must actively manage eventual consistency across Kafka ordering, search indexing, and cache invalidation.
September 22, 2026
by Dhruv Goel
· 1,969 Views · 2 Likes
article thumbnail
When Production Stops Moving: Running Claude Code Across a Distributed Enterprise Integration Team
Learn how Claude Code helps enterprise teams build, troubleshoot, and manage service integrations with MCP, CI/CD, automated reviews, and stronger governance.
September 21, 2026
by Balaji Venkatasubramaniyar DZone Core CORE
· 2,713 Views · 2 Likes
article thumbnail
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest
Stable operation IDs prevent iOS retries from duplicating agent tools across LangGraph, MCP Tasks, Kafka, and App Attest.
September 21, 2026
by Uthej Mopathi DZone Core CORE
· 1,978 Views · 3 Likes
article thumbnail
Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
Learn how forgotten internet-facing assets expand your attack surface and how continuous asset discovery, inventory, and ownership can reduce security risks.
September 21, 2026
by Igboanugo David Ugochukwu DZone Core CORE
· 1,548 Views · 2 Likes
article thumbnail
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration
MCP, Kafka, and REST APIs are not the same: this comparison maps each to the right layer of your agentic AI architecture.
September 18, 2026
by Kai Wähner DZone Core CORE
· 3,832 Views · 1 Like
article thumbnail
The New API Contract Is Probabilistic: Building Reliable Systems Around Unreliable Model Outputs
AI model outputs are unpredictable, so developers must use validation, testing, monitoring, and safe fallbacks to build reliable systems around them.
September 17, 2026
by Micheal Chukwube
· 2,346 Views · 1 Like
article thumbnail
The Trinity of Modern Data Architecture: Process Intelligence, Event-Driven Integration, and Trusted Agentic AI
Process intelligence, event-driven integration, and trusted agentic AI must be designed as one converged architecture for real business value.
September 16, 2026
by Kai Wähner DZone Core CORE
· 2,817 Views · 1 Like
article thumbnail
How to Perform Response Verification in REST-Assured Java for API Testing: Part 2
Master REST-Assured response verification in Java with Hamcrest Matchers, JSON assertions, API validations, and real-world examples.
September 11, 2026
by Faisal Khatri DZone Core CORE
· 2,977 Views · 4 Likes
article thumbnail
Prompt Caching: Overriding Tokenization for Faster and More Cost-Effective AI
Prompt caching allows AI systems to reuse the processing of unchanged token sequences, resulting in faster inference, lower latency, and reduced costs.
September 11, 2026
by Ravi Ranjan Shahi
· 3,588 Views · 2 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×