DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Coding

Also known as the build stage of the SDLC, coding focuses on the writing and programming of a system. The Zones in this category take a hands-on approach to equip developers with the knowledge about frameworks, tools, and languages that they can tailor to their own build needs.

Functions of Coding

Frameworks

Frameworks

A framework is a collection of code that is leveraged in the development process by providing ready-made components. Through the use of frameworks, architectural patterns and structures are created, which help speed up the development process. This Zone contains helpful resources for developers to learn about and further explore popular frameworks such as the Spring framework, Drupal, Angular, Eclipse, and more.

Java

Java

Java is an object-oriented programming language that allows engineers to produce software for multiple platforms. Our resources in this Zone are designed to help engineers with Java program development, Java SDKs, compilers, interpreters, documentation generators, and other tools used to produce a complete application.

JavaScript

JavaScript

JavaScript (JS) is an object-oriented programming language that allows engineers to produce and implement complex features within web browsers. JavaScript is popular because of its versatility and is preferred as the primary choice unless a specific function is needed. In this Zone, we provide resources that cover popular JS frameworks, server applications, supported data types, and other useful topics for a front-end engineer.

Languages

Languages

Programming languages allow us to communicate with computers, and they operate like sets of instructions. There are numerous types of languages, including procedural, functional, object-oriented, and more. Whether you’re looking to learn a new language or trying to find some tips or tricks, the resources in the Languages Zone will give you all the information you need and more.

Tools

Tools

Development and programming tools are used to build frameworks, and they can be used for creating, debugging, and maintaining programs — and much more. The resources in this Zone cover topics such as compilers, database management systems, code editors, and other software tools and can help ensure engineers are writing clean code.

Latest Premium Content
Trend Report
Platform Engineering and DevOps
Platform Engineering and DevOps
Trend Report
Developer Experience
Developer Experience
Refcard #291
Code Review Core Practices
Code Review Core Practices
Refcard #400
Java Application Containerization and Deployment
Java Application Containerization and Deployment

DZone's Featured Coding Resources

AI Agents Leaked 13,000 Screenshots: Why Enterprise Approval Controls Failed

AI Agents Leaked 13,000 Screenshots: Why Enterprise Approval Controls Failed

By Tim Freestone
Thirteen thousand internal screenshots from 343 technology companies are sitting in public GitHub repositories because a coding agent could not attach an image to a private pull request. Cybernews reported that the exposed material includes customer records, billing data, payment system screens, and unreleased product features. The agents appear to have completed the task using access they already held, but the route they took exposed data publicly. That last detail is the story. When a private repository could not render the image, the agents created public ones, mostly under employees’ personal accounts. No policy forbade it in a form software could act on, and no control stood between the task and the result. It would be comforting to call this an outlier. A survey of more than 900 executives and technical practitioners says it is closer to the norm. Gravitee’s State of AI Agent Security 2026 report found that only 14.4% of organizations have every AI agent go live with full security and IT approval, while 82% of executives say they feel confident their existing policies protect them from unauthorized agent actions. The sequence is deploy first, approve later, investigate after that. The survey comes from an API management vendor, so treat the figures as directional. Directional is enough. Nobody is reporting that agents wait politely for a security review, and the screenshot leak is what that looks like when the agents are good at their jobs. The gap between AI policy and enforcement Look beneath the confidence. A policy is not a control, and a count of incidents tells you how often something went wrong without telling you whether you could explain it to a regulator. The harder question is a plain one. Can you prove, on demand, who authorized this agent, which data it touched, and under what rule? The same survey puts a number on how far most organizations are from that answer. On average, only 47.1% of an organization’s AI agents are actively monitored or secured. Run that through an audit. An agent that is not actively monitored can leave critical gaps in the record. An agent without its own identity cannot be tied to the person who delegated the work, and agents that share credentials cannot say which of them acted. Each gap removes a link between an action and an accountable human, and an investigator needs every link. If the honest answer is “we would need to reconstruct it,” you do not have a security gap so much as an evidence gap. A detection gap belongs to the SOC. An evidence gap belongs to the CISO and the Chief Compliance Officer together, because one must produce the record and the other must stand behind it when a regulator’s clock starts. When an AI security gap becomes a compliance problem The clock runs in days. Kiteworks Data Security and Compliance Risk: 2026 Annual Survey Report found that 50% of organizations cannot produce a complete AI data access audit record within one business day. Notification windows, customer contract terms, and assessor timelines do not wait for an evidence package that takes weeks. I call the gap between confidence and enforcement governance theater, and I say it without contempt for the executives involved. They are reasoning from the best information they have: a written policy. Nobody has shown them whether that policy is technically enforced at the point where an agent reaches for data. A rule that no system enforces is a statement of intent, and auditors will read it as such. Regulators do not regulate models. HIPAA, PCI DSS, and the financial regulators do not ask whether a clinician or a program disclosed the data, and a screenshot of a billing console in a public repository raises the same questions either way. The obligation attaches to the data, and the evidence requirement attaches with it. That is also why a system prompt is not a compliance control. An instruction telling a model to stay away from certain material can be bypassed or changed, and an assessor will not accept “the agent was told not to” as proof of access control. Enforcement must sit with the data, independent of the model, and it must leave a record. What must change before the next agent ships Start by naming an owner for every agent. Not the team that built it and not the vendor whose model it calls, but one accountable person who can say what the agent may write, publish, and share, and who delegated the work. Ownership of AI security is unsettled in most organizations, and that vacuum is the cheapest gap to close. Next, give agents their own identities and tie each one to the human who authorized it. Humans and agents are two classes of identity that belong under one governance model, with one policy and one audit trail. The goal is not independence for agents. It is attribution for everything they do, with authority scoped to each action rather than inherited whole from a developer’s session. Then inventory every place an agent can write. List each destination an agent might use to make a human see an artifact, record who owns the account and whether the default is public, and refuse or human-gate anything world-readable outside your control. Treat screenshots and recordings as data, because they carry customer records and tokens and slip past rules written for documents. Finally, run the test your auditor will run. Pick a recent agent action and ask your team to produce the complete evidence package covering authorization, data accessed, policy applied, and the log that proves it. Time the answer. If it takes days, that number is your real exposure, and it will persuade a board faster than any survey. The organizations that close the approval gap will not be the ones with the longest policy binders. They will be the ones who can answer the auditor’s question before anyone asks it. Editor’s note: This article originally appeared on our sister publication, TechRepublic. More
Decoding the “Black Box”: Evaluating Agent Tool Chains in Production

Decoding the “Black Box”: Evaluating Agent Tool Chains in Production

By Gaurav Bhardwaj
One of the most interesting parts of my role as a cloud solution architect is working directly with customers as they move from experimentation to production. The conversations change significantly at that point. During an early proof of concept, the questions are usually foundational: Can the model answer the question?Can the agent call the API?Can we connect our enterprise data?Can we build the experience? But when customers start thinking about production, the questions become much harder: Can I trust the agent to take an action?How do I know it followed the business process?How can I prove that it used the right tool?What happens if it gets the right answer but takes the wrong path? I'll use a simplified travel scenario throughout this article. Imagine building an agent that can upgrade an airline passenger when certain business conditions are met. A user asks: “Can you upgrade Alex Johnson's SEA-to-JFK flight AA245 to business class if he is eligible?” The agent responds: “Alex is eligible, and I have successfully submitted the business-class upgrade request.” At first glance, everything looks good. The answer is clear. The request appears to have been completed. The demo works. But I found myself asking a different question: “What did the agent actually do before giving us that answer?” That question became much more interesting than the answer itself. From Evaluating Answers to Evaluating Behavior For a traditional generative AI application, evaluating the final response makes sense. Was it relevant? Grounded? Coherent? Accurate? Those questions don't disappear when we build agents, but agents introduce another dimension. An agent can understand an intent, decide which tool to use, generate tool arguments, execute the tool, inspect the result, choose another tool, take an action, and eventually generate a response. Microsoft Foundry describes this same challenge: production-ready agent applications need evaluation not only of the final output, but also of the quality and efficiency of the workflow that produced it. For our travel scenario, this might be the intended workflow: User request ➔ Look up customer ➔ Check upgrade eligibility ➔ Create upgrade request ➔ Return confirmation Now imagine the agent instead does this: User request ➔ Create upgrade request ➔ Look up customer ➔ Check eligibility ➔ Return confirmation The final answer could still be perfect. But the workflow is absolutely not. That is what I mean by the agentic black box. A Small Version of the Customer Problem When working through an architecture question, I try to make the problem as small as possible first to separate the core issue from enterprise complexity. I created three simple Python functions. The first finds the customer: Python import json def lookup_customer(name: str) -> str: """Find a customer using their name.""" customers = { "Alex Johnson": { "customer_id": "CUST-1001", "loyalty_tier": "Gold" } } customer = customers.get(name) if not customer: return json.dumps({"error": "Customer not found"}) return json.dumps(customer) The second determines whether that customer is eligible for an upgrade: Python def check_upgrade_eligibility(customer_id: str, flight_number: str) -> str: """Check whether the customer can receive an upgrade.""" if customer_id == "CUST-1001": return json.dumps({ "customer_id": customer_id, "flight_number": flight_number, "eligible": True, "reason": "Gold member with available upgrade inventory" }) return json.dumps({ "customer_id": customer_id, "flight_number": flight_number, "eligible": False }) Finally, the tool that performs the business action: Python def create_upgrade_request(customer_id: str, flight_number: str, target_cabin: str) -> str: """Create an airline upgrade request.""" return json.dumps({ "request_id": "UPG-9001", "customer_id": customer_id, "flight_number": flight_number, "target_cabin": target_cabin, "status": "submitted" }) In a real environment, each function could represent a Microsoft Foundry Agent, an Azure Function, an MCP Server, a Logic App, or an enterprise system. But from the agent's perspective, the important question remains: Which tool should I call, with what parameters, and at what point in the workflow? Turning Business Rules into Agent Instructions The customer requirement sounds straightforward: “Do not create an upgrade until eligibility has been confirmed.” That sentence is doing something very important — it is defining a business control. I can expose my Python functions to a Foundry agent as function tools and establish the Foundry project client. Python from typing import Any, Callable, Set from azure.ai.projects.models import FunctionTool, ToolSet from azure.ai.projects import AIProjectClient from azure.identity import DefaultAzureCredential import os user_functions: Set[Callable[..., Any]] = { lookup_customer, check_upgrade_eligibility, create_upgrade_request, } functions = FunctionTool(user_functions) toolset = ToolSet() toolset.add(functions) project_client = AIProjectClient( endpoint=os.environ["AZURE_AI_PROJECT"], credential=DefaultAzureCredential(), ) Now we create the agent, converting a business rule into expected agent behavior: Python agent = project_client.agents.create_agent( model=os.environ["MODEL_DEPLOYMENT_NAME"], name="travel-upgrade-agent", instructions=""" You help airline customers request flight upgrades. For every upgrade request: 1. Look up the customer first. 2. Check the customer's upgrade eligibility. 3. Only if the customer is eligible, create an upgrade request. 4. Never create an upgrade request before eligibility is confirmed. 5. Clearly explain the result to the user. """, toolset=toolset, ) Running the Customer Scenario We can create a thread, submit a test request, and execute the run: Python thread = project_client.agents.threads.create() message = project_client.agents.messages.create( thread_id=thread.id, role="user", content="Can you upgrade Alex Johnson's SEA-to-JFK flight AA245 to business class if he is eligible?", ) run = project_client.agents.runs.create_and_process( thread_id=thread.id, agent_id=agent.id, ) If we inspect the output and stop there, we have demonstrated that the agent can work. But we haven't demonstrated that it worked correctly. If an agent can create a ticket, modify a reservation, or provision infrastructure, I want visibility into the trajectory. Did it select the right tool? Did it send correct parameters? Did it take the next action at the right time? Using AIAgentConverter A Foundry execution produces multiple messages and interactions. Parsing all of that manually into an evaluation schema is tedious. That is where AIAgentConverter is useful. It converts a Foundry thread and run into the inputs expected by supported evaluators. Python import json from azure.ai.evaluation import AIAgentConverter converter = AIAgentConverter(project_client) converted_data = converter.convert( thread.id, run.id, ) print(json.dumps(converted_data, indent=2, default=str)) This was the point where the evaluation problem clicked for me. Instead of evaluating only the final string, I now have a normalized representation of the agent interaction that evaluators can inspect. Evaluation Question #1: Did the agent understand the user? "Upgrade the flight if he is eligible" is subtly different from "Upgrade the flight." Using IntentResolutionEvaluator, we can measure whether the agent correctly identified that condition. Python import os from azure.ai.evaluation import IntentResolutionEvaluator model_config = { "azure_deployment": os.environ["AZURE_DEPLOYMENT_NAME"], "api_key": os.environ["AZURE_OPENAI_API_KEY"], "azure_endpoint": os.environ["AZURE_OPENAI_ENDPOINT"], "api_version": os.environ["AZURE_API_VERSION"], } intent_evaluator = IntentResolutionEvaluator( model_config=model_config, threshold=3, ) intent_result = intent_evaluator(**converted_data) Evaluation Question #2: Did it call the right tools? Next, we inspect tool behavior using ToolCallAccuracyEvaluator. Python from azure.ai.evaluation import ToolCallAccuracyEvaluator tool_evaluator = ToolCallAccuracyEvaluator( model_config=model_config, threshold=3, ) tool_result = tool_evaluator(**converted_data) Why does this matter? If the agent selected the correct function but supplied the customer's name ("Alex Johnson") where the API expected an ID ("CUST-1001"), it fails. A language model could still generate an extremely convincing final response to cover this up. Agent behavior is part of the software surface we need to test. Tool Accuracy Is Not the Same as Tool Order Suppose the agent makes these valid calls: lookup_customer ➔ create_upgrade_request ➔ check_upgrade_eligibility. The tools and parameters are correct, but the sequence violates our business process. That is why I wouldn't use Tool Call Accuracy alone. For sequencing, Foundry provides Task Navigation Efficiency, which compares the actual sequence against an expected sequence using three modes: exact_match: Requires the exact same content and order. (Useful for strict payment workflows).in_order_match: Permits extra exploratory steps while preserving the expected order of mandatory steps.any_order_match: Expected steps can occur in any order. (Useful for research tasks). This flexibility proves that evaluation matching isn't just an AI decision—it's a business-process decision. One Successful Demo Is Not Enough A successful demonstration proves the agent worked once. Production readiness asks: How consistently does it work across model upgrades, API changes, and prompt tweaks? AIAgentConverter can prepare thread data for batch evaluation: Python filename = os.path.join(os.getcwd(), "agent_evaluation_data.jsonl") converter.prepare_evaluation_data( thread_ids=[thread_id_1, thread_id_2, thread_id_3, thread_id_4], filename=filename, ) evaluators = { "intent_resolution": IntentResolutionEvaluator(model_config=model_config), "tool_call_accuracy": ToolCallAccuracyEvaluator(model_config=model_config), } from azure.ai.evaluation import evaluate results = evaluate( data=filename, evaluation_name="travel-agent-regression", evaluators=evaluators, azure_ai_project=os.environ["AZURE_AI_PROJECT"], ) My Favorite Evaluation Cases Come From Failures When a customer discovers an edge case — like the agent submitting an action before validating eligibility — I don't just fix the prompt. I turn it into a permanent regression case: JSON { "query": "Upgrade Alex Johnson's flight if he is eligible.", "expected_actions": [ "lookup_customer", "check_upgrade_eligibility", "create_upgrade_request" ] } Over time, your evaluation dataset becomes a history of: "Things we have learned that this agent must never get wrong again." How I Now Think About Agent Evaluation My mental model for architecture discussions has shifted to this layered approach: Plain Text USER INTENT | v +---------+ | AGENT | +----+----+ | +-------------+-------------+ | | | v v v INTENT PROCESS RESPONSE | | | | +------+------+ | | | | | | v v v v v Understand Tool Input Order Quality Choice Intent resolution: Did it understand what the user wanted?Tool selection: Did it choose the right tool?Tool input accuracy: Were the parameters correct?Tool call success: Did the execution succeed?Tool output utilization: Did it correctly use the returned result?Task navigation efficiency: Did it follow the sequence?Response quality: Was the final answer useful and grounded? Where Microsoft Agent Framework Fits While AIAgentConverter is documented in Foundry's classic Agent Service workflow, newer applications built using the Microsoft Agent Framework can integrate Foundry evaluation more directly through FoundryEvals. A simplified current pattern looks like this: Python import os from azure.identity.aio import AzureCliCredential from agent_framework import Agent, evaluate_agent from agent_framework.azure import FoundryChatClient from agent_framework.foundry import FoundryEvals async def evaluate_travel_agent(): credential = AzureCliCredential() chat_client = FoundryChatClient( project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"], model=os.environ.get("FOUNDRY_MODEL", "gpt-4o"), credential=credential, ) agent = Agent( client=chat_client, name="travel-upgrade-agent", instructions=( "Help customers with flight upgrades. " "Always verify eligibility before submitting an upgrade." ), tools=[lookup_customer, check_upgrade_eligibility, create_upgrade_request], ) query = "Upgrade Alex Johnson's AA245 flight to business class if he is eligible." response = await agent.run(query) evaluators = FoundryEvals( client=chat_client, evaluators=[ FoundryEvals.INTENT_RESOLUTION, FoundryEvals.TOOL_CALL_ACCURACY, FoundryEvals.TASK_NAVIGATION_EFFICIENCY, ], ) results = await evaluate_agent( agent=agent, responses=response, queries=[query], evaluators=evaluators, ) for result in results: print(f"Status: {result.status}") print(f"Passed: {result.passed}/{result.total}") print(f"Report: {result.report_url}") (Note: Agent evaluation SDKs are evolving rapidly. Validate against the current Microsoft Learn documentation and your installed SDK version before production implementation.) The Real Customer Question Was Trust Looking back, what started this entire line of thinking wasn't really an SDK question. It was a customer asking, implicitly: “How comfortable should I be allowing this agent to take action in my business?” And I realized I could not answer that question simply by looking at the final response. I needed to understand: Plain Text What did it understand? What did it call? What did it send? What came back? What did it do next? Did it follow the process? Did it obey the business rules? That is why trajectory evaluation has become such an important part of how I think about agent architecture. The Path Is Part of the Product For a chatbot, the final answer may be the primary product. For an agent, I increasingly think: The path is part of the product. When an agent starts calling APIs, modifying systems, creating transactions, triggering workflows, or taking enterprise actions, evaluating only the final response is no longer enough. We need to evaluate behavior. For me, that is the real value behind capabilities such as: JSON evaluation_dimensions = [ "Intent Resolution", "Tool Call Accuracy", "Tool Selection", "Tool Input Accuracy", "Tool Output Utilization", "Tool Call Success", "Task Navigation Efficiency"] And it is why I think tools such as AIAgentConverter, the Azure AI Evaluation SDK, Microsoft Foundry evaluators, and Microsoft Agent Framework deserve a place in the architecture conversation much earlier than the final production-readiness review. Because before I tell a customer: “Yes, I think this agent is ready,” I want to be able to answer one additional question: “Do we know what it actually did?” That, to me, is where evaluating agents becomes much more interesting than simply evaluating answers. More
Building High-Performance Time-Series Applications With Java and QuestDB
Building High-Performance Time-Series Applications With Java and QuestDB
By Otavio Santana DZone Core CORE
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
By Jubin Soni, FBCS DZone Core CORE
Building IoT Time-Series Applications With Java and Apache IoTDB
Building IoT Time-Series Applications With Java and Apache IoTDB
By Otavio Santana DZone Core CORE
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices

The Illusion of a Single Database In a traditional monolithic application, maintaining data consistency is straightforward. If you need to create a new order and update warehouse inventory, you wrap the logic inside a single database transaction: Java @Transactional public void placeOrder(OrderRequest request) { orderRepository.save(request.toOrder()); inventoryRepository.decrementStock(request.getItemId(), request.getQuantity()); } If the inventory update throws an exception, the relational database rolls back the entire transaction. Either both operations succeed, or neither does. In a distributed microservices architecture, that safety net disappears. When your OrderService writes a record to a local PostgreSQL database and immediately publishes an event to an Apache Kafka cluster to notify the InventoryService, you are dealing with two completely independent, non-atomic systems. The Catastrophic Failure Modes When you attempt to write to a local database and publish an event within the same business method, one of two failures will inevitably occur: Scenario A (Database First, Network Fails) [Save to Database: SUCCESS] ──> [Network / Kafka Outage: FAILS] Result: Order exists in database, but downstream services are never notified. Scenario B (Publish First, Database Fails) [Publish to Kafka: SUCCESS] ──> [Database Unique Constraint Violation: FAILS] Result: Downstream services charge payment or pack inventory for an order that was never saved. Distributed two-phase commit (2PC) protocols are notoriously slow, fragile, and rarely supported across modern cloud-native message brokers. To achieve guaranteed consistency without blocking throughput, the industry-standard architecture is the Transactional Outbox Pattern. The Blueprint: The Transactional Outbox Pattern Instead of trying to speak to two external systems at once, the microservice performs all its operations within a single, local database boundary. Plain Text [Incoming Request] │ ▼ ┌───────────────────────────────────────────────────────────┐ │ Local ACID Transaction │ │ ├── 1. Insert into orders table │ │ └── 2. Insert into outbox_events table │ └───────────────────────────────────────────────────────────┘ │ ▼ [Outbox Relay / Change Data Capture (CDC)] │ ▼ [Message Broker: Kafka Topic] Atomic local write: The application saves the domain entity (orders) and a corresponding event payload into an outbox_events table inside the exact same local @Transactional boundary.Asynchronous relay: An independent background process reads the outbox table and publishes the messages to Kafka.Acknowledgment: Once Kafka acknowledges receipt, the relay marks the outbox event as published or removes the row. 1. Database Schema for the Outbox Define a dedicated outbox table designed for high-throughput polling and sequential reads: SQL CREATE TABLE outbox_events ( id UUID PRIMARY KEY, aggregate_type VARCHAR(255) NOT NULL, aggregate_id VARCHAR(255) NOT NULL, event_type VARCHAR(255) NOT NULL, payload JSONB NOT NULL, created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP, processed BOOLEAN DEFAULT FALSE, processed_at TIMESTAMP WITH TIME ZONE ); CREATE INDEX idx_outbox_unprocessed ON outbox_events (created_at) WHERE processed = FALSE; 2. The Application Layer: Atomic Persistence The Spring service writes both the domain entity and the outbox event in one atomic operation: Java package com.example.outbox.service; import com.example.outbox.dto.OrderRequest; import com.example.outbox.entity.Order; import com.example.outbox.entity.OutboxEvent; import com.example.outbox.repository.OrderRepository; import com.example.outbox.repository.OutboxRepository; import com.fasterxml.jackson.databind.ObjectMapper; import org.springframework.stereotype.Service; import org.springframework.transaction.annotation.Transactional; import java.time.Instant; import java.util.UUID; @Service public class OrderService { private final OrderRepository orderRepository; private final OutboxRepository outboxRepository; private final ObjectMapper objectMapper; public OrderService(OrderRepository orderRepository, OutboxRepository outboxRepository, ObjectMapper objectMapper) { this.orderRepository = orderRepository; this.outboxRepository = outboxRepository; this.objectMapper = objectMapper; } @Transactional public void createOrder(OrderRequest request) { // 1. Persist domain entity Order order = new Order(UUID.randomUUID(), request.getCustomerId(), request.getTotalAmount()); orderRepository.save(order); // 2. Persist outbox event inside the exact same transaction try { String jsonPayload = objectMapper.writeValueAsString(order); OutboxEvent outbox = new OutboxEvent( UUID.randomUUID(), "ORDER", order.getId().toString(), "ORDER_CREATED", jsonPayload, Instant.now(), false ); outboxRepository.save(outbox); } catch (Exception e) { throw new RuntimeException("Failed to serialize outbox event payload", e); } } } 3. The Relay Layer: Reliable Kafka Publishing An asynchronous background worker polls the unprocessed records in the outbox and delivers them to the broker: Java package com.example.outbox.relay; import com.example.outbox.entity.OutboxEvent; import com.example.outbox.repository.OutboxRepository; import org.slf4j.Logger; import org.slf4j.LoggerFactory; import org.springframework.kafka.core.KafkaTemplate; import org.springframework.scheduling.annotation.Scheduled; import org.springframework.stereotype.Component; import org.springframework.transaction.annotation.Transactional; import java.time.Instant; import java.util.List; @Component public class OutboxMessageRelay { private static final Logger log = LoggerFactory.getLogger(OutboxMessageRelay.class); private final OutboxRepository outboxRepository; private final KafkaTemplate<String, String> kafkaTemplate; public OutboxMessageRelay(OutboxRepository outboxRepository, KafkaTemplate<String, String> kafkaTemplate) { this.outboxRepository = outboxRepository; this.kafkaTemplate = kafkaTemplate; } @Scheduled(fixedDelayString = "${app.outbox.poll-interval-ms:2000}") @Transactional public void publishPendingEvents() { List<OutboxEvent> pendingEvents = outboxRepository.findTop50ByProcessedFalseOrderByCreatedAtAsc(); for (OutboxEvent event : pendingEvents) { try { // Publish using aggregateId as the Kafka partition key to preserve message ordering kafkaTemplate.send("orders-events", event.getAggregateId(), event.getPayload()) .whenComplete((result, ex) -> { if (ex == null) { event.setProcessed(true); event.setProcessedAt(Instant.now()); outboxRepository.save(event); log.info("Successfully relayed outbox event: {}", event.getId()); } else { log.error("Failed to relay event to Kafka: {}", event.getId(), ex); } }); } catch (Exception e) { log.error("Synchronous dispatch failure for outbox event: {}", event.getId(), e); break; // Halt batch progression to preserve order } } } } 4. Production Engineering Guardrails Preserve Partition Ordering Always use the aggregate_id (such as order_id) as the partition key when publishing the Kafka message. This ensures all state transitions for a single business entity land on the exact same Kafka partition and are processed in strict sequence. Polling vs. Log-Based CDC (Debezium) Scheduled polling is easy to set up and ideal for small-to-medium systems. For high-volume enterprise platforms processing thousands of writes per second, replace database polling with Change Data Capture (CDC) engines like Debezium. Debezium reads the database write-ahead log (WAL) directly, streaming changes to Kafka with zero query overhead on the operational database. Idempotency on the Consumer The Transactional Outbox pattern guarantees at-least-once delivery. If the relay publishes an event to Kafka but crashes before marking the outbox row as processed, it may re-send the message upon restart. Downstream consumers must maintain an idempotency check (e.g., tracking processed message IDs in Redis) to discard duplicates. Architectural Strategy Matrix Dimension Dual-Write (@Transactional + Kafka) Distributed 2PC Transactional Outbox Pattern Data Consistency Broken (Silent inconsistencies) Strong Strong (Eventual consistency) System Latency Low High (Blocking locks) Ultra-low local execution Broker Resilience Fragile (Network crashes drop events) Low High (Decoupled publishing) Operational Simplicity Deceptively simple Complex Straightforward Summary In distributed systems, atomicity cannot cross network boundaries. When you attempt to update a local database and publish a message to an event bus inside the same method, failure is a mathematical certainty over time. By shifting to the Transactional Outbox Pattern, you leverage the battle-tested ACID guarantees of your relational database to capture domain state and outbound events simultaneously. This eliminates the dual-write anti-pattern, guarantees at-least-once delivery, and builds a dependable bridge between relational transactions and event-driven architecture.

By Rahul Tewari
OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule
OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

OpenSearch is an open-source, distributed search and analytics suite derived as a fork of Elasticsearch and maintained under the Apache 2.0 license. When it comes to memory configuration, the guidance is often reduced to a few rules of thumb: swapoff -a, vm.swappiness=1, or bootstrap.memory_lock, and allocating 50% of available memory to the JVM heap while leaving the rest for Lucene and the filesystem page cache, OpenSearch off-heap caches, network buffers, and other system needs. These recommendations are repeated throughout documentation, blog posts, and operational guides, yet their origins and the mechanisms that justify these specific values are rarely examined. Undoubtedly, they provide a reasonable and safe starting point or a safe upper bound in most of the cases, but a safe default is not necessarily an optimal configuration. All of this raises even more questions. How do these defaults affect cluster performance? What is the optimal JVM heap ratio? Does memory given up by the JVM actually become filesystem page cache, and at what point does that trade-off stop paying off? How do read/write latency correlate with the heap ratio? These questions become particularly important in resource-constrained environments and in the cloud, where long-term contracts may make existing instances significantly cheaper, making horizontal or vertical scaling a difficult decision. In this article, we'll try to answer these questions through benchmarking. This is Act 1 of a two-act series. Act 1 focuses on identifying the cause of the latency problems we observed with the current defaults. Act 2 will explore what other heap-ratio values might look like for read/write loads. Along the way, I'll share the tools and commands used throughout the investigation, making this article a practical reference as well for you and for myself when I inevitably need to retrace the investigation months later. Knowledge Context The story also crosses several boundaries, such as the Kernel VM, the JVM, and Lucene. So, it’s important to outline the concepts mentioned in this part of the article beforehand, both for the context and, optionally, to enrich the AI context if you'd like to summarize everything. AreaWhere memory livesWhy it matters hereLinux page cacheFile-backed RAMLucene relies heavily on it for index data; under memory pressure, these pages can be reclaimed and read again later.Linux swapDisk-backed anonymous memoryAnonymous process memory can be swapped out under pressure. vm.swappiness influences this decision but does not prohibit it.Linux PSIKernel pressure signalShows time tasks spend stalled due to CPU, memory, or I/O pressure. We'll use I/O PSI while investigating latency.JVM heapAnonymous memoryControlled by Xms/Xmx; contains Java objects and several OpenSearch data structures.JVM native memoryAnonymous/file-backed memory outside XmxIncludes code cache, metaspace, stacks, direct buffers, and native allocations. Heap metrics do not account for all of it.OpenSearch cachesHeap/off-heap, depending on cacheTheir sizes may depend on heap size, which becomes important when we change jvm_heap_ratio.OpenSearch indexing bufferHeapIts size depends on heap and therefore becomes an important variable in Act 2.Lucene mmapFile-backed/page cacheLucene index files mapped into the process do not consume JVM heap; resident pages compete for physical RAM. Environment I used Aiven for OpenSearch on Azure, with cluster metrics exported to Thanos. The OpenSearch Benchmark metrics don't provide everything we need to answer our questions, particularly host-level metrics such as Linux PSI and swap activity. Exporting the cluster metrics to Thanos allows us to use PromQL queries later to retrieve the additional metrics needed for the investigation. The cluster consists of 3 nodes: CPUAMD EPYC 7763v (Milan)vCPU / RAM2 vCPU, 8 GiBDisk Size175 GiB per nodeAzure Regionazure-westeuropeAzure DiskPremiumV2_LRSAzure SKUStandard_D2as_v5 OpenSearch Version3.6.0JDKjava-21-openjdk-headlessGCG1GC Act 1. The Latency and an Extra GB The symptom: elevated query latency across the cluster, first reported by the customer after a kernel and Azure image upgrade. The load pattern on OpenSearch itself remained unchanged, as did the cluster configuration and settings. The monitoring panels give us the first clue. In the screenshots below, the green vertical line marks the moment of the upgrade. After that point, the page cache grows by roughly a gigabyte, while I/O PSI, previously close to zero, starts showing significant spikes. Nothing crashed, no alert fired, and from the JVM point of view everything looked normal. So where did that extra gigabyte of page cache come from? Nothing was actually freed. It moved. The interesting part isn't just that memory went to swap; it's which memory. When you have thousands of running clusters, there is always a small fraction of them operating close to the edge: relatively stable, yet sensitive enough that even a small change can noticeably affect performance. Like a star nearing the end of its lifetime, they may look stable right up until something disturbs the balance. The immediate cause of the page-cache change was identified fairly quickly: Azure applies tuning parameters that differ from the Linux kernel defaults, and those parameters were not applied by the older image. Once applied, the larger buffers and read_ahead increased the filesystem cache footprint, putting additional pressure on anonymous memory and eventually pushing some of it to swap. But rather than stopping there, let's use this incident as an opportunity to experiment with the heap ratio and make the behavior of OpenSearch instances explicit and less dependent on such environmental changes. Evidence It Is on Swap; None of It Locked First, find what is going on on a node itself: Shell PID=$(pgrep -f 'org.opensearch.bootstrap.OpenSearch') grep -E 'VmRSS|RssAnon|RssFile|RssShmem|VmSwap|VmLck' /proc/$PID/status Shell VmRSS: 4366408 kB # resident RssAnon: 3329896 kB # heap + anonymous native RssFile: 1036496 kB # resident mmap'd Lucene pages RssShmem: 16 kB VmSwap: 2679496 kB # on swap VmLck: 0 kB # bootstrap.memory_lock=fasle, none locked Shell grep -E 'MemFree|MemAvailable|Cached|SwapFree' /proc/meminfo Shell MemFree: 258216 kB MemAvailable: 3479264 kB Cached: 3403196 kB SwapCached: 854396 kB SwapFree: 4745444 kB The Swap Device Is dm-crypt Then confirm there is somewhere for it to go, and on what kind of device: Shell swapon --show Plain Text NAME TYPE SIZE USED PRIO /dev/dm-4 partition 8G 3.5G -1 The dm-* swap device is the interesting detail to catch and to keep in mind. This is an encrypted device, so once a page is requested it could drive more I/O -> more dm-crypt allocations -> more high-order pressure. A self-reinforcing loop and a good example of read amplification. Paging Is Live, Not Historical The next logical step is to check whether it is live paging or just a stale historical tail. The vmstat 1 5 the Linux Virtual Memory Statistics Tool should give us an exact answer for this, where non-zero swap blocks in and out (marked as si, so): Shell vmstat 1 5 Plain Text procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 1 0 3620928 159716 5864 3704212 113 80 3361 577 3147 12 8 7 84 1 0 0 0 0 3625240 189036 5860 3705260 0 5228 24268 5228 4841 4072 14 12 69 4 0 0 0 0 3625240 166196 5860 3710696 0 0 32 0 3572 2556 21 5 74 0 0 0 0 0 3625240 162668 5860 3716200 8 0 164 0 3729 2619 20 7 73 0 0 0 12 0 3625312 147692 5860 3724288 0 140 9207 524 3679 4024 22 14 63 1 0 0 Non-zero si/so in 4 of 5 samples show the live paging process, swpd also climbing across five seconds, which is good proof. The Swapped Pages Are Anonymous, Not File-Backed Let's also check swap memory consumption for each of the process's mappings, to confirm that swap is heap-related: Shell PID=$(pgrep -f org.opensearch.bootstrap.OpenSearch) awk '/^[0-9a-f]/{h=$0} /^Swap:/{if($2>0)print $2" kB "h}' /proc/$PID/smaps | sort -rn | head -20 Plain Text 1160380 kB 708400000-7ffe00000 rw-p 00000000 00:00 0 62296 kB 7f30c0000000-7f30c3f4b000 rw-p 00000000 00:00 0 61452 kB 7f1fc8000000-7f1fcbc03000 rw-p 00000000 00:00 0 61448 kB 7f1fa8000000-7f1fabc02000 rw-p 00000000 00:00 0 61444 kB 7f2f38000000-7f2f3bc01000 rw-p 00000000 00:00 0 61444 kB 7f24b4000000-7f24b7c01000 rw-p 00000000 00:00 0 61444 kB 7f22ac000000-7f22afc01000 rw-p 00000000 00:00 0 61444 kB 7f2238000000-7f223bc01000 rw-p 00000000 00:00 0 61444 kB 7f216c000000-7f216fc01000 rw-p 00000000 00:00 0 61444 kB 7f2168000000-7f216bc01000 rw-p 00000000 00:00 0 61444 kB 7f2148000000-7f214bc01000 rw-p 00000000 00:00 0 61444 kB 7f1fec000000-7f1fefc01000 rw-p 00000000 00:00 0 61444 kB 7f1fe8000000-7f1febc01000 rw-p 00000000 00:00 0 61444 kB 7f1fcc000000-7f1fcfc01000 rw-p 00000000 00:00 0 61444 kB 7f1fac000000-7f1fafc01000 rw-p 00000000 00:00 0 60772 kB 7f1fc0000000-7f1fc3c14000 rw-p 00000000 00:00 0 59952 kB 7f311e000000-7f3123f41000 rw-p 00000000 00:00 0 59340 kB 7f30b4000000-7f30b7c9b000 rw-p 00000000 00:00 0 57180 kB 7f3110000000-7f3113e3e000 rw-p 00000000 00:00 0 43084 kB 7f3130800000-7f3133d70000 rwxp 00000000 00:00 0 The Largest Swapped Region Is But why is what we are seeing above a heap-related area? There are a few clues for that. The region size 0x708400000 - 0x7ffe00000 is exactly 4,154,458,112 bytes = 3,962 MiB as we use -Xms == -Xmx and the whole thing is committed at startup, and nothing else in a JVM process is a single contiguous ~4 GB anonymous rw-p mapping. Second, It's the lowest mapping in the address space, smaps_rollup [rollup] line starts at exactly 708400000: Shell cat /proc/$PID/smaps_rollup Shell 708400000-7ffd3bb56000 ---p 00000000 00:00 0 [rollup] Private_Dirty: 3213788 kB Swap: 2680532 kB SwapPss: 2679412 kB Locked: 0 kB The JIT Code Cache Is Swapped Too Decoding the top swapped regions: 1160380 kB at 708400000 – is the JVM heap, the compressed‑oops heap base and matches the [rollup] start from smaps_rollup.The dozens of 61444 kB regions – these areas are probably related to native/off‑heap: Netty, JNI, Lucene native, etc.43084 kB marked rwxp – the JIT code cache, also swapped out, a bad sign. Together, these regions account for almost exactly the ~2.5 GB of swapped memory we observed earlier: cold heap regions, native/off-heap allocations, code cache, and possibly thread stacks. Practically, this means two consequences: GC can amplify swap latency. 1 GB of the JVM heap was swapped out. G1 does not necessarily touch all of those pages during a mixed collection, but any GC phase that accesses a swapped page incurs a major fault and has to bring it back through the dm-crypt device. Hence, short GC work can produce significantly longer pauses.A swapped page can also contain executable code. The next call into a swapped-out compiled method can trigger a major fault before the code can run. The resulting latency may land on an otherwise random request and be difficult to attribute directly to GC, index I/O, or the query itself. Evidence of Sustained Anon Memory Churn workingset_refault_anon counts anonymous memory refault events after reclaim; it does not count unique pages. Together with pswpin and pswpout, it shows how much anonymous memory paging has accumulated since boot. Shell grep -E 'workingset_(refault|activate)_anon|pswpin|pswpout' /proc/vmstat Plain Text workingset_refault_anon 89061837 workingset_activate_anon 1856069 pswpin 86690374 pswpout 61618603 These counters are cumulative, so fetching them at 10-minute intervals clearly shows that this wasn't a one-time eviction. There was sustained process: nearly 89 million anonymous pages were repeatedly swapped out and faulted back in. vm.swappiness = 1 and GC Amplification In this story vm.swappiness=1 was set since the cluster's inception. It does what Linux defines it to do, but it doesn't provide the protection we wanted. I suspect this matters particularly in the most resource-constrained deployments. How do we know that? The entire result above is a counterexample. swappiness biases the kernel's choice between reclaiming file-backed and anonymous pages. It does not prevent anonymous pages from being swapped out. Even at 1, this can still happen under sustained memory pressure. On a resource-constrained node whose index is several times larger than its RAM, pressure on the filesystem cache is not an exceptional condition; this is the normal operating state. This has two important consequences: It is not self-healing. Nothing proactively pages anonymous memory back in on a schedule. A swapped-out page returns to RAM only when it is accessed again, and cold memory, as it's defined, may remain untouched for a long time. As a result, the cold JVM memory can remain in swap indefinitely.It may remain invisible until something touches it. At steady state, a cold tail of the heap can remain in swap without producing obvious symptoms. The problem becomes visible when those pages are touched again, causing major page faults and potentially amplifying GC and request latency. Key Takeaways So, the root cause of the latency problems is an oversized heap combined with page cache pressure (the cold heap tail has been swapped out). vm.swappiness=1 did not protect the heap. It biases what gets reclaimed; it does not prevent anonymous memory from being swapped out.vm.swappiness=1 should not be relied on with the other defaults in production. An oversized heap can lead to GC amplification that is difficult to detect.Most of the swapped-out memory wasn't heap at all, but malloc arenas and the JIT code cache, none of it inside Xmx, so heap metrics didn't show it.Swap on dm-crypt exacerbates the issue, resulting in longer GC pauses and random request latency spikes. dm-crypt may be unavoidable in production due to security requirements. The fix isn't another swappiness tweak. It's two things: stop committing heap you don't use, and make the heap you do commit non-evictable. Follow Up In Act 2, we'll answer the remaining questions raised at the beginning of this article and look more closely at the trade-off introduced by bootstrap.memory_lock. This setting makes the heap resident and swap-immune, but it also turns jvm_heap_ratio from a soft default into a permanent memory commitment. The question then becomes: what heap ratio best suits different read and write workloads? There is one more complication to mention in advance: in a resource-constrained environment, merge storms can distort benchmark results, making an otherwise good heap ratio appear poor. See the screenshot below.

By Maxim Muzafarov
Grounding AI Agents in Governed Data
Grounding AI Agents in Governed Data

Today, every vendor offering BI solutions has incorporated a chat box. Whether you use Copilot or some other natural-language interface that connects you to a data warehouse, just ask a question in simple words, and it will generate SQL automatically. While this is conducive to productivity in other industries, in banking it presents an opportunity for a new attack. It is not enough to simply say that wrong SQL can be produced. It’s that an ungoverned text-to-SQL layer may join tables it shouldn’t, return columns that should have been masked. As a result of a lack of oversight, a marketing analyst could receive a query containing raw account numbers, since none of the components of the stack told the system to do otherwise. Prompt-level guardrails (“please don’t show PII”) are not a security control. They’re just a suggestion, and a model under adversarial pressure ot just a confusing prompt will ignore a suggestion. The issue isn't just about providing a more intelligent prompt; rather, it's about placing the artificial intelligence assistant on the same layer of information as human analysts, meaning an environment where the database manages the relevant security details as per row and column criteria instead of relying on technology. Consequently, if the analyst does not have access to the specific column, there is no justification for the AI assistant to have access to it as well. The diagram below (Figure 1) shows the steps taken to build that layer in BigQuery: a validated semantic layer that allows human dashboards and AI-generated queries to be connected to the same quality definitions. This would allow the assistant to leverage existing security rather than creating it. Figure 1. The Semantic Layer Resolver Step 1: Stop Letting Anyone (Human or AI) Query Raw Tables The initial phase is architectural, not related to AI: there is no query made by a person or a system involving the base tables. Instead, each of the metrics that are accessible to consumers is defined at least once in a BigQuery view, and its calculation logic is embedded in that view. SQL -- Certified metric: Risk-Weighted Assets, defined once, queried everywhere CREATE VIEW analytics.risk_weighted_assets AS SELECT exposure.customer_id, exposure.region, exposure.exposure_class, exposure.outstanding_balance, risk_weights.weight_pct, ROUND(exposure.outstanding_balance * risk_weights.weight_pct / 100, 2) AS rwa_amount, CURRENT_TIMESTAMP() AS calculated_at FROM finance.exposures AS exposure JOIN reference.basel_risk_weights AS risk_weights ON exposure.exposure_class = risk_weights.exposure_class WHERE exposure.status = 'ACTIVE'; The view of "risk-weighted assets" created through a dashboard, a notebook, and an LLM agent is identical. There is no alternative version in a researcher’s spreadsheet, nor can an AI agent "helpfully" recreate the calculation based on exposure tables but use incorrect risk weightings. Step 2: Enforce Security at the Data Layer, Not the Application Layer It is important to ensure that BigQuery includes row-level and column-level security and associates it with the table. This means the principle will work irrespective of the entity making the query. SQL -- Row-level security: a regional analyst only ever sees their region's rows CREATE ROW ACCESS POLICY regional_filter ON analytics.risk_weighted_assets GRANT TO ('group:[email protected]') FILTER USING (region = 'EMEA'); Column masking works the same way, through policy tags rather than per-report logic: YAML # Dataplex policy tag: applied once, enforced everywhere the column is queried taxonomy: financial-pii policyTags: - displayName: "customer-account-number" description: "Masked for all roles except fraud-investigation" - displayName: "customer-ssn" description: "Masked for all roles except compliance-audit" Once a policy tag is applied to a column, a user, or any AI agent acting under that user's identity, who doesn’t have the appropriate fine-grained reader role, will receive either a null value or a hashed value. There is no mistake that the model has to circumvent; it’s simply a different result. This is what makes querying with AI safe, since whatever the query for the AI is, it cannot reveal anything that the column policy prohibits. Step 3: Give the Grounding Layer Metadata to Query Against It is impossible for an LLM to adhere to rules of governance it knows nothing about. Accordingly, a metadata directory is necessary for the semantic layer that contains a description of each certified metric with enough detail for the agent to turn an inquiry posed in natural language into the correct interpretation and filtering process, not simply provide it with a raw schema dump. JSON { "metric_id": "risk_weighted_assets", "display_name": "Risk-Weighted Assets", "view": "analytics.risk_weighted_assets", "owner": "[email protected]", "sensitivity": "internal", "allowed_dimensions": ["region", "exposure_class", "customer_id"], "definition": "Balance times Basel risk weight, summed by class.", "lineage": ["finance.exposures", "reference.basel_risk_weights"], "last_certified": "2026-06-01" } This record is the thing the AI agent actually reads. The document specifies which view will be interrogated, lists the dimensions available for filtering results, and identifies who to contact if something goes wrong. It is worth mentioning that in this record there is no schema given for the finance exposes table, which leaves the model nothing to "discover" about. Step 4: Route Natural-Language Requests Through the Semantic Layer, Not the Warehouse When the certified metrics with their metadata have been obtained, the resolution process consists of transforming the user's natural-language question into a query that uses an allowed view rather than directly referring to the underlying schema. Python class SemanticLayerResolver: def __init__(self, metric_catalog, bq_client): self.catalog = metric_catalog # metric_id -> metadata, from Step 3 self.bq_client = bq_client def resolve(self, nl_request: str, user_identity: str) -> QueryResult: # 1. Map the request to a certified metric, never to a raw table. # A constrained classifier over self.catalog.keys() works better # here than open-ended text-to-SQL against the full warehouse. metric = self.match_metric(nl_request) if metric is None: return QueryResult.refuse("No certified metric found.") # 2. Extract filters, restricted to the metric's allowed_dimensions. filters = self.extract_filters( nl_request, metric["allowed_dimensions"] ) # 3. Build SQL against the certified view only. sql = self.build_query(metric["view"], filters) # 4. Execute as the requesting user, so BigQuery's row/column # security applies exactly as it would for a human query. result = self.bq_client.query(sql, user=user_identity) # 5. Attach lineage and certification metadata to the answer, # so "what the AI said" is auditable like any report. return QueryResult( data=result, metric_id=metric["metric_id"], lineage=metric["lineage"], certified_at=metric["last_certified"], ) The important line is step 4: the query is executed under the requesting user instead of using a shared service account. Hence, all the downstream access control mechanisms are automatically applied. There is no need for a separate permission system for the resolver because it has no access rights that exceed the rights of the requesting user. Step 5: Audit Every AI-Generated Query Like You Would a Human's The governance teams will not agree on a system based on the suggestion of " having faith in the model." What they approve is proof in every case where a resolver has provided information, just as is done when an individual writes a report. Python def log_ai_query(user_identity, nl_request, result: QueryResult): audit_log.write({ "user": user_identity, "request": nl_request, "metric_id": result.metric_id, "lineage": result.lineage, "policy_version": result.certified_at, "row_count": result.row_count, "timestamp": now(), }) One financial institution successfully applied this approach. What used to be a lengthy project in which one would have to analyze whether an AI assistant could access customer information has been transformed into something evaluated right away: the assistant can perform the same functions as a human worker. The financial institution was also measuring the new trend of using a certified semantic layer, not only in regard to the AI being discussed. Conflicts over defining metrics across different business lines practically vanished when the organization no longer had to create a separate “AI-compliant” data model. The Real Insight: Governance Is What Makes AI Fast, Not What Slows It Down It’s easy to assume that the best approach to deal with the LLM and sensitive data combination is to include a review step in which a human sits in on every step of the process, or another model is deployed to analyze the first model’s outputs before they are used. This is not only unscalable, but it also misses the point. Another way is to make sure that the insecure path cannot be taken, rather than simply being shunned. If the data layer implements row-level security, column masking, and certified metric definitions, you can confirm that an AI agent querying the data cannot generate queries that reveal any data previously available to someone with the same role. As a result, there is no need to verify output against constraints, since they were already included in the model. This shift is suggested by this pattern. Governed self-service, the architecture that permits a business analyst to carry out data initiatives safely in the absence of ticket submission, also creates a secure basis for AI-enhanced analysis. But it wasn't the main purpose. It is just a coincidence that it has worked out this way.

By Jeevan reddy Geereddy
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI

Regularly, someone reopens the same argument. XML or JSON or YAML, as if one has to win and the others lose. It usually comes up in a context like data contracts, where a team has to pick a format and defend it. The framing is wrong. These formats were built for different jobs in different eras, and the more useful question is which one fits the job in front of you. So here is the history, the trade-offs, and where each one still wins, including what changes now that LLMs and agents read and write structured data too. XML, JSON, and YAML at a Glance These three formats are different ways to represent structured data. XML is verbose and rigorous. JSON is compact and universal. YAML is readable and config-friendly. None is strictly best. Each won a different era and a different job, and validation became its own layer, led today by JSON Schema. Key Takeaways XML led enterprise integration for two decades and now lives mostly in legacy systems; JSON won web APIs; YAML won cloud-native config.XML, JSON, and YAML serialize data. JSON Schema and XSD validate it. They are different layers, not competitors.YAML has no schema language of its own. It borrows JSON Schema, which is how Kubernetes and similar tools validate YAML.JSON Schema now underpins LLM tool calling and structured outputs, which puts it at the center of agentic AI.Authored in YAML, validated by a schema, enforced in the pipeline: data contracts are the clearest example of a wider pattern in data governance and orchestration tools. Serialization vs. Validation: Two Different Jobs XML, JSON, and YAML are serialization formats. You author data in them, and a parser reads them back. JSON Schema and XML Schema (XSD) are validation languages. They describe what valid data looks like, and a validator checks a document against that description. So comparing YAML with JSON Schema is not a fair fight. One is a format you write. The other is a contract you check against. Keep that split in mind. Most of the real story is about how the two layers interact. A Short History of XML, JSON, and YAML Each format rose with a shift in how we built systems. XML came first, standardized by the W3C in 1998 with roots in SGML. It became the backbone of enterprise integration. SOAP, WSDL, and the early ESB and SOA stacks all spoke XML. It was verbose but rigorous, and it shipped with a full schema system in XSD. The ecosystem also grew heavy. The sprawling WS-* stack of SOAP extensions became so complex that many engineers came to call it WS-* hell, which is part of why lighter approaches eventually took over. JSON came out of the JavaScript world in the early 2000s. Douglas Crockford formalized it from JavaScript object literals, and json.org went up in 2002. As REST and AJAX replaced SOAP for web APIs, JSON replaced XML as the default wire format. It was lighter, easier to read, and mapped directly to data structures in most languages. JSON Schema followed later as a separate community effort. YAML appeared in 2001 as "YAML Ain't Markup Language," designed to be human-friendly first. Since version 1.2 in 2009, it is a superset of JSON. It found its home in the cloud-native era. Kubernetes, Terraform, Ansible, CI/CD pipelines, GitOps. Anywhere humans hand-write configuration that lives in Git. One detail matters for later. Each format handled schema differently. XML built it in with XSD. JSON bolted it on with JSON Schema. YAML never built one and borrowed JSON Schema instead. Other Formats Worth Knowing: TOML, HCL, Protobuf, and Config Languages A comparison limited to three formats would feel a decade out of date. The landscape is wider now. TOML is simple and serves as the config format for Rust's Cargo and Python's Poetry. HCL is HashiCorp's language for Terraform. In the streaming world, Protobuf and Avro take a schema-first approach and serialize to compact binary, which is why they sit under Kafka and gRPC. There is also a newer category built to fix YAML's weaknesses: configuration languages. CUE, Pkl from Apple, and KCL from the CNCF add expressions, validation, and reuse on top of the data model, then render plain YAML or JSON as output. They are not serialization formats. They are programs that generate configuration. For teams drowning in thousands of lines of near-duplicate YAML, they are worth a look. The rest of this post stays on XML, JSON, and YAML, since they are still the three you choose between most days. How XML, JSON, and YAML Differ in Practice The differences show up the moment you write them by hand. XML wraps everything in opening and closing tags and supports attributes, namespaces, and comments. It is precise and self-describing. It is also heavy. A small payload turns into a wall of angle brackets. JSON uses braces, brackets, and quoted keys. It is compact and unambiguous, and every major language parses it natively. Its one notable omission is comments. The spec does not allow them, which is a real constraint for anything humans need to annotate. YAML uses indentation instead of brackets and braces. It supports comments, multi-line strings, anchors for reuse, and multiple documents in one file. It reads closer to how people think about nested data. The cost is that whitespace carries meaning, so structure is easy to break. Pros and Cons of XML, JSON, and YAML XML's strength is rigor: namespaces, mature validation with XSD, XPath for querying, and decades of tooling. Its weakness is weight and friction. Few people enjoy writing it by hand, and it feels dated for new web APIs. JSON's strength is ubiquity and simplicity. It is the lingua franca of web APIs; it maps cleanly to data structures, and it parses fast everywhere. Its weaknesses are the lack of comments and the absence of native validation in the base spec. YAML's strength is readability. It is friendly to engineers and non-engineers, it diffs cleanly in Git, and it supports comments and reuse. The trade-offs come from the same design. Significant whitespace makes it fragile, so one wrong indentation can break the file, and loose typing causes surprises, like the Norway problem where the country code NO once parsed as the boolean false. YAML 1.2 and StrictYAML help, but the tension stays. The same whitespace that makes YAML readable makes it fragile. One caveat on YAML's downsides. They mostly bite humans. When tools generate and validate the files, as Kubernetes operators and config languages like CUE or Pkl do, the fragility matters far less. In fact, the sweet spot is machines generating and editing while humans mainly read and review, which plays to YAML's strengths. XML vs. JSON vs. YAML: A Comparison Table The table below sums up how XML, JSON, and YAML compare across the dimensions that matter most in practice, from readability and verbosity to schema support and failure modes. DimensionXMLJSONYAMLHuman readabilityLowMediumHighVerbosityHighMediumLowCommentsYesNoYesNative schemaXSD, built inJSON Schema, add-onNone, borrows JSON SchemaValidation maturityVery matureMature, now dominantMature, via JSON SchemaIDE toolingMatureMatureMatureType safetyStrong with XSDBasic, strong with schemaWeak, coercion surprisesGit-diff friendlinessPoorGoodExcellentLearning curveSteepEasyEasy to start, subtle trapsTypical useSOAP, documents, enterpriseWeb APIs, data exchangeConfig, contracts, pipelinesMain failure modeBloat and complexityNo comments, no base validationIndentation and type coercion Does YAML Have a Schema? YAML has no schema language of its own. It uses JSON Schema. Because YAML 1.2 maps onto the same data model as JSON, a JSON Schema validator can validate a YAML document without modification. The format you author in and the language that validates it are decoupled, and the decoupling is a feature. The mechanics are simple. A tool publishes a JSON Schema describing its YAML structure. Your IDE applies that schema as you type, giving autocompletion, inline validation, and error checking. In VS Code this runs through the YAML language server. SchemaStore acts as a public registry of JSON Schemas that editors auto-apply to hundreds of known config files, from Kubernetes manifests to GitHub Actions workflows. Kubernetes is the largest example. Custom resources validate against OpenAPI structural schemas, which are a JSON Schema dialect generated from the underlying Go types. CI and workflow tools like CircleCI follow the same idea, publishing a JSON Schema for their YAML that editors use for validation and autocompletion. The pattern is consistent. The code is the source of truth; it emits the schema, and the schema validates the YAML. So XSD was XML's built-in answer. JSON Schema is JSON's bolt-on answer. YAML's answer is to reuse JSON Schema. The result is that JSON Schema became the shared validation layer for both JSON and YAML. JSON Schema and Agentic AI: Tool Calling, Structured Outputs, and MCP Structured data formats used to be a backend concern. Now they sit at the center of how AI systems work, and JSON Schema is the format doing the work. When an LLM calls a tool, the tool is defined by a JSON Schema. When you ask a model for structured output, you hand it a JSON Schema and the model fills it in. Several providers go further with constrained decoding, which restricts the model token by token so the output cannot violate the schema. OpenAI's Structured Outputs guarantees schema compliance this way, and Google's Gemini and others offer similar structured-output modes. The Model Context Protocol (MCP), the emerging standard for connecting models to tools, adopted JSON Schema 2020-12 as its default dialect for tool inputs and outputs in 2025. The shift underneath is the interesting part. Schemas have always enforced structure, since XSD already validated and rejected SOAP messages at runtime. What is new is where the enforcement sits. The same JSON Schema that validates a config file now also shapes a language model's output as it generates, token by token. The contract moved from checking data after the fact to steering how it is produced. Notice the division of labor. Humans author agent and workflow config in YAML, because it is readable. Machines exchange JSON, because it is precise. JSON Schema validates both. Data Contracts and Governance: Where Format Choice Matters This is where the choice stops being academic. Data contracts are where formats, schemas, and governance meet. A data contract is a formal agreement between the team that produces a dataset and the teams that consume it. It defines fields, types, allowed values, freshness, ownership, and quality rules. The shift it represents is governance moving out of documents and into code. A contract in a Confluence page is documentation. A contract in version-controlled YAML, checked in CI/CD, is an enforceable control. The tooling has converged on YAML for the same reasons that make YAML good for config. The Open Data Contract Standard, at version 3.1.0 under the Linux Foundation's Bitol project, defines contracts in YAML and ships a JSON Schema so editors can validate them. Soda's SodaCL expresses data quality checks in YAML and runs them in pipelines, with a cloud layer that adds stakeholder approval. dbt embeds model contracts and tests in YAML. Great Expectations validates against a declarative spec. Different tools, same pattern. Author the contract in YAML, validate it with a schema, enforce it in the pipeline. The pattern is a familiar one. Databases use schemas to keep bad data out of storage. Streaming platforms use schema registries to keep bad data out of event streams. Analytical datasets now use contracts to do the same thing, one layer up. Data contracts are the clearest case, but the same model runs through orchestration and governance tooling. Workflow engines like Argo and Kestra define pipelines in YAML that a JSON Schema validates. Policy-as-code tools like Open Policy Agent and Kyverno keep rules in version control and enforce them in CI/CD or at deploy time. Author in YAML, validate with a schema, enforce in the pipeline. The format choice and the validation layer are the same story at every level. Which Format Should You Use? A Quick Decision Guide XML fits when you need namespaces, document-centric markup, or integration with SOAP and legacy enterprise systems that already speak it. JSON is the default for web APIs and machine-to-machine exchange, where precision and universal parsing matter more than human authoring. Reach for YAML for configuration and contracts that humans write and review in Git, where readability and comments earn their keep. For validation, reach for JSON Schema in almost every modern case, including for your YAML. Reach for XSD when you are already in the XML world. And when your YAML starts repeating itself across hundreds of files, look at a configuration language like CUE, Pkl, or KCL before the duplication gets worse. The same logic shows up across the modern data stack. Integration pipelines and APIs move JSON. Process mining still reads XML-based event logs like XES while newer event streams carry JSON. Data contracts and platform config are authored in YAML and validated by JSON Schema. Different layer, different format, same principle. Stop Asking Which Format Is Best XML, JSON, and YAML were never really competing for the same job. XML won enterprise integration for two decades, then its weight and WS-* complexity pushed teams toward lighter options, so today it lives mostly in legacy systems. JSON won the web API era and became the default wire format. YAML won the cloud-native config era. Validation sits on its own layer: XSD for XML and JSON Schema for JSON and YAML, and JSON Schema now also underpins how agents and AI systems exchange structured data. The useful question is not which format is best. It is which layer you are working in. Author where humans read. Exchange where machines parse. Validate everywhere. Get those three right and the format debate mostly takes care of itself.

By Kai Wähner DZone Core CORE
Building Time-Series Applications With Java and InfluxDB
Building Time-Series Applications With Java and InfluxDB

InfluxDB is essential for applications that analyze continuously changing data. In IoT, this includes tracking temperature, pressure, energy use, or machine telemetry over time. Financial systems use similar models for market prices, exchange rates, trading activity, portfolio values, and risk metrics. The key requirement is the ability to ingest large volumes of timestamped data and efficiently query current, historical, and evolving values. This versatility makes InfluxDB valuable beyond traditional monitoring. Enterprises use it for observability, infrastructure metrics, logistics, industrial systems, customer activity, fraud detection, transaction trends, and business KPIs. When time is central to data storage and queries, a dedicated time-series database simplifies architecture and enables more intuitive queries. Understanding InfluxDB InfluxDB is designed to preserve data value by keeping its relationship to time. Instead of treating timestamps as standard columns, InfluxDB organizes data by time-based measurements, attributes, and values. Each record represents an observation at a specific moment, such as a sensor’s temperature, API latency, or asset price. This structure supports direct queries, including retrieving the latest value, assessing recent trends, or calculating averages over defined intervals. This structure is ideal for workloads that produce continuous data. IoT devices generate millions of measurements, infrastructure platforms track ongoing metrics, and financial applications monitor prices and transactions. In these cases, data is usually appended, not updated, and queries often focus on recent values, time ranges, aggregations, or trends. As a result, InfluxDB is well suited for modern architectures. Distributed applications, cloud platforms, microservices, connected devices, and real-time business systems all generate continuous temporal data. Beyond storage, applications have to efficiently query and aggregate this data to assess current conditions and understand system evolution. The key point is that InfluxDB is not only an “IoT database.” InfluxDB is not limited to IoT use cases. It can serve as the temporal component in a more extensive polyglot architecture. For example, a relational database may manage transactional data, a document database may handle flexible aggregates, and InfluxDB can store the history of measurements, signals, and operational changes. When business needs require insight into how something evolves over time, this time-based specialization provides a significant architectural advantage. Hands-On: Java With InfluxDB Starting with Eclipse JNoSQL 1.1.18, Java applications can work with time-series databases through a consistent programming model. Let’s see this in practice with InfluxDB 3. For this example, InfluxDB will run locally without authentication to simplify setup. This approach is suitable for development and demonstrations, but --without-auth must not be used in production. Start InfluxDB 3 Core with Docker: Shell docker run -d \ --name influxdb-instance \ -p 8181:8181 \ influxdb:3.11.0-core \ influxdb3 serve \ --node-id jnosql \ --object-store memory \ --without-auth Once the server is running, create the database used by the application: Java docker exec influxdb-instance \ influxdb3 create database \ --host http://localhost:8181 \ metrics Setting Up the Java Application Eclipse JNoSQL uses Jakarta technologies such as CDI and JSON-B, along with Eclipse MicroProfile Config for externalized configuration. These APIs are available in runtimes including Open Liberty, Payara, Helidon, and Quarkus. Beyond the standard JNoSQL dependencies, add the InfluxDB driver: XML <dependency> <groupId>org.eclipse.jnosql.databases</groupId> <artifactId>jnosql-influxdb</artifactId> <version>${jnosql.version}</version> </dependency> Then configure the connection through microprofile-config.properties: Properties files jnosql.timeseries.database=metrics jnosql.influxdb.url=http://localhost:8181 jnosql.influxdb.token=jnosql-influxdb-test-token Since authentication is disabled for this local instance, the token serves only as a required placeholder for the driver configuration. In production, these values should be externalized and overridden using MicroProfile Config, aligning with Twelve-Factor App principles. Modeling Time-Series Data In this example, account transactions are modeled over time: Java @Entity public class AccountTransaction { @Id private Instant id; @Column private String account; @Column private BigDecimal amount; @Column private String currency; @Column private TransactionStatus status; // constructors, getters, and setters } The transaction status can be represented with a simple enum: Java public enum TransactionStatus { APPROVED, DECLINED, PENDING } The key detail is the Instant identifier. Each transaction records a specific point in time, enabling queries for the latest state and historical navigation. Using TimeSeriesTemplate You can now insert transactions and query them using TimeSeriesTemplate: Java public class App { public static void main(String[] args) { var firstTransaction = new AccountTransaction( Instant.parse("2026-09-20T08:00:00Z"), "account-42", new BigDecimal("79.90"), "EUR", TransactionStatus.APPROVED ); var secondTransaction = new AccountTransaction( Instant.parse("2026-09-20T09:00:00Z"), "account-42", new BigDecimal("24.50"), "EUR", TransactionStatus.APPROVED ); var latestTransaction = new AccountTransaction( Instant.parse("2026-09-20T10:15:00Z"), "account-42", new BigDecimal("120.00"), "EUR", TransactionStatus.DECLINED ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { TimeSeriesTemplate template = container.select(TimeSeriesTemplate.class).get(); template.insert(firstTransaction); template.insert(secondTransaction); template.insert(latestTransaction); var currentStatus = template .select(AccountTransaction.class) .where("account") .eq("account-42") .orderBy("id") .desc() .limit(1) .singleResult(); System.out.println( "Current account status: " + currentStatus ); var history = template .select(AccountTransaction.class) .where("account") .eq("account-42") .orderBy("id") .desc() .skip(1) .limit(10) .result(); System.out.println("Transaction history:"); history.forEach(System.out::println); } } } These queries address two common time-series scenarios: retrieving the latest observation for the account and obtaining its recent history, excluding the current record. Using Jakarta Data The same model can also be exposed through a Jakarta Data repository: Java @Repository public interface AccountTransactionRepository extends BasicRepository<AccountTransaction, Instant> { List<AccountTransaction> findByAccountOrderByIdDesc( String account, Limit limit); } The application code is now repository-oriented: Java public class App2 { public static void main(String[] args) { var firstTransaction = new AccountTransaction( Instant.parse("2026-09-20T08:00:00Z"), "account-42", new BigDecimal("79.90"), "EUR", TransactionStatus.APPROVED ); var secondTransaction = new AccountTransaction( Instant.parse("2026-09-20T09:00:00Z"), "account-42", new BigDecimal("24.50"), "EUR", TransactionStatus.APPROVED ); var latestTransaction = new AccountTransaction( Instant.parse("2026-09-20T10:15:00Z"), "account-42", new BigDecimal("120.00"), "EUR", TransactionStatus.DECLINED ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { AccountTransactionRepository repository = container.select(AccountTransactionRepository.class).get(); repository.save(firstTransaction); repository.save(secondTransaction); repository.save(latestTransaction); var currentStatus = repository .findByAccountOrderByIdDesc( "account-42", Limit.of(1) ) .stream() .findFirst(); System.out.println( "Current account status: " + currentStatus ); var history = repository .findByAccountOrderByIdDesc( "account-42", Limit.range(2, 10) ); System.out.println("Recent transaction history:"); history.forEach(System.out::println); } } } Notably, the application expresses time-oriented business queries without relying directly on the InfluxDB client API. Familiar Java mapping and repository concepts are retained, while InfluxDB delivers specialized storage and query capabilities. Where InfluxDB Fits in an Enterprise Architecture While InfluxDB is commonly associated with monitoring or IoT use cases, its role in enterprise architecture is broader. It works best as a specialized time-series store that supplements, rather than replaces, other databases. Relational databases can continue to serve as systems of record for customers, orders, and transactions, while messaging platforms such as Kafka distribute events. InfluxDB then stores time-based operational metrics, telemetry, prices, business indicators, or application signals from these systems. This separation simplifies the overall architecture. Transactional systems are designed for preserving consistency, relationships, and state changes, while time-series databases excel at handling continuous data, recent-state queries, historical ranges, and time-based aggregations. Combining both workloads in a single database may work initially, but as temporal data grows, it frequently leads to complex indexing, partitioning, and retention strategies. For example, an e-commerce platform may store orders and payments in a relational database, while using InfluxDB for checkout latency, payment approval rates, inventory changes, and orders-per-minute. A financial platform can keep transactional records in its core database and use InfluxDB for exchange-rate history, portfolio measurements, or operational risk indicators. Similarly, an industrial platform might store machine metadata in a relational system and use InfluxDB for temperature, pressure, and vibration readings. The architectural value lies in specialization. InfluxDB is most effective when applications need to answer questions like “what is happening now?”, “what changed during this period?”, or “how is this metric trending?” without placing all temporal responsibilities on the transactional database. Conclusion InfluxDB is a strong fit when applications need to work with continuously changing data, recent state, and historical context. With Eclipse JNoSQL 1.1.18, Java developers can use InfluxDB through familiar Jakarta APIs instead of depending directly on database-specific client code, making time-series workloads easier to integrate into modern enterprise applications.

By Otavio Santana DZone Core CORE
Wasm Inside Neo4j: Building the Example That Didn't Exist
Wasm Inside Neo4j: Building the Example That Didn't Exist

In a recent DZone article, Running Sentiment Analysis Inside Neo4j With a Java Plugin, we explored several approaches to running sentiment analysis inside the Neo4j database engine. One of those approaches — embedding a Wasm runtime inside a Java UDF — was described like this: Theoretically, we could embed a Wasm runtime such as wasmtime inside a Java UDF and execute the VADER Wasm module from within Neo4j, getting Wasm's sandbox guarantees inside Neo4j's plugin model. It's technically feasible but no published working example appears to exist and the complexity cost is high relative to the alternatives. An interesting idea to watch, but not practical today. This article builds that working example. We'll show how to embed a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity score map callable directly from Cypher. We'll cover the tools and inspection techniques needed to understand what the Wasm compiler generates and why the Java calling convention looks the way it does. The full source code is available on GitHub. What We're Building We're embedding a wasmtime Wasm runtime inside a Neo4j Java UDF using wasmtime-java, a community JNI binding for the Wasmtime runtime. It's not an official Bytecode Alliance product, but it ships prebuilt native libraries for all major platforms and is sufficient for this proof-of-concept. A Rust function compiled to WebAssembly rides inside the plugin JAR alongside the Java code. When Cypher calls the UDF, Java initializes the Wasm runtime, loads the binary, and invokes the Rust function — all inside the Neo4j JVM process with no external API calls and no network round-trips. Note: This article was tested specifically against wasmtime-java 0.19.0. The API used here is version-specific; newer releases or alternative JVM Wasm runtimes may expose different interfaces and calling conventions. Prerequisites You'll need the following installed if you wish to follow along. We're using Apple Silicon (ARM64) as our development platform, so we'll note where the setup differs from other platforms. Java We're using OpenJDK 21 (tested with 21.0.12.1). Install it using your platform's package manager or download it directly from adoptium.net. On macOS via Homebrew: Shell brew install openjdk@21 On Ubuntu/Debian: Shell sudo apt install openjdk-21-jdk On Windows, download and run the installer from Adoptium. Confirm your Java version: Shell java -version You should see a Java 21 runtime. If you're on Apple Silicon, also confirm you're running a native ARM64 JVM with: Shell uname -m You should see arm64. Not running under ARM64 will likely break the wasmtime-java JNI library loading. Maven We're using Maven 3.9.6. On Apple Silicon, be cautious about installing Maven via Homebrew as, at the time of writing, the Homebrew Maven formula pulls in OpenJDK 26 as a dependency, which conflicts with a Java 21 installation. If your package manager installs an incompatible JDK alongside Maven, verify the runtime with mvn -version and configure JAVA_HOME as necessary. Installing Maven manually is the safest approach: Shell cd ~ curl -O https://archive.apache.org/dist/maven/maven-3/3.9.6/binaries/apache-maven-3.9.6-bin.tar.gz tar xzf apache-maven-3.9.6-bin.tar.gz Then add Maven to your PATH and make it persist across terminal sessions: Shell echo 'export PATH="$HOME/apache-maven-3.9.6/bin:$PATH"' >> ~/.zshrc source ~/.zshrc On Linux, add the same line to ~/.bashrc instead.On Windows, download the zip from maven.apache.org and add the bin folder to your system PATH via System Properties. Confirm Maven is using Java 21: Shell mvn -version You should see Java version: 21 in the output. Rust We're using Rust 1.96.0. Install via rustup if not already present: Shell curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh On Windows, download and run rustup-init.exe from rustup.rs. To pin to the specific Rust version we tested with: Shell rustup toolchain install 1.96.0 rustup default 1.96.0 Then add the WASI target: Shell rustup target add wasm32-wasip1 This target works identically across macOS, Linux, and Windows. WABT The WebAssembly Binary Toolkit gives us wasm-objdump for inspecting Wasm binaries. We tested with version 1.0.41. On macOS: Shell brew install wabt On Ubuntu/Debian: Shell sudo apt install wabt On Windows, download the latest release from github.com/WebAssembly/wabt/releases. wit-bindgen This is the interface types generator for WebAssembly. There are two distinct version numbers to be aware of: the wit-bindgen-cli command-line tool and the wit-bindgen Rust crate used as a dependency inside the Wasm module. These can differ. We tested with CLI version 0.59.0 and Rust crate version 0.40.0 (specified in Cargo.toml). The generated binary identifies the crate version through the export name cabi_realloc_wit_bindgen_0_40_0. Install the pinned CLI version via Cargo on all platforms: Shell cargo install wit-bindgen-cli --version 0.59.0 Confirm it's installed: Shell wit-bindgen --version Neo4j Desktop We're using Neo4j Desktop with a local database instance. Download from Neo4j for Desktop. The pom.xml in this article is pinned to Neo4j 2026.07.0 — update the neo4j.version property to match your own Desktop installation. wasmtime-java Platform Support The wasmtime-java library ships prebuilt JNI native libraries for: macOS aarch64macOS x86_64Linux aarch64Linux x86_64Windows x86_64 No additional setup is needed, as Maven pulls the correct native library for your platform automatically. Version Summary For reference, here are all the component versions used in this article: ComponentVersionOpenJDK21.0.12.1Maven3.9.6Rust1.96.0WABT1.0.41wit-bindgen CLI0.59.0wit-bindgen crate0.40.0vader_sentiment crate0.1.1wasmtime-java0.19.0Neo4j2026.07.0 Getting the Code Clone the repository before following along. All source files are provided so you don't need to create them manually. Shell cd ~ git clone --filter=blob:none --sparse https://github.com/VeryFatBoy/neo4j.git cd neo4j git sparse-checkout set wasm-udf mv wasm-udf ../wasm-udf cd ../wasm-udf Project Structure Before creating any files, here's the final layout we're building toward. There are two separate projects: A Rust crate that compiles to Wasm.A Maven project that hosts the Neo4j UDF. First, the Rust crate: Plain Text sentimentable/ ├── Cargo.toml ├── src/ │ └── lib.rs └── wit/ └── sentimentable.wit Second, the Maven project: Plain Text neo4j-wasm-udf/ ├── pom.xml └── src/ └── main/ ├── java/ │ └── com/example/ │ ├── WasmUDF.java │ └── SentimentUDF.java └── resources/ ├── add.wat ├── add.wasm └── sentimentable.wasm The Wasm binaries in resources/ are bundled into the plugin JAR at build time. The Rust crate and Maven project are kept separate, and the Wasm binary is the handoff point between them. The project layout is also shown in Figure 1. Figure 1. Two-Project Layout How the Wasm Plumbing Works Before diving into the code, it's worth understanding the three layers that make this possible. Core Wasm and WASI WebAssembly defines a portable binary format and a stack-based execution model. On its own, it only understands numbers, such as integers and floats. When a Wasm module needs system capabilities, like memory allocation or I/O, it uses WASI (WebAssembly System Interface), a standardized set of system calls that a host runtime implements. Our Rust code targets wasm32-wasip1, which means it compiles to Wasm with WASI preview 1 system calls. The wasmtime runtime implements those calls on the host side. wasmtime-java This library wraps the wasmtime Wasm runtime in a JNI binding, making it callable from Java. It ships prebuilt native libraries for all major platforms, so adding it as a Maven dependency is all that's needed — no separate wasmtime installation required. The Java API lets us load a Wasm binary, set up a WASI context, and call exported functions directly. wit-bindgen and the String ABI Core WebAssembly functions operate on Wasm value types such as integers and floats. WIT (WebAssembly Interface Types) and the Component Model provide higher-level interface types such as strings, tuples, and records; wit-bindgen generates the lowering and lifting code needed to represent those types at the Wasm boundary. For strings, it uses a pointer-and-length convention: the caller allocates memory inside the Wasm module using a generated cabi_realloc function, writes the string bytes there and passes the memory address and byte length as two integers. The Rust code reads the string from that address. For return values, the lowering strategy depends on the type, which we'll see when we inspect the generated binary. With those three pieces in place, the calling chain looks like this: Plain Text Cypher query -> Neo4j routes to @UserFunction -> Java initializes wasmtime engine + WASI context -> Java allocates string in Wasm memory -> Java calls exported Wasm function -> Rust executes VADER scoring -> Java reads result from Wasm memory -> Java returns Map<String, Double> to Neo4j -> Neo4j returns result to Cypher Graphically, the calling chain is also shown in Figure 2. Figure 2. Calling Chain We built up to this through two simpler stepping-stone cases: Case 1: A trivial integer addition to prove the chain works.Case 2: A single compound score to introduce string passing and WASI. The full walkthrough of both, including the WasmUDF.java implementation is in a technical report on the GitHub repo. The Maven Project XML <project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd"> <modelVersion>4.0.0</modelVersion> <groupId>com.example</groupId> <artifactId>neo4j-wasm-udf</artifactId> <version>1.0-SNAPSHOT</version> <packaging>jar</packaging> <properties> <maven.compiler.source>21</maven.compiler.source> <maven.compiler.target>21</maven.compiler.target> <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding> <neo4j.version>2026.07.0</neo4j.version> </properties> <dependencies> <dependency> <groupId>org.neo4j</groupId> <artifactId>neo4j</artifactId> <version>${neo4j.version}</version> <scope>provided</scope> </dependency> <dependency> <groupId>io.github.kawamuray.wasmtime</groupId> <artifactId>wasmtime-java</artifactId> <version>0.19.0</version> </dependency> </dependencies> <build> <plugins> <plugin> <artifactId>maven-compiler-plugin</artifactId> <configuration> <source>21</source> <target>21</target> </configuration> </plugin> <plugin> <groupId>org.apache.maven.plugins</groupId> <artifactId>maven-shade-plugin</artifactId> <version>3.5.1</version> <executions> <execution> <phase>package</phase> <goals><goal>shade</goal></goals> <configuration> <artifactSet> <excludes> <exclude>org.neo4j:*</exclude> </excludes> </artifactSet> <shadedArtifactAttached>false</shadedArtifactAttached> </configuration> </execution> </executions> </plugin> </plugins> </build> </project> Two things worth noting here: org.neo4j:neo4j is declared as provided scope — Neo4j is already present in the database JVM at runtime, so we exclude it from the bundled JAR.We use maven-shade-plugin rather than maven-jar-plugin to produce a fat JAR that bundles wasmtime-java and its native libraries alongside our code. Update the neo4j.version property to match your own Neo4j Desktop installation. Case 3: Full Polarity Map VADER produces four scores: compound, positive, negative, and neutral. In this case, we update the Rust function to return all four and the Java UDF to return them as a Map<String, Double> — matching the return shape of the Java VADER UDF from the previous article. The sentimentable.wit File We change the return type from a single f32 to a tuple of four f32 values: Plain Text package local:sentimentable; world sentimentable { export sentimentable: func(input: string) -> tuple<f32, f32, f32, f32>; } We use a tuple rather than a named record. Both would work, but a tuple is simpler on the Java side — we read four consecutive f32 values from memory at known offsets without needing to decode field names. The lib.rs File Rust wit_bindgen::generate!({ world: "sentimentable", }); struct Component; impl Guest for Component { fn sentimentable(input: String) -> (f32, f32, f32, f32) { lazy_static::lazy_static! { static ref ANALYZER: vader_sentiment::SentimentIntensityAnalyzer<'static> = vader_sentiment::SentimentIntensityAnalyzer::new(); } let scores = ANALYZER.polarity_scores(input.as_str()); ( *scores.get("compound").unwrap_or(&0.0) as f32, *scores.get("pos").unwrap_or(&0.0) as f32, *scores.get("neg").unwrap_or(&0.0) as f32, *scores.get("neu").unwrap_or(&0.0) as f32, ) } } export!(Component); Build: Shell cd ~/wasm-udf/sentimentable cargo build --target wasm32-wasip1 --release Inspecting the Binary Let's check the exports first: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "^Export" -A 6 Four exports should be present: memory, sentimentable, cabi_realloc, and cabi_realloc_wit_bindgen_0_40_0. Now let's find the type signature of the sentimentable function. Find the sig index: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "func\[9\]" | head -1 Then look it up: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "type\[9\]" You should see: Plain Text - type[9] (i32, i32) -> i32 The signature is (i32, i32) -> i32 . This is the key difference between returning a single scalar and returning a tuple: wit-bindgen uses a direct f32 return for a single value, but switches to an indirect result pointer when returning a tuple. What's written at that pointer is four f32 values (16 bytes) at consecutive 4-byte offsets. The Java side reads all four. This illustrates an important distinction between the WIT interface definition and the generated core Wasm ABI. The WIT signature and the Wasm-level signature are different layers: wit-bindgen lowers WIT types to a core Wasm ABI, and the lowering strategy depends on the return type. A single scalar such as f32 is returned directly as a Wasm value. A tuple is returned indirectly through linear memory, with the caller receiving a pointer to where the values were written. The Java calling code must match the generated ABI rather than the WIT definition, which is why inspecting the binary with wasm-objdump before writing the Java wrapper is essential. Figure 3 shows the memory layout. Figure 3. Memory Layout. Figure 4 compares Cases 2 and 3. Figure 4. Case 2 vs. Case 3 ABI Comparison The SentimentUDF.java file Java package com.example; import io.github.kawamuray.wasmtime.Engine; import io.github.kawamuray.wasmtime.Func; import io.github.kawamuray.wasmtime.Linker; import io.github.kawamuray.wasmtime.Memory; import io.github.kawamuray.wasmtime.Module; import io.github.kawamuray.wasmtime.Store; import io.github.kawamuray.wasmtime.WasmFunctions; import io.github.kawamuray.wasmtime.WasmValType; import io.github.kawamuray.wasmtime.wasi.WasiCtx; import io.github.kawamuray.wasmtime.wasi.WasiCtxBuilder; import org.neo4j.procedure.Description; import org.neo4j.procedure.Name; import org.neo4j.procedure.UserFunction; import java.io.InputStream; import java.nio.ByteBuffer; import java.nio.ByteOrder; import java.nio.charset.StandardCharsets; import java.util.HashMap; import java.util.Map; public class SentimentUDF { @UserFunction("com.example.wasm.sentiment") @Description("Scores text using VADER sentiment analysis compiled to Wasm. Returns compound, positive, negative, neutral.") public Map<String, Double> sentiment(@Name("text") String text) throws Exception { if (text == null || text.isBlank()) { return Map.of("compound", 0.0, "positive", 0.0, "negative", 0.0, "neutral", 1.0); } byte[] wasmBytes; try (InputStream is = SentimentUDF.class.getResourceAsStream("/sentimentable.wasm")) { if (is == null) throw new RuntimeException("sentimentable.wasm not found in resources"); wasmBytes = is.readAllBytes(); } WasiCtx wasi = new WasiCtxBuilder().inheritStdout().inheritStderr().build(); try (Store<Void> store = Store.withoutData(wasi); Engine engine = store.engine(); Module module = Module.fromBinary(engine, wasmBytes); Linker linker = new Linker(engine)) { WasiCtx.addToLinker(linker); linker.module(store, "", module); Memory memory = linker.get(store, "", "memory").get().memory(); Func reallocFn = linker.get(store, "", "cabi_realloc").get().func(); WasmFunctions.Function4<Integer, Integer, Integer, Integer, Integer> realloc = WasmFunctions.func(store, reallocFn, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32); byte[] inputBytes = text.getBytes(StandardCharsets.UTF_8); int len = inputBytes.length; int strPtr = realloc.call(0, 0, 1, len); ByteBuffer buf = memory.buffer(store); buf.position(strPtr); buf.put(inputBytes); Func sentimentFn = linker.get(store, "", "sentimentable").get().func(); WasmFunctions.Function2<Integer, Integer, Integer> scoreFn = WasmFunctions.func(store, sentimentFn, WasmValType.I32, WasmValType.I32, WasmValType.I32); int resultPtr = scoreFn.call(strPtr, len); // read four f32 values at 4-byte offsets: compound, pos, neg, neu ByteBuffer resultBuf = memory.buffer(store); resultBuf.order(ByteOrder.LITTLE_ENDIAN); float compound = resultBuf.getFloat(resultPtr); float positive = resultBuf.getFloat(resultPtr + 4); float negative = resultBuf.getFloat(resultPtr + 8); float neutral = resultBuf.getFloat(resultPtr + 12); Map<String, Double> result = new HashMap<>(); result.put("compound", (double) compound); result.put("positive", (double) positive); result.put("negative", (double) negative); result.put("neutral", (double) neutral); return result; } } } The return type is (Map<String, Double> ), the null guard returning a neutral map and the four getFloat() reads at consecutive 4-byte offsets from the result pointer. Build and Deploy Copy the Wasm binary, build and deploy: Shell cp ~/wasm-udf/sentimentable/target/wasm32-wasip1/release/sentimentable.wasm \ ~/wasm-udf/neo4j-wasm-udf/src/main/resources/ cd ~/wasm-udf/neo4j-wasm-udf mvn -q clean package cp target/neo4j-wasm-udf-1.0-SNAPSHOT.jar \ ~/Library/Application\ Support/neo4j-desktop/Application/Data/dbmss/<your-dbms-id>/plugins/ Stop Neo4j, restart it, and run the verification queries. Positive sentence: Cypher RETURN com.example.wasm.sentiment('The movie was great') AS scores; Result: JSON { "compound": 0.624893307685852, "positive": 0.577464759349823, "negative": 0.0, "neutral": 0.4225352108478546 } Capitalization test: Cypher RETURN com.example.wasm.sentiment('The movie was GREAT!') AS scores; Result: JSON { "compound": 0.7290259003639221, "positive": 0.6307692527770996, "negative": 0.0, "neutral": 0.3692307770252228 } Empty string guard: Cypher RETURN com.example.wasm.sentiment('') AS scores; Result: JSON { "compound": 0.0, "positive": 0.0, "negative": 0.0, "neutral": 1.0 } All three cases behave correctly. Summary We set out to build the working example that our previous article said didn't exist. Here's what we showed. We embedded a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity map — compound, positive, negative and neutral — matching the return shape of the Java VADER UDF from the previous article. The wit-bindgen tuple ABI writes four f32 values to consecutive memory addresses; the Java side reads them back with a LITTLE_ENDIAN ByteBuffer. All four scores are correct, capitalization sensitivity works, and the empty string guard returns a sensible neutral map. The result is a workable integration pattern rather than a universal replacement for a native Java implementation. With the per-call initialization used in this proof of concept, the approach is best suited to low-frequency workloads where Wasm isolation and portability justify the additional complexity. For high-throughput workloads, the natural next step is to benchmark and reuse the Wasmtime engine and compiled module while keeping execution state appropriately isolated between calls. In the next article, we'll look at running TypeSafe AI's Jev inside Neo4j for calibrated sentiment decisions. Stay tuned! The full source code is available on GitHub.

By Akmal Chaudhri DZone Core CORE
Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus
Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus

Enterprise context — Acme FinServ. SOC 2 CC7 (system monitoring) requires that Acme can detect and investigate anomalous activity. When an agent-driven workflow touches customer data at 2 AM, "we have logs somewhere" is not an answer an auditor accepts. The distributed trace built in this part is the forensic evidence trail: a single trace ID that ties the Goose prompt to every agentgateway policy decision and every Quarkus tool call, so a post-incident review can reconstruct exactly which agent did what, in what order, and how long each governed hop took. The Core Problem In Part 1, we built a Quarkus MCP tool server. In Part 2, we secured it with agentgateway's JWT authentication, RBAC, and ExtMCP guardrails. The architecture works — but when something goes wrong in production, you're flying blind. Agentic workflows are fundamentally different from traditional request-response APIs. A single user prompt like "Debug customer CUST-4091" triggers a multi-round-trip loop: Goose calls tools/list to discover available toolsThe LLM selects getCustomerStatus and Goose sends tools/callThe LLM reads the response, sees primaryRegion: US-EAST-1, and chains a second tools/call to getZoneHealthLogsThe LLM correlates both results and generates a diagnostic summary Each of these hops crosses process boundaries: Goose → agentgateway → Quarkus. Without distributed tracing, you see four isolated HTTP requests in your access logs. You cannot tell they belong to the same agentic workflow. When step 3 takes 12 seconds instead of 200ms, you have no waterfall to pinpoint whether the latency came from agentgateway policy evaluation, Quarkus bean validation, or a slow downstream call. This creates black holes in telemetry dashboards — the exact gap that autonomous agents exploit to degrade silently. The Solution: W3C Trace Context Across All Three Layers The fix is standard distributed tracing, applied to the MCP transport layer: agentgateway exports spans for every proxied MCP request and propagates traceparent headers to the backend.Quarkus with quarkus-opentelemetry picks up the incoming traceparent, creates child spans for tool execution and bean validation, and exports them to the same Jaeger instance.Jaeger correlates both sides into a single trace waterfall — one view from agent prompt to tool result. Prerequisites Everything from Parts 1 and 2, plus: Podman – for running Jaeger (podman compose) Verify Podman is available: Shell podman --version Step 1: Launching the Observability Backend We use Jaeger v2 as both the OTLP collector and the trace UI. A single container accepts traces from agentgateway on port 4317 (OTLP gRPC) and from Quarkus on port 4318 (OTLP HTTP), and serves the query UI on port 16686. Shell cd part3-observability podman compose up -d This starts Jaeger v2 with OTLP collection enabled by default. Verify it's running: Shell curl -sf http://localhost:16686/ > /dev/null && echo "Jaeger UI is ready" Open http://localhost:16686 — you'll see an empty Jaeger UI. We'll populate it with MCP traces in the following steps. Production Alternative: Grafana Tempo For production deployments, replace Jaeger with Grafana Tempo backed by object storage (S3/GCS). The OTLP endpoint stays the same — only the compose.yml changes. Grafana provides richer dashboards, alerting, and long-term trace retention. Step 2: Enabling OpenTelemetry in Quarkus Add the quarkus-opentelemetry extension to Part 1's pom.xml: Properties files <dependency> <groupId>io.quarkus</groupId> <artifactId>quarkus-opentelemetry</artifactId> </dependency> Configure the exporter in application.properties: Properties files # OpenTelemetry quarkus.otel.service.name=customer-tools quarkus.otel.exporter.otlp.traces.endpoint=http://localhost:4318 quarkus.otel.exporter.otlp.traces.protocol=http/protobuf quarkus.otel.traces.sampler=always_on quarkus.otel.traces.suppress-non-application-uris=false PropertyPurposeservice.nameIdentifies this service in Jaeger's service dropdowntraces.endpointOTLP HTTP receiver — Jaeger's port 4318 (base URL only; Quarkus appends /v1/traces)traces.protocolhttp/protobuf — Quarkus uses its Vert.x-based HTTP exportertraces.sampleralways_on — sample every span (reduce in production)suppress-non-application-urisfalse — include MCP endpoint spans (they'd be filtered otherwise) When no OTLP collector is running (Parts 1 and 2 without Jaeger), Quarkus logs a connection warning, but the MCP server works normally. When the collector IS running (Part 3), traces flow automatically. Zero code changes to the MCP tools. Rebuild Part 1: Shell cd part1-quarkus-mcp mvn package -DskipTests What Quarkus Auto-Instruments With quarkus-opentelemetry on the classpath and the SDK enabled, Quarkus automatically creates spans for: LayerSpan NameWhat It CapturesHTTP serverPOST /mcpInbound MCP request with method, status, latencyCDI beansCustomerServiceTools.getCustomerStatusTool execution time within the MCP handlerBean ValidationHibernateValidatorParameter validation before tool logic runsREST clientOutbound HTTP callsAny downstream API calls (future extensions) No @WithSpan annotations needed. The Quarkus OpenTelemetry extension instruments the reactive pipeline automatically. Step 3: Configuring W3C Trace Context in agentgateway agentgateway supports native OpenTelemetry trace export. Add a tracing block to the gateway configuration: Properties files config: adminAddr: localhost:15000 tracing: otlpEndpoint: http://localhost:4317 otlpProtocol: grpc randomSampling: 1.0 FieldPurposeotlpEndpointOTLP receiver — Jaeger's port 4317otlpProtocolgrpc for OTLP/gRPC (also supports http)randomSamplingSample 100% of traces (reduce to 0.01–0.1 in production) How Trace Propagation Works When agentgateway receives an MCP request: Creates a root span for the proxy operation (e.g., agentgateway.mcp.proxy)Injects a traceparent header into the forwarded request to Quarkus:traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01Quarkus reads the traceparent, creates a child span under the same trace ID, and records tool executionBoth spans export to Jaeger via OTLP, where they appear as a single correlated trace This is standard W3C Trace Context propagation — the same mechanism used across all OpenTelemetry-instrumented services. Configuration Files Part 3 provides two agentgateway configurations: ConfigUse Caseconfig-traced.yamlTracing only — proxy + OTLP export, no security layersconfig-traced-guardrails.yamlTracing + ExtMCP guardrails — observe the guardrail evaluation spans too Step 4: Running the Interactive Demo Start all services with the one-command script: Shell cd part3-observability ./start-all.sh The script starts Jaeger, Quarkus (with OTel enabled), and agentgateway (with trace export), then launches the demo SPA on :8890. Open the MCP Observability Console at http://localhost:8890/index.html and walk through the three demo steps: Initialize – Establishes an MCP session through agentgateway. The architecture diagram animates the trace propagation: root span creation in agentgateway, traceparent injection, child span in Quarkus, and OTLP export to Jaeger.List Tools – Discovers all 5 tools through the traced proxy. The trace waterfall panel shows the agentgateway proxy span and the Quarkus HTTP span side by side with timing.Multi-Tool Workflow – Simulates Goose's multi-turn reasoning: getCustomerStatus (finds region US-EAST-1) → getZoneHealthLogs (checks zone health) → getSLACompliance (correlates SLA metrics). Each step generates a full trace with waterfall visualization. The stat tiles track traces generated, spans collected, and Jaeger status. Click Open Jaeger to view the real trace waterfalls in the Jaeger UI at http://localhost:16686. Step 5: Generating Traces via CLI To generate additional traces manually, simulate a multi-turn agentic workflow: Shell # Step 1: Initialize MCP session export MCP_SESSION_ID=$(curl -s -D - http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}' \ | grep -i "mcp-session-id:" | sed 's/.*: //' | tr -d '\r') # Step 2: Discover tools curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 3: Agent calls getCustomerStatus (first tool invocation) curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"getCustomerStatus","arguments":{"customerId":"CUST-4091"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 4: Agent chains getZoneHealthLogs based on the region from step 3 curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"getZoneHealthLogs","arguments":{"zoneId":"US-EAST-1"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 5: Agent fetches SLA compliance for correlation curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":5,"method":"tools/call","params":{"name":"getSLACompliance","arguments":{"serviceId":"api-gateway"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . Each of these requests generates a trace that flows through agentgateway into Quarkus and lands in Jaeger. Step 6: Visualizing the Trace Waterfall in Jaeger Open http://localhost:16686 in your browser. Finding Traces In the Service dropdown, select customer-tools (Quarkus) or agentgatewayClick Find TracesClick on any trace to open the waterfall view Reading the Waterfall A typical tools/call trace shows the following span hierarchy: Shell agentgateway.mcp.proxy [12ms] └─ POST /mcp [8ms] ← Quarkus HTTP server └─ CustomerServiceTools.getCustomerStatus [2ms] ← CDI tool execution SpanServiceWhat It Tells Youagentgateway.mcp.proxyagentgatewayTotal proxy overhead including policy evaluationPOST /mcpcustomer-toolsQuarkus HTTP handling time for the MCP requestgetCustomerStatuscustomer-toolsPure tool execution time (business logic) What to Look For Proxy overhead: The gap between the agentgateway span and the Quarkus span shows network + policy evaluation time. If this grows, check guardrail server latency.Validation time: Bean Validation spans appear before tool execution. Regex-heavy patterns like ^CUST-[0-9]{4,8}$ are fast, but complex validators on large payloads can add latency.Multi-turn correlation: When Goose chains multiple tool calls (e.g., getCustomerStatus → getZoneHealthLogs), each appears as a separate trace. The mcp-session-id tag lets you filter all traces belonging to one agent session.Error traces: Failed validations (invalid customer ID format) or guardrail rejections (blocked poison payloads) produce error spans with exception details. Connecting Goose for Real Traces Launch Goose pointed at agentgateway and prompt a multi-tool workflow: Shell goose session "Debug customer CUST-4091 — check their account status, then pull health logs for their region and SLA compliance for api-gateway." This generates a burst of correlated traces in Jaeger showing Goose's multi-turn tool orchestration from the proxy layer down to individual tool execution spans. What We Achieved Starting from the secured architecture in Part 2, we added full observability without changing any MCP tool code: LayerWhat We AddedConfig ChangeQuarkusquarkus-opentelemetry dependencypom.xml + application.propertiesagentgatewaytracing block in config YAMLconfig-traced.yamlObservability backendJaeger all-in-one via Podman Composecompose.yml The entire stack runs locally with a single ./start-all.sh command and produces end-to-end trace waterfalls in Jaeger. Production Considerations ConcernLocal (this tutorial)ProductionTrace backendJaeger all-in-one (in-memory)Grafana Tempo + object storageSampling rate100% (default: 1.0)1-10% or adaptive samplingTrace retentionContainer lifetimeDays/weeks in durable storageAlertingManual Jaeger inspectionGrafana alerting on span latency/error rateMetricsTraces onlyAdd Prometheus + quarkus-micrometer for RED metrics Coming Up in Part 4 With tracing in place, you can now see every MCP tool call flowing through the system. In Part 4, we will move beyond single-agent tool calls to multi-agent orchestration — using the Agent-to-Agent (A2A) protocol to coordinate autonomous agents that can delegate work, enforce governance via AGENTS.md, and call back into our MCP tool services.

By Daniel Oh DZone Core CORE
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

In my previous article, I walked through running coding agents inside Docker Sandboxes on a local machine. We installed the sbx CLI, started with a small project, and covered the commands needed to run, stop, and remove a sandbox. This time, I want to take that same workflow off the laptop. Docker added cloud sandboxes in version 0.42.0. You can now use sbx --cloud to run an agent on Docker-managed infrastructure instead of using your machine for the sandbox’s compute. The command is simple to use. The part that is worth understanding is how you get your code into that environment, work with the agent, and bring the changes back locally. That is what we will do here. Nothing complicated; we will start with a small Python project, one coding task, and a cloud sandbox. We will remove the sandbox when we are done with the work. What Changes With a Cloud Sandbox? The sbx CLI still runs in your terminal. With --cloud, supported commands target Docker’s cloud service rather than your local sandbox environment. For example: PowerShell sbx ls Lists your local sandboxes. PowerShell sbx --cloud ls Lists your cloud sandboxes. That distinction matters throughout this walkthrough. If you forget --cloud, you are not asking about the same environment. Cloud sandboxes also have separate credentials and network policies. Do not assume that an agent login or network policy you configured locally is already available in the cloud. For this example, we will copy individual files explicitly. That keeps it easy to see what we send to the sandbox and what we bring back. Before You Start You will need: An updated sbx CLI with cloud support, introduced in version 0.42.0.A Docker account with an active Docker Agentic Platform plan for cloud compute.Authentication for the coding agent you want to use. This walkthrough uses Claude.Python 3 available in the sandbox image for the example. Note: The free sbx CLI does not mean cloud compute is free. Docker bills cloud compute based on usage, and your model provider bills inference separately. Check your account’s pricing before starting. Also, use a small sample project first. Running remotely means sending code off your machine. For company repositories, make sure that is allowed before uploading anything. The host-side commands below use PowerShell. Paths inside the cloud sandbox use Linux-style paths. Step 1: Sign In and Configure the Agent First, check your installed version: PowerShell sbx version If you are still using an older version from the previous walkthrough, update it before continuing. Sign in to Docker: PowerShell sbx login For Claude, Docker documents a cloud OAuth flow: PowerShell sbx --cloud secret set anthropic --oauth Complete the provider sign-in with an account that has the required access. Notice the --cloud flag here, too. These credentials are stored for cloud use, separately from your local sandbox credentials. There is no reason to put a token in our Python files or paste it into an agent prompt. Step 2: Create a Small Project Let us give the agent something specific to fix. Create a project folder: PowerShell New-Item -ItemType Directory -Path .\cloud-sandbox-demo Set-Location .\cloud-sandbox-demo Inside it, create a file named slug.py: Python def make_slug(text): return text.lower().replace(" ", "-") This converts "Docker Sandboxes" into "docker-sandboxes". It works for that input, but it does not handle whitespace very well. Leading spaces become leading hyphens. Repeated spaces become repeated hyphens. Tabs are not handled at all. Now create test_slug.py: Python import unittest from slug import make_slug class SlugTests(unittest.TestCase): def test_two_words(self): self.assertEqual(make_slug("Docker Sandboxes"),"docker-sandboxes") if __name__ == "__main__": unittest.main() We have one passing case and a clear improvement to make. The point is not that this function needs cloud compute. It is small enough that we can focus on the sandbox workflow without spending half the article explaining an application. Step 3: Start a Cloud Sandbox Run the following command: PowerShell sbx --cloud run --detached --name cloud-demo --ttl 1h claude This creates a cloud sandbox and starts the agent without attaching your terminal to it. The flags in the above command are for doing useful things: --cloud selects the cloud environment.--detached returns control to your terminal.--name cloud-demo gives the sandbox a recognizable name.--ttl 1h requests a one-hour lifetime. Important: The documented default action when the TTL expires is deletion. Treat this as a disposable environment, and copy your work out before the deadline. The command prints a sandbox ID. You can also find it with: PowerShell sbx --cloud ls Copy that ID into a PowerShell variable: PowerShell $sandbox = "PASTE_YOUR_SANDBOX_ID_HERE" Use the real ID returned by Docker, not the placeholder above. Keep using this terminal for the remaining commands. One detail to remember is a detached cloud run creates a new sandbox. It is not the command to run repeatedly when you want to reconnect to the same one. Step 4: Copy the Project Into the Sandbox Create a directory for our example: PowerShell sbx --cloud exec $sandbox mkdir -p /workspace/demo The mkdir command runs inside the Linux sandbox, not on Windows. Now copy the two files: PowerShell sbx --cloud cp .\slug.py "${sandbox}:/workspace/demo/slug.py" sbx --cloud cp .\test_slug.py "${sandbox}:/workspace/demo/test_slug.py" The ${sandbox} syntax is intentional. In PowerShell, it separates the variable name from the colon used in Docker’s SANDBOX:PATH format. This is also why I am copying individual files rather than uploading the entire folder. We do not need a virtual environment, local configuration, or an accidentally included .env file for this task. Run the existing test inside the sandbox: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v If your selected image does not include Python 3, add it inside the sandbox before continuing. The existing test only covers two words separated by one space. Passing it does not mean the whitespace handling is correct yet. Step 5: Give the Agent a Narrow Task Attach to the running cloud sandbox: PowerShell sbx --cloud attach $sandbox Now give Claude a concrete task: Plain Text Work on the Python project in /workspace/demo. Update make_slug so that: - The output remains lowercase. - Leading and trailing whitespace is removed. - Consecutive whitespace becomes a single hyphen. - Spaces, tabs, and newlines are handled consistently. - Empty input returns an empty string. Add unit tests for these cases using unittest. Keep the existing test. Do not add third-party dependencies or modify files outside this project. Run the tests and summarize which files you changed. This is much more useful than asking the agent to “improve the project.” We have told it what the function should do, which edge cases matter, and how much freedom it has. There is no reason for it to introduce a framework or reorganize the project. The prompt is task guidance, though — not a security policy. File access, network access, and credentials still need the appropriate sandbox controls. Once the agent finishes, use Ctrl + backslash to detach and return to your local terminal. Detaching does not stop the cloud sandbox. Step 6: Run the Tests and Bring the Changes Back Run the test command again from your terminal: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v This executes inside the cloud sandbox. It is not running against your original local files. For this task, a straightforward implementation could look like: Python def make_slug(text): return "-".join(text.lower().split()) Calling split() without a separator handles consecutive whitespace and removes leading and trailing whitespace. Joining those words with a hyphen gives us the requested behavior. The agent may arrive at a different implementation. Read it rather than assuming that passing tests makes every change worth keeping. Create a separate folder for the returned files: PowerShell New-Item -ItemType Directory -Path .\review Copy the modified files into it: PowerShell sbx --cloud cp "${sandbox}:/workspace/demo/slug.py" .\review\slug.py sbx --cloud cp "${sandbox}:/workspace/demo/test_slug.py" .\review\test_slug.py Your original files are still untouched. If you have Git installed, compare the versions: PowerShell git diff --no-index -- .\slug.py .\review\slug.py git diff --no-index -- .\test_slug.py .\review\test_slug.py You can also compare them in your editor. Look at the tests as closely as the implementation. Did the agent actually add cases for tabs and newlines? Did it keep the original test? Did it add anything unrelated? For a real repository, I would bring the changes into a working branch and use the normal review process. The sandbox changes where the agent works. It does not replace code review. What About Web Applications? Our Python example does not start a server. If you use a web project instead, cloud sandboxes can expose an application through a public HTTPS URL. For an application already listening on sandbox port 3000: PowerShell sbx --cloud ports $sandbox --publish 3000 sbx --cloud ports $sandbox Use the URL returned by Docker. This is different from publishing a local port such as localhost:3000. In cloud mode, the command accepts the sandbox port, and Docker assigns the public URL. Note: Publicly reachable is not the same as private. Do not expose an unauthenticated admin page, secrets, or sensitive test data. Remove the exposure when you no longer need it: PowerShell sbx --cloud ports $sandbox --unpublish 3000 Step 7: Clean Up the Cloud Sandbox Before cleanup, make sure the files you want to keep are on your machine. If you want to pause rather than delete, Docker documents cloud stop as preserving the sandbox’s memory and disk state: PowerShell sbx --cloud stop $sandbox Do not assume that preserved resources have no cost. Check your plan’s billing terms. For this small exercise, we have already copied the results out, so we can remove the sandbox: PowerShell sbx --cloud rm $sandbox Confirm the removal when prompted, then list your cloud sandboxes: PowerShell sbx --cloud ls There is an important difference from my earlier article: sbx --cloud rm --all is intentionally disabled. Cloud cleanup requires explicit sandbox identifiers. That is a useful safeguard. A cloud credential may have access to more than the one environment you were experimenting with. A Few Things That Can Slow You Down If the agent cannot authenticate, check its cloud credentials. A successful local session does not prove that cloud authentication is configured. If it cannot reach a service, check the cloud network policy. Do not immediately open access to everything just to make an error disappear. If your local files have not changed, remember the workflow we used: we copied files into the cloud and copied the results back. Those copies are not a live synchronization mechanism. And if you are coming back to a running sandbox, use attach. Repeating the detached creation command gives you another sandbox, not another connection to the original one. Conclusion What I like about this addition is that it keeps the workflow familiar. We are still using sbx, still giving the agent a specific project, and still deciding what work to keep. The difference is where that work happens. Start small. Send only the files the agent needs, give it one clear task, and bring the results back into your normal development process. Once that feels comfortable, move on to a larger repository or a task that actually benefits from remote compute. And copy the changes back before the sandbox expires. A useful fix is not very useful if the only copy disappears with the environment.

By Naga Santhosh Reddy Vootukuri DZone Core CORE
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI

Quick Summary Both Embabel and LangGraph4j let a Java developer build multi-step AI agents without leaving the JVM.Embabel hands the framework a goal and a bag of typed actions, and lets a planner decide the order on its own.LangGraph4j asks the developer to draw the exact graph of nodes and edges by hand.We will understand both philosophies through a real Embabel agent, a small Kolkata street-crossing example, a comparison table, and finally a bigger question — is Java catching up with Python in enterprise AI work? Where the Story Starts If you are a Java developer today, you are watching two worlds collide. On one side are large language models, which grew up almost entirely in Python. On the other side is enterprise Java, which has spent twenty-five years learning to build systems that banks, insurance companies, and hospitals can actually trust. Two frameworks are now trying to bring these two worlds together on the JVM: Embabel and LangGraph4j. Both help you build an "agent" — a piece of software that uses an LLM to complete a task in several steps, rather than in one single prompt. But the way they think about "steps" is completely different. That difference is what this article is about. Meeting Embabel Through a Real Piece of Code The best way to understand Embabel is to look at actual code, not a slide. Here is a small agent that gives retirement planning advice, written the Embabel way. Java package com.example.demo; import com.embabel.agent.api.annotation.AchievesGoal; import com.embabel.agent.api.annotation.Action; import com.embabel.agent.api.annotation.Agent; import com.embabel.agent.api.common.OperationContext; import com.embabel.agent.domain.io.UserInput; import java.util.Arrays; import java.util.List; @Agent(name = "RetirementPlannerAgent", description = "This agent provides retirement planning advice.") public class RetirementPlannerAgent { record RetirementUserInput(int presentAge, int targetRetirementAge, double annualIncome) { } record RetirementPlanAdvice(String advice) { } record RetirementPlanAdvices(List<RetirementPlanAdvice> advices) { } enum RiskToleranceLevel { LOW, MEDIUM, HIGH } @Action(description = "Identify the present age, target retirement age, and annual income of the user from the user message.") public RetirementUserInput identifyRetirementUserInput(UserInput userInput, OperationContext context) { String content = userInput.getContent(); return context.ai().withDefaultLlm() .creating(RetirementUserInput.class) .fromPrompt(""" Identify the present age, target retirement age, and annual income of the user from the following message and return them as a JSON object. User message: %s """.formatted(content)); } @Action(description = "Identify the risk tolerance level of the user based on the present age, target retirement age, and annual income.") public RiskToleranceLevel identifyRiskToleranceLevel(RetirementUserInput retirementUserInput, OperationContext context) { return context.ai().withDefaultLlm() .creating(RiskToleranceLevel.class) .fromPrompt(""" Identify the risk tolerance level of the user based on the following information and return it as a JSON object. Permitted values for risk tolerance level are: %s Present age: %d Target retirement age: %d Maximum years to retirement: %d Annual income: %.2f """.formatted(Arrays.toString(RiskToleranceLevel.values()), retirementUserInput.presentAge(), retirementUserInput.targetRetirementAge(), (retirementUserInput.targetRetirementAge() - retirementUserInput.presentAge()), retirementUserInput.annualIncome())); } @Action(description = "Provide retirement plan advice based on the user's risk tolerance level.") @AchievesGoal(description = "Provide retirement plan advice based on the user's risk tolerance level.") public RetirementPlanAdvices provideRetirementPlanAdvice(RiskToleranceLevel riskToleranceLevel, OperationContext context) { String systemPrompt = """ You are a retirement planning advisor. Based on the user's risk tolerance level, provide a list of retirement plan advices. """; return context.ai().withDefaultLlm() .creating(RetirementPlanAdvices.class) .fromPrompt(""" %s User's risk tolerance level: %s """.formatted(systemPrompt, riskToleranceLevel.name())); } } Now look closely at what is missing from this code. There is no method called runAgent() that calls identifyRetirementUserInput(), then identifyRiskToleranceLevel(), then provideRetirementPlanAdvice(), in that order. Nowhere did the developer type out the sequence. Instead, each @Action simply states two things: What type it needs as input (its precondition).What type it produces as output (its effect). provideRetirementPlanAdvice needs a RiskToleranceLevel. identifyRiskToleranceLevel happens to produce a RiskToleranceLevel from a RetirementUserInput. And identifyRetirementUserInput produces that RetirementUserInput from the raw UserInput. Embabel's planner looks at all this at runtime and works out, on its own, that this is the only order in which the goal (@AchievesGoal) can be reached. This is Embabel's whole philosophy in one sentence: give the framework a goal and a set of typed building blocks, and let it plan. The Same Investment Advisory Workflow, Wired by Hand in LangGraph4j Now let us build the exact same three-step advisory flow — read the user's numbers, work out risk tolerance, give advice — the LangGraph4j way. Here, the developer draws the graph explicitly, and the shared state is a plain key-value map (an AgentState) rather than Soham's strongly typed records. Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.NodeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.List; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; /** * @author Soham Sengupta * @since 2026-09-13 * @description The retirement/investment advisory workflow from the * Embabel example above, this time wired explicitly as a LangGraph4j * graph. Every step, and the order between the steps, is declared * here by the developer - there is no planner discovering it. */ public class RetirementPlannerGraph { // Step 1: pull the present age, target retirement age, and income out of free text. static class IdentifyUserInputNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String userMessage = state.<String>value("userMessage").orElseThrow(); // In a real system: call the LLM here (say, via langchain4j) and parse // presentAge / targetRetirementAge / annualIncome out of userMessage. return Map.of( "presentAge", 32, "targetRetirementAge", 60, "annualIncome", 1200000.0); } } // Step 2: classify how much investment risk this user can reasonably take. static class IdentifyRiskToleranceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { int presentAge = state.<Integer>value("presentAge").orElseThrow(); int targetRetirementAge = state.<Integer>value("targetRetirementAge").orElseThrow(); // In a real system: call the LLM here with these values and ask it to // return LOW, MEDIUM, or HIGH as the risk tolerance level. String riskTolerance = (targetRetirementAge - presentAge) > 20 ? "HIGH" : "MEDIUM"; return Map.of("riskTolerance", riskTolerance); } } // Step 3: turn the risk tolerance into a concrete list of investment advice. static class ProvideAdviceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String riskTolerance = state.<String>value("riskTolerance").orElseThrow(); // In a real system: call the LLM here to draft actual advice - suitable // Indian investment instruments for this riskTolerance level, and so on. List<String> advice = List.of("Suggested investment mix for a " + riskTolerance + " risk profile."); return Map.of("advice", advice); } } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("identifyUserInput", node_async(new IdentifyUserInputNode())) .addNode("identifyRiskTolerance", node_async(new IdentifyRiskToleranceNode())) .addNode("provideAdvice", node_async(new ProvideAdviceNode())) .addEdge(START, "identifyUserInput") .addEdge("identifyUserInput", "identifyRiskTolerance") .addEdge("identifyRiskTolerance", "provideAdvice") .addEdge("provideAdvice", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke( Map.of("userMessage", "I am 32, want to retire at 60, and earn 12 lakh a year.")); result.ifPresent(state -> System.out.println(state.data())); } } Two things stand out next to the Embabel version. First, there is a main method here that explicitly lists identifyUserInput -> identifyRiskTolerance -> provideAdvice as edges — in the Embabel version, that sequence was never written down anywhere; it was worked out by the planner. Second, the shared state (AgentState) is just a bag of string keys and values, read back out with state.value("presentAge"), instead of Soham's own strongly typed RetirementUserInput and RiskToleranceLevel. For this particular workflow, which happens to be a strict straight line with no branching, LangGraph4j's graph is refreshingly easy to read top to bottom. The real difference shows up once branching enters the picture, which is exactly where our next example — crossing a Kolkata road — comes in. A Bit of History: From Servlet to Spring, Now From Spring AI to Embabel To understand why Embabel is built this way, it helps to know who built it. Embabel comes from Rod Johnson — the same person who created the Spring Framework more than two decades ago. Back in the early 2000s, enterprise Java was drowning in heavy J2EE application servers and Enterprise Java Beans. Rod was solving real problems in the finance industry at the time, found the existing tools too heavy, and wrote a book and a framework that simplified things a great deal. That framework became Spring, and it changed how an entire generation of Java developers worked. In 2025, Rod did something similar again, this time for AI agents. When he introduced Embabel to the Java community, he framed it using a comparison that Java developers will find very familiar: Spring AI is to Embabel roughly what the plain old Servlet API once was to Spring MVC. Spring AI gives you the low-level plumbing — talking to a model, building a prompt, calling a tool. Embabel sits one level above that, giving you the actual application framework — goals, actions, planning, and a proper domain model — the same kind of jump in abstraction that Spring itself brought to raw Servlets and EJBs, twenty years back. Embabel (pronounced "Em-BAY-bel") is written mainly in Kotlin, but as you can see from the retirement planner code above, it feels completely natural to use from plain Java. It is also built to sit closely with Spring, which is exactly why an existing Spring shop can pick it up without much friction. GOAP: The Planning Engine Hiding Inside Embabel The planning idea inside Embabel is not new — it is borrowed from video games, and it is called GOAP, short for Goal-Oriented Action Planning. GOAP was built to make game characters (think of soldiers in an old shooter game) decide, on their own, a believable sequence of actions to reach a goal, instead of following a fixed script. GOAP needs three things: A state of the world as it stands right now.A set of actions, each with a precondition (what must be true to run it) and an effect (what becomes true after it runs).A goal, which is simply a desired state. A search algorithm (usually the well-known A* algorithm) then works out the cheapest chain of actions that gets you from where you are to where you want to be. Embabel uses exactly this idea, but instead of asking the LLM to "think step by step" about the plan (which is often unreliable) or asking the developer to hard-code the entire flow (which is rigid), it asks a deterministic planner to search over your own typed Java or Kotlin methods. The precondition of an action is simply the input type it needs. The effect is simply the output type it produces. Your domain classes — RetirementUserInput, RiskToleranceLevel, RetirementPlanAdvices — literally become the "world state" the planner reasons about. No LLM guesswork is involved in deciding the order; the LLM is only used inside each action, for the part it is actually good at — understanding and generating language. Understanding the GOAP + OOAD Loop, With a Kolkata Traffic Signal All this can feel a bit abstract, so let us make it concrete with something every Kolkata resident understands very well — crossing a busy road. Picture Soham standing at a signal with his four-year-old son, Kit, holding his hand. It is a typical Kolkata crossing — buses, yellow taxis, and autos, and the signal, while present, is not always fully obeyed. Sometimes there is a traffic constable standing in the middle of the road, waving vehicles through by hand, overriding the signal completely. Step one — model the world as an object (this is the OOAD part). In Object-Oriented Analysis and Design, we are trained to represent a real-world situation as a class with clearly named fields. Here, the "world" Soham is observing can be written as one simple record: record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } Step two — define the goal. The goal is not "the signal is green." The goal is "Soham and Kit have reached the other side, safely." Reaching a green signal is only useful if it actually leads there. Step three — define the actions, each with its precondition and effect (this is the GOAP part). Notice that none of these actions know about each other. Each one only knows what it needs and what it produces: holdChildHand – produces handHeld = true. This is usually the very first thing a responsible parent does, well before even looking at the signal.waitForSafeSignal – needs the current RoadSituation, and produces signalGreen = true once either the light turns green or a traffic constable is present and waving pedestrians across.confirmTrafficClear – because a green signal in Kolkata does not always mean an auto will not sneak through, this action looks both ways and produces trafficClear = true.crossTheRoad (the @AchievesGoal action) – only runs once handHeld, trafficClear, and (signalGreen or policeOnDuty) are all true. Step four — let the planner loop. This is the actual "GOAP + OOAD loop": the planner looks at the current RoadSituation object, picks whichever action's precondition is already satisfied and whose effect moves the world closer to the goal, executes it, updates the RoadSituation, and checks again if the goal is reached. It keeps looping — plan, act, update state, re-check — until crossTheRoad finally fires. The beautiful part is what happens on a day when the signal itself is not working — a fairly common event in Kolkata, especially during a power cut or during Puja season when the police fully take charge of a crossing. If tomorrow you add one more action, say waitForPoliceWave, which also produces a "safe to proceed" fact, the planner will simply discover this new path on its own the next time it runs. Nobody needs to redraw anything, because nobody drew anything explicit in the first place. The Same Scenario as Embabel Code Here is a simplified sketch of the above scenario, written in the same style as the retirement planner. Treat it as a teaching example rather than a compiled, production-ready class: Java /** * @author Soham Sengupta * @since 2026-09-13 * @description A small agent that plans how Soham and his four-year-old * son Kit can safely cross a busy Kolkata road. Written purely to show * how Embabel's GOAP-style planner reasons over typed domain objects * (OOAD) to reach a goal, without the developer wiring the order by hand. */ @Agent(name = "StreetCrossingAgent", description = "Plans a safe road crossing for a parent and a young child.") public class StreetCrossingAgent { record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } record CrossingPlan(String narrative) { } @Action(description = "Hold Kit's hand before anything else - the first rule of the road.") public RoadSituation holdChildHand(UserInput userInput) { return new RoadSituation(false, false, true, false); } @Action(description = "Wait till the signal turns green, or till a traffic constable waves pedestrians across.") public RoadSituation waitForSafeSignal(RoadSituation situation, OperationContext context) { boolean safeToMove = situation.signalGreen() || situation.policeOnDuty(); return new RoadSituation(safeToMove, situation.trafficClear(), situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Look right, then left, then right again - the signal alone is not a guarantee in Kolkata traffic.") public RoadSituation confirmTrafficClear(RoadSituation situation) { return new RoadSituation(situation.signalGreen(), true, situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Cross only when hand is held, the way is clear, and the signal or constable allows it.") @AchievesGoal(description = "Soham and Kit have reached the other side of the road safely.") public CrossingPlan crossTheRoad(RoadSituation situation) { return new CrossingPlan( "Soham held Kit's hand tight, waited for the green man, checked both sides once more, then crossed together."); } } Now compare this with how the same scenario would look in LangGraph4j's philosophy — as an explicit graph you draw yourself, branches and all: Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.AsyncEdgeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; public class StreetCrossingGraph { // Stand-ins for a real signal sensor and a quick look both ways. private static boolean checkSignal() { return true; } private static boolean lookBothWays() { return true; } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("holdHand", node_async(state -> Map.of("handHeld", true))) .addNode("waitForSignal", node_async(state -> { boolean policeOnDuty = state.<Boolean>value("policeOnDuty").orElse(false); return Map.of("signalGreen", checkSignal() || policeOnDuty); })) .addNode("checkTraffic", node_async(state -> Map.of("trafficClear", lookBothWays()))) .addNode("cross", node_async(state -> Map.of("crossed", true))) .addEdge(START, "holdHand") .addEdge("holdHand", "waitForSignal") .addConditionalEdges("waitForSignal", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("signalGreen").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "checkTraffic", "retry", "waitForSignal")) .addConditionalEdges("checkTraffic", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("trafficClear").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "cross", "retry", "checkTraffic")) .addEdge("cross", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke(Map.of("policeOnDuty", false)); result.ifPresent(state -> System.out.println(state.data())); } } Notice the two addConditionalEdges calls — this is how LangGraph4j handles a branch: an edge action returns a label ("proceed" or "retry"), and a small map resolves that label to the actual next node. This is exactly the graph the developer must draw by hand, action by action and branch by branch, for a scenario that Embabel's planner worked out on its own. As a flowchart, that graph looks like this: Both pieces of code reach the same goal. But in Embabel, nobody drew this flowchart — the planner found it. In LangGraph4j, this flowchart is the code. If a new real-world case turns up tomorrow, Embabel's planner can absorb it automatically as long as the new action's types fit; LangGraph4j needs a human to open the graph and add a new node or edge. Comparing the Two Philosophies Aspect Embabel LangGraph4j Core idea Give a goal and typed actions; a GOAP/A* planner works out the order Developer explicitly wires nodes and edges into a graph Mental model "What do I want, and what building blocks do I have?" "What are my steps, and how do they branch?" Control flow Discovered at runtime by the planner Declared upfront by the developer Role of the LLM Used only inside actions, never for deciding sequence Can be used inside nodes; sequence is still fixed by the graph Adapting to a new case Often automatic, if a new action's types fit the gap Needs a human to add a new node or edge Tracing "why this order" Needs the planner's own logging/tooling to see the chosen path Very direct — the graph is already the flowchart Roots Kotlin-first, Java-friendly, close to Spring A faithful Java port of Python's LangGraph, works with Langchain4j and Spring AI Maturity (as of late 2026) Young, pre-1.0, moving fast Older and more widely adopted, with a large existing community When to Reach for Which Reach for Embabel when: Your goal is clear, but the exact path to it can honestly vary depending on the situation.You want a deterministic, non-LLM planner deciding the order, not the LLM guessing it.You are already deep in the Spring ecosystem and like strongly typed domain models.New cases keep appearing over time, and you would rather add one new action than redraw a graph. Reach for LangGraph4j when: You already know the exact stages of your workflow — say, a well-understood pipeline of retrieve, rerank, generate, and validate.You want the flow to be visible as an actual graph, easy to explain to a non-technical stakeholder.Your team is porting an existing Python LangGraph pipeline and wants the Java version to mirror it closely.You value a larger, more mature community with more examples to learn from, at least for now. Neither approach is "better" in an absolute sense. Embabel bets on planning; LangGraph4j bets on explicitness. Pick the one that matches how well you actually know your workflow in advance. To Conclude: Java Still Has a Say in Enterprise AI Python remains, without question, the home of AI research — the notebooks, the training loops, the enormous ecosystem of machine learning libraries were built there first, and will likely stay there. Nobody sensible is arguing Java should train the next large language model. But training a model is only one part of the story. The other part — the much bigger part, in terms of sheer lines of code running in the real world — is taking an already-trained model and safely wiring it into systems that already exist: a bank's core banking platform, an insurance company's policy engine, a hospital's records system. The overwhelming majority of that existing code, in most large enterprises, is written in Java and Spring, not Python. That is Java's home ground, built up over more than two decades. This is exactly the ground both Embabel and LangGraph4j are fighting on. Java also tends to run this kind of orchestration work faster than Python at execution time, which matters once you are calling these agents thousands of times a day inside a live enterprise system. And when your applications are already written in Java, keeping the AI layer in Java too — rather than routing every call out to a separate Python service — often turns out to be the simpler, safer choice. So, the real contest in enterprise AI is perhaps not "who trains the smarter model" — Python wins that one comfortably. It is "who can be trusted to make that model's decisions reliably inside a bank's core system, an insurance engine, or a hospital record system." That is precisely the kind of trust Java has spent two decades earning. With Rod Johnson effectively writing a sequel to his own Spring story, and with LangGraph4j bringing a proven Python pattern faithfully onto the JVM, Java is not sitting out this wave of AI. It has simply chosen to fight the battle it already knows how to win. This piece focused on Embabel's goal-and-planner philosophy. Embabel also has other ideas worth a separate deep-dive later — like its approach to agentic search and enterprise memory. A hands-on, step-by-step guide to setting up Embabel from scratch will follow as a companion piece. Here's the link to the source code: https://github.com/trainerpb/embabel-hello-world/tree/feature/revision.

By Soham Sengupta
How Go Maps Work: From Buckets to Swiss Tables
How Go Maps Work: From Buckets to Swiss Tables

Hey Mates! “How does a map work in Go?” is one of my favorite interview questions. It sounds simple, but it opens up a conversation about hashing, collisions, memory layout, and why two implementations with the same average O(1) lookup complexity can behave quite differently. With Go 1.24, that conversation got more interesting: the map implementation switched to Swiss Tables. Let’s walk through how the old implementation worked, what changed, and how the new design finds your keys. TL;DR Go 1.24, released in February 2025, completely replaced the internal implementation of the built-in `map` with a design based on Google's Swiss Tables. The syntax and the behavior guaranteed by the Go specification did not change, so existing programs required no migration. In microbenchmarks, some map operations became up to 60% faster. Across the Go team's full-application benchmark suite, geometric mean CPU time improved by about 1.5%. Datadog reported roughly 70% less memory for one exceptionally large map—an impressive case study, not a universal promise. Why Replace map at All? map is one of the most frequently used data structures in Go. It appears in caches, configuration, indexes, and every `map[string]any` we would rather not discuss. The old implementation served Go well for more than a decade. Hash-table research did not stop, however. At CppCon 2017, Google engineers Sam Benzaquen, Alkis Evlogimenos, Matt Kulukundis, and Roman Perepelitsa presented a new cache-friendly design. It became known as Swiss Tables, after the Google Zürich office where the team worked. In 2018, Google released an implementation in the C++ Abseil library. The design then spread across ecosystems: C++: absl::flat_hash_map in Abseil;Rust: the standard HashMap is built on hashbrown, a Swiss Tables implementation, since Rust 1.36;Go: first through third-party packages such as dolthub/swiss and cockroachdb/swiss, then in the built-in map starting with Go 1.24. The route into Go was collaborative. Community members built early prototypes. Peter Mattis of CockroachDB combined those ideas with solutions for Go-specific requirements in cockroachdb/swiss. The Go 1.24 runtime implementation is heavily based on that work. Before we get to the “Swiss” part, let us quickly review hash tables. If collisions and load factors are already familiar, skip to section 3. Hash Tables From First Principles Imagine a theater coat check where coats are retrieved by surname. You say “Smith,” and the attendant applies a simple rule to decide which section to search first, section 17, perhaps. They do not scan the entire room; they go directly to one small area and inspect a few tags. The key is the surname.The value is the coat.The hashfunction turns a key into the starting section. The same key always produces the same result within a particular map.A slot stores one key/value pair. Lookup is O(1) on average because the hash takes us to a small part of the table instead of forcing us to scan every entry. This is an average-case property, not an unconditional guarantee. The unavoidable problem is a collision: there are finitely many locations, so different keys eventually choose the same starting point. Two classic strategies handle this: Chaining. Section 17 holds a list of key/value pairs. Lookup walks that small list. Hans Peter Luhn of IBM described this approach in 1953.Open addressing. Slot 17 is occupied, so try another slot, then another, until a suitable one is found. The order of locations is called the probe sequence. Open addressing was used in 1954 and formally published in 1957. Both ideas are about 70 years old. The interesting part is how modern implementations make them friendly to modern CPUs. The Old Implementation: Go 1.23 and Earlier The old Go map was a hybrid. It used fixed-size buckets, while excess collisions were handled with chains of overflow buckets. In spirit, it was closer to chaining. A map had an hmap header pointing to an array of buckets. Each `bmap` bucket contained exactly 8 slots: Plain Text flowchart LR subgraph HMAP["hmap header"] direction TB C["count: number of entries"] B["B: log2 bucket count"] P["buckets: array pointer"] end P --> ARR["bucket array: 2^B buckets"] ARR --> BKT subgraph BKT["bmap bucket: 8 slots"] direction TB TH["tophash: 8 filter bytes"] K["8 keys"] V["8 values"] OV["overflow pointer"] end OV --> OB1["overflow bucket"] OB1 --> OB2["another overflow bucket"] Three details matter: tophash: Eight bytes at the start of the bucket, one per slot. Each byte contains the top bits of that key's hash. Comparing a byte is cheaper than comparing a full string key, so it acts as a fast filter.Keys and values are stored separately: Eight keys followed by eight values. This reduces alignment padding.Overflow buckets: When the primary bucket cannot hold another entry, the runtime allocates another bucket and links it into a chain. To find grape, the runtime roughly did this: Calculate `hash("grape")`.Use the low `B` bits to select a bucket.Check the bucket's eight `tophash` bytes one at a time.When a byte matches, compare the full key. If the key matches, return the value.If the bucket has no match, follow its pointer to the overflow bucket and repeat.If the chain ends, the key is absent. The old map grew when average occupancy exceeded 6.5 entries per 8-slot bucket, a load factor of 81.25%, or when too many overflow buckets accumulated. The number of primary buckets doubled. Crucially, growth was incremental. Go did not move the entire old array at once. Each write evacuated a little more data. One unlucky insertion therefore did not have to copy a gigabyte-sized map in a single pause. Pointer chasing. Every overflow hop reads another part of the heap and increases the risk of a CPU-cache miss. A long overflow chain can therefore be expensive.Serial metadata checks. The runtime inspected `tophash` byte by byte and slot by slot. Eight slots could mean eight loop iterations and several branches.Memory overhead. Overflow buckets and their pointers cost memory. Raising the load factor much above 81% made overflow chains more common and lookup slower. The garbage collector could also have more pointers to scan. The goal was clear: fewer pointers, better locality, and less serial work. Swiss Tables provide exactly that. Meet Swiss Tables A Swiss Table is open addressing adapted to modern CPUs. Three decisions drive the design: Split the hash into an address and a short fingerprint.Pack the metadata for a group of slots into one machine word.Compare the fingerprints of 8 slots in parallel, then compare full keys only for the candidates. A 64-bit hash is divided into two unequal pieces: Plain Text 64-bit key hash ┌─────────────────────────────────────────────┬─────────────┐ │ h1: upper 57 bits │ h2: 7 bits │ │ chooses where probing begins │ fingerprint │ └─────────────────────────────────────────────┴─────────────┘ h1 selects the initial group.h2 is a 7-bit fingerprint used as a cheap filter before a full-key comparison. In the coat-check analogy, h1 chooses a section, and h2 is a short mark on each tag that quickly rules out unrelated coats. Slots are arranged in groups of 8. Each group has a 64-bit control word, one byte per slot: Plain Text control word: 64 bits = 8 bytes ┌────┬────┬────┬────┬────┬────┬────┬────┐ │ ∅ │ 15 │ 27 │ 5C │ ∅ │ 3A │ 71 │ † │ └────┴────┴────┴────┴────┴────┴────┴────┘ │ │ │ │ └─ occupied, h2 = 0x15 └─ deleted tombstone └─ empty Each byte describes its slot: Plain Text | Byte value | Meaning | | `0b1000_0000` (`0x80`) | **empty** slot | | `0b1111_1110` (`0xFE`) | **deleted** slot, or tombstone | | `0b0xxx_xxxx` | occupied; the low seven bits contain h2 | Occupied slots always have a zero high bit, while empty and deleted slots have a one. That encoding enables efficient parallel tests. Suppose we are looking for h2 = 0x27. Conceptually, the operation looks like this: Plain Text probe word: 27 │ 27 │ 27 │ 27 │ 27 │ 27 │ 27 │ 27 == │ == │ == │ == │ == │ == │ == │ == control word: ∅ │ 15 │ 27 │ 5C │ ∅ │ 3A │ 71 │ † result: 0 │ 0 │ 1 │ 0 │ 0 │ 0 │ 0 │ 0 ↑ candidate slot 2 On amd64, Go recognizes this operation as an intrinsic and uses SIMD instructions. Other architectures have a portable SWAR implementation — SIMD Within A Register — that processes all eight bytes using arithmetic on one 64-bit word. The following simplified code is optional. The important result is a bit mask of slots whose h2 may match: Go func matchH2(ctrl uint64, h2 uint8) uint64 { broadcast := uint64(h2) * 0x0101010101010101 x := ctrl ^ broadcast return (x - 0x0101010101010101) &^ x & 0x8080808080808080 } There are two reasons a candidate is not yet a result: h2 has only seven bits, so an occupied slot has a 1/128 chance of sharing the same fingerprint by accident;The portable bit trick can produce a rare extra candidate because of subtraction borrow. Neither affects correctness. Go always performs a full-key comparison before returning a value. If the initial group has no matching key, open addressing continues at group granularity. Go uses a quadratic probe sequence. Three rules are enough to understand lookup: Matching h2 bytes produce candidate slots whose full keys are checked;If the group contains an empty slot, stop — the key is absent;If the group has no match and no empty slot, visit the next group in the probe sequence. One probe handles the metadata for eight slots, so densely populated tables can still be searched efficiently. Because group metadata is cheaper to inspect, Swiss Tables can remain more densely populated. Go allows an average load up to 7/8 = 87.5%, compared with 81.25% in the old implementation. More useful entries in the same backing storage usually means less memory per key. Go's Extension: A Directory of Tables Go could not simply port Abseil's implementation. Two language and runtime requirements needed additional design work. A conventional Swiss Table grows all at once: allocate a larger array and move every entry. For a gigabyte-sized map, the insertion that triggers growth would suffer a noticeable pause. Go is widely used for latency-sensitive services, and its old maps already bounded growth work per insertion. A large Go map is a directory of independent Swiss Tables, not one unbounded table. Each table covers part of the hash space and has a maximum capacity of 1024 slots, or 128 groups: Plain Text flowchart TD M["map[K]V"] --> DIR["directory: array of table pointers"] DIR --> T0["table 0: up to 1024 slots"] DIR --> T1["table 1: up to 1024 slots"] DIR --> TN["table N"] T1 --> G0["group 0: control word + 8 slots"] T1 --> G1["group 1: control word + 8 slots"] T1 --> GK["up to 128 groups"] A variable number of upper hash bits selects the table. This technique is a form of extendible hashing. Growth is local. When a table below the limit fills, only that table grows. Once a table reaches the limit, it splits into two. The amount of growth work caused by one insertion is therefore bounded by one table of at most 1024 slots: roughly at most 896 live entries at the 7/8 load threshold (not the entire map). Small maps get a special fast path. A map that starts small and never exceeds 8 entries lives directly in a single group with no table directory. A map that has already grown, or was created with a large `hint`, does not shrink back into this representation after deletions. Unlike many hash tables, Go explicitly permits modifying a map during iteration: an entry deleted before it is reached must not be produced;an entry updated before it is reached must produce its latest value;a newly inserted entry may or may not be produced. Growth reshuffles storage, so an iterator cannot simply walk the current array. Go's iterator retains the old table to determine traversal order. Before returning an entry, it consults the current table to confirm that the key still exists and to obtain the latest value. According to the Go team, iteration is the most complex part of the implementation. Deletion and Tombstones With open addressing, deleting an entry cannot always turn its slot into an ordinary empty slot. Lookup stops at an empty slot. Creating one in the middle of another key's probe sequence could therefore make a live key unreachable. A tombstone, encoded as deleted (`0xFE`), means: “an entry used to be here; continue probing.” Insertions may reuse tombstones. Go avoids a tombstone when the group already contains another empty slot. Any lookup reaching that group would stop there anyway, so turning the deleted slot into empty cannot break a probe sequence. Accumulated tombstones disappear during a later grow or split. Live entries are copied into new groups and deleted slots are not. Go 1.24 does not implement a separate same-size grow for this purpose. And Performance Numbers According to the Go team and early production reports: Microbenchmarks: some map operations are up to 60% faster than in Go 1.23. Results vary widely, and a few edge cases regress.Full-application benchmarks: the Go team's suite showed about 1.5% geometric-mean CPU-time improvement for the whole application. A specific program may behave differently.Memory: Datadog measured roughly 70% less map memory for one very large map with about 3.5 million entries in a high-traffic environment. This favorable case benefited from denser storage, no overflow buckets, and growth that did not retain two enormous bucket arrays at once. Savings were much smaller in another environment. Large maps often benefit more because overflow chains and cache misses hurt the old design more strongly. The exact result depends on key and value sizes and on the mix of reads, writes, deletes, hits, and misses. Measure your own workload. Takeaways The Go Swiss Tables story is a good example of changing a foundational component carefully: Start with a proven design already used by Abseil and Rust's `hashbrown`.Adapt it to Go's invariants: bounded growth latency and modification during iteration.Preserve the public API and specified behavior. Hash tables are seven decades old, yet a better match between data layout and modern hardware can still remove CPU time and memory across a large ecosystem. “Solved” problems often have room for another good engineering pass. References [Faster Go maps with Swiss Tables] — the primary Go team article[Go 1.24 Release Notes] — runtime changes and `GOEXPERIMENT=noswissmap`[Go 1.24 `internal/runtime/maps` source] — original source[Go 1.26 runtime source] — the release in which the old map implementation was removed[Abseil Swiss Tables Design Notes] — original sourceMatt Kulukundis at CppCon 2017Datadog's workload-specific memory case study[`hashbrown`] — the Swiss Tables implementation behind Rust's standard `HashMap`

By Ilia Ivankin

The Latest Coding Topics

article thumbnail
Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A
Multi-agent systems are common. Here we build a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol.
October 9, 2026
by Purnanga Borah
· 351 Views · 1 Like
article thumbnail
Why Your GPU Fleet Is Both Full and Idle: Building a Capacity Orchestration Layer
Managing GPU capacity manually produces fleets that are simultaneously fully allocated and poorly utilized. A capacity orchestration layer can fix the problem.
October 9, 2026
by Ankit Sinha
· 365 Views · 1 Like
article thumbnail
Supercharging AI Agents with Azure Context: A Hands-On Guide to Azure MCP
Here’s a step-by-step guide for cloud engineers on how to bridge local AI assistants with live Azure infrastructure using the Model Context Protocol.
October 8, 2026
by Ammar Ekbote
· 556 Views · 1 Like
article thumbnail
AI Agents Leaked 13,000 Screenshots: Why Enterprise Approval Controls Failed
A reported leak of 13,000 screenshots shows how AI agents can bypass weak approval and audit controls even when organizations have written policies.
October 7, 2026
by Tim Freestone
· 752 Views · 1 Like
article thumbnail
Decoding the “Black Box”: Evaluating Agent Tool Chains in Production
Evaluating AI agents means checking the full trajectory, not just the answer — intent, tool calls, order, and accuracy — using Microsoft Foundry's evaluation tools.
October 7, 2026
by Gaurav Bhardwaj
· 674 Views · 1 Like
article thumbnail
Building High-Performance Time-Series Applications With Java and QuestDB
QuestDB combines fast time-series ingestion with a familiar SQL model. With Eclipse JNoSQL 1.1.18, Java developers can use it through TimeSeriesTemplate and Jakarta Data.
October 7, 2026
by Otavio Santana DZone Core CORE
· 773 Views · 1 Like
article thumbnail
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
This guide walks through custom Azure ML model training and deployment to a Managed Online Endpoint and connecting it to a Microsoft Foundry agent as a function tool.
October 6, 2026
by Jubin Soni, FBCS DZone Core CORE
· 4,968 Views · 3 Likes
article thumbnail
Building IoT Time-Series Applications With Java and Apache IoTDB
Apache IoTDB targets device-oriented time-series workloads. Eclipse JNoSQL 1.1.18 lets Java devs use it with TimeSeriesTemplate and Jakarta Data repositories.
October 6, 2026
by Otavio Santana DZone Core CORE
· 1,005 Views · 3 Likes
article thumbnail
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices
Stop dual-write data inconsistencies. Learn how to architect the Transactional Outbox Pattern using Java, Spring Boot, and PostgreSQL for reliable Kafka events.
October 6, 2026
by Rahul Tewari
· 1,407 Views · 1 Like
article thumbnail
OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule
The story has two parts. In Act 1, we trace elevated latency back to a JVM heap, showing how anonymous memory can still end up in swap even with vm.swappiness=1.
October 6, 2026
by Maxim Muzafarov
· 1,023 Views · 1 Like
article thumbnail
Grounding AI Agents in Governed Data
Stop trusting prompts alone to protect PII. Put AI agents behind the same governed semantic layer as your human analysts.
October 6, 2026
by Jeevan reddy Geereddy
· 830 Views · 1 Like
article thumbnail
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI
XML, JSON, and YAML compared: history, trade-offs, and where each wins, plus why JSON Schema is now the contract layer for AI agents.
October 5, 2026
by Kai Wähner DZone Core CORE
· 778 Views · 2 Likes
article thumbnail
Building Time-Series Applications With Java and InfluxDB
InfluxDB brings time-oriented storage to Java applications, while Eclipse JNoSQL 1.1.18 simplifies integration through TimeSeriesTemplate and Jakarta Data repositories.
October 5, 2026
by Otavio Santana DZone Core CORE
· 742 Views · 1 Like
article thumbnail
Wasm Inside Neo4j: Building the Example That Didn't Exist
A Rust VADER sentiment analyzer compiled to WebAssembly, embedded inside a Neo4j Java UDF, and callable directly from Cypher using wasmtime-java.
October 2, 2026
by Akmal Chaudhri DZone Core CORE
· 1,183 Views · 2 Likes
article thumbnail
Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus
Add W3C Trace Context propagation across Goose, agentgateway, and Quarkus to turn opaque agentic tool loops into fully observable distributed traces in Jaeger.
October 2, 2026
by Daniel Oh DZone Core CORE
· 1,013 Views · 2 Likes
article thumbnail
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
In this article, we will discuss how to run your coding agents in the cloud using sbx. Cloud compute is usage-billed, so keep track of your sandboxes accordingly.
October 2, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 1,257 Views · 1 Like
article thumbnail
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
The architectural divide between state-machine rigidity and agentic flexibility in financial systems, comparing stateful, multi-agent workflows.
October 1, 2026
by Soham Sengupta
· 1,161 Views · 1 Like
article thumbnail
How Go Maps Work: From Buckets to Swiss Tables
This article is for Go developers who write `m := make(map[string]int)` every day but have never looked under the hood. We will use a coat-check analogy.
October 1, 2026
by Ilia Ivankin
· 943 Views · 1 Like
article thumbnail
The Silent Container Death: A TCP Dial That Never Times Out
A pod goes into CrashLoopBackOff. You pull the logs expecting a stack trace, a panic, an error string — and then nothing. No error. No exit message. Magic.
September 30, 2026
by Alexander Fo
· 1,658 Views · 1 Like
article thumbnail
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
Learn about graph databases by building an F1 teammate network from real Formula 1 data and using Cypher to connect Max Verstappen to Juan Manuel Fangio.
September 30, 2026
by Jeremy Morgan
· 1,085 Views · 2 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×