DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Tools

Development and programming tools are used to build frameworks, and they can be used for creating, debugging, and maintaining programs — and much more. The resources in this Zone cover topics such as compilers, database management systems, code editors, and other software tools and can help ensure engineers are writing clean code.

icon
Latest Premium Content
Trend Report
Kubernetes in the Enterprise
Kubernetes in the Enterprise
Refcard #366
Advanced Jenkins
Advanced Jenkins
Refcard #378
Apache Kafka Patterns and Anti-Patterns
Apache Kafka Patterns and Anti-Patterns

DZone's Featured Tools Resources

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

By Naga Santhosh Reddy Vootukuri DZone Core CORE
In my previous article, I walked through running coding agents inside Docker Sandboxes on a local machine. We installed the sbx CLI, started with a small project, and covered the commands needed to run, stop, and remove a sandbox. This time, I want to take that same workflow off the laptop. Docker added cloud sandboxes in version 0.42.0. You can now use sbx --cloud to run an agent on Docker-managed infrastructure instead of using your machine for the sandbox’s compute. The command is simple to use. The part that is worth understanding is how you get your code into that environment, work with the agent, and bring the changes back locally. That is what we will do here. Nothing complicated; we will start with a small Python project, one coding task, and a cloud sandbox. We will remove the sandbox when we are done with the work. What Changes With a Cloud Sandbox? The sbx CLI still runs in your terminal. With --cloud, supported commands target Docker’s cloud service rather than your local sandbox environment. For example: PowerShell sbx ls Lists your local sandboxes. PowerShell sbx --cloud ls Lists your cloud sandboxes. That distinction matters throughout this walkthrough. If you forget --cloud, you are not asking about the same environment. Cloud sandboxes also have separate credentials and network policies. Do not assume that an agent login or network policy you configured locally is already available in the cloud. For this example, we will copy individual files explicitly. That keeps it easy to see what we send to the sandbox and what we bring back. Before You Start You will need: An updated sbx CLI with cloud support, introduced in version 0.42.0.A Docker account with an active Docker Agentic Platform plan for cloud compute.Authentication for the coding agent you want to use. This walkthrough uses Claude.Python 3 available in the sandbox image for the example. Note: The free sbx CLI does not mean cloud compute is free. Docker bills cloud compute based on usage, and your model provider bills inference separately. Check your account’s pricing before starting. Also, use a small sample project first. Running remotely means sending code off your machine. For company repositories, make sure that is allowed before uploading anything. The host-side commands below use PowerShell. Paths inside the cloud sandbox use Linux-style paths. Step 1: Sign In and Configure the Agent First, check your installed version: PowerShell sbx version If you are still using an older version from the previous walkthrough, update it before continuing. Sign in to Docker: PowerShell sbx login For Claude, Docker documents a cloud OAuth flow: PowerShell sbx --cloud secret set anthropic --oauth Complete the provider sign-in with an account that has the required access. Notice the --cloud flag here, too. These credentials are stored for cloud use, separately from your local sandbox credentials. There is no reason to put a token in our Python files or paste it into an agent prompt. Step 2: Create a Small Project Let us give the agent something specific to fix. Create a project folder: PowerShell New-Item -ItemType Directory -Path .\cloud-sandbox-demo Set-Location .\cloud-sandbox-demo Inside it, create a file named slug.py: Python def make_slug(text): return text.lower().replace(" ", "-") This converts "Docker Sandboxes" into "docker-sandboxes". It works for that input, but it does not handle whitespace very well. Leading spaces become leading hyphens. Repeated spaces become repeated hyphens. Tabs are not handled at all. Now create test_slug.py: Python import unittest from slug import make_slug class SlugTests(unittest.TestCase): def test_two_words(self): self.assertEqual(make_slug("Docker Sandboxes"),"docker-sandboxes") if __name__ == "__main__": unittest.main() We have one passing case and a clear improvement to make. The point is not that this function needs cloud compute. It is small enough that we can focus on the sandbox workflow without spending half the article explaining an application. Step 3: Start a Cloud Sandbox Run the following command: PowerShell sbx --cloud run --detached --name cloud-demo --ttl 1h claude This creates a cloud sandbox and starts the agent without attaching your terminal to it. The flags in the above command are for doing useful things: --cloud selects the cloud environment.--detached returns control to your terminal.--name cloud-demo gives the sandbox a recognizable name.--ttl 1h requests a one-hour lifetime. Important: The documented default action when the TTL expires is deletion. Treat this as a disposable environment, and copy your work out before the deadline. The command prints a sandbox ID. You can also find it with: PowerShell sbx --cloud ls Copy that ID into a PowerShell variable: PowerShell $sandbox = "PASTE_YOUR_SANDBOX_ID_HERE" Use the real ID returned by Docker, not the placeholder above. Keep using this terminal for the remaining commands. One detail to remember is a detached cloud run creates a new sandbox. It is not the command to run repeatedly when you want to reconnect to the same one. Step 4: Copy the Project Into the Sandbox Create a directory for our example: PowerShell sbx --cloud exec $sandbox mkdir -p /workspace/demo The mkdir command runs inside the Linux sandbox, not on Windows. Now copy the two files: PowerShell sbx --cloud cp .\slug.py "${sandbox}:/workspace/demo/slug.py" sbx --cloud cp .\test_slug.py "${sandbox}:/workspace/demo/test_slug.py" The ${sandbox} syntax is intentional. In PowerShell, it separates the variable name from the colon used in Docker’s SANDBOX:PATH format. This is also why I am copying individual files rather than uploading the entire folder. We do not need a virtual environment, local configuration, or an accidentally included .env file for this task. Run the existing test inside the sandbox: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v If your selected image does not include Python 3, add it inside the sandbox before continuing. The existing test only covers two words separated by one space. Passing it does not mean the whitespace handling is correct yet. Step 5: Give the Agent a Narrow Task Attach to the running cloud sandbox: PowerShell sbx --cloud attach $sandbox Now give Claude a concrete task: Plain Text Work on the Python project in /workspace/demo. Update make_slug so that: - The output remains lowercase. - Leading and trailing whitespace is removed. - Consecutive whitespace becomes a single hyphen. - Spaces, tabs, and newlines are handled consistently. - Empty input returns an empty string. Add unit tests for these cases using unittest. Keep the existing test. Do not add third-party dependencies or modify files outside this project. Run the tests and summarize which files you changed. This is much more useful than asking the agent to “improve the project.” We have told it what the function should do, which edge cases matter, and how much freedom it has. There is no reason for it to introduce a framework or reorganize the project. The prompt is task guidance, though — not a security policy. File access, network access, and credentials still need the appropriate sandbox controls. Once the agent finishes, use Ctrl + backslash to detach and return to your local terminal. Detaching does not stop the cloud sandbox. Step 6: Run the Tests and Bring the Changes Back Run the test command again from your terminal: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v This executes inside the cloud sandbox. It is not running against your original local files. For this task, a straightforward implementation could look like: Python def make_slug(text): return "-".join(text.lower().split()) Calling split() without a separator handles consecutive whitespace and removes leading and trailing whitespace. Joining those words with a hyphen gives us the requested behavior. The agent may arrive at a different implementation. Read it rather than assuming that passing tests makes every change worth keeping. Create a separate folder for the returned files: PowerShell New-Item -ItemType Directory -Path .\review Copy the modified files into it: PowerShell sbx --cloud cp "${sandbox}:/workspace/demo/slug.py" .\review\slug.py sbx --cloud cp "${sandbox}:/workspace/demo/test_slug.py" .\review\test_slug.py Your original files are still untouched. If you have Git installed, compare the versions: PowerShell git diff --no-index -- .\slug.py .\review\slug.py git diff --no-index -- .\test_slug.py .\review\test_slug.py You can also compare them in your editor. Look at the tests as closely as the implementation. Did the agent actually add cases for tabs and newlines? Did it keep the original test? Did it add anything unrelated? For a real repository, I would bring the changes into a working branch and use the normal review process. The sandbox changes where the agent works. It does not replace code review. What About Web Applications? Our Python example does not start a server. If you use a web project instead, cloud sandboxes can expose an application through a public HTTPS URL. For an application already listening on sandbox port 3000: PowerShell sbx --cloud ports $sandbox --publish 3000 sbx --cloud ports $sandbox Use the URL returned by Docker. This is different from publishing a local port such as localhost:3000. In cloud mode, the command accepts the sandbox port, and Docker assigns the public URL. Note: Publicly reachable is not the same as private. Do not expose an unauthenticated admin page, secrets, or sensitive test data. Remove the exposure when you no longer need it: PowerShell sbx --cloud ports $sandbox --unpublish 3000 Step 7: Clean Up the Cloud Sandbox Before cleanup, make sure the files you want to keep are on your machine. If you want to pause rather than delete, Docker documents cloud stop as preserving the sandbox’s memory and disk state: PowerShell sbx --cloud stop $sandbox Do not assume that preserved resources have no cost. Check your plan’s billing terms. For this small exercise, we have already copied the results out, so we can remove the sandbox: PowerShell sbx --cloud rm $sandbox Confirm the removal when prompted, then list your cloud sandboxes: PowerShell sbx --cloud ls There is an important difference from my earlier article: sbx --cloud rm --all is intentionally disabled. Cloud cleanup requires explicit sandbox identifiers. That is a useful safeguard. A cloud credential may have access to more than the one environment you were experimenting with. A Few Things That Can Slow You Down If the agent cannot authenticate, check its cloud credentials. A successful local session does not prove that cloud authentication is configured. If it cannot reach a service, check the cloud network policy. Do not immediately open access to everything just to make an error disappear. If your local files have not changed, remember the workflow we used: we copied files into the cloud and copied the results back. Those copies are not a live synchronization mechanism. And if you are coming back to a running sandbox, use attach. Repeating the detached creation command gives you another sandbox, not another connection to the original one. Conclusion What I like about this addition is that it keeps the workflow familiar. We are still using sbx, still giving the agent a specific project, and still deciding what work to keep. The difference is where that work happens. Start small. Send only the files the agent needs, give it one clear task, and bring the results back into your normal development process. Once that feels comfortable, move on to a larger repository or a task that actually benefits from remote compute. And copy the changes back before the sandbox expires. A useful fix is not very useful if the only copy disappears with the environment. More
Wasm Inside Neo4j: Building the Example That Didn't Exist

Wasm Inside Neo4j: Building the Example That Didn't Exist

By Akmal Chaudhri DZone Core CORE
In a recent DZone article, Running Sentiment Analysis Inside Neo4j With a Java Plugin, we explored several approaches to running sentiment analysis inside the Neo4j database engine. One of those approaches — embedding a Wasm runtime inside a Java UDF — was described like this: Theoretically, we could embed a Wasm runtime such as wasmtime inside a Java UDF and execute the VADER Wasm module from within Neo4j, getting Wasm's sandbox guarantees inside Neo4j's plugin model. It's technically feasible but no published working example appears to exist and the complexity cost is high relative to the alternatives. An interesting idea to watch, but not practical today. This article builds that working example. We'll show how to embed a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity score map callable directly from Cypher. We'll cover the tools and inspection techniques needed to understand what the Wasm compiler generates and why the Java calling convention looks the way it does. The full source code is available on GitHub. What We're Building We're embedding a wasmtime Wasm runtime inside a Neo4j Java UDF using wasmtime-java, a community JNI binding for the Wasmtime runtime. It's not an official Bytecode Alliance product, but it ships prebuilt native libraries for all major platforms and is sufficient for this proof-of-concept. A Rust function compiled to WebAssembly rides inside the plugin JAR alongside the Java code. When Cypher calls the UDF, Java initializes the Wasm runtime, loads the binary, and invokes the Rust function — all inside the Neo4j JVM process with no external API calls and no network round-trips. Note: This article was tested specifically against wasmtime-java 0.19.0. The API used here is version-specific; newer releases or alternative JVM Wasm runtimes may expose different interfaces and calling conventions. Prerequisites You'll need the following installed if you wish to follow along. We're using Apple Silicon (ARM64) as our development platform, so we'll note where the setup differs from other platforms. Java We're using OpenJDK 21 (tested with 21.0.12.1). Install it using your platform's package manager or download it directly from adoptium.net. On macOS via Homebrew: Shell brew install openjdk@21 On Ubuntu/Debian: Shell sudo apt install openjdk-21-jdk On Windows, download and run the installer from Adoptium. Confirm your Java version: Shell java -version You should see a Java 21 runtime. If you're on Apple Silicon, also confirm you're running a native ARM64 JVM with: Shell uname -m You should see arm64. Not running under ARM64 will likely break the wasmtime-java JNI library loading. Maven We're using Maven 3.9.6. On Apple Silicon, be cautious about installing Maven via Homebrew as, at the time of writing, the Homebrew Maven formula pulls in OpenJDK 26 as a dependency, which conflicts with a Java 21 installation. If your package manager installs an incompatible JDK alongside Maven, verify the runtime with mvn -version and configure JAVA_HOME as necessary. Installing Maven manually is the safest approach: Shell cd ~ curl -O https://archive.apache.org/dist/maven/maven-3/3.9.6/binaries/apache-maven-3.9.6-bin.tar.gz tar xzf apache-maven-3.9.6-bin.tar.gz Then add Maven to your PATH and make it persist across terminal sessions: Shell echo 'export PATH="$HOME/apache-maven-3.9.6/bin:$PATH"' >> ~/.zshrc source ~/.zshrc On Linux, add the same line to ~/.bashrc instead.On Windows, download the zip from maven.apache.org and add the bin folder to your system PATH via System Properties. Confirm Maven is using Java 21: Shell mvn -version You should see Java version: 21 in the output. Rust We're using Rust 1.96.0. Install via rustup if not already present: Shell curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh On Windows, download and run rustup-init.exe from rustup.rs. To pin to the specific Rust version we tested with: Shell rustup toolchain install 1.96.0 rustup default 1.96.0 Then add the WASI target: Shell rustup target add wasm32-wasip1 This target works identically across macOS, Linux, and Windows. WABT The WebAssembly Binary Toolkit gives us wasm-objdump for inspecting Wasm binaries. We tested with version 1.0.41. On macOS: Shell brew install wabt On Ubuntu/Debian: Shell sudo apt install wabt On Windows, download the latest release from github.com/WebAssembly/wabt/releases. wit-bindgen This is the interface types generator for WebAssembly. There are two distinct version numbers to be aware of: the wit-bindgen-cli command-line tool and the wit-bindgen Rust crate used as a dependency inside the Wasm module. These can differ. We tested with CLI version 0.59.0 and Rust crate version 0.40.0 (specified in Cargo.toml). The generated binary identifies the crate version through the export name cabi_realloc_wit_bindgen_0_40_0. Install the pinned CLI version via Cargo on all platforms: Shell cargo install wit-bindgen-cli --version 0.59.0 Confirm it's installed: Shell wit-bindgen --version Neo4j Desktop We're using Neo4j Desktop with a local database instance. Download from Neo4j for Desktop. The pom.xml in this article is pinned to Neo4j 2026.07.0 — update the neo4j.version property to match your own Desktop installation. wasmtime-java Platform Support The wasmtime-java library ships prebuilt JNI native libraries for: macOS aarch64macOS x86_64Linux aarch64Linux x86_64Windows x86_64 No additional setup is needed, as Maven pulls the correct native library for your platform automatically. Version Summary For reference, here are all the component versions used in this article: ComponentVersionOpenJDK21.0.12.1Maven3.9.6Rust1.96.0WABT1.0.41wit-bindgen CLI0.59.0wit-bindgen crate0.40.0vader_sentiment crate0.1.1wasmtime-java0.19.0Neo4j2026.07.0 Getting the Code Clone the repository before following along. All source files are provided so you don't need to create them manually. Shell cd ~ git clone --filter=blob:none --sparse https://github.com/VeryFatBoy/neo4j.git cd neo4j git sparse-checkout set wasm-udf mv wasm-udf ../wasm-udf cd ../wasm-udf Project Structure Before creating any files, here's the final layout we're building toward. There are two separate projects: A Rust crate that compiles to Wasm.A Maven project that hosts the Neo4j UDF. First, the Rust crate: Plain Text sentimentable/ ├── Cargo.toml ├── src/ │ └── lib.rs └── wit/ └── sentimentable.wit Second, the Maven project: Plain Text neo4j-wasm-udf/ ├── pom.xml └── src/ └── main/ ├── java/ │ └── com/example/ │ ├── WasmUDF.java │ └── SentimentUDF.java └── resources/ ├── add.wat ├── add.wasm └── sentimentable.wasm The Wasm binaries in resources/ are bundled into the plugin JAR at build time. The Rust crate and Maven project are kept separate, and the Wasm binary is the handoff point between them. The project layout is also shown in Figure 1. Figure 1. Two-Project Layout How the Wasm Plumbing Works Before diving into the code, it's worth understanding the three layers that make this possible. Core Wasm and WASI WebAssembly defines a portable binary format and a stack-based execution model. On its own, it only understands numbers, such as integers and floats. When a Wasm module needs system capabilities, like memory allocation or I/O, it uses WASI (WebAssembly System Interface), a standardized set of system calls that a host runtime implements. Our Rust code targets wasm32-wasip1, which means it compiles to Wasm with WASI preview 1 system calls. The wasmtime runtime implements those calls on the host side. wasmtime-java This library wraps the wasmtime Wasm runtime in a JNI binding, making it callable from Java. It ships prebuilt native libraries for all major platforms, so adding it as a Maven dependency is all that's needed — no separate wasmtime installation required. The Java API lets us load a Wasm binary, set up a WASI context, and call exported functions directly. wit-bindgen and the String ABI Core WebAssembly functions operate on Wasm value types such as integers and floats. WIT (WebAssembly Interface Types) and the Component Model provide higher-level interface types such as strings, tuples, and records; wit-bindgen generates the lowering and lifting code needed to represent those types at the Wasm boundary. For strings, it uses a pointer-and-length convention: the caller allocates memory inside the Wasm module using a generated cabi_realloc function, writes the string bytes there and passes the memory address and byte length as two integers. The Rust code reads the string from that address. For return values, the lowering strategy depends on the type, which we'll see when we inspect the generated binary. With those three pieces in place, the calling chain looks like this: Plain Text Cypher query -> Neo4j routes to @UserFunction -> Java initializes wasmtime engine + WASI context -> Java allocates string in Wasm memory -> Java calls exported Wasm function -> Rust executes VADER scoring -> Java reads result from Wasm memory -> Java returns Map<String, Double> to Neo4j -> Neo4j returns result to Cypher Graphically, the calling chain is also shown in Figure 2. Figure 2. Calling Chain We built up to this through two simpler stepping-stone cases: Case 1: A trivial integer addition to prove the chain works.Case 2: A single compound score to introduce string passing and WASI. The full walkthrough of both, including the WasmUDF.java implementation is in a technical report on the GitHub repo. The Maven Project XML <project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd"> <modelVersion>4.0.0</modelVersion> <groupId>com.example</groupId> <artifactId>neo4j-wasm-udf</artifactId> <version>1.0-SNAPSHOT</version> <packaging>jar</packaging> <properties> <maven.compiler.source>21</maven.compiler.source> <maven.compiler.target>21</maven.compiler.target> <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding> <neo4j.version>2026.07.0</neo4j.version> </properties> <dependencies> <dependency> <groupId>org.neo4j</groupId> <artifactId>neo4j</artifactId> <version>${neo4j.version}</version> <scope>provided</scope> </dependency> <dependency> <groupId>io.github.kawamuray.wasmtime</groupId> <artifactId>wasmtime-java</artifactId> <version>0.19.0</version> </dependency> </dependencies> <build> <plugins> <plugin> <artifactId>maven-compiler-plugin</artifactId> <configuration> <source>21</source> <target>21</target> </configuration> </plugin> <plugin> <groupId>org.apache.maven.plugins</groupId> <artifactId>maven-shade-plugin</artifactId> <version>3.5.1</version> <executions> <execution> <phase>package</phase> <goals><goal>shade</goal></goals> <configuration> <artifactSet> <excludes> <exclude>org.neo4j:*</exclude> </excludes> </artifactSet> <shadedArtifactAttached>false</shadedArtifactAttached> </configuration> </execution> </executions> </plugin> </plugins> </build> </project> Two things worth noting here: org.neo4j:neo4j is declared as provided scope — Neo4j is already present in the database JVM at runtime, so we exclude it from the bundled JAR.We use maven-shade-plugin rather than maven-jar-plugin to produce a fat JAR that bundles wasmtime-java and its native libraries alongside our code. Update the neo4j.version property to match your own Neo4j Desktop installation. Case 3: Full Polarity Map VADER produces four scores: compound, positive, negative, and neutral. In this case, we update the Rust function to return all four and the Java UDF to return them as a Map<String, Double> — matching the return shape of the Java VADER UDF from the previous article. The sentimentable.wit File We change the return type from a single f32 to a tuple of four f32 values: Plain Text package local:sentimentable; world sentimentable { export sentimentable: func(input: string) -> tuple<f32, f32, f32, f32>; } We use a tuple rather than a named record. Both would work, but a tuple is simpler on the Java side — we read four consecutive f32 values from memory at known offsets without needing to decode field names. The lib.rs File Rust wit_bindgen::generate!({ world: "sentimentable", }); struct Component; impl Guest for Component { fn sentimentable(input: String) -> (f32, f32, f32, f32) { lazy_static::lazy_static! { static ref ANALYZER: vader_sentiment::SentimentIntensityAnalyzer<'static> = vader_sentiment::SentimentIntensityAnalyzer::new(); } let scores = ANALYZER.polarity_scores(input.as_str()); ( *scores.get("compound").unwrap_or(&0.0) as f32, *scores.get("pos").unwrap_or(&0.0) as f32, *scores.get("neg").unwrap_or(&0.0) as f32, *scores.get("neu").unwrap_or(&0.0) as f32, ) } } export!(Component); Build: Shell cd ~/wasm-udf/sentimentable cargo build --target wasm32-wasip1 --release Inspecting the Binary Let's check the exports first: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "^Export" -A 6 Four exports should be present: memory, sentimentable, cabi_realloc, and cabi_realloc_wit_bindgen_0_40_0. Now let's find the type signature of the sentimentable function. Find the sig index: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "func\[9\]" | head -1 Then look it up: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "type\[9\]" You should see: Plain Text - type[9] (i32, i32) -> i32 The signature is (i32, i32) -> i32 . This is the key difference between returning a single scalar and returning a tuple: wit-bindgen uses a direct f32 return for a single value, but switches to an indirect result pointer when returning a tuple. What's written at that pointer is four f32 values (16 bytes) at consecutive 4-byte offsets. The Java side reads all four. This illustrates an important distinction between the WIT interface definition and the generated core Wasm ABI. The WIT signature and the Wasm-level signature are different layers: wit-bindgen lowers WIT types to a core Wasm ABI, and the lowering strategy depends on the return type. A single scalar such as f32 is returned directly as a Wasm value. A tuple is returned indirectly through linear memory, with the caller receiving a pointer to where the values were written. The Java calling code must match the generated ABI rather than the WIT definition, which is why inspecting the binary with wasm-objdump before writing the Java wrapper is essential. Figure 3 shows the memory layout. Figure 3. Memory Layout. Figure 4 compares Cases 2 and 3. Figure 4. Case 2 vs. Case 3 ABI Comparison The SentimentUDF.java file Java package com.example; import io.github.kawamuray.wasmtime.Engine; import io.github.kawamuray.wasmtime.Func; import io.github.kawamuray.wasmtime.Linker; import io.github.kawamuray.wasmtime.Memory; import io.github.kawamuray.wasmtime.Module; import io.github.kawamuray.wasmtime.Store; import io.github.kawamuray.wasmtime.WasmFunctions; import io.github.kawamuray.wasmtime.WasmValType; import io.github.kawamuray.wasmtime.wasi.WasiCtx; import io.github.kawamuray.wasmtime.wasi.WasiCtxBuilder; import org.neo4j.procedure.Description; import org.neo4j.procedure.Name; import org.neo4j.procedure.UserFunction; import java.io.InputStream; import java.nio.ByteBuffer; import java.nio.ByteOrder; import java.nio.charset.StandardCharsets; import java.util.HashMap; import java.util.Map; public class SentimentUDF { @UserFunction("com.example.wasm.sentiment") @Description("Scores text using VADER sentiment analysis compiled to Wasm. Returns compound, positive, negative, neutral.") public Map<String, Double> sentiment(@Name("text") String text) throws Exception { if (text == null || text.isBlank()) { return Map.of("compound", 0.0, "positive", 0.0, "negative", 0.0, "neutral", 1.0); } byte[] wasmBytes; try (InputStream is = SentimentUDF.class.getResourceAsStream("/sentimentable.wasm")) { if (is == null) throw new RuntimeException("sentimentable.wasm not found in resources"); wasmBytes = is.readAllBytes(); } WasiCtx wasi = new WasiCtxBuilder().inheritStdout().inheritStderr().build(); try (Store<Void> store = Store.withoutData(wasi); Engine engine = store.engine(); Module module = Module.fromBinary(engine, wasmBytes); Linker linker = new Linker(engine)) { WasiCtx.addToLinker(linker); linker.module(store, "", module); Memory memory = linker.get(store, "", "memory").get().memory(); Func reallocFn = linker.get(store, "", "cabi_realloc").get().func(); WasmFunctions.Function4<Integer, Integer, Integer, Integer, Integer> realloc = WasmFunctions.func(store, reallocFn, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32); byte[] inputBytes = text.getBytes(StandardCharsets.UTF_8); int len = inputBytes.length; int strPtr = realloc.call(0, 0, 1, len); ByteBuffer buf = memory.buffer(store); buf.position(strPtr); buf.put(inputBytes); Func sentimentFn = linker.get(store, "", "sentimentable").get().func(); WasmFunctions.Function2<Integer, Integer, Integer> scoreFn = WasmFunctions.func(store, sentimentFn, WasmValType.I32, WasmValType.I32, WasmValType.I32); int resultPtr = scoreFn.call(strPtr, len); // read four f32 values at 4-byte offsets: compound, pos, neg, neu ByteBuffer resultBuf = memory.buffer(store); resultBuf.order(ByteOrder.LITTLE_ENDIAN); float compound = resultBuf.getFloat(resultPtr); float positive = resultBuf.getFloat(resultPtr + 4); float negative = resultBuf.getFloat(resultPtr + 8); float neutral = resultBuf.getFloat(resultPtr + 12); Map<String, Double> result = new HashMap<>(); result.put("compound", (double) compound); result.put("positive", (double) positive); result.put("negative", (double) negative); result.put("neutral", (double) neutral); return result; } } } The return type is (Map<String, Double> ), the null guard returning a neutral map and the four getFloat() reads at consecutive 4-byte offsets from the result pointer. Build and Deploy Copy the Wasm binary, build and deploy: Shell cp ~/wasm-udf/sentimentable/target/wasm32-wasip1/release/sentimentable.wasm \ ~/wasm-udf/neo4j-wasm-udf/src/main/resources/ cd ~/wasm-udf/neo4j-wasm-udf mvn -q clean package cp target/neo4j-wasm-udf-1.0-SNAPSHOT.jar \ ~/Library/Application\ Support/neo4j-desktop/Application/Data/dbmss/<your-dbms-id>/plugins/ Stop Neo4j, restart it, and run the verification queries. Positive sentence: Cypher RETURN com.example.wasm.sentiment('The movie was great') AS scores; Result: JSON { "compound": 0.624893307685852, "positive": 0.577464759349823, "negative": 0.0, "neutral": 0.4225352108478546 } Capitalization test: Cypher RETURN com.example.wasm.sentiment('The movie was GREAT!') AS scores; Result: JSON { "compound": 0.7290259003639221, "positive": 0.6307692527770996, "negative": 0.0, "neutral": 0.3692307770252228 } Empty string guard: Cypher RETURN com.example.wasm.sentiment('') AS scores; Result: JSON { "compound": 0.0, "positive": 0.0, "negative": 0.0, "neutral": 1.0 } All three cases behave correctly. Summary We set out to build the working example that our previous article said didn't exist. Here's what we showed. We embedded a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity map — compound, positive, negative and neutral — matching the return shape of the Java VADER UDF from the previous article. The wit-bindgen tuple ABI writes four f32 values to consecutive memory addresses; the Java side reads them back with a LITTLE_ENDIAN ByteBuffer. All four scores are correct, capitalization sensitivity works, and the empty string guard returns a sensible neutral map. The result is a workable integration pattern rather than a universal replacement for a native Java implementation. With the per-call initialization used in this proof of concept, the approach is best suited to low-frequency workloads where Wasm isolation and portability justify the additional complexity. For high-throughput workloads, the natural next step is to benchmark and reuse the Wasmtime engine and compiled module while keeping execution state appropriately isolated between calls. In the next article, we'll look at running TypeSafe AI's Jev inside Neo4j for calibrated sentiment decisions. Stay tuned! The full source code is available on GitHub. More
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
By Jeremy Morgan
The Silent Container Death: A TCP Dial That Never Times Out
The Silent Container Death: A TCP Dial That Never Times Out
By Alexander Fo
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
By Jerzy Kopaczewski
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture

Kafka-compatible ingestion is showing up everywhere now, including in analytics platforms that have nothing to do with running Kafka. Databricks added Kafka-compatible APIs to its Zerobus Ingest service. Snowflake went a step further at its Summit and announced Datastream, a native, fully Kafka-compatible streaming service. In both cases, existing Kafka producers can stream straight into the platform with no code changes, just a config change. This is good news for the ecosystem. It also confirms a point I have made for years: the Kafka API has become the de facto standard for moving events around, the same way the Amazon S3 API became the standard for object storage. But the trend hides an important distinction. Using the Kafka API for ingestion into an analytics platform is a very different thing from building an event-driven enterprise architecture on Kafka. And in practice, most companies do both: they run Kafka as the operational backbone, and one of the most common sink connectors in that setup feeds the lakehouse. This post is about that difference, ingestion versus architecture, why it matters, and why the two approaches are complementary rather than competing. What Is Apache Kafka? Messaging, Storage, Connect, and Streams Most people know Kafka as messaging. Pub-sub. A way to send events from A to B in real time. That is the part everyone learns first, and it is real, but it is only a quarter of the picture. Apache Kafka (the open source framework) is four things in one platform: Messaging is the real-time publish-and-subscribe layer that decouples producers from consumers. Storage is the durable, replayable commit log. A Kafka topic is not a transient buffer that forgets data the moment it is read. Events are persisted, ordered, and can be replayed by any consumer at any time. This is what makes Kafka a system of record for events, not just a pipe. Kafka Connect is the integration framework. Hundreds of connectors move data into and out of Kafka from databases, message queues, SaaS applications, cloud storage, and yes, lakehouses. Kafka Streams and stream processing handle continuous processing of data while it is in motion. You filter, join, aggregate, and enrich events as they flow, rather than landing everything first and processing it later in batch. The point is simple: Kafka is a platform, not a message queue. That combination of storage, integration, and processing on top of messaging is exactly what lets it serve as a backbone for an entire enterprise, not just a transport between two systems. The persistent log also decouples producers from consumers in time, so each system reads the same events at its own pace, whether real-time, batch, or request-response. That is what makes Kafka a tool for data consistency across the whole estate, not only real-time speed and scale. Kafka Protocol vs. Kafka Framework: What 'Kafka-Compatible' Really Means Here is the distinction that explains the whole "Kafka is everywhere" trend, and it is one people constantly blur. The Kafka API, also called the Kafka protocol, is the open wire protocol licensed under Apache 2.0. It defines how clients talk to brokers: the requests, the binary format, the rules. It is open, and anyone can implement against it. The Kafka framework is the actual open-source software: the brokers, Kafka Connect, Kafka Streams. It is one implementation of that protocol, the original one. Because the protocol is open, other systems can speak it without using the framework underneath. Snowflake is refreshingly direct about this with Datastream: it is compatible with the Kafka wire protocol but does not use Kafka under the hood. The technology is pure Snowflake. That is the protocol-versus-framework split in a single sentence, stated by the vendor. And it is precisely why the Kafka API became the de facto standard for event streaming. But there is a crucial caveat. Protocol-compatible does not mean feature-complete. Many systems that advertise "Kafka-compatible" implement only the messaging core, and even there they often miss features. They often skip the storage semantics, connectors, stream processing, and exactly-once guarantees. So when a product says it supports the Kafka API, that tells you the on-ramp works. It does not tell you that you have a streaming platform underneath. The devil is in the details, and the details are usually everything except the basic produce-and-consume path. Two Jobs for the Kafka API: Ingestion vs. Architecture Once you separate the protocol from the platform, two very different use cases come into focus. The first is Kafka as the operational backbone. Event-driven microservices. Real-time applications. The central nervous system that connects operational and analytical systems across the business. Kafka as the strategic integration layer that replaces the old ESB, ETL, and iPaaS middleware. In this world, data is processed in motion, consumed by many systems at once, and the workloads are mission-critical, stateful, and continuous. The second is Kafka as an ingestion path into analytics. Here the Kafka API is a convenient, standardized on-ramp to get events into a warehouse or lakehouse, where they are processed analytically, often in micro-batches. One producer, one destination, analytics at the end. The first is an event-driven architecture. The second is a smarter pipe into an analytics platform. And here is the part that matters most: most companies do not pick one. They do both. They deploy Kafka for the event-driven architecture, run their operational workloads on it, and then one of the most common sink connectors in that whole setup is the one feeding the lakehouse. Kafka runs the live business, and Snowflake or Databricks gets a continuous feed of those same events for analytics, reporting, and AI. The two are complementary, not competing. Streaming Platform vs. Lakehouse Ingestion: Confluent, Databricks, Snowflake This is where the recent vendor announcements fit in, and it helps to see the two approaches side by side. On one side are the data streaming platforms. Whether you run open-source Apache Kafka yourself or work with a vendor such as Confluent, the goal is the same: build the full platform around the event-driven architecture. Operational and analytical workloads, Connect for integration, stream processing for data in motion, and the durable log as a system of record. There is a rich and fast-moving market here, with strong options well beyond the obvious names. For an overview of how it is evolving, see my Data Streaming Landscape 2026. On the other side are the analytics platforms adding Kafka-compatible ingestion. Databricks Zerobus and Snowflake Datastream are two clear examples. Their goal is narrower and perfectly reasonable: get producer data into their platform without requiring a separate Kafka cluster or a separate streaming vendor. Snowflake is explicit that Datastream is purpose-built for teams that want to replace their Kafka infrastructure with a native Snowflake service, landing topics directly as governed Snowflake or Iceberg tables. Worth noting on maturity: Databricks has rolled out its Kafka-compatible APIs in beta, with the rest of Zerobus generally available, while Snowflake Datastream is still heading into private preview. Both implement the protocol as an ingest interface into their own platform, not as a general-purpose streaming backbone. Both approaches are valid. They solve different problems. When to Use Kafka vs. Lakehouse Ingestion If all you need is to land events in the lakehouse, then Kafka-compatible ingestion built directly into the analytics platform is a fine choice. There is nothing wrong with it. It is one less system to run, one less vendor to manage, and the config-change migration story is real. No need to overthink it. But be clear-eyed about what most enterprises actually do. They do not use Kafka only for ingestion into analytics. They use it for operational use cases, for a strategic event-driven enterprise architecture, and as the integration platform that replaces legacy middleware. For that, an analytics platform's ingest service is not a substitute. It was never designed to be one. It is the on-ramp, or at most a replacement for the on-ramp, not the backbone. So choose based on the use case, not the headline. If you only need the pipe into the lakehouse, take the pipe. If you are building the operational nervous system of your business, you need the platform, and the lakehouse becomes one important consumer of it rather than the center of it. More vendors adopting the Kafka API is a win for everyone. It confirms the protocol is the standard. The only thing to stay sharp on is which problem you are solving: ingestion into analytics, or the event-driven enterprise architecture that runs the business.

By Kai Wähner DZone Core CORE
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code

AI-generated software needs provenance that survives beyond the chat window. A code review can show what changed, but it rarely shows which model produced a fragment, what prompt and repository context influenced it, which agent or tool executed the change, which intent was being implemented, or whether the recorded history was altered later. Provenance fills that gap by treating generation as a supply-chain event rather than an ephemeral interaction. The core idea is already established in adjacent standards where W3C PROV models provenance through entities, activities, and agents, while SLSA records how software artifacts were produced so downstream consumers can verify expected processes and inputs. Generation Provenance as Engineering Metadata The first design rule is to separate authorship from provenance. Provenance answers where code came from and how it was produced, and it does not, by itself, determine legal ownership. The U.S. Copyright Office states that generative-AI output is copyrightable only where sufficient human-authored expressive elements exist, and that prompting alone is not enough. Employment agreements, contributor agreements, licenses, and jurisdictional law still govern ownership questions. Provenance instead supplies evidence for attribution, review, audit, and accountability. A minimal record should bind the generated artifact to the model provider, model identifier and revision, agent identity and version, prompt digest, context digests, execution trace, repository commit, intent digest, timestamp, and approving human or service identity. Hosted model aliases can change over time, so a provider-returned model identifier or immutable deployment revision is preferable to a friendly model name alone. The record should also contain cryptographic digests for generated files so later edits cannot silently inherit stale provenance. JSON { "artifact": "src/billing/CancelService.java", "sha256": "7e91...c42a", "commit": "9f3c1ad", "model": {"provider": "acme-ai", "id": "code-model", "revision": "2026-08-14"}, "agent": {"id": "repo-agent", "version": "3.7.2"}, "prompt": "sha256:18ab...90ef", "context": ["git:9f3c1ad^", "sha256:44c2...bb10"], "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1...", "approvedBy": "team:payments-reviewers" } Git commit trailers provide a low-friction place to attach pointers because Git supports structured token-value trailers at the end of commit messages. The commit should store references and digests rather than sensitive prompts themselves. Plain Text AI-Provenance: sha256:5df0...a992 AI-Model: acme-ai/code-model@2026-08-14 AI-Agent: [email protected] Prompt-Digest: sha256:18ab...90ef Context-Digest: sha256:44c2...bb10 AI-Trace: urn:uuid:2cf1... Intent-Digest: sha256:a771...0d61 File-level attribution can use a compact pointer rather than duplicating the full record. A generated region can carry a comment such as // ai-provenance: urn:gen:2cf1..., while the referenced sidecar record maps that generation to file hashes and, when needed, line ranges. This keeps source readable and prevents model metadata from becoming scattered, inconsistent comments. From SBOM to Generation BOM AI-generated code needs an additional description of the generation process. CycloneDX already supports source code, machine-learning models, component provenance, formulation describing how objects were created, and citations that attribute supplied information to entities or processes. Its ML-BOM capability records models, datasets, configurations, and provenance, while the earlier model-card framework established the broader practice of documenting model identity, intended use, evaluation, and limitations. A generation BOM can therefore be implemented as a small signed sidecar linked to the repository and release artifact, rather than inventing a second source-control system. JSON { "bomFormat": "GenerationBOM", "specVersion": "0.1", "subject": {"path": "src/billing/CancelService.java", "sha256": "7e91...c42a"}, "generator": {"agent": "[email protected]", "model": "acme-ai/code-model@2026-08-14"}, "inputs": {"prompt": "sha256:18ab...90ef", "context": ["sha256:44c2...bb10"]}, "intent": "sha256:a771...0d61", "trace": "urn:uuid:2cf1..." } The intent digest is especially important. A prompt records instructions presented to a model, but an intent contract records the behavior that must remain true after generation. Such a contract can contain permitted change scope, protected behaviors, security constraints, and acceptance criteria. Provenance then connects the produced code not merely to an AI request, but to a reviewable engineering objective. This mirrors data-lineage systems such as OpenLineage, which associate runs, jobs, datasets, and extensible facets so downstream analysis can reconstruct how an output was produced. Artifact hashes alone cannot establish reproducibility when generation depends on mutable infrastructure. Provenance should therefore bind execution parameters such as model configuration, decoding settings, tool versions, retrieval indexes, and policy revisions. Capturing these values converts provenance from a historical label into a verifiable reconstruction boundary for later audits and incident analysis. Tamper Evidence and CI Enforcement Metadata becomes trustworthy only when alteration is detectable. SLSA explicitly treats provenance authenticity and digital-signature verification as mechanisms for detecting tampering, and recommends approaches that improve compromise detection, including transparency logs. Sigstore provides signing with short-lived identity-bound certificates and records signing events in Rekor, an append-only transparency log. A provenance file can be signed as a blob during CI: Shell cosign sign-blob \ --bundle generation-provenance.sigstore.json \ generation-provenance.json Verification should occur before merge or release, not after an incident. SLSA similarly emphasizes that provenance has little value unless a consumer verifies it against expected properties. Shell provctl verify \ --commit "$GIT_COMMIT" \ --require-model \ --require-context-digest \ --require-intent \ --require-signature \ --max-unattributed-lines 0 A practical gate should reject AI-marked changes when the artifact digest no longer matches, required model or agent fields are absent, the provenance signature fails, the intent contract is missing, or the trace cannot be resolved. Human-edited code should not be forced into artificial AI attribution; instead, the policy should distinguish generated, transformed, and manually authored regions. Execution traces can preserve tool calls, retrievals, test runs, and agent steps and are designed to make software-supply-chain steps transparent by recording what happened, by whom, and in what order, and their runtime-trace predicate can describe system events associated with a supply-chain step. Runtime verification closes another gap. Provenance can prove which generation path produced a deployment, but not that the resulting behavior remains correct under production conditions. Release telemetry should therefore link runtime incidents back to commit, provenance record, model revision, and intent contract. That correlation turns an AI-related defect from an unstructured forensic exercise into a query over lineage. Accountability Without Capturing Everything Capturing every prompt verbatim is usually the wrong default. Prompts and retrieved context may contain credentials, personal data, proprietary code, customer information, or licensed material. A safer design stores encrypted source material in an access-controlled evidence store and places digests, object references, retention class, and classification labels in Git-visible provenance. High-sensitivity environments can retain only keyed digests and approved summaries where reproduction is less important than proof of correspondence. Storage and performance costs also require boundaries. Full agent traces can be large, while line-level metadata can become noisy after refactoring. The durable unit should normally be a generation event bound to artifact digests and commits, with finer-grained ranges reserved for high-risk code. Developer ergonomics matter equally, as provenance capture should be automatic in IDE agents, repository bots, and CI runners rather than dependent on manual form filling. Regulation strengthens the case for disciplined records without creating a universal rule that every AI-generated source line must carry a label. The EU AI Act requires general-purpose AI model providers to maintain technical documentation and copyright-compliance policies, while NIST SP 800-218A extends secure development practices specifically for generative AI across the software lifecycle. These frameworks reinforce documentation, traceability, and governance, but a code-provenance system should be treated as engineering evidence rather than a substitute for legal analysis. AI-generated code should enter a repository with the same expectations applied to any other supply-chain artifact: origin, inputs, process, identity, integrity, and approval must be recoverable later. The strongest implementation is not a comment saying that AI was used, but a signed provenance chain linking model and agent identity, prompt and context digests, execution trace, intent contract, commit, generated artifact, review decision, and runtime evidence. Teams adopting AI-assisted development should make that chain automatic, verify it in CI, protect sensitive evidence separately, and fail closed for unattributed high-risk changes. That converts provenance from documentation into an enforceable engineering control and makes accountability possible long after the generation session has disappeared.

By Uthej Mopathi DZone Core CORE
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data

Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq

By Jubin Soni, FBCS DZone Core CORE
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure

Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.

By Akhil Madineni DZone Core CORE
Building a Product Recommendation Engine With Neo4j — No ML Library Required
Building a Product Recommendation Engine With Neo4j — No ML Library Required

When many developers think about recommendation engines, they think of machine learning: collaborative filtering models, matrix factorization, embedding vectors, and training pipelines. What surprises many people is that you can build a genuinely useful recommendation system with nothing more than a graph database and several Cypher queries. No scikit-learn, no TensorFlow, no model training. Just the natural structure of the data doing the work. In this article, we'll build a product recommendation engine on top of Neo4j Aura using two Jupyter notebooks. The first generates a realistic synthetic dataset and loads it into Aura. The second runs four recommendation queries directly in Cypher and visualizes the results with Plotly. Everything runs locally in a Python virtual environment against a free cloud Neo4j instance. The full source code is available on GitHub. Why Graphs Are a Natural Fit for Recommendations The core intuition behind most recommendation approaches is relationship: this customer bought that product, those products appear together in the same order, this product shares attributes with that one. In a relational database, capturing these relationships means multiple self-joins across large tables. A query like "find products bought by customers who also bought what this customer bought" quickly becomes difficult to write and expensive to execute at scale. In a graph, that same question is a traversal. We follow edges from a customer to the products they purchased, hop across to other customers who share those products, and collect what else those customers bought. The query is short, the intent is clear, and the graph engine is optimized for exactly this kind of path-following work. Prerequisites AuraDB is Neo4j's fully managed cloud database. A free tier is available with no credit card required. Sign up at Get Started for Free.Create a new AuraDB Free instance.When the instance is created, download or note the credentials — the connection URI, username, and password.Once the instance is running, open the Query tab and connect to the instance.Confirm it's empty with MATCH (n) RETURN count(n) which should return 0 A virtual environment is highly recommended. For example: Shell python3 -m venv ~/recommendation-engine-env source ~/recommendation-engine-env/bin/activate Before starting Jupyter, export the connection details as environment variables in your shell: Shell export NEO4J_URI="neo4j+s://xxxx.databases.neo4j.io" export NEO4J_USERNAME="your_username_here" export NEO4J_PASSWORD="your_password_here" The Graph Model Before we write any code, let's define the graph. We have four node types and three relationship types. Nodes Customer – id, name, email, city, country.Product – id, name, description, price.Category – name (e.g., Electronics, Clothing, Books).Tag – name (e.g. "wireless", "eco-friendly", "premium"). Relationships (:Customer)-[:PURCHASED {order_id, quantity, order_date}]->(:Product) — order metadata lives on the relationship rather than a separate Order node, which keeps our Cypher clean.(:Product)-[:BELONGS_TO]->(:Category)(:Product)-[:TAGGED_WITH]->(:Tag) The decision to put order_id, quantity and order_date on the PURCHASED relationship is worth discussing. It means a single customer can have multiple PURCHASED relationships to the same product (each with a different order_id) and we can group by order_id to find products that appeared together in the same basket — which is exactly what our co-purchase query needs. Figure 1 illustrates exactly this point, as we have a customer, two products, and the same order_id. Figure 1. Shared order_id enables co-purchase queries Notebook 1: Data Generation and Loading Rather than sourcing an external dataset, we'll generate synthetic data using Faker. This keeps the notebook fully self-contained, and readers can run it as-is without downloading anything. We'll generate 2,000 customers, 500 products across 15 categories, and 20,000 orders. Each order is a basket of several products sharing the same order_id — this is the key design decision that makes the frequently-bought-together query work. With an average basket of 3 products, we end up with around 60,000 PURCHASED relationships in the graph. Realistic Product Names Faker's default catch_phrase() method produces output like "Proactive exuding encoding" — readable enough for a demo but not really useful in an article. Instead, we define a PRODUCT_VOCAB dictionary keyed by category, each containing lists of adjectives, nouns, use cases, and benefit statements. A product name is then a simple combination, as follows: Python def make_product_name(category): vocab = PRODUCT_VOCAB[category] adj = random.choice(vocab["adjectives"]) noun = random.choice(vocab["nouns"]) return f"{adj} {noun}" def make_product_description(category, name): vocab = PRODUCT_VOCAB[category] use_case = random.choice(vocab["use_cases"]) benefit = random.choice(vocab["benefits"]) return f"The {name} is designed for {use_case}. {benefit}." This gives us names like "Wireless Noise-Canceling Earbuds," "Organic Ground Coffee" and "Ergonomic Lumbar Support Cushion" — realistic enough to make the recommendation output meaningful. Basket-Based Order Generation Each order picks a random customer, generates a unique order_id, and samples several products into a basket. We then flatten the basket into individual order lines, each carrying the shared order_id: Python orders = [] for _ in range(NUM_ORDERS): order_id = str(uuid.uuid4()) customer = random.choice(customers) order_date = (start_date + timedelta(days=random.randint(0, 730))).strftime("%Y-%m-%d") basket = random.sample(products, k=random.randint(2, 4)) for product in basket: orders.append({ "order_id": order_id, "customer_id": customer["id"], "product_id": product["id"], "quantity": random.randint(1, 5), "order_date": order_date }) Loading Into Aura Data loading is in batches of 100 using MERGE statements. To show progress during the load, we'll use tqdm as ~60,000 order lines can take several minutes, and the progress bars make it easy to see what's happening: Python with driver.session() as session: customer_batches = range(0, len(customers), BATCH_SIZE) for i in tqdm(customer_batches, desc="Loading customers", unit="batch", colour="#1f77b4"): session.execute_write(load_customers, customers[i:i+BATCH_SIZE]) product_batches = range(0, len(products), BATCH_SIZE) for i in tqdm(product_batches, desc="Loading products ", unit="batch", colour="#1f77b4"): session.execute_write(load_products, products[i:i+BATCH_SIZE]) for product_id, tags in tqdm(product_tags.items(), desc="Loading tags ", unit="product", colour="#1f77b4"): session.execute_write(load_tags, product_id, tags) order_batches = range(0, len(orders), BATCH_SIZE) for i in tqdm(order_batches, desc="Loading orders ", unit="batch", colour="#1f77b4"): session.execute_write(load_orders, orders[i:i+BATCH_SIZE]) A verification query at the end confirms the counts. The Four Recommendation Queries Notebook 2 runs four Cypher queries against the loaded graph, each implementing a different recommendation strategy. Before running any query, we fetch a stable seed customer, product, and category: Python with driver.session() as session: customer = session.run(""" MATCH (c:Customer) RETURN c.id AS customer_id, c.name AS customer_name ORDER BY c.name ASC LIMIT 1 """).single() product = session.run(""" MATCH (p:Product)<-[r:PURCHASED]-() RETURN p.id AS product_id, p.name AS product_name, count(r) AS order_count ORDER BY order_count DESC LIMIT 1 """).single() top_cat = session.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category) RETURN cat.name AS category, count(p) AS total ORDER BY total DESC LIMIT 1 """).single() We pick the alphabetically first customer for consistency, the most-purchased product to ensure co-purchase data exists, and the category with the most products for the trending query. This makes the notebook reproducible across runs. Query 1: Collaborative Filtering The classic "customers who bought this also bought" approach. We find customers who share at least one purchased product with the seed customer, then collect what else those customers bought — excluding anything the seed customer already purchased. Python def collaborative_filtering(tx, customer_id, limit=5): result = tx.run(""" MATCH (target:Customer {id: $customer_id})-[:PURCHASED]->(p:Product) <-[:PURCHASED]-(other:Customer)-[:PURCHASED]->(rec:Product) WHERE NOT (target)-[:PURCHASED]->(rec) RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(other) AS score ORDER BY score DESC, id ASC LIMIT $limit """, customer_id=customer_id, limit=limit) return result.data() The score is the number of other customers whose purchasing overlap with our target customer also led them to buy the recommended product. A higher score means more customers in the overlap group bought it, making it a stronger signal. In Cypher, the traversal reads almost like the description: start at the target customer, follow PURCHASED edges to products, hop to other customers who bought the same products, then follow their PURCHASED edges to new products. Query 2: Frequently Bought Together This query finds products that appeared in the same order as the seed product. The key is matching on order_id across two PURCHASED relationships from the same customer: Python def frequently_bought_together(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})<-[r1:PURCHASED]-(c:Customer) -[r2:PURCHASED]->(other:Product) WHERE r1.order_id = r2.order_id AND other.id <> $product_id RETURN other.id AS id, other.name AS product, other.price AS price, count(c) AS frequency ORDER BY frequency DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The WHERE r1.order_id = r2.order_id clause is what makes this work. It constrains the traversal to only consider cases where both products were part of the same order, not just bought by the same customer at different times. frequency counts how many distinct customers placed an order containing both products together. Query 3: Content-Based Filtering Rather than looking at purchase behavior, this query finds products similar to the seed product based on shared tags. The more tags two products have in common, the more similar they are: Python def content_based(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})-[:TAGGED_WITH]->(t:Tag) <-[:TAGGED_WITH]-(rec:Product) WHERE rec.id <> $product_id RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(t) AS shared_tags ORDER BY shared_tags DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The traversal goes outward from the seed product through its tags, then back inward to any other product that shares those same tags. count(t) gives the number of shared tags, which serves as a simple but effective similarity score. This approach works without any purchase history, making it useful for recommending products to new customers or for newly listed products with no order data yet. Query 4: Trending in Category This query finds the most purchased products in the top category within a fixed date window. In our case, this is from 2024-10-01 onwards: Python def trending_in_category(tx, category_name, cutoff="2024-10-01", limit=5): result = tx.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category {name: $category_name}) MATCH (:Customer)-[r:PURCHASED]->(p) WHERE date(r.order_date) >= date($cutoff) RETURN p.id AS id, p.name AS product, p.price AS price, count(r) AS purchases ORDER BY purchases DESC, id ASC LIMIT $limit """, category_name=category_name, cutoff=cutoff, limit=limit) return result.data() We use date() conversion on the stored string order_date to enable date comparison. count(r) counts individual PURCHASED relationships rather than distinct customers, so a customer who bought the same product multiple times within the window is counted each time — reflecting genuine demand volume rather than unique buyer count. Notebook 2: Results Each query outputs a table followed by a Plotly horizontal bar chart. Here are the results for our seed data. Collaborative Filtering Figure 2 returns five products. The top recommendation is An Introduction to Public Speaking, driven by the number of customers whose purchasing overlap with Aaron Boyd also led them to buy it. Heavy-Duty Cable Management Box and Educational Coding Robot follow closely, showing that the overlap group bought broadly across categories rather than clustering in one area. Figure 2. Collaborative filtering Frequently Bought Together Figure 3 shows products co-purchased with the Durable Grooming Brush in the same order basket. The top results — Waterproof Hammock and Natural Body Lotion at frequency 4, followed by Adjustable Lumbar Support Cushion, Sugar-Free Collagen Powder and Slim-Fit Hiking Vest at frequency 3 — show which products most commonly appeared alongside the seed product in the same order. The cross-category spread here (Beauty, Outdoor, Clothing, Health, Office) is a feature of random synthetic data; in a real system, we'd expect more category clustering. Figure 3. Frequently bought together Content-Based Filtering Figure 4 finds products sharing the most tags with the seed product. All five results share 2 tags with the Durable Grooming Brush — Smart Mechanical Keyboard, Waterproof Toiletry Bag, Ergonomic Whiteboard, Cold-Pressed Hot Sauce, and Durable Dumbbell Pair. The cross-category reach (Sports, Food & Drink, Office, Travel, Electronics) illustrates the tag graph doing its job: shared attributes like "durable" or "waterproof" create similarity links that cross category boundaries, which is useful for surface-level discovery recommendations. Figure 4. Content-based filtering Trending in Category Figure 5 shows the top 5 products in Toys — the category with the most products in our graph — with purchase counts from 2024-10-01 onwards. Battery-Free Coding Robot leads, followed by Battery-Free Building Blocks Set, Interactive Remote Control Car, Wooden Magnetic Drawing Board, and Creative Puzzle Game. The scores are tight here, which makes sense because within a single category over a fixed time window, popular products tend to cluster around similar purchase volumes. Figure 5. Trending Summary We've built a working product recommendation engine using nothing but Neo4j, Cypher, and a few Python libraries. No ML framework, no training data, no model deployment. The four queries cover the most common recommendation patterns in production systems: Collaborative filteringCo-purchase analysisContent similarityTrending detection The graph model is the foundation that makes this possible. Storing orders as relationships with properties means co-purchase queries are a natural traversal rather than a complex join. Adding tags as nodes means similarity queries are just path-matching. Because everything lives in the same graph, we can also combine these approaches. For example, filtering collaborative filtering results by tag similarity using a single extended Cypher query. The full source code is available on GitHub.

By Akmal Chaudhri DZone Core CORE
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults

Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report

By Naga Santhosh Reddy Vootukuri DZone Core CORE
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control

Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Kubernetes environments can drift, accumulate one-off fixes, and diverge across teams until a routine deploy breaks or a cost spike forces a review. This checklist gives platform, SRE, and engineering teams a way to keep clusters, deployments, and automation manageable as Kubernetes operations scale and teams grow. It covers standards, observability, releases, access, drift, and cost. Review it before promoting a service to production and revisit it as your environments shift. Cluster Standards and Environment Discipline Most teams run more than one Kubernetes cluster, and those clusters diverge over time as they are upgraded and modified independently. At that point, a fix or runbook that works on one cluster can’t be trusted to work on another. Keeping the fleet operable requires every cluster to run a supported Kubernetes version and follow the same approved platform settings and policies. Document centrally controlled settings (e.g., Kubernetes versions, networking, admission policies) separately from service-team settings (e.g., pod resource requests, autoscaling, ConfigMaps)Standardize namespace, labeling, and resource-quota conventions so workloads are identified and bounded consistently across clustersMaintain each cluster’s baseline configuration in version control; use reconciliation to apply it and correct untracked changesMaintain an approved Kubernetes version range across environments; track each cluster against this range and upgrade it before its current version reaches end of supportRecord any cluster setting that differs from the standard baseline, including the justification, approver, expiry date, and whether it must be restored or reapproved control areawhat to standardizeminimum evidence Kubernetes version Supported version range and upgrade cadence Version inventory showing every cluster within the supported range Cluster baseline Networking, ingress, and baseline policies Declarative config in version control, reconciled to live state Namespaces and quotas Naming, labels, and resource quotas Quota and label audit across clusters Exceptions Approved deviations from baseline Record of approved deviations with justification and expiry date Deployment Consistency and Release Safety Kubernetes makes it easy to ship a change to production several ways: a CI/CD pipeline, a Helm upgrade run by hand, or a kubectl apply straight from a laptop. Each runs different checks, but the manual ones skip the tests and approvals that a pipeline would enforce. A repeatable release path applies the same gates every time and provides a reliable way to recover when a deployment fails. Require every service to follow the same approved deployment path from commit to production, with consistent release steps and controls across teams and environmentsPromote the same versioned, immutable artifact through every environment without rebuilding it at each stageRequire every change to clear the same automated gates (e.g., tests, policy checks, health checks) before reaching productionRoll out production changes in stages (e.g., canary release, percentage-based traffic shift); automatically stop or roll back when predefined health criteria are not metFor every production change, require a rollback, feature disablement, or recovery path that has been tested before releaseFor each deployment, assign an owner accountable for monitoring it through release and triggering rollback on failureRecord every production deployment with its artifact version, approver, and timestamp so the active release stays auditable Observability and Operational Readiness A Kubernetes cluster keeps workloads running by restarting and rescheduling them, so a service can keep failing without the failure ever becoming obvious. A pod stuck in CrashLoopBackOff or failing its readiness probe can remain unhealthy for hours, and if it emits no metrics or logs of its own, there’s nothing to tell you what went wrong. Catching that early depends on each service surfacing its own signals rather than waiting for the cluster to show something is wrong. Require every new service to ship with a minimum observability baseline before production: metrics, structured logs, traces, and liveness and readiness probesDefine service health signals (e.g., latency, traffic, errors, saturation), each with a threshold and assigned team that responds when it is breachedStandardize structured logging and trace context so a request can be followed end to endRoute every alert to an on-call rotation or runbook; retire alerts no one acts onMaintain a quarterly reviewed runbook for each service, including known failure modes, escalation contacts, and recovery stepsSet minimum retention periods for metrics, logs, and traces, with documented justification and explicit approval for shorter retention periodsRun a post-incident review after every major outage; apply findings to update runbooks, alerts, and service baselines Access Controls and Automation Guardrails A Kubernetes cluster usually serves many teams and workloads through a single shared control plane. A role with too much access, for example, can affect them all at once. And when the credential is shared, there’s no way to tell later who actually made the change. Access that stays narrow and tied to a single identity keeps a mistake or a compromised account from impacting the whole cluster. Use namespace-scoped RBAC roles with only the required permissions; grant cluster-wide administrator access only through logged, justified, time-limited exceptionsGive each automation its own scoped service account so automated and privileged actions trace to a distinct identity instead of shared credentialsReserve break-glass access for emergency production changes, with time limits and post-use reviewUse admission policies to reject workloads with unsigned images, privileged containers, or settings barred by platform standardsRecord the actor, target, and timestamp for every privileged or automated action in the Kubernetes audit log; regularly review for activity that does not match an approved change or access requestUse short-lived, automatically rotated ServiceAccount tokens for workloads; revoke credentials and RBAC bindings when a person, workload, or automated process is decommissioned Drift and Failure Management Over time, a Kubernetes cluster’s live state can drift from the configuration stored in version control. This could be due to a hotfix applied directly to a live resource during an incident or an incomplete rollout that leaves the cluster partially updated. If those differences are not fixed, a subsequent deployment may conflict with the live state or overwrite a manual change, and version control may no longer accurately reflect what is running in the cluster. Use automated checks to compare live cluster state with the version-controlled baseline at defined intervals; record each mismatch and notify the team responsible for the affected resourceSet risk-based remediation deadlines for detected drift, requiring teams to restore the baseline or approve a time-limited exception for the changed configuration before the deadlineLog every manual production change and resolve it within a defined period by updating the baseline or reverting the live resource to its declared stateSet an SLO and error budget for each service, identify the team tracking budget use, and pause feature work to prioritize reliability fixes when the budget is exhaustedRun root-cause reviews for recurring failures and apply findings to update baselines, policies, and admission checks instead of patching each instanceTest failure scenarios (e.g., pod disruption, node loss, dependency outages) on a defined schedule, confirm services recover as expected, and track remediation for any gaps example drift patternwhat usually reveals it Manual live-resource change Reconciliation diff against declared state Version or baseline skew Scheduled cluster inventory audit Expired break-glass fix Exception register entry past its window Repeated failure patched one service at a time Same root cause across incident reviews Cost Awareness and Resource Discipline In Kubernetes, resource requests for CPU and memory determine how much cluster capacity is reserved for a workload. Teams may size these requests for peak demand and leave them unchanged even when normal usage is much lower. Across many workloads, this unused capacity adds up and can cause the cluster to run more nodes than actual demand requires, increasing infrastructure costs. Set CPU and memory requests based on representative usage data; set limits where appropriate based on workload behavior and reliability requirementsReview workloads whose requests exceed observed use by a defined threshold, accounting for traffic patterns and reliability needsRequire cost-allocation labels for each workload by team and namespace; correct unallocated spend and missing or inaccurate labelsReclaim idle and orphaned resources (e.g., unused volumes, stale namespaces, oversized nodes) on a monthly cadenceSet autoscaling thresholds based on demand and reliability requirements; periodically review settings that fall outside the approved rangeRegularly review sustained overprovisioning or low utilization; reduce excess capacity or record why it must be retained when avoidable cost exceeds a set threshold Closing Run this checklist before a service enters production and at regular intervals afterward. Repeat it when clusters are upgraded, team responsibilities change, or services are added or retired. Resolve failed checks and revisit approved exceptions before they expire. Unresolved configuration drift can accumulate across environments until teams begin to treat it as the intended baseline. This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report

By Abhishek Gupta DZone Core CORE
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.

The first warning sign wasn't an outage. It was a boring pull request. We changed one App Service setting. It was the sort of change that should have resulted in a small plan and a quick review. Instead, Terraform refreshed networking, private endpoints, DNS, Key Vaults, storage accounts, app services, and monitoring before showing what would actually change. Nothing was broken; that was the point. Terraform did exactly what it was designed to do: account for everything represented in state before calculating change. The problem was that our Terraform state had become a single, platform-sized boundary that every small change had to pass through, and one no team could fully own. If you have run a landing zone as a single Terraform configuration, you have probably had a version of that pull request. The instinct afterward is to blame size: the configuration has grown too large, so break it up. That instinct is wrong, or at least incomplete. Size is uncomfortable, but coupling is what actually hurts. Nothing in the change touched networking, DNS, or those key vaults. They were dragged into the plan because everything was bound together through one state. At first, that coupling just means slow plans and noisy reviews. Later, it raises a harder question: who actually owns this? Where the Coupling Shows Up Start with the plan. In a monolith, Terraform has to account for everything represented in the state before it can tell you what changed. You can target a single resource, but that is an escape hatch, not a way to run a platform. So the wait scales with the size of the estate, not your change. Both a one-line edit and a fifty-resource migration get stuck behind the same refresh before the diff appears. Provider upgrades show the same problem. A single root configuration pins one set of provider versions, so you cannot move networking to a newer azurerm version and leave everything else behind. Every upgrade becomes all-or-nothing, which means it keeps losing to smaller, safer priorities. Ours sat on azurerm 2.97 and only moved to the 4.x line once the upgrade could no longer be put off. The monolith had made the jump too big to schedule any sooner. The bigger concern is blast radius. One state file, one lock, one plan. A bad apply, a corrupted state, a destroy that catches more than you aimed at: whatever goes wrong can reach more of the platform than the change was ever meant to touch, because nothing in the layout is there to contain it. The dependency graph suffers too. Unrelated resources get sequenced together just because they share a graph. A network change might wait on unrelated compute, DNS on policy. The graph ends up reflecting accidental grouping rather than real dependencies. The result is clear. There is no small change. You cannot ship a DNS record or a new Key Vault without running the entire configuration through plan and apply. Every change is a platform change, carrying platform risk and requiring review, no matter how minor. These look like separate problems, but all come from the same design choice: too many unrelated concerns tied into one Terraform boundary. Where Coupling Becomes Ownership It is easy to call these operational annoyances: slow plans, awkward upgrades, risky applies, the tax you pay for a big configuration. But the same coupling appears in review and approval, where it stops being just an operational problem. Once too many concerns share the same state, pipeline, and approval path, the question is no longer only "how long did the plan take?" It becomes "who is accountable for the boundary this change is crossing?" Take private connectivity. A single private endpoint on Azure isn't handled by just one team. The application team owns the service behind it. The platform team manages the landing zone, subnet, and endpoint placement. Private DNS zones might be managed centrally or by another team. Security or governance may require the service to be private. How these map to teams varies, but in a monolith, everything ends up in the same state, pipeline, and plan. So "who owns this?" rarely has a clear answer. However you split teams, they are coupled through a single configuration that none can truly own. When the application team changes its service, the same config still carries platform connectivity and governance controls. You cannot draw ownership along your real organizational boundaries, because the code does not have them. Both slow plans and unclear ownership trace back to the same issue: shared concerns treated as if they belong to just one team. Figure 1: When Terraform boundaries stop matching ownership boundaries. The monolith gives Terraform one boundary. Organizations have several. The pain comes when small changes have to cross boundaries that no team fully owns. Reach for the Coupling, Not the Size The reflex now is to split the state and move on. But splitting a landing zone poorly can be worse than leaving it alone. If you split along the wrong lines, you trade one blast radius for tangled cross-state dependencies. You also lose the single plan that at least showed the whole graph in one place. For example, splitting private endpoints into one state and private DNS zones into another may look clean on paper. But if different teams deploy them without a clear agreement, every new endpoint becomes a coordination headache, not a smaller change. Moving files into separate folders does nothing if the same pipeline, credentials, and approval path still govern everything. Decomposition should follow actual coupling, not just line count. So the next question is not "how many states should we create?" It is "which boundaries are real enough for teams to own, deploy, and recover independently?" If your Terraform monolith hurts, do not start by counting files or resources. Look at what is actually being coupled. Slow plans and unclear ownership are both signs that your Terraform boundaries no longer match your real ownership boundaries.

By Naveen Kalapala
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest
When an iOS Retry Executes an Agent Twice: Building Effectively-Once Tool Workflows With LangGraph, MCP Tasks, Kafka, and App Attest

A mobile request can fail without the server-side work failing. An iOS app may time out, lose the response after a POST has reached the service, or retry after connectivity changes while the original execution is still progressing. Apple explicitly distinguishes safe retry behavior by HTTP method and notes that URLSession can retry requests in some connection-loss cases, waitsForConnectivity can also cause the system to continue a request when connectivity returns. The dangerous state is therefore not “request failed,” but “completion is unknown.” If that request starts an agent that charges an account, reserves inventory, sends a message, or invokes an MCP tool, a second submission can become a second side effect. The Retry Boundary Is the Real Transaction Boundary “Exactly once” is too strong for a workflow crossing an iPhone, HTTP, an agent runtime, an MCP server, Kafka, a database, and an external API. Kafka can provide exactly-once guarantees within defined Kafka processing boundaries, but those guarantees do not atomically include arbitrary remote tool effects. The practical target is effectively-once behavior, and retries are expected, but every effect is guarded by a stable operation identity and converges on one committed outcome. Kafka’s idempotent producer suppresses duplicate records caused by producer retries, while transactional producers can atomically publish across Kafka partitions; the producer documentation also limits idempotence guarantees to a producer session and requires read_committed consumers for end-to-end transactional visibility. The operation identity must exist before the first network attempt. An iOS client can create an operationId when an action becomes durable local intent, persist it, and reuse it across transport retries. Transport material such as a server challenge may change, but the business ID must not. The server treats (subjectId, operationId) as a uniqueness boundary and stores a canonical payload hash with it. PostgreSQL unique constraints enforce row uniqueness, while INSERT ... ON CONFLICT provides an atomic conflict path under concurrency. SQL INSERT INTO agent_operation(subject_id, operation_id, payload_hash, status) VALUES (:subject, :operationId, :payloadHash, 'ACCEPTED') ON CONFLICT (subject_id, operation_id) DO NOTHING; A conflict with the same payload hash returns the existing operation; a different hash rejects key reuse. The record should exist before agent execution starts, and the accepted response should expose the durable operation identity. Let LangGraph Resume Without Repeating Effects LangGraph persistence is useful precisely because durable execution can replay code. With a checkpointer, LangGraph saves state at super-step boundaries; if execution resumes after a failure, an affected node can run again from the beginning. Official guidance consequently requires idempotent node logic, and task results can be checkpointed so completed task work can be reused during resume instead of recomputed. Replaying from an earlier checkpoint can also re-trigger later LLM calls and API requests. A stable business operation should therefore map to a stable LangGraph thread, while every effectful tool boundary receives the same operation ID. Python config = {"configurable": {"thread_id": operation_id} result = graph.invoke( {"operation_id": operation_id, "command": command}, config ) Checkpointing reduces recomputation but does not replace downstream idempotency. A reservation can succeed before its task result is durably checkpointed. LangGraph’s functional API therefore recommends idempotent tasks because an incomplete task can execute again during resume. Python @task def reserve_inventory(operation_id, sku, quantity): return mcp.call_tool("reserve_inventory", { "operationId": operation_id, "sku": sku, "quantity": quantity }) The significant property in this snippet is not the decorator. The important part is that the business identity crosses the graph boundary and reaches the tool implementation. A downstream inventory service can then use that identity to return a previously committed reservation rather than creating another one. MCP Tasks Are Durable Handles, Not Deduplication Keys The current MCP Tasks design is especially relevant to long-running agent tools. In the July 28, 2026 protocol revision, Tasks moved into the io.modelcontextprotocol/tasks extension. A server can return a durable task handle, and the client can poll with tasks/get, provide input with tasks/update, or request cancellation with tasks/cancel. The task is durably created before its handle is returned, which allows polling after a disconnect. That durability solves result retrieval after task creation, but it does not by itself deduplicate the request that creates the task. The task ID is server-generated. If the server creates task A, the response disappears, and the original tools/call is sent again, a naïve implementation can create task B. Therefore, the business operationId must be part of the tool arguments or equivalent application metadata, and task creation must first look up an existing operation. This follows directly from MCP’s server-generated task-ID model combined with retry ambiguity at the HTTP boundary. The MCP server can return an existing task handle for the same authenticated subject, operation ID, and payload hash, and later return the stored terminal result. Cancellation should also be idempotent because MCP defines it as cooperative rather than a guarantee that underlying work stops immediately. Keep Kafka Guarantees Inside Kafka Kafka is most valuable after the operation has been claimed. A database transaction can persist operation state with an outbox row carrying the same ID. Kafka producer idempotence protects against duplicates caused by producer retries, while consumers can still use the operation ID for application-level deduplication. Kafka transactions can atomically cover Kafka writes, but they do not extend over an MCP server or payment API. The event contract should preserve causality rather than inventing a new identity at each hop. JSON { "operationId": "8E7B6D9E-...", "type": "AgentToolCompleted", "tool": "reserve_inventory", "status": "SUCCEEDED" } A consumer can enforce uniqueness on (consumerName, operationId, eventType) or make the state transition conditional. Kafka delivery guarantees and application idempotency then reinforce each other instead of being treated as interchangeable. Bind Retry Identity to App Attest Without Blocking Legitimate Retries App Attest addresses a different failure mode: whether a request comes from a legitimate app instance and whether signed request material has been replayed or altered. Apple’s current guidance uses a server-provided challenge for assertions and requires the server to validate a strictly increasing assertion counter; that counter is specifically an anti-replay signal. Assertions are generated locally on the device after key attestation. The App Attest assertion must not become the business idempotency token. A legitimate retry should obtain fresh challenge material and generate a fresh assertion while retaining the original operation ID. The data hashed for the assertion can bind the server challenge, operation ID, and canonical payload hash together. Swift let payloadHash = SHA256.hash(data: body) let clientData = challenge + operationID.data + Data(payloadHash) let clientDataHash = Data(SHA256.hash(data: clientData)) let assertion = try await service.generateAssertion( keyID, clientDataHash: clientDataHash ) Apple recommends server-controlled challenges, server-side validation, and assertion-counter tracking as assertions are generated on demand without a round trip to Apple’s servers. The server verifies App Attest, checks that the challenge binds the operation ID and payload, then performs the idempotency lookup. A fresh assertion can retry the same operation; a replayed assertion fails anti-replay validation; an altered payload fails the hash check. Effectively-Once Behavior Is a Composition Property Reliable agent execution does not come from asking iOS to retry less often or from labeling a Kafka pipeline “exactly once.” It comes from carrying one durable business identity across every retry and every boundary, claiming that identity atomically before execution, making LangGraph effects idempotent under resume, using MCP Tasks as durable result handles rather than creation-time deduplication keys, restricting Kafka’s exactly-once guarantees to Kafka’s transactional domain, and using App Attest to prove request integrity without confusing anti-replay state with business deduplication. When those boundaries align, a lost mobile response can cause another HTTP attempt, another graph invocation, or another poll, but it does not cause another business effect. That is the operational meaning of effectively once.

By Uthej Mopathi DZone Core CORE
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration
MCP vs REST/HTTP API vs Kafka: The Architect's Guide to Agentic AI Integration

Every major AI vendor now supports the Model Context Protocol. The framing is almost always the same: MCP is the universal connector for AI agents in the enterprise. That framing sets up a false choice. MCP, REST/HTTP APIs, and Apache Kafka are not alternatives. They solve different problems at different layers of the architecture. Treating them as competing options produces systems that are fragile exactly where they need to be reliable. These three technologies can and do coexist in the same architecture. The question is not which one to pick. It is which one belongs where, and what the tradeoffs are when more than one could technically do the job. This article maps that decision: what each technology is built for, where the boundaries are, and where the genuine gray areas lie. 1. What Is MCP and What Is It Built For? Anthropic introduced the Model Context Protocol in November 2024 as an open standard for connecting AI assistants to external tools and data sources. Before MCP, every AI model required a custom connector to each external system. Three models, ten systems: thirty custom integrations to build and maintain. MCP collapses that to one standard interface. Any compliant client talks to any compliant server without prior coordination. OpenAI adopted MCP in March 2025. Google DeepMind confirmed support in April 2025. By December 2025, MCP had reached over 97 million monthly SDK downloads across Python, TypeScript, Java, Kotlin, C#, and Swift. Anthropic donated the protocol to the Agentic AI Foundation under the Linux Foundation, with AWS, Google, Microsoft, Bloomberg, and OpenAI as platinum members. MCP is no longer a developer experiment. Signals of enterprise maturity are arriving quickly: AI agents paying for API access autonomously, cross-SDK interoperability between Anthropic and OpenAI converging on MCP Resources, composable enterprise workflows where agents read tool signatures and compose cross-system flows without predefined paths, and an official MCP Registry launched in late 2025 as the community-driven server directory. The 2026 roadmap focuses on scalable transport, agent-to-agent communication, governance maturation, and enterprise readiness covering audit trails and SSO-integrated authentication. MCP handles tool access: how an agent calls an external capability. It does not handle agent-to-agent coordination, which is the domain of protocols like Google's Agent-to-Agent (A2A). MCP and A2A are complementary and address different layers of agentic architecture. The moment MCP is asked to do more than tool access, the architecture starts to break. Security Maturity Is Still Catching Up With Adoption Most incidents disclosed in 2025 and early 2026 are implementation failures, not protocol flaws. An Endor Labs analysis of 2,614 MCP implementations found 82% use file system operations prone to path traversal and 67% use APIs related to code injection. Enterprise-grade authentication with OAuth 2.1 and SAML/OIDC is on the 2026 roadmap but still in progress. The practical controls for today: apply least privilege, limit MCP server access to only the systems and data each tool requires, and monitor tool definitions for unexpected changes. 2. MCP vs. REST/HTTP API MCP and REST/HTTP APIs serve different consumers and should not be treated as interchangeable. REST is an architectural style built on HTTP, widely adopted but with no fixed conventions for discovery, error formats, or method naming. Well-designed REST APIs backed by OpenAPI specifications work well for direct, programmatic data access when a native SDK or versioned API already exists and teams know how to operate it. MCP enforces consistency at the interface level because the consumer is an AI model that cannot tolerate creative API interpretation. MCP standardizes how a tool is called. It does not standardize what the tool returns, how fresh that data is, or whether two agents calling the same tool simultaneously see the same state. For direct data access to vector stores, databases, or business application APIs, a well-governed REST API, native SDK, or Kafka Connect integration is almost always the better choice: lower latency, no protocol overhead, mature tooling. For giving AI agents standardized, discoverable access to a broader set of tools across vendors and frameworks, MCP is the right layer. The two are complementary, not competing. Tool Design Matters as Much as the Protocol Choice One important nuance on tool design: mapping one-to-one from existing APIs to MCP tools rarely works well. What matters is tool granularity, smart metadata, and thoughtful assembly of the MCP layer. An MCP server that exposes well-structured, semantically rich tools lets an AI agent reason about capabilities and compose workflows. This is reminiscent of the composability questions from the enterprise SOA (Service-oriented Architecture) era. SOA promised flexible service composition but delivered integration chaos when governance, metadata quality, and service granularity were treated as afterthoughts. MCP faces the same risk. The protocol is sound; what determines success is the discipline applied to how tools are defined, documented, and assembled. What MCP Does Not Do What MCP does not do matters as much as what it does. It does not manage data, guarantee message delivery, enforce governance, or guarantee consistency across systems. It is an interface layer, not a data pipeline. That boundary becomes even clearer when looking at what Kafka does, which is structurally different from both MCP and REST. 3. Apache Kafka: Event Broker, Decoupling, and the Backbone Role Operational data is the live data that runs business processes: order states, inventory levels, transaction records, customer accounts, risk scores. It originates in systems like SAP, Salesforce, Oracle, and mainframes, and it changes continuously. Kafka is architecturally different from both HTTP and MCP in one way that matters most: it decouples producers and consumers through a persistent, ordered, append-only log. With HTTP or MCP, the caller and the callee are coupled at request time. Every integration is point-to-point. If the target system is slow or unavailable, the caller is directly affected. Kafka breaks that coupling entirely. A producer writes an event once. Any number of consumers read it independently, at their own pace, using their own communication paradigm. One consumer processes records in real time. Another runs nightly batch analytics over the same events. A third powers a stream processing pipeline. A fourth writes results to a data lake via Apache Iceberg. All of them consume the same underlying data product. None of them affects the others. Kafka supports three consumption patterns from a single event stream: streaming, request-response, and batch. The event exists once; each consumer is independent. This is the pub/sub event broker model, and it is what makes Kafka the integration backbone between operational and analytical systems. The diagram below shows this decoupling: a single Kafka topic serving real-time applications, HTTP-based consumers, batch analytics, and MCP agent interfaces simultaneously. Stream Processing With Kafka Streams and Apache Flink Stream processing is a core complement to Apache Kafka, extending the platform from event transport into real-time data processing and decisioning. Kafka Streams is a lightweight Java library embedded in applications. It is well-suited for streaming ETL and simple to medium stateful stream processing without requiring a separate cluster. It integrates closely with existing JVM-based services. Apache Flink is a distributed stream processing engine designed for more complex workloads. It supports Java, Python, and SQL APIs, making it accessible to both application developers and data engineers. Flink runs as a dedicated cluster or in managed environments and is built for high-scale scenarios such as multi-stream joins, event-time processing, large state management, exactly-once semantics, Complex Event Processing (CEP), real-time analytics, and AI model inference. Both approaches extend Kafka with processing capabilities. The choice depends on workload complexity, required deployment model, and preferred programming language, not on replacing Kafka’s role as the event streaming backbone. A detailed comparison is available in the post Apache Kafka and Apache Flink: A Match Made in Heaven. Operational and Analytical Integration, Including the Data Lakehouse Kafka is not only for operational data integration. It serves as the ingestion layer into data lakes, feeds real-time analytical pipelines, enables stream processing with embedded AI models, and connects business applications bidirectionally. A governed data streaming platform provides schema registry, lineage tracking, role-based access control, and exactly-once delivery semantics across all of that. It serves both operational and analytical use cases and acts as the bridge between those two worlds. For how streaming and the lakehouse converge via Apache Iceberg, see Data Streaming Meets Lakehouse. Kafka's append-only commit log is the foundation of data consistency across the enterprise. Every downstream consumer sees the same data in the same order. That is not just a performance feature. It is what prevents the architecture where every system has its own version of the truth. 4. The Tradeoffs: It Is Not Black and White The choice between MCP, REST/HTTP APIs, and Kafka is rarely clean. All three can play a role in the same architecture. REST/HTTP APIs work well for operational data access when volume is moderate and a well-governed API already exists. A REST API backed by a Kafka-derived serving layer can return consistent, current data. The API is the interface; the streaming platform is what makes the data trustworthy behind it. A financial services firm exposing account balances via REST is not doing it wrong, as long as those balances are derived from a governed, consistent data source rather than pulled directly from a source system on every request. Kafka becomes the clear choice when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when the same events need to feed operational applications, analytical pipelines, and AI agents simultaneously. MCP fits best when access is supplementary, loosely coupled, and low-frequency. A support agent looking up a ServiceNow ticket before drafting a response, or a sales assistant pulling the latest slide deck from Google Drive before a call, are good fits. The key test is simple: does it matter if the data the agent receives is a few seconds or minutes old? If yes, MCP should not own that responsibility. If no, MCP is the right interface. SAP: Clean Separation Between ERP Integration and Developer Tooling The boundary between MCP and REST is not a choice between two equivalent options for the same integration. SAP is the clearest example of a clean separation. SAP exposes extensive REST and OData APIs for ERP integration: order management, finance, supply chain, procurement, and HR data flowing bidirectionally between SAP and other enterprise systems. SAP's MCP servers serve an entirely different purpose: developer tooling for ABAP code generation, CAP application development, UI5 and Fiori assistance, and operational tasks like transport validation and incident management. An architect connecting SAP order events to downstream systems uses OData and Kafka Connect. A developer asking an AI coding assistant to generate ABAP code uses the SAP MCP server. Different consumers, different use cases, different data. No overlap. Salesforce and ServiceNow: Same Data, Different Consumer Salesforce and ServiceNow follow a different pattern. Their MCP servers wrap the same underlying REST APIs and expose the same underlying data, but for a different consumer. A developer-written integration calls the Salesforce REST API directly with known endpoints and hardcoded logic. An AI agent calls the Salesforce MCP server, which wraps that same API to make it discoverable and stateful for an agent that cannot read documentation or manage its own session state. The data is identical. The access path differs based on who is consuming it. This is not a free choice between equivalent options. It is the same system serving two different client types through two different interface layers. REST vs. Kafka for Operational Data: The Harder Call The harder boundary is between REST and Kafka for operational data. Both can technically serve it, and that is where the real architectural decision lies. REST is simpler to start with but introduces point-to-point coupling, integration spaghetti at scale, and consistency risks when the same data needs to reach multiple consumers. Kafka is more complex to operate but provides the decoupling, consistency, and governance that enterprise architectures require when the same data needs to reach many consumers reliably. The two are not mutually exclusive. A common and well-proven pattern combines both: Kafka handles the event backbone, decoupling, and consistency, while a REST layer sits on top for synchronous request-response access, API management integration, or compatibility with systems that cannot speak the native Kafka protocol. This is particularly common in mobile applications, legacy system integration, and API gateway architectures. For a detailed look at how REST and Kafka complement each other in practice, see Request-Response with REST/HTTP vs. Data Streaming with Apache Kafka. 5. Decision Framework: MCP, REST/HTTP, or Kafka? Choosing between MCP, REST/HTTP, and Kafka is not a single decision but a set of tradeoffs that depend on data volume, consumer type, consistency requirements, and what is already in production. The comparison table below makes those tradeoffs concrete across eight dimensions. When to Use Which: A Guide to the Decision Tree The decision tree below walks through the same logic as a series of questions, routing to the right choice based on the integration's actual requirements. Use MCP when the integration is supplementary and tool-like: Slack, Google Drive, ServiceNow tickets, internal knowledge bases. The agent needs context to act, not a stream of events to react to. Eventual consistency is acceptable. Apply least privilege, monitor tool definitions for changes, and isolate MCP servers from production systems. Use a REST/HTTP API or native SDK when a well-documented API or SDK already exists and the engineering team knows how to operate it. The access pattern is direct, moderate-volume, and latency-sensitive. REST is also a reasonable choice for operational data when the backend is a governed Kafka-derived serving layer and consistency properties are inherited, not assumed. Use Apache Kafka when data is high-volume or high-velocity, when multiple consumers need the same events, when ordering and exactly-once delivery matter, or when governance, lineage, and auditability are non-negotiable. Kafka is also the right choice when the same data needs to feed operational applications, real-time analytics, data lakes, and AI agents simultaneously. Use the real-time context engine when an AI agent needs current, consistent operational context for autonomous decisions. Kafka and Flink govern the data. MCP provides the agent interface. The consistency guarantee comes from the streaming layer, not from MCP. The practical question is not which protocol to choose. It is whether the data architecture underneath the agents can be trusted. Agents making autonomous decisions about inventory, risk, or customer service are only as reliable as the data they act on. 6. Where MCP and Kafka Work Together: The Real-Time Context Engine There is one pattern where MCP and data streaming complement each other directly: the real-time context engine. Kafka and Flink process and govern the data: ingesting from operational systems, applying transformations and filters, producing real-time materialized views. Those views are then exposed to AI agents through a standardized MCP interface. The streaming platform owns the data, its freshness, and its consistency guarantees. MCP owns the interface to the agent. Neither layer bleeds into the other's responsibility. Data consistency is not delegated to MCP. The streaming platform enforces it upstream before the MCP interface comes into play. The agent calls a tool and receives context that is current, governed, and consistent, not because MCP guarantees it, but because the streaming platform does. Any compliant AI agent, whether Claude, ChatGPT, Amazon Bedrock, LlamaIndex, or CrewAI, can call the context engine and receive current context from operational systems without needing to understand Kafka topics, Flink jobs, or schema evolution. An agent routing shipments from yesterday's inventory, approving transactions against a risk score from three hours ago, or reading an account balance that has not propagated: none of these is reliable. A real-time context engine eliminates this class of error at the source, reduces hallucinations, lowers inference cost, and anchors decisions to current operational reality. From Data Freshness to Agent Governance Enterprise readiness for this pattern also depends on how agents are governed once deployed. Trust, control, and accountability become central once agents start chaining decisions across domains. The context engine is the data layer of that answer. Governance of the agents themselves, covering what they are permitted to do, under what conditions, and with what audit trail, is the other half. This is the dimension enterprise buyers are actively evaluating when selecting agent orchestration platforms. The diagram below shows how the three layers fit together: the streaming platform as the data backbone, the context engine as the governed serving layer, and MCP as the clean interface to agents. 7. Conclusion: One Protocol, One Job MCP has earned its place in the enterprise architecture stack. What it has not yet earned is the role of universal integration layer, and understanding that distinction is what this article has been about. The broader architecture this sits inside connects three interdependent pillars. Event-driven data integration, with Kafka as the backbone, moves data reliably between operational and analytical systems and delivers governed data products to every consumer. Process intelligence is the orchestration layer that determines which decisions to automate, in what sequence, and under what conditions, giving agentic workflows the structure and governance they need to be trustworthy. Trusted agentic AI is where MCP plays its role: the standardized, governed interface through which agents access external tools and context, anchored to real data by the streaming layer beneath it. For a vendor-by-vendor analysis of trust and lock-in across the major AI platforms, see the Enterprise Agentic AI Landscape 2026. For a deeper look at how the three pillars fit together as an enterprise architecture framework, see The Trinity of Modern Data Architecture: Process Intelligence, Event-Driven Integration, and Trusted Agentic AI. One protocol, one job. That is the right way to use MCP.

By Kai Wähner DZone Core CORE

Monthly Top Tools Experts

expert thumbnail

Abhishek Gupta

Principal PM, Azure Cosmos DB,
Microsoft

I mostly work on open-source technologies including distributed data systems, Kubernetes and Go
expert thumbnail

Yitaek Hwang

Software Engineer,
NYDIG

The Latest Tools Topics

article thumbnail
Multi-Agent Orchestration on AWS With AgentCore Runtime and A2A
Multi-agent systems are common. Here we build a small multi-agent system on Amazon Bedrock AgentCore Runtime using the Agent-to-Agent (A2A) protocol.
October 9, 2026
by Purnanga Borah
· 369 Views · 1 Like
article thumbnail
Supercharging AI Agents with Azure Context: A Hands-On Guide to Azure MCP
Here’s a step-by-step guide for cloud engineers on how to bridge local AI assistants with live Azure infrastructure using the Model Context Protocol.
October 8, 2026
by Ammar Ekbote
· 571 Views · 1 Like
article thumbnail
AI Agents Leaked 13,000 Screenshots: Why Enterprise Approval Controls Failed
A reported leak of 13,000 screenshots shows how AI agents can bypass weak approval and audit controls even when organizations have written policies.
October 7, 2026
by Tim Freestone
· 769 Views · 1 Like
article thumbnail
Decoding the “Black Box”: Evaluating Agent Tool Chains in Production
Evaluating AI agents means checking the full trajectory, not just the answer — intent, tool calls, order, and accuracy — using Microsoft Foundry's evaluation tools.
October 7, 2026
by Gaurav Bhardwaj
· 686 Views · 1 Like
article thumbnail
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
This guide walks through custom Azure ML model training and deployment to a Managed Online Endpoint and connecting it to a Microsoft Foundry agent as a function tool.
October 6, 2026
by Jubin Soni, FBCS DZone Core CORE
· 4,983 Views · 3 Likes
article thumbnail
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices
Stop dual-write data inconsistencies. Learn how to architect the Transactional Outbox Pattern using Java, Spring Boot, and PostgreSQL for reliable Kafka events.
October 6, 2026
by Rahul Tewari
· 1,412 Views · 1 Like
article thumbnail
Wasm Inside Neo4j: Building the Example That Didn't Exist
A Rust VADER sentiment analyzer compiled to WebAssembly, embedded inside a Neo4j Java UDF, and callable directly from Cypher using wasmtime-java.
October 2, 2026
by Akmal Chaudhri DZone Core CORE
· 1,185 Views · 2 Likes
article thumbnail
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
In this article, we will discuss how to run your coding agents in the cloud using sbx. Cloud compute is usage-billed, so keep track of your sandboxes accordingly.
October 2, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 1,257 Views · 1 Like
article thumbnail
The Silent Container Death: A TCP Dial That Never Times Out
A pod goes into CrashLoopBackOff. You pull the logs expecting a stack trace, a panic, an error string — and then nothing. No error. No exit message. Magic.
September 30, 2026
by Alexander Fo
· 1,658 Views · 1 Like
article thumbnail
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
Learn about graph databases by building an F1 teammate network from real Formula 1 data and using Cypher to connect Max Verstappen to Juan Manuel Fangio.
September 30, 2026
by Jeremy Morgan
· 1,085 Views · 2 Likes
article thumbnail
AWS 7R Migration Strategies: A Decision Framework for Engineering Teams
Learn how to classify workloads, choose the right migration path, and avoid the traps that turn 6-week projects into 6-month ones.
September 30, 2026
by Jerzy Kopaczewski
· 960 Views · 2 Likes
article thumbnail
Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs Architecture
Databricks and Snowflake speak the Kafka protocol, but Kafka for lakehouse ingestion is not Kafka as an event-driven architecture.
September 30, 2026
by Kai Wähner DZone Core CORE
· 2,945 Views · 1 Like
article thumbnail
Git Blame Isn’t Enough: Building Verifiable Provenance for AI-Generated Code
AI-generated code needs verifiable provenance linking intent, context, models, edits, approvals, commits, and artifacts across the software lifecycle.
September 30, 2026
by Uthej Mopathi DZone Core CORE
· 1,201 Views · 3 Likes
article thumbnail
A Deep Dive into the Microsoft Foundry Document Intelligence SDK: From PDF to Structured Data
A hands-on guide to Microsoft Foundry Document Intelligence SDK for extracting text, structured fields, and document data from PDFs for RAG and AI pipelines.
September 29, 2026
by Jubin Soni, FBCS DZone Core CORE
· 6,301 Views · 3 Likes
article thumbnail
From Giant Prompts to On-Demand Skills: Build an Extensible AI Agent With Progressive Disclosure
Progressive disclosure replaces giant prompts with lightweight skill summaries and on-demand instructions, keeping AI agents focused and extensible.
September 28, 2026
by Akhil Madineni DZone Core CORE
· 1,152 Views · 3 Likes
article thumbnail
Building a Product Recommendation Engine With Neo4j — No ML Library Required
A graph models customers, products, categories and tags, making collaborative filtering, co-purchase analysis, content similarity, and trending queries graph traversals.
September 28, 2026
by Akmal Chaudhri DZone Core CORE
· 902 Views · 1 Like
article thumbnail
Building a Practical Cloud-Native Golden Path: A Guide to Kubernetes-Based Service Delivery, Self-Service, and Developer-Friendly Defaults
Golden paths standardize software delivery with self-service workflows, deployment guardrails, and observability while preserving team autonomy.
September 25, 2026
by Naga Santhosh Reddy Vootukuri DZone Core CORE
· 1,627 Views · 1 Like
article thumbnail
Kubernetes Operations Playbook: The Essentials for Keeping Scale, Complexity, and Drift Under Control
Kubernetes operations can drift as teams scale. Use this checklist to standardize clusters, releases, observability, access, reliability, and cost.
September 23, 2026
by Abhishek Gupta DZone Core CORE
· 2,567 Views · 1 Like
article thumbnail
Your Terraform Monolith Isn't Too Big. It's Tightly Coupled.
Terraform monoliths hurt when one state couples too many resources and owners. Split around ownership boundaries, not size.
September 22, 2026
by Naveen Kalapala
· 2,178 Views · 1 Like
article thumbnail
Architecting for <1s Latency: Managing Eventual Consistency in Distributed Search Platforms
To maintain sub-second search freshness, logistics systems must actively manage eventual consistency across Kafka ordering, search indexing, and cache invalidation.
September 22, 2026
by Dhruv Goel
· 1,988 Views · 2 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×