Classification Never Left. It Just Got a New Home in LLMs.
Building High-Performance Time-Series Applications With Java and QuestDB
Agentic AI Threat Intelligence Essentials
Getting Started With Agentic AI for SecOps
Hello DZone community! We’re refreshing our newsletters to help you follow the topics that matter most to your work, explore new ideas, and stay connected with the developer community. Our Zone newsletters are becoming six focused newsletters, with new names and related topics that were developed with input from some of our fantastic community members, and we’re excited to bring them to your inbox: Beyond the Rows: Big data and databasesMind the Model: AI for developers and engineersShip & Scale: Cloud, DevOps, performance, and AgileThe Attack Surface: SecurityDistributed by Design: Microservices, integration, and IoTCode & Craft: Java and web development Each topical newsletter will arrive twice a month, with editions scheduled on Tuesdays and Thursdays. Meet DZone Digest We’re also bringing DZone Daily and DZone Weekly together into DZone Digest, arriving every Wednesday. It will be your weekly roundup of articles and insights from across DZone. What Else Is Changing? All newsletters got a refreshed design, making room for the content you care about and more opportunities to discover upcoming events. We’re also opening newsletter subscriptions to everyone, including developers who aren’t DZone members. Our goal is to make DZone’s newsletters more useful, with clearer topic choices and a regular cadence that helps you keep learning. We’d love to hear from you: Which newsletter are you most interested in, and what topics would you like us to cover? Share your thoughts in the comments. You can subscribe to them here. (If you’re subscribed to any of our Zone newsletters, you’ll begin receiving the updated newsletter covering your topics. If you’re subscribed to DZone Daily or DZone Weekly, you’ll now receive DZone Digest.)
Large AI evaluations rarely fail at convenient boundaries. A batch may contain tens of thousands of prompts, retrieval cases, tool-use scenarios, or judge-model comparisons, and each case can involve expensive network calls plus result persistence. Restarting the entire batch after a worker crash wastes inference spend and can change the meaning of the run when model outputs are nondeterministic. Temporal provides durable orchestration, but durability alone does not create application-level checkpoints. A reliable controller needs an explicit recovery boundary: completed evaluation cases stay completed, retries restart from a durable cursor, and the workflow remains small enough to replay efficiently. Put the Retry Boundary Around Recoverable Work The key design choice is the unit that gets retried. Temporal requires Workflow code to remain deterministic, while Activities are the place for external API calls and other nondeterministic work. That separation fits evaluation systems naturally: the Workflow owns lifecycle and policy, while an Activity calls the model, loads test cases, writes results, and advances progress. Temporal describes Activities as the failure-prone, side-effecting layer and supports retries and heartbeats for long-running work. A naive Workflow can schedule one Activity for the whole evaluation and rely only on an Activity retry. Without checkpoint logic, a retry re-enters the Activity from its method boundary, making case zero the recovery point. Scheduling one Activity per case avoids that problem but can create a large Workflow Event History when the dataset is large. A better middle ground is a checkpoint-aware Activity that processes a bounded slice of cases and records progress after each committed result. The Workflow can keep retry behavior explicit rather than hiding it inside HTTP-client loops: Java ActivityOptions options = ActivityOptions.newBuilder() .setStartToCloseTimeout(Duration.ofHours(2)) .setHeartbeatTimeout(Duration.ofSeconds(30)) .setRetryOptions(RetryOptions.newBuilder() .setInitialInterval(Duration.ofSeconds(2)) .setBackoffCoefficient(2.0) .setMaximumInterval(Duration.ofMinutes(1)) .setMaximumAttempts(8) .build()) .build(); EvaluationSummary summary = Workflow.newActivityStub(EvaluationActivities.class, options) .evaluate(spec); StartToCloseTimeout bounds one Activity attempt, while the shorter heartbeat timeout lets Temporal notice a dead or disconnected worker much earlier. Temporal documents that a heartbeat timeout marks the Activity attempt as failed when heartbeats stop and allows another attempt to be scheduled according to the retry policy. Treat the Checkpoint as a Commit Record A cursor such as nextIndex = 3840 is useful only when every result before that cursor is durably committed. The safe ordering is result write, checkpoint advance, then heartbeat. Reversing that order creates a lost-result window: a heartbeat could claim that case 3839 finished even though the corresponding result never reached durable storage. When result rows and checkpoints share a transactional database, both persistence operations can be committed atomically. The evaluation identity also has to be immutable. A resumable run should pin the dataset snapshot, model identifier, inference parameters, prompt template, judge configuration, and scoring code version. Otherwise, a resumed process can silently mix results produced under different semantics. The checkpoint key should therefore be based on that immutable run identity rather than only on a human-readable batch name. A compact Activity implementation can reconcile Temporal's last heartbeat with an external authoritative checkpoint: Java public EvaluationSummary evaluate(EvalSpec spec) { ActivityExecutionContext ctx = Activity.getExecutionContext(); EvalCheckpoint heartbeat = ctx .getHeartbeatDetails(EvalCheckpoint.class) .orElse(EvalCheckpoint.start()); EvalCheckpoint durable = checkpoints.load(spec.runId()); int start = Math.max(heartbeat.nextIndex(), durable.nextIndex()); for (int index = start; index < spec.caseCount(); index++) { EvalCase testCase = corpus.load(spec.datasetVersion(), index); EvalResult result = evaluator.run(spec.modelVersion(), testCase); results.upsert(spec.runId(), testCase.id(), result); EvalCheckpoint saved = checkpoints.advanceIfGreater(spec.runId(), index + 1, testCase.id()); ctx.heartbeat(saved); } return results.summarize(spec.runId()); } The external checkpoint is authoritative because Temporal heartbeats are a liveness and progress mechanism, not a transactional commit log for evaluation output. Temporal notes that heartbeats can be throttled by the Worker, and progress recorded immediately before a failure is available to the next attempt only if that heartbeat reached the Temporal Service before the Worker crashed. The Math.max reconciliation tolerates a lagging heartbeat, while advanceIfGreater prevents an older attempt from moving the durable cursor backward. Make Duplicate Execution Harmless Checkpointing reduces repeated work but does not eliminate it. A worker can persist a model result and die before advancing the checkpoint, so the next attempt may execute the same case again. A timed-out attempt can also remain alive briefly from the perspective of an external system. The result store must therefore make duplicate case writes safe. An idempotency key derived from runId and caseId provides a clean boundary. The persistence operation should insert once or perform a deterministic upsert instead of appending a second record. If the model provider supports request-level idempotency, the same stable key can be propagated there; otherwise, the local result store still prevents duplicate scoring records. The invariant is that repeating a case changes no already-committed state. Checkpoint granularity then becomes an economic decision. Heartbeating after every case minimizes replay but increases heartbeat traffic. Heartbeating every small group reduces coordination overhead but increases the maximum repeated work after failure. Temporal Workers may throttle heartbeat delivery, so application correctness must never depend on every heartbeat being observed. External progress commits can still happen at case granularity even when heartbeats are less frequent. Heartbeats also provide a natural cancellation channel. Temporal delivers Activity cancellation through heartbeat calls, so a long evaluation loop that heartbeats regularly can stop promptly rather than continuing expensive inference after cancellation. The checkpoint remains intact, making an operator-initiated retry or later continuation predictable. Keep Long Evaluations Replayable Checkpoint-aware Activities solve worker failure, but very large or continuously running evaluations can still grow Workflow history. Temporal's Continue-As-New mechanism starts a new Workflow Execution with fresh Event History while carrying forward the latest relevant state; the Workflow ID remains the same and the Run ID changes. Temporal recommends the mechanism when history becomes large or when long-lived executions need a fresh execution boundary. That mechanism works best with small Workflow state. A continuation should carry identifiers and cursors, not thousands of model responses. Large artifacts belong in a database or object store, with the Workflow retaining only stable references such as runId, dataset version, current shard, and checkpoint token. Temporal has separately warned against keeping excessive data in Workflow state and recommends external storage for large data. For a sharded evaluation, each Activity can own a deterministic range such as cases 20,000 through 24,999. After a shard completes, the Workflow records only the shard result reference and schedules the next range. A continuation boundary can be taken after a suitable number of shards: Java if (Workflow.getInfo().isContinueAsNewSuggested()) { ContinueState next = new ContinueState( spec.runId(), nextShard, aggregateRef); Workflow.continueAsNew(spec, next); } Temporal's Java documentation exposes isContinueAsNewSuggested() so application code can checkpoint state at a safe point before history limits become a concern. This creates two complementary recovery layers: heartbeats and external checkpoints resume work inside an Activity, while Continue-As-New controls the lifetime and replay cost of the Workflow itself. Conclusion A reliable AI evaluation controller should treat recovery as a data-consistency problem rather than a generic retry problem. Temporal supplies durable execution, retry policies, heartbeat-based failure detection, cancellation delivery, and fresh-history continuation, but the application still defines what “completed” means. Persisting each case idempotently, advancing a monotonic external checkpoint only after that persistence succeeds, and heartbeating the resulting cursor turns a failed worker into a small replay window instead of a full-batch restart. Keeping immutable evaluation semantics attached to the run and using Continue-As-New for long histories preserves both correctness and operational efficiency. The result is a controller that resumes the evaluation at the last trustworthy commit point rather than paying again for work that has already finished.
Apache IoTDB is well-suited for environments where time-series data from connected devices and industrial systems is generated continuously and requires efficient querying. Typical use cases include monitoring temperature, pressure, vibration, energy usage, machine status, and device telemetry. This approach also applies to manufacturing, smart infrastructure, fleet monitoring, utilities, and edge computing. IoTDB stands out for its focus on large-scale time-series workloads from devices and industrial systems. It is purpose-built for high-frequency data ingestion, historical analysis, and time-based queries. This makes it ideal for applications that require insight into both the current state and historical trends of devices or processes. Why Apache IoTDB Matters Apache IoTDB excels at managing large volumes of device-generated data over time. It addresses issues beyond storage, such as continuous data ingestion, efficient organization by device and timestamp, and fast queries for both recent and historical values. This makes IoTDB well-suited for systems with temporal, device-oriented data models. Common use cases include industrial IoT, smart factories, energy monitoring, connected vehicles, building automation, predictive maintenance, and edge computing. For example, manufacturers can track machinery data, utilities can analyze energy consumption across many meters, and fleet operators can monitor location and engine telemetry. In each scenario, IoTDB enables applications to determine current status, review device behavior over time, and detect measurements outside expected ranges. Hands-On: Java With Apache IoTDB We will build a simple Java application using Apache IoTDB. For simplicity, IoTDB will run locally in Docker. Shell docker run -d \ --name iotdb-instance \ -p 6667:6667 \ -e dn_rpc_address=0.0.0.0 \ apache/iotdb:2.0.11-standalone Next, create the database using IoTDB’s Table SQL dialect: Shell docker exec -it iotdb-instance \ /iotdb/sbin/start-cli.sh \ -h 127.0.0.1 \ -p 6667 \ -u root \ -pw root \ -sql_dialect table \ -e "CREATE DATABASE IF NOT EXISTS jnosql" For this local example, we use the default root/root credentials and disable client redirection. These settings are suitable for single-node development only. In production, use secure credentials and configure redirection based on your IoTDB cluster topology. Configuring the Application Next, configure the connection in microprofile-config.properties: Properties files jnosql.timeseries.database=jnosql jnosql.iotdb.host=localhost jnosql.iotdb.port=6667 jnosql.iotdb.username=root jnosql.iotdb.password=root jnosql.iotdb.enable.redirection=false Eclipse JNoSQL is built on Jakarta APIs such as CDI, JSON-B, and Eclipse MicroProfile Config. These APIs are supported by popular runtimes including Helidon, Quarkus, Open Liberty, and other Jakarta EE-compatible environments. Then add the Apache IoTDB driver: XML <dependency> <groupId>org.eclipse.jnosql.databases</groupId> <artifactId>jnosql-iotdb</artifactId> <version>${jnosql.version}</version> </dependency> Modeling Sensor Data A sensor-oriented model aligns well with IoTDB: Java @Entity public class SensorReading { @Id private Instant id; @Column private String sensor; @Column private double temperature; @Column private double humidity; // constructors, getters, and setters } The Instant field records the timestamp of the reading. The other fields capture the measurements at that time. The entity can also be exposed through Jakarta Data: Java @Repository public interface SensorReadingRepository extends BasicRepository<SensorReading, Instant> { List<SensorReading> findBySensorOrderByIdDesc( String sensor, Limit limit); } Using TimeSeriesTemplate With the infrastructure in place, insert several readings and retrieve both the latest value and recent history: Java "sensor-01", 22.1, 46.5 ); var latestReading = new SensorReading( Instant.parse("2026-09-20T10:15:00Z"), "sensor-01", 23.6, 48.2 ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { TimeSeriesTemplate template = container.select(TimeSeriesTemplate.class).get(); template.insert(firstReading); template.insert(secondReading); template.insert(latestReading); var currentReading = template .select(SensorReading.class) .where("sensor") .eq("sensor-01") .orderBy("id") .desc() .limit(1) .singleResult(); System.out.println( "Current sensor reading: " + currentReading ); var history = template .select(SensorReading.class) .where("sensor") .eq("sensor-01") .orderBy("id") .desc() .skip(1) .limit(10) .result(); System.out.println("Recent sensor history:"); history.forEach(System.out::println); } } } The first query answers a common IoT question: what is the latest reading from this sensor? The second retrieves its recent history, excluding the latest observation. Using Jakarta Data The same use case can be implemented using the repository: Java public class App2 { public static void main(String[] args) { var firstReading = new SensorReading( Instant.parse("2026-09-20T08:00:00Z"), "sensor-01", 21.4, 45.0 ); var secondReading = new SensorReading( Instant.parse("2026-09-20T09:00:00Z"), "sensor-01", 22.1, 46.5 ); var latestReading = new SensorReading( Instant.parse("2026-09-20T10:15:00Z"), "sensor-01", 23.6, 48.2 ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { SensorReadingRepository repository = container.select(SensorReadingRepository.class).get(); repository.save(firstReading); repository.save(secondReading); repository.save(latestReading); var currentReading = repository .findBySensorOrderByIdDesc( "sensor-01", Limit.of(1) ) .stream() .findFirst(); System.out.println( "Current sensor reading: " + currentReading ); var history = repository .findBySensorOrderByIdDesc( "sensor-01", Limit.range(2, 10) ); System.out.println("Recent sensor history:"); history.forEach(System.out::println); } } } Both approaches model the domain in terms of latest state and historical observations. This is where a time-series database like Apache IoTDB is more effective than using a general-purpose database with timestamps as regular fields. Why Apache IoTDB Is More Than Sensor Storage While Apache IoTDB is often associated with IoT, its capabilities go well beyond basic sensor data storage. It excels when organizations must manage large volumes of device-centric data over time while retaining relationships among measurements, devices, and timestamps. As a result, IoTDB is valuable for industrial systems, utilities, transportation, smart infrastructure, manufacturing, and edge computing. For example, in predictive maintenance, a machine may continuously report vibration, temperature, pressure, energy consumption, and operating condition. The true value is not only found in the latest readings, but in studying how these measurements change together over time. Maintenance systems may compare recent data with historical trends, spot anomalies before failures, or correlate changes across multiple devices. This requires more than simple sensor data storage. This approach also benefits energy systems, where smart meters, solar panels, batteries, and grid equipment generate continuous measurement streams that require time-based analysis. In transportation, vehicles produce ongoing data such as speed, location, fuel consumption, battery status, and engine telemetry. Similarly, smart buildings rely on time-series data from HVAC systems, occupancy sensors, and energy meters. IoTDB is especially effective when system architecture is organized around devices and their measurements. It can serve as the historical data layer for operational systems, while other technologies manage transactional or business information. For example, a relational database might store details about a machine, its owner, contracts, or maintenance schedules, while IoTDB manages the extensive measurement data the machine generates. This separation lets enterprise applications treat telemetry as a primary workload, rather than forcing high-frequency device data into models meant for business entities and transactions. Conclusion Apache IoTDB is a strong fit for applications that need to ingest and analyze device-generated data over time, especially in IoT, industrial, energy, and edge scenarios. With Eclipse JNoSQL 1.1.18, Java developers can access these capabilities through familiar Jakarta APIs, keeping the application model consistent while still leveraging IoTDB’s time-series specialization.
OpenSearch is an open-source, distributed search and analytics suite derived as a fork of Elasticsearch and maintained under the Apache 2.0 license. When it comes to memory configuration, the guidance is often reduced to a few rules of thumb: swapoff -a, vm.swappiness=1, or bootstrap.memory_lock, and allocating 50% of available memory to the JVM heap while leaving the rest for Lucene and the filesystem page cache, OpenSearch off-heap caches, network buffers, and other system needs. These recommendations are repeated throughout documentation, blog posts, and operational guides, yet their origins and the mechanisms that justify these specific values are rarely examined. Undoubtedly, they provide a reasonable and safe starting point or a safe upper bound in most of the cases, but a safe default is not necessarily an optimal configuration. All of this raises even more questions. How do these defaults affect cluster performance? What is the optimal JVM heap ratio? Does memory given up by the JVM actually become filesystem page cache, and at what point does that trade-off stop paying off? How do read/write latency correlate with the heap ratio? These questions become particularly important in resource-constrained environments and in the cloud, where long-term contracts may make existing instances significantly cheaper, making horizontal or vertical scaling a difficult decision. In this article, we'll try to answer these questions through benchmarking. This is Act 1 of a two-act series. Act 1 focuses on identifying the cause of the latency problems we observed with the current defaults. Act 2 will explore what other heap-ratio values might look like for read/write loads. Along the way, I'll share the tools and commands used throughout the investigation, making this article a practical reference as well for you and for myself when I inevitably need to retrace the investigation months later. Knowledge Context The story also crosses several boundaries, such as the Kernel VM, the JVM, and Lucene. So, it’s important to outline the concepts mentioned in this part of the article beforehand, both for the context and, optionally, to enrich the AI context if you'd like to summarize everything. AreaWhere memory livesWhy it matters hereLinux page cacheFile-backed RAMLucene relies heavily on it for index data; under memory pressure, these pages can be reclaimed and read again later.Linux swapDisk-backed anonymous memoryAnonymous process memory can be swapped out under pressure. vm.swappiness influences this decision but does not prohibit it.Linux PSIKernel pressure signalShows time tasks spend stalled due to CPU, memory, or I/O pressure. We'll use I/O PSI while investigating latency.JVM heapAnonymous memoryControlled by Xms/Xmx; contains Java objects and several OpenSearch data structures.JVM native memoryAnonymous/file-backed memory outside XmxIncludes code cache, metaspace, stacks, direct buffers, and native allocations. Heap metrics do not account for all of it.OpenSearch cachesHeap/off-heap, depending on cacheTheir sizes may depend on heap size, which becomes important when we change jvm_heap_ratio.OpenSearch indexing bufferHeapIts size depends on heap and therefore becomes an important variable in Act 2.Lucene mmapFile-backed/page cacheLucene index files mapped into the process do not consume JVM heap; resident pages compete for physical RAM. Environment I used Aiven for OpenSearch on Azure, with cluster metrics exported to Thanos. The OpenSearch Benchmark metrics don't provide everything we need to answer our questions, particularly host-level metrics such as Linux PSI and swap activity. Exporting the cluster metrics to Thanos allows us to use PromQL queries later to retrieve the additional metrics needed for the investigation. The cluster consists of 3 nodes: CPUAMD EPYC 7763v (Milan)vCPU / RAM2 vCPU, 8 GiBDisk Size175 GiB per nodeAzure Regionazure-westeuropeAzure DiskPremiumV2_LRSAzure SKUStandard_D2as_v5 OpenSearch Version3.6.0JDKjava-21-openjdk-headlessGCG1GC Act 1. The Latency and an Extra GB The symptom: elevated query latency across the cluster, first reported by the customer after a kernel and Azure image upgrade. The load pattern on OpenSearch itself remained unchanged, as did the cluster configuration and settings. The monitoring panels give us the first clue. In the screenshots below, the green vertical line marks the moment of the upgrade. After that point, the page cache grows by roughly a gigabyte, while I/O PSI, previously close to zero, starts showing significant spikes. Nothing crashed, no alert fired, and from the JVM point of view everything looked normal. So where did that extra gigabyte of page cache come from? Nothing was actually freed. It moved. The interesting part isn't just that memory went to swap; it's which memory. When you have thousands of running clusters, there is always a small fraction of them operating close to the edge: relatively stable, yet sensitive enough that even a small change can noticeably affect performance. Like a star nearing the end of its lifetime, they may look stable right up until something disturbs the balance. The immediate cause of the page-cache change was identified fairly quickly: Azure applies tuning parameters that differ from the Linux kernel defaults, and those parameters were not applied by the older image. Once applied, the larger buffers and read_ahead increased the filesystem cache footprint, putting additional pressure on anonymous memory and eventually pushing some of it to swap. But rather than stopping there, let's use this incident as an opportunity to experiment with the heap ratio and make the behavior of OpenSearch instances explicit and less dependent on such environmental changes. Evidence It Is on Swap; None of It Locked First, find what is going on on a node itself: Shell PID=$(pgrep -f 'org.opensearch.bootstrap.OpenSearch') grep -E 'VmRSS|RssAnon|RssFile|RssShmem|VmSwap|VmLck' /proc/$PID/status Shell VmRSS: 4366408 kB # resident RssAnon: 3329896 kB # heap + anonymous native RssFile: 1036496 kB # resident mmap'd Lucene pages RssShmem: 16 kB VmSwap: 2679496 kB # on swap VmLck: 0 kB # bootstrap.memory_lock=fasle, none locked Shell grep -E 'MemFree|MemAvailable|Cached|SwapFree' /proc/meminfo Shell MemFree: 258216 kB MemAvailable: 3479264 kB Cached: 3403196 kB SwapCached: 854396 kB SwapFree: 4745444 kB The Swap Device Is dm-crypt Then confirm there is somewhere for it to go, and on what kind of device: Shell swapon --show Plain Text NAME TYPE SIZE USED PRIO /dev/dm-4 partition 8G 3.5G -1 The dm-* swap device is the interesting detail to catch and to keep in mind. This is an encrypted device, so once a page is requested it could drive more I/O -> more dm-crypt allocations -> more high-order pressure. A self-reinforcing loop and a good example of read amplification. Paging Is Live, Not Historical The next logical step is to check whether it is live paging or just a stale historical tail. The vmstat 1 5 the Linux Virtual Memory Statistics Tool should give us an exact answer for this, where non-zero swap blocks in and out (marked as si, so): Shell vmstat 1 5 Plain Text procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 1 0 3620928 159716 5864 3704212 113 80 3361 577 3147 12 8 7 84 1 0 0 0 0 3625240 189036 5860 3705260 0 5228 24268 5228 4841 4072 14 12 69 4 0 0 0 0 3625240 166196 5860 3710696 0 0 32 0 3572 2556 21 5 74 0 0 0 0 0 3625240 162668 5860 3716200 8 0 164 0 3729 2619 20 7 73 0 0 0 12 0 3625312 147692 5860 3724288 0 140 9207 524 3679 4024 22 14 63 1 0 0 Non-zero si/so in 4 of 5 samples show the live paging process, swpd also climbing across five seconds, which is good proof. The Swapped Pages Are Anonymous, Not File-Backed Let's also check swap memory consumption for each of the process's mappings, to confirm that swap is heap-related: Shell PID=$(pgrep -f org.opensearch.bootstrap.OpenSearch) awk '/^[0-9a-f]/{h=$0} /^Swap:/{if($2>0)print $2" kB "h}' /proc/$PID/smaps | sort -rn | head -20 Plain Text 1160380 kB 708400000-7ffe00000 rw-p 00000000 00:00 0 62296 kB 7f30c0000000-7f30c3f4b000 rw-p 00000000 00:00 0 61452 kB 7f1fc8000000-7f1fcbc03000 rw-p 00000000 00:00 0 61448 kB 7f1fa8000000-7f1fabc02000 rw-p 00000000 00:00 0 61444 kB 7f2f38000000-7f2f3bc01000 rw-p 00000000 00:00 0 61444 kB 7f24b4000000-7f24b7c01000 rw-p 00000000 00:00 0 61444 kB 7f22ac000000-7f22afc01000 rw-p 00000000 00:00 0 61444 kB 7f2238000000-7f223bc01000 rw-p 00000000 00:00 0 61444 kB 7f216c000000-7f216fc01000 rw-p 00000000 00:00 0 61444 kB 7f2168000000-7f216bc01000 rw-p 00000000 00:00 0 61444 kB 7f2148000000-7f214bc01000 rw-p 00000000 00:00 0 61444 kB 7f1fec000000-7f1fefc01000 rw-p 00000000 00:00 0 61444 kB 7f1fe8000000-7f1febc01000 rw-p 00000000 00:00 0 61444 kB 7f1fcc000000-7f1fcfc01000 rw-p 00000000 00:00 0 61444 kB 7f1fac000000-7f1fafc01000 rw-p 00000000 00:00 0 60772 kB 7f1fc0000000-7f1fc3c14000 rw-p 00000000 00:00 0 59952 kB 7f311e000000-7f3123f41000 rw-p 00000000 00:00 0 59340 kB 7f30b4000000-7f30b7c9b000 rw-p 00000000 00:00 0 57180 kB 7f3110000000-7f3113e3e000 rw-p 00000000 00:00 0 43084 kB 7f3130800000-7f3133d70000 rwxp 00000000 00:00 0 The Largest Swapped Region Is But why is what we are seeing above a heap-related area? There are a few clues for that. The region size 0x708400000 - 0x7ffe00000 is exactly 4,154,458,112 bytes = 3,962 MiB as we use -Xms == -Xmx and the whole thing is committed at startup, and nothing else in a JVM process is a single contiguous ~4 GB anonymous rw-p mapping. Second, It's the lowest mapping in the address space, smaps_rollup [rollup] line starts at exactly 708400000: Shell cat /proc/$PID/smaps_rollup Shell 708400000-7ffd3bb56000 ---p 00000000 00:00 0 [rollup] Private_Dirty: 3213788 kB Swap: 2680532 kB SwapPss: 2679412 kB Locked: 0 kB The JIT Code Cache Is Swapped Too Decoding the top swapped regions: 1160380 kB at 708400000 – is the JVM heap, the compressed‑oops heap base and matches the [rollup] start from smaps_rollup.The dozens of 61444 kB regions – these areas are probably related to native/off‑heap: Netty, JNI, Lucene native, etc.43084 kB marked rwxp – the JIT code cache, also swapped out, a bad sign. Together, these regions account for almost exactly the ~2.5 GB of swapped memory we observed earlier: cold heap regions, native/off-heap allocations, code cache, and possibly thread stacks. Practically, this means two consequences: GC can amplify swap latency. 1 GB of the JVM heap was swapped out. G1 does not necessarily touch all of those pages during a mixed collection, but any GC phase that accesses a swapped page incurs a major fault and has to bring it back through the dm-crypt device. Hence, short GC work can produce significantly longer pauses.A swapped page can also contain executable code. The next call into a swapped-out compiled method can trigger a major fault before the code can run. The resulting latency may land on an otherwise random request and be difficult to attribute directly to GC, index I/O, or the query itself. Evidence of Sustained Anon Memory Churn workingset_refault_anon counts anonymous memory refault events after reclaim; it does not count unique pages. Together with pswpin and pswpout, it shows how much anonymous memory paging has accumulated since boot. Shell grep -E 'workingset_(refault|activate)_anon|pswpin|pswpout' /proc/vmstat Plain Text workingset_refault_anon 89061837 workingset_activate_anon 1856069 pswpin 86690374 pswpout 61618603 These counters are cumulative, so fetching them at 10-minute intervals clearly shows that this wasn't a one-time eviction. There was sustained process: nearly 89 million anonymous pages were repeatedly swapped out and faulted back in. vm.swappiness = 1 and GC Amplification In this story vm.swappiness=1 was set since the cluster's inception. It does what Linux defines it to do, but it doesn't provide the protection we wanted. I suspect this matters particularly in the most resource-constrained deployments. How do we know that? The entire result above is a counterexample. swappiness biases the kernel's choice between reclaiming file-backed and anonymous pages. It does not prevent anonymous pages from being swapped out. Even at 1, this can still happen under sustained memory pressure. On a resource-constrained node whose index is several times larger than its RAM, pressure on the filesystem cache is not an exceptional condition; this is the normal operating state. This has two important consequences: It is not self-healing. Nothing proactively pages anonymous memory back in on a schedule. A swapped-out page returns to RAM only when it is accessed again, and cold memory, as it's defined, may remain untouched for a long time. As a result, the cold JVM memory can remain in swap indefinitely.It may remain invisible until something touches it. At steady state, a cold tail of the heap can remain in swap without producing obvious symptoms. The problem becomes visible when those pages are touched again, causing major page faults and potentially amplifying GC and request latency. Key Takeaways So, the root cause of the latency problems is an oversized heap combined with page cache pressure (the cold heap tail has been swapped out). vm.swappiness=1 did not protect the heap. It biases what gets reclaimed; it does not prevent anonymous memory from being swapped out.vm.swappiness=1 should not be relied on with the other defaults in production. An oversized heap can lead to GC amplification that is difficult to detect.Most of the swapped-out memory wasn't heap at all, but malloc arenas and the JIT code cache, none of it inside Xmx, so heap metrics didn't show it.Swap on dm-crypt exacerbates the issue, resulting in longer GC pauses and random request latency spikes. dm-crypt may be unavoidable in production due to security requirements. The fix isn't another swappiness tweak. It's two things: stop committing heap you don't use, and make the heap you do commit non-evictable. Follow Up In Act 2, we'll answer the remaining questions raised at the beginning of this article and look more closely at the trade-off introduced by bootstrap.memory_lock. This setting makes the heap resident and swap-immune, but it also turns jvm_heap_ratio from a soft default into a permanent memory commitment. The question then becomes: what heap ratio best suits different read and write workloads? There is one more complication to mention in advance: in a resource-constrained environment, merge storms can distort benchmark results, making an otherwise good heap ratio appear poor. See the screenshot below.
The Illusion of a Single Database In a traditional monolithic application, maintaining data consistency is straightforward. If you need to create a new order and update warehouse inventory, you wrap the logic inside a single database transaction: Java @Transactional public void placeOrder(OrderRequest request) { orderRepository.save(request.toOrder()); inventoryRepository.decrementStock(request.getItemId(), request.getQuantity()); } If the inventory update throws an exception, the relational database rolls back the entire transaction. Either both operations succeed, or neither does. In a distributed microservices architecture, that safety net disappears. When your OrderService writes a record to a local PostgreSQL database and immediately publishes an event to an Apache Kafka cluster to notify the InventoryService, you are dealing with two completely independent, non-atomic systems. The Catastrophic Failure Modes When you attempt to write to a local database and publish an event within the same business method, one of two failures will inevitably occur: Scenario A (Database First, Network Fails) [Save to Database: SUCCESS] ──> [Network / Kafka Outage: FAILS] Result: Order exists in database, but downstream services are never notified. Scenario B (Publish First, Database Fails) [Publish to Kafka: SUCCESS] ──> [Database Unique Constraint Violation: FAILS] Result: Downstream services charge payment or pack inventory for an order that was never saved. Distributed two-phase commit (2PC) protocols are notoriously slow, fragile, and rarely supported across modern cloud-native message brokers. To achieve guaranteed consistency without blocking throughput, the industry-standard architecture is the Transactional Outbox Pattern. The Blueprint: The Transactional Outbox Pattern Instead of trying to speak to two external systems at once, the microservice performs all its operations within a single, local database boundary. Plain Text [Incoming Request] │ ▼ ┌───────────────────────────────────────────────────────────┐ │ Local ACID Transaction │ │ ├── 1. Insert into orders table │ │ └── 2. Insert into outbox_events table │ └───────────────────────────────────────────────────────────┘ │ ▼ [Outbox Relay / Change Data Capture (CDC)] │ ▼ [Message Broker: Kafka Topic] Atomic local write: The application saves the domain entity (orders) and a corresponding event payload into an outbox_events table inside the exact same local @Transactional boundary.Asynchronous relay: An independent background process reads the outbox table and publishes the messages to Kafka.Acknowledgment: Once Kafka acknowledges receipt, the relay marks the outbox event as published or removes the row. 1. Database Schema for the Outbox Define a dedicated outbox table designed for high-throughput polling and sequential reads: SQL CREATE TABLE outbox_events ( id UUID PRIMARY KEY, aggregate_type VARCHAR(255) NOT NULL, aggregate_id VARCHAR(255) NOT NULL, event_type VARCHAR(255) NOT NULL, payload JSONB NOT NULL, created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP, processed BOOLEAN DEFAULT FALSE, processed_at TIMESTAMP WITH TIME ZONE ); CREATE INDEX idx_outbox_unprocessed ON outbox_events (created_at) WHERE processed = FALSE; 2. The Application Layer: Atomic Persistence The Spring service writes both the domain entity and the outbox event in one atomic operation: Java package com.example.outbox.service; import com.example.outbox.dto.OrderRequest; import com.example.outbox.entity.Order; import com.example.outbox.entity.OutboxEvent; import com.example.outbox.repository.OrderRepository; import com.example.outbox.repository.OutboxRepository; import com.fasterxml.jackson.databind.ObjectMapper; import org.springframework.stereotype.Service; import org.springframework.transaction.annotation.Transactional; import java.time.Instant; import java.util.UUID; @Service public class OrderService { private final OrderRepository orderRepository; private final OutboxRepository outboxRepository; private final ObjectMapper objectMapper; public OrderService(OrderRepository orderRepository, OutboxRepository outboxRepository, ObjectMapper objectMapper) { this.orderRepository = orderRepository; this.outboxRepository = outboxRepository; this.objectMapper = objectMapper; } @Transactional public void createOrder(OrderRequest request) { // 1. Persist domain entity Order order = new Order(UUID.randomUUID(), request.getCustomerId(), request.getTotalAmount()); orderRepository.save(order); // 2. Persist outbox event inside the exact same transaction try { String jsonPayload = objectMapper.writeValueAsString(order); OutboxEvent outbox = new OutboxEvent( UUID.randomUUID(), "ORDER", order.getId().toString(), "ORDER_CREATED", jsonPayload, Instant.now(), false ); outboxRepository.save(outbox); } catch (Exception e) { throw new RuntimeException("Failed to serialize outbox event payload", e); } } } 3. The Relay Layer: Reliable Kafka Publishing An asynchronous background worker polls the unprocessed records in the outbox and delivers them to the broker: Java package com.example.outbox.relay; import com.example.outbox.entity.OutboxEvent; import com.example.outbox.repository.OutboxRepository; import org.slf4j.Logger; import org.slf4j.LoggerFactory; import org.springframework.kafka.core.KafkaTemplate; import org.springframework.scheduling.annotation.Scheduled; import org.springframework.stereotype.Component; import org.springframework.transaction.annotation.Transactional; import java.time.Instant; import java.util.List; @Component public class OutboxMessageRelay { private static final Logger log = LoggerFactory.getLogger(OutboxMessageRelay.class); private final OutboxRepository outboxRepository; private final KafkaTemplate<String, String> kafkaTemplate; public OutboxMessageRelay(OutboxRepository outboxRepository, KafkaTemplate<String, String> kafkaTemplate) { this.outboxRepository = outboxRepository; this.kafkaTemplate = kafkaTemplate; } @Scheduled(fixedDelayString = "${app.outbox.poll-interval-ms:2000}") @Transactional public void publishPendingEvents() { List<OutboxEvent> pendingEvents = outboxRepository.findTop50ByProcessedFalseOrderByCreatedAtAsc(); for (OutboxEvent event : pendingEvents) { try { // Publish using aggregateId as the Kafka partition key to preserve message ordering kafkaTemplate.send("orders-events", event.getAggregateId(), event.getPayload()) .whenComplete((result, ex) -> { if (ex == null) { event.setProcessed(true); event.setProcessedAt(Instant.now()); outboxRepository.save(event); log.info("Successfully relayed outbox event: {}", event.getId()); } else { log.error("Failed to relay event to Kafka: {}", event.getId(), ex); } }); } catch (Exception e) { log.error("Synchronous dispatch failure for outbox event: {}", event.getId(), e); break; // Halt batch progression to preserve order } } } } 4. Production Engineering Guardrails Preserve Partition Ordering Always use the aggregate_id (such as order_id) as the partition key when publishing the Kafka message. This ensures all state transitions for a single business entity land on the exact same Kafka partition and are processed in strict sequence. Polling vs. Log-Based CDC (Debezium) Scheduled polling is easy to set up and ideal for small-to-medium systems. For high-volume enterprise platforms processing thousands of writes per second, replace database polling with Change Data Capture (CDC) engines like Debezium. Debezium reads the database write-ahead log (WAL) directly, streaming changes to Kafka with zero query overhead on the operational database. Idempotency on the Consumer The Transactional Outbox pattern guarantees at-least-once delivery. If the relay publishes an event to Kafka but crashes before marking the outbox row as processed, it may re-send the message upon restart. Downstream consumers must maintain an idempotency check (e.g., tracking processed message IDs in Redis) to discard duplicates. Architectural Strategy Matrix Dimension Dual-Write (@Transactional + Kafka) Distributed 2PC Transactional Outbox Pattern Data Consistency Broken (Silent inconsistencies) Strong Strong (Eventual consistency) System Latency Low High (Blocking locks) Ultra-low local execution Broker Resilience Fragile (Network crashes drop events) Low High (Decoupled publishing) Operational Simplicity Deceptively simple Complex Straightforward Summary In distributed systems, atomicity cannot cross network boundaries. When you attempt to update a local database and publish a message to an event bus inside the same method, failure is a mathematical certainty over time. By shifting to the Transactional Outbox Pattern, you leverage the battle-tested ACID guarantees of your relational database to capture domain state and outbound events simultaneously. This eliminates the dual-write anti-pattern, guarantees at-least-once delivery, and builds a dependable bridge between relational transactions and event-driven architecture.
Most enterprise AI programs do not fail because the model is too weak. They stall because the data underneath the model is fragmented, delayed, poorly documented, or too expensive to access repeatedly. The common response is to propose a complete platform replacement. That sounds clean on a diagram and becomes dangerous in production. Existing warehouses often support financial reporting, operational dashboards, manufacturing analytics, and regulatory processes that cannot pause while a new AI platform is assembled. The better strategy is to build an AI-ready data layer around stable business contracts. The organization modernizes how data is stored, governed, observed, and served without forcing every existing consumer to migrate at once. AI Readiness Is a Data Contract Problem An AI-ready platform needs more than raw data in inexpensive storage. It needs trusted definitions, reproducible history, fresh operational signals, discoverable lineage, and predictable query behavior. A model trained on an ambiguous customer identifier or an unstable product hierarchy will produce unstable results regardless of model quality. Start by identifying the business entities that must remain consistent across old and new systems. Typical examples include customer, product, supplier, order, invoice, material, and account. Define a canonical contract for each entity before choosing the final storage engine. The contract should specify field names, data types, ownership, accepted values, freshness expectations, and compatibility rules. It should also separate business meaning from physical implementation. A field can move from a legacy warehouse to an open table format without forcing downstream users to learn a new definition. Here is a simplified contract expressed as YAML: YAML entity: product_component owner: supply_chain_data primary_key: product_id, component_id, effective_from freshness: maximum_delay_minutes: 30 fields: product_id: {type: string, nullable: false} component_id: {type: string, nullable: false} quantity: {type: decimal, nullable: false} effective_from: {type: timestamp, nullable: false} effective_to: {type: timestamp, nullable: true} compatibility: additive_columns: allowed destructive_changes: require_new_version This small document does something important. It gives legacy reports, data pipelines, and AI applications the same definition to depend on. Build an Abstraction Layer Before Moving Consumers The riskiest migration pattern moves data and consumers at the same time. When a report changes after cutover, the team cannot easily tell whether the problem came from extraction, transformation, business logic, or presentation. Instead, create a stable semantic or compatibility layer between consumers and physical tables. Existing reports continue reading familiar columns while the implementation behind the view changes gradually. SQL CREATE VIEW analytics.product_component_current AS SELECT product_id, component_id, CAST(quantity AS DECIMAL(18, 4)) AS quantity, effective_from, effective_to FROM modern_layer.product_component WHERE is_current = TRUE; The view is intentionally boring. That is a strength. It preserves a contract while engineers replace ingestion, storage, and transformation components behind it. During transition, the same interface can point to the legacy source, the modern source, or a reconciled combination. Consumers migrate when the new path is proven, not when the infrastructure team finishes installing it. Use Layering to Separate Ingestion From Business Meaning A practical architecture separates raw ingestion, normalized data, and business-ready models. The names are less important than the boundaries. The ingestion layer preserves source fidelity and arrival metadata. The normalized layer resolves types, keys, duplicates, and schema differences. The business layer applies reusable definitions for reporting, features, and AI retrieval. This separation prevents source-system changes from leaking directly into AI applications. It also allows the same governed business model to support batch analytics, streaming decisions, feature engineering, and retrieval-augmented generation. Open table formats can help because they support schema evolution, snapshot history, and rollback. The Apache Iceberg documentation explains how column additions, renames, and partition changes can occur as metadata operations without rewriting every historical file. Those capabilities are useful, but they do not replace contracts. A technically valid schema change can still break business meaning. Run Both Paths and Reconcile Continuously Dual running is not wasted infrastructure. It is how teams prove that a modern data layer is safe. For a defined period, execute legacy and modern pipelines from the same source data. Compare row counts, key coverage, financial totals, null rates, duplicate rates, and business-specific invariants. Do not rely only on aggregate equality, because two incorrect datasets can produce the same total. Python def compare_snapshots(legacy, modern): checks = { "row_count": legacy.count() == modern.count(), "key_coverage": legacy.keys() == modern.keys(), "amount_total": abs(legacy.sum("amount") - modern.sum("amount")) < 0.01, "duplicate_keys": modern.duplicate_count() == 0, } failed = name for name, passed in checks.items() if not passed if failed: raise ValueError(f"Reconciliation failed: {failed}") Real implementations need tolerance rules, exception handling, and audit records, but the principle remains simple. A migration is complete only when correctness is demonstrated repeatedly across normal operations, period close, late-arriving data, and recovery scenarios. Make Lineage and Observability Part of the Product AI systems often combine data from many pipelines. When an answer changes, teams need to know which source, transformation, or model version caused it. Capture lineage at execution time rather than asking engineers to document it later. The OpenLineage specification defines interoperable metadata around datasets, jobs, and runs. Whether a team adopts that standard or another approach, the important point is to connect every published dataset to its inputs, code version, execution, owner, and quality results. Monitor the data layer with service-level objectives. Useful signals include freshness delay, failed contract checks, schema drift, incomplete partitions, reconciliation differences, query latency, and cost per workload. Infrastructure uptime alone is not enough. A pipeline can be running while delivering yesterday's data or silently dropping a critical field. Add Real-Time Access Only Where the Decision Requires It AI readiness is often confused with making everything real time. That creates unnecessary cost and operational complexity. Classify datasets by decision latency. Fraud detection or equipment monitoring may need event-level updates. Product recommendations may accept a few minutes of delay. Financial reporting may prioritize completeness and controlled closing over speed. When streaming is justified, design for replay, idempotency, and explicit processing guarantees. The Apache Kafka Streams documentation describes transactional and idempotent processing for exactly-once behavior within supported read-process-write flows. Teams still need to test external side effects and recovery paths rather than assuming one configuration solves end-to-end correctness. Control Cost Through Workload Isolation Legacy warehouses often mix ingestion, transformation, dashboards, experiments, and ad hoc queries in one shared resource pool. AI adds expensive feature generation, embedding creation, and large scans to that competition. Separate workloads by purpose and apply budgets, concurrency limits, caching, and retention policies independently. Store reusable features and business models once instead of recomputing them in every notebook. Track cost by dataset and workload so teams can see whether freshness or model accuracy justifies the additional compute. Predictable cost is part of the data contract. A dataset that is technically available but economically impractical to query is not AI-ready. Modernize by Proving One Business Slice Do not begin with the entire enterprise. Choose one domain with meaningful AI potential and stable business ownership. Product structures, customer identity, inventory, or service events are common candidates. Build the canonical contract, abstraction layer, modern pipeline, reconciliation suite, lineage, and cost controls for that slice. Keep existing reports working. Then connect one AI use case to the same governed layer. The result becomes a reusable migration pattern. Future domains inherit working templates for contracts, quality checks, dual runs, observability, and cutover. Modernization accelerates because the organization is no longer debating the architecture from scratch. An AI-ready data layer is not a separate platform waiting for the enterprise to catch up. It is a controlled evolution of the enterprise data system itself. The safest path keeps trusted analytics running while gradually replacing the foundations beneath them.
Today, every vendor offering BI solutions has incorporated a chat box. Whether you use Copilot or some other natural-language interface that connects you to a data warehouse, just ask a question in simple words, and it will generate SQL automatically. While this is conducive to productivity in other industries, in banking it presents an opportunity for a new attack. It is not enough to simply say that wrong SQL can be produced. It’s that an ungoverned text-to-SQL layer may join tables it shouldn’t, return columns that should have been masked. As a result of a lack of oversight, a marketing analyst could receive a query containing raw account numbers, since none of the components of the stack told the system to do otherwise. Prompt-level guardrails (“please don’t show PII”) are not a security control. They’re just a suggestion, and a model under adversarial pressure ot just a confusing prompt will ignore a suggestion. The issue isn't just about providing a more intelligent prompt; rather, it's about placing the artificial intelligence assistant on the same layer of information as human analysts, meaning an environment where the database manages the relevant security details as per row and column criteria instead of relying on technology. Consequently, if the analyst does not have access to the specific column, there is no justification for the AI assistant to have access to it as well. The diagram below (Figure 1) shows the steps taken to build that layer in BigQuery: a validated semantic layer that allows human dashboards and AI-generated queries to be connected to the same quality definitions. This would allow the assistant to leverage existing security rather than creating it. Figure 1. The Semantic Layer Resolver Step 1: Stop Letting Anyone (Human or AI) Query Raw Tables The initial phase is architectural, not related to AI: there is no query made by a person or a system involving the base tables. Instead, each of the metrics that are accessible to consumers is defined at least once in a BigQuery view, and its calculation logic is embedded in that view. SQL -- Certified metric: Risk-Weighted Assets, defined once, queried everywhere CREATE VIEW analytics.risk_weighted_assets AS SELECT exposure.customer_id, exposure.region, exposure.exposure_class, exposure.outstanding_balance, risk_weights.weight_pct, ROUND(exposure.outstanding_balance * risk_weights.weight_pct / 100, 2) AS rwa_amount, CURRENT_TIMESTAMP() AS calculated_at FROM finance.exposures AS exposure JOIN reference.basel_risk_weights AS risk_weights ON exposure.exposure_class = risk_weights.exposure_class WHERE exposure.status = 'ACTIVE'; The view of "risk-weighted assets" created through a dashboard, a notebook, and an LLM agent is identical. There is no alternative version in a researcher’s spreadsheet, nor can an AI agent "helpfully" recreate the calculation based on exposure tables but use incorrect risk weightings. Step 2: Enforce Security at the Data Layer, Not the Application Layer It is important to ensure that BigQuery includes row-level and column-level security and associates it with the table. This means the principle will work irrespective of the entity making the query. SQL -- Row-level security: a regional analyst only ever sees their region's rows CREATE ROW ACCESS POLICY regional_filter ON analytics.risk_weighted_assets GRANT TO ('group:[email protected]') FILTER USING (region = 'EMEA'); Column masking works the same way, through policy tags rather than per-report logic: YAML # Dataplex policy tag: applied once, enforced everywhere the column is queried taxonomy: financial-pii policyTags: - displayName: "customer-account-number" description: "Masked for all roles except fraud-investigation" - displayName: "customer-ssn" description: "Masked for all roles except compliance-audit" Once a policy tag is applied to a column, a user, or any AI agent acting under that user's identity, who doesn’t have the appropriate fine-grained reader role, will receive either a null value or a hashed value. There is no mistake that the model has to circumvent; it’s simply a different result. This is what makes querying with AI safe, since whatever the query for the AI is, it cannot reveal anything that the column policy prohibits. Step 3: Give the Grounding Layer Metadata to Query Against It is impossible for an LLM to adhere to rules of governance it knows nothing about. Accordingly, a metadata directory is necessary for the semantic layer that contains a description of each certified metric with enough detail for the agent to turn an inquiry posed in natural language into the correct interpretation and filtering process, not simply provide it with a raw schema dump. JSON { "metric_id": "risk_weighted_assets", "display_name": "Risk-Weighted Assets", "view": "analytics.risk_weighted_assets", "owner": "[email protected]", "sensitivity": "internal", "allowed_dimensions": ["region", "exposure_class", "customer_id"], "definition": "Balance times Basel risk weight, summed by class.", "lineage": ["finance.exposures", "reference.basel_risk_weights"], "last_certified": "2026-06-01" } This record is the thing the AI agent actually reads. The document specifies which view will be interrogated, lists the dimensions available for filtering results, and identifies who to contact if something goes wrong. It is worth mentioning that in this record there is no schema given for the finance exposes table, which leaves the model nothing to "discover" about. Step 4: Route Natural-Language Requests Through the Semantic Layer, Not the Warehouse When the certified metrics with their metadata have been obtained, the resolution process consists of transforming the user's natural-language question into a query that uses an allowed view rather than directly referring to the underlying schema. Python class SemanticLayerResolver: def __init__(self, metric_catalog, bq_client): self.catalog = metric_catalog # metric_id -> metadata, from Step 3 self.bq_client = bq_client def resolve(self, nl_request: str, user_identity: str) -> QueryResult: # 1. Map the request to a certified metric, never to a raw table. # A constrained classifier over self.catalog.keys() works better # here than open-ended text-to-SQL against the full warehouse. metric = self.match_metric(nl_request) if metric is None: return QueryResult.refuse("No certified metric found.") # 2. Extract filters, restricted to the metric's allowed_dimensions. filters = self.extract_filters( nl_request, metric["allowed_dimensions"] ) # 3. Build SQL against the certified view only. sql = self.build_query(metric["view"], filters) # 4. Execute as the requesting user, so BigQuery's row/column # security applies exactly as it would for a human query. result = self.bq_client.query(sql, user=user_identity) # 5. Attach lineage and certification metadata to the answer, # so "what the AI said" is auditable like any report. return QueryResult( data=result, metric_id=metric["metric_id"], lineage=metric["lineage"], certified_at=metric["last_certified"], ) The important line is step 4: the query is executed under the requesting user instead of using a shared service account. Hence, all the downstream access control mechanisms are automatically applied. There is no need for a separate permission system for the resolver because it has no access rights that exceed the rights of the requesting user. Step 5: Audit Every AI-Generated Query Like You Would a Human's The governance teams will not agree on a system based on the suggestion of " having faith in the model." What they approve is proof in every case where a resolver has provided information, just as is done when an individual writes a report. Python def log_ai_query(user_identity, nl_request, result: QueryResult): audit_log.write({ "user": user_identity, "request": nl_request, "metric_id": result.metric_id, "lineage": result.lineage, "policy_version": result.certified_at, "row_count": result.row_count, "timestamp": now(), }) One financial institution successfully applied this approach. What used to be a lengthy project in which one would have to analyze whether an AI assistant could access customer information has been transformed into something evaluated right away: the assistant can perform the same functions as a human worker. The financial institution was also measuring the new trend of using a certified semantic layer, not only in regard to the AI being discussed. Conflicts over defining metrics across different business lines practically vanished when the organization no longer had to create a separate “AI-compliant” data model. The Real Insight: Governance Is What Makes AI Fast, Not What Slows It Down It’s easy to assume that the best approach to deal with the LLM and sensitive data combination is to include a review step in which a human sits in on every step of the process, or another model is deployed to analyze the first model’s outputs before they are used. This is not only unscalable, but it also misses the point. Another way is to make sure that the insecure path cannot be taken, rather than simply being shunned. If the data layer implements row-level security, column masking, and certified metric definitions, you can confirm that an AI agent querying the data cannot generate queries that reveal any data previously available to someone with the same role. As a result, there is no need to verify output against constraints, since they were already included in the model. This shift is suggested by this pattern. Governed self-service, the architecture that permits a business analyst to carry out data initiatives safely in the absence of ticket submission, also creates a secure basis for AI-enhanced analysis. But it wasn't the main purpose. It is just a coincidence that it has worked out this way.
Testing how a WebRTC application behaves when the network fails is usually done by hand: open two browser windows, turn off Wi-Fi, count to ten, turn it back on, watch what happens. It works, after a fashion. It is also unrepeatable, untimed, impossible to run unattended, and useless for comparing two implementations — the interruption is never quite the same twice, and nothing records what actually occurred. I hit this while testing recovery behavior across browser engines. I needed the same interruption, applied at the same point in the connection lifecycle, repeated dozens of times, producing machine-readable output. Getting there took longer than expected, mostly because the obvious approach doesn't do what it appears to. This article describes the approach that worked, the failure modes that shaped it, and the parts worth reusing. The harness is open source; the code, raw trial records, and environment metadata are linked at the end. The Obvious Approach Is Misleading Playwright exposes BrowserContext.setOffline(true), which is the natural first reach: JavaScript await context.setOffline(true); It does something — just not the thing you want. setOffline operates at the network layer Playwright controls, so it reliably kills HTTP requests and WebSocket connections. Your signaling channel dies immediately, which looks convincing in the logs. What it does not reliably do is break an established peer connection whose media path runs over loopback or the local network. ICE keeps exchanging traffic and connectionState stays connected. If you are testing signaling recovery, this is a legitimate tool. If you are testing what an application does when the media path fails, you are testing nothing, and the dead signaling channel makes it look like you are. There is a second-order problem. Because setOffline blocks the page's WebSocket, any signaling that has to survive the outage — an ICE restart offer, for instance — cannot travel over the browser's own connection. In my harness, this forced signaling to be bridged through the Node process driving the test rather than the page, so that offers and answers could still cross between peers while one of them was offline from the browser's point of view. That is a workaround for a tool limitation, not a property of WebRTC, and it is worth knowing before you build around it. Make the Media Path Something You Control The alternative is to stop trying to break the network and instead force all traffic through a process you own. Setting iceTransportPolicy: "relay" in the RTCConfiguration restricts candidate gathering to relay candidates only — host and server-reflexive candidates are excluded from the pool entirely (MDN). Point that relays at a coturn instance on the local machine, and the entire media path now depends on one process you can signal. JavaScript const pc = new RTCPeerConnection({ iceTransportPolicy: "relay", iceServers: [{ urls: [ "turn:127.0.0.1:65050?transport=udp", "turn:127.0.0.1:65050?transport=tcp" ], username: "harness", credential: "harnesssecret" }] }); Interrupting the connection then becomes process control: Shell kill -STOP $TURN_PID # relay stops forwarding kill -CONT $TURN_PID # relay resumes SIGSTOP rather than SIGTERM matters here. The process is suspended, not terminated, so it stops relaying immediately while retaining its port bindings and allocation state. SIGCONT resumes it in place, with no restart and no re-binding race. The property that makes this useful for experiments is that the outage is *parameterized*. "Restore the relay three seconds after the connection reports failed" becomes a variable rather than a stopwatch and good intentions. Running the same interruption at one, three, and five seconds across two browsers is then just a loop. Choose the Method Once and Record It A harness that silently falls back between interruption methods produces data you cannot interpret. If trial 7 paused a relay and trial 8 called setOffline, the comparison is meaningless — and you will not know unless the selection is recorded. Select once at startup, log what was available alongside what was chosen, and write the selection into every trial record: JSON { "selected": "host_coturn_sigstop", "fallbackUsed": false, "attempted": [ { "available": true, "method": "host_coturn_stop_forward", "reason": "turnserver found at /opt/homebrew/opt/coturn/bin/turnserver" }, { "available": false, "method": "os_wifi_power", "reason": "ALLOW_WIFI_TOGGLE not set to 1" }, { "available": true, "method": "playwright_offline", "reason": "always available (may not break loopback WebRTC)" } ], "iceTransportPolicy": "relay" } The os_wifi_power entry deserves a comment. Toggling the actual interface via networksetup -setairportpower is closer to a real network event than pausing a relay, and is therefore a better test in principle. It is gated behind an explicit environment variable because a test suite that disables your machine's Wi-Fi without asking is a poor citizen, particularly in CI. Preflight, and Fail Loudly A long matrix run that breaks on trial 3 and produces 57 rows of garbage is worse than one that refuses to start. Before any measured trials, verify that each browser can reach every state the experiment depends on, and record what was observed: JSON { "engine": "chromium", "browserVersion": "151.0.7922.34", "steps": [ { "step": "ice_disconnected", "ok": true, "waitedMs": 5029 }, { "step": "ice_failed", "ok": true, "waitedMs": 9992 }, { "step": "full_reconnect", "ok": true, "elapsedMs": 145 } ] } The preflight also aborts hard on one specific error class: SDP negotiation failures, m-line ordering errors, and InvalidAccessError. These indicate the harness is broken rather than the connection under test, and allowing them through contaminates every downstream result. A run that does not finish with zero of these should be discarded. The preflight numbers are themselves a result. Under this interruption method, Chrome reached disconnected at 5.0 s and failed at 10.0 s; Firefox took 11.2 s and 19.9 s. Both engines are working within the consent-freshness bounds described in RFC 7675, which sets a 30-second consent expiry with checks roughly every five seconds — but the specific thresholds differ, and that difference is only visible because the interruption is byte-for-byte identical across engines. Manual Wi-Fi toggling cannot produce that comparison. Attribute Recovery to the Right Connection This is the detail that determines whether the results mean anything, and it is easy to omit. When testing whether a connection recovers, "the session is connected again" is not the same claim as "the original connection recovered." A freshly built RTCPeerConnection reaching connected is indistinguishable, in connectionState alone, from the original one returning. Counting both as recovery measures nothing. Tag every peer connection at construction and compare identity at the moment of recovery: JavaScript metrics.originalPeer = (pcInstanceId === iceRestartPcInstanceId); An epoch counter addresses the related hazard: callbacks from a torn-down connection fire after its replacement exists, and without a guard they write into the current trial's record. JavaScript if (epochAtStart !== pcEpoch) return; // stale callback, ignore Neither guard is exotic. Both are the difference between a dataset and a pile of numbers. Not Every Failure Test Needs a Network One of the two experiments in my harness tests whether a superseded recovery cycle can proceed into a destructive rebuild. That is a concurrency question, not a connectivity one, so it requires no interruption at all — an artificial delay forces the overlap window deterministically. The result runs identically anywhere, produces the same outcome every time, and completes in seconds. Before building network machinery, it is worth checking which of your failure scenarios are actually about the network. What the Setup Produces The stack is Node with Playwright driving two browser contexts, a small WebSocket signaling server, and coturn as the controllable relay. Each trial emits a JSON file with a timestamped event log covering connectionState, iceConnectionState, signalingState, restart and rebuild boundaries, cycle identifiers, and final outcome — plus a CSV summary row. Alongside those, the run captures an environment record: OS and architecture, Node and Playwright versions, both browser versions, ICE configuration, and the selected interruption method. That last file matters more than it sounds. The distance between "I tested this, and it worked" and a result someone else can check lies almost entirely in whether the environment was recorded next to the numbers. Two Things I Would Do Differently The environment record did not capture a git commit hash, so published results cannot be tied to an exact code revision. An obvious gap in hindsight, and the first thing I would add. More substantively: pausing a TURN relay is not a network interface change. It tests relay failure specifically. The engine timings above may not transfer to a genuine Wi-Fi-to-cellular handoff, where interface teardown, address changes, and gathering behavior all differ. The method buys reproducibility at the cost of realism, and that trade should be stated rather than discovered by a reader. Running It Yourself The harness is at github.com/jaynirmal15/webrtc-recovery-harness under an MIT license. The raw trial records, environment metadata, and preflight output from the runs described above are in the results/ directory. If you are testing recovery behavior in your own application, the interruption machinery is the part worth lifting.
When model providers announced 1-million-token and multi-million-token context windows, the software engineering world celebrated. The immediate narrative was simple and appealing: document chunking is dead, complex RAG pipelines are obsolete, and developers can now dump entire codebases, legal libraries, or multi-year enterprise datasets into a single model call. This breakthrough led engineering teams into a new trap: assuming that a model's capacity to accept context equals its ability to reason effectively over that context. In production, relying on massive context windows as a substitute for intelligent information retrieval leads to severe performance degradation, runaway infrastructure bills, and unpredictable hallucinations. Operating large-scale LLM architectures taught me a hard truth: A larger context window gives a model more surface area to get confused. Before you expand your context length, you must optimize your context quality. The Hidden Trap: "Needle in a Haystack" and Attention Degradation In traditional database systems, doubling the size of a query payload doesn't degrade the accuracy of the returned records. Relational algebra operates on exact logic. Large language models do not query data; they process spatial attention. As prompt lengths scale into hundreds of thousands of tokens, the self-attention mechanisms inside the transformer architecture begin to struggle. Key information buried deep in the middle of a massive prompt often suffers from "Lost in the Middle" syndrome, where the model strongly attends to tokens at the very beginning and very end of the prompt while ignoring crucial details placed in between. Diagram of Attention Degradation in Large Context Windows When an application fails under a massive context window, the model rarely throws an error. Instead, it silently ignores conflicting constraints, merges unrelated concepts, or produces plausible-sounding responses derived from irrelevant sections of your input. Step 1: Measure Prompt Noise-to-Signal Ratio In simple LLM integrations, teams measure throughput and token count. When scaling large-context applications, you must measure context density, the proportion of tokens directly relevant to the user's intent versus the filler text passed into the prompt. During a recent enterprise project, I built a utility to compute contextual relevance scores before submitting payloads to ultra-long context models: Python import logging from typing import List logger = logging.getLogger("ContextOptimizer") class ContextDensityAnalyzer: def __init__(self, key_terms: List[str]): self.key_terms = [term.lower() for term in key_terms] def analyze_density(self, document_text: str) -> dict: total_words = len(document_text.split()) if total_words == 0: return {"density_score": 0.0, "total_words": 0} # Calculate frequency of target domain terms in payload matched_terms = sum( document_text.lower().count(term) for term in self.key_terms ) density_score = round(matched_terms / total_words, 4) logger.info(f"Analyzed {total_words} words. Context Density: {density_score}") return { "density_score": density_score, "total_words": total_words, "status": "PASS" if density_score > 0.015 else "HIGH_NOISE" } # Example Usage analyzer = ContextDensityAnalyzer(key_terms=["quarterly revenue", "compliance", "EBITDA"]) payload = "..." # Large retrieved document payload metrics = analyzer.analyze_density(payload) By filtering out low-density documents prior to prompt assembly, I reduced token volume by 65% while simultaneously increasing precision on factual extraction tasks. Step 2: The Latency Penalty of Quadratic and Pre-Fill Processing In standard REST APIs, payload size marginally impacts network transport time. In transformer architectures, processing large context windows introduces a massive latency penalty during the pre-fill phase (time-to-first-token). While time-to-first-token (TTFT) for a 4K token prompt might take 300 milliseconds, pre-filling a 200K token prompt can take 8 to 15 seconds before the model generates a single word. Python import time def evaluate_prefill_latency(client, model: str, context_text: str, user_query: str): full_prompt = f"Context:\n{context_text}\n\nQuestion: {user_query}" start_time = time.time() # Measure time to first token response_stream = client.chat.completions.create( model=model, messages=[{"role": "user", "content": full_prompt}], stream=True ) ttft = None for chunk in response_stream: if chunk.choices[0].delta.content: ttft = time.time() - start_time break # Captured Time-To-First-Token logger.info(f"Model: {model} | TTFT: {ttft:.2f} seconds") return ttft If your application requires real-time user engagement (such as customer support or interactive co-pilots), high TTFT caused by bloated context windows will ruin the user experience long before the model finishes its output. Step 3: Compare Cost-to-Accuracy Across Context Tiers Passing massive amounts of unstructured text into a model for every query creates an exponential cost structure. What engineering teams miss is that the relationship between context length and accuracy is non-linear: doubling the context window doubles the cost but rarely doubles accuracy. Strategy Token Load Avg Latency (TTFT) Accuracy Rate Relative Cost Naive Context Dump 128,000+ tokens ~6.5 seconds 72% (Lost in Middle) 10x Baseline Hybrid RAG + Reranking 8,000 tokens ~0.8 seconds 91% (High Precision) 1x Baseline Hierarchical Summarization 16,000 tokens ~1.4 seconds 86% (Broad Context) 2x Baseline Table of Performance Comparison Across Context Optimization Strategies Optimizing for cost requires evaluating whether structured pre-retrieval (like semantic search paired with cross-encoder reranking) produces a better result at a fraction of the token cost. Step 4: Hybrid Architecture: Context Windows + Precision RAG The most resilient enterprise AI systems don't choose between large context windows and RAG; they use large context windows inside a structured retrieval architecture. Instead of dumping an entire database into the model, use vector retrieval and reranking to select the top relevant passages, and leverage the expanded context window exclusively to hold multi-turn conversation history and rich system instruction contracts. Python def assemble_intelligent_context(user_query: str, vector_db, reranker) -> str: # Step 1: Broad retrieval raw_docs = vector_db.similarity_search(user_query, k=25) # Step 2: Rerank to extract dense, high-signal passages ranked_docs = reranker.rank(query=user_query, documents=raw_docs, top_n=5) # Step 3: Assemble compact, structured context structured_context = "\n---\n".join([doc.page_content for doc in ranked_ranked_docs]) return f"RELEVANT CONTEXT:\n{structured_context}\n\nUSER QUERY: {user_query}" This hybrid approach ensures that the context window is populated only with dense, actionable information, preventing attention degradation and keeping latency low. Diagram of a Context Optimization Pipeline for LLM Inference Context Quality as the Engine for Enterprise AI Scalability The AI industry will continue pushing context limits from millions to tens of millions of tokens. But raw capacity is an infrastructure feature, not an architecture strategy. Relying on massive context windows as a crutch for poor data pipeline design is the modern equivalent of storing an entire relational database in server memory because you don't want to build indexes. Before you scale up your context window size, invest in context quality, intelligent chunking, and strict relevance filtering. When you feed your models high-density, low-noise prompts, you don't just get cheaper and faster applications; you build a system that executes predictably at scale. Conclusion Bigger context windows are an impressive engineering feat, but they are not a silver bullet for enterprise AI systems. As context size expands, the trade-offs in attention accuracy, latency, and operational cost become impossible to ignore. Real enterprise performance isn't achieved by seeing how much data a model can swallow in a single request; it is achieved by engineering precise, high-density context pipelines that deliver the exact right information at the right time. True intelligence in production starts with discipline, not volume.
How It Started This started from a plain problem: I kept hitting token limits at work. I was using Claude Code for real engineering work, and I was burning through budget faster than I wanted. The obvious question was: can I cut that down without hurting the quality of what the agent produces? Not "just use a cheaper model and hope." Something more deliberate — a set of tools that each attack a different part of the token bill. How much context gets read. How much gets re-read. How verbose the agent's own output is. How it finds its way around a codebase in the first place. That question turned into a side project: token-optimization-stack, a public repo with setup docs for tools that reduce token spend. And token-stack-benchmarks, a benchmark harness to actually test whether any of it worked. I'm writing this up because the project ended, a few weeks in, in a place I didn't expect. Not with a working stack and a savings number. Instead, with proof that the token-savings numbers I was looking at were actively misleading — and a cost problem that made the whole thing stop making sense before I could even publish a result. I think both of those are more useful to write about than a clean win would have been. What the Stack Looked Like Early On The first version of the stack had five tools in it: Graphify – turns a codebase into a queryable knowledge graph. The agent can ask "what calls this" instead of reading files to find out.Serena – lets the agent navigate and edit code by symbol, instead of raw file reads and text edits.Headroom – advertised as transparent context compression. Its docs said it needed "no behavioral changes" once installed.LiteLLM – a routing layer. The idea: send easy subtasks to a cheap model and hard ones to an expensive model.Caveman – compresses the agent's own output. Terser replies, compressed subagent output, less back-and-forth. Two of those five didn't survive contact with a real benchmark. Headroom didn't do what its own docs implied. The only way to register it without wrapping the whole claude command in a separate launcher is headroom init claude. That just adds an on-demand MCP tool — something the agent can call, not something that compresses context automatically. Running headroom doctor confirmed nothing was actually being routed through it unless you also ran a separate proxy process with an ANTHROPIC_BASE_URL override. That's a much heavier setup than "no behavioral changes" suggested. On top of that, its mcp serve command crashed against a current MCP SDK — it needed an old, pinned mcp<2 dependency just to start. LiteLLM had a different problem. Its usage-based routing is a load-balancing strategy across provider endpoints, not the complexity-based, per-task routing I actually wanted. And mechanically, Claude Code sends one fixed model for an entire session — there's no way to swap models mid-task based on how hard a step is. The tool I wanted didn't exist yet, at least not in this shape. I removed both rather than keep them in as unverified claims. What was left — Graphify, Serena, a compression/caching layer called LeanCTX, and Caveman — became the actual stack I tested. Why I Had to Stop Not because the idea was wrong. That came later. The experiment itself stopped making financial sense. The rigorous version of this test — real tasks from SWE-bench Verified and Multi-SWE-bench, sixteen repos, five versions of the stack, three repeats each — works out to about 4,800 agent runs. I never got close to that. Instead I ran a much cheaper pilot: 31 tasks, 2 versions of the stack, one repeat, on the cheapest model I had (claude-haiku-4-5, medium effort). Even that only partly finished — 11 of 31 task pairs — and it already cost about $5.60 in raw API spend. That's before EC2 costs, Docker builds, or the multi-day slog of getting this running cleanly on both EC2 and an Apple Silicon Mac. Scale that same per-run cost up to the full test and you pass $1,200 in API spend — on the cheapest model available, before a single result is even trustworthy. Sonnet costs 2x what Haiku does, on both input and output tokens ($2/$10 per million tokens vs. Haiku's $1/$5). So switching to it to get a trustworthy result would push the same test past $2,400. And Haiku wasn't trustworthy: it got zero correct fixes on the Java tasks, and it broke two of the three Python tasks the plain baseline had already solved. Is this savings number — or this correctness failure — actually about the stack? Or is it about the fact that I'm running everything on the cheapest model I could afford to run 4,800 times of? I didn't have a good answer. That's where I stopped. What I Actually Found, for What It's Worth Even the partial pilot data was worth sharing, because it directly contradicts what a token-savings-only view would have told me. On the three Python tasks the plain baseline agent solved correctly, adding the full stack did this: On one task, the agent treated a clear, self-contained bug report as if it were ambiguous. It asked a one-line clarifying question on its very first turn, then just stopped — num_turns: 1, no error, nothing left to score. The plain baseline took 30 turns on the exact same prompt and fixed the bug. Read only off the token dashboard, this run showed 97% fewer tokens used — the single best-looking number in the whole pilot, and it came from the one run that did no work at all.On a second task, the stack produced a real patch. It applied cleanly. The target test still failed.The third task stayed correct — but used more tokens than the baseline, not fewer. I'm not treating "2 of 3" as a rate. Three tasks are too small a sample to turn into a percentage. But something else holds, even at this size: the token numbers and the correctness numbers pointed in opposite directions, and the worst result in the batch produced the best-looking number. That doesn't need a bigger sample to be true — it happened, on a real task, and it's exactly the kind of failure a token-savings-only report can't catch. I'd also expect this to get worse on a weak model, not better. Haiku has less room to recover once a terser style takes away its ability to push back or think through whether a task is really ambiguous. A stronger model might ask the same question but keep working anyway — or not need to ask at all. I didn't get to test that. It's a specific, checkable prediction for whoever picks this up next, not just a guess. Token and cost savings numbers, without a real correctness check against the actual test suite, aren't just incomplete — they can point in exactly the wrong direction. And the biggest, flashiest savings number is a plausible place for that to happen, not an unlikely one. The fix is simple: report cost per solved task, not cost per task. Under that measure, the 97%-savings run isn't a win with an asterisk. It's infinitely expensive, because it solved zero tasks. That one change closes the trap — a dashboard built around it can't turn a silent failure into a headline number. Pair that with something even cheaper to check: turn count. A run that takes 1 turn when the baseline took 30 is a giant red flag, one that no token dashboard shows on its own. And unlike correctness scoring, checking it costs nothing — no test suite, no scoring setup, no Docker. It's already sitting in the same log that produced the token numbers. If You Want to Pick This Up I'm not going to keep running this. Not because I think the question is answered — I just can't afford to answer it properly right now. If you want to take it further, both repos are public: token-optimization-stack – the stack itself: setup scripts and docs for Graphify, Serena, LeanCTX, and Caveman, plus the benchmark methodology this pilot followed.token-stack-benchmarks – the test harness: a Dockerized runner for each version of the stack, task sampling, and the SWE-bench/Multi-SWE-bench scoring scripts. Also everything I ran into getting a Linux-shaped harness to work on both EC2 and an Apple Silicon Mac — case-sensitivity bugs, CPU architecture mismatches, a new Python version breaking a scoring dependency, and more. Contributions and forks are welcome. So are "here's why your pilot was wrong" pull requests. A few concrete places to start: The broken-patch case is still a mystery. The empty-patch failure has a clear cause now (see above). This one doesn't. On the second Python task, a real patch applied cleanly and still didn't fix the bug. Nothing in the logs explains why the stack produced a wrong-but-plausible answer instead of a right one. That's the harder failure mode, and nobody's looked into it yet.Running the full 5-arm test would confirm a real suspect, not just a guess. The 97% run failed because of behavior, not because context got lost. That points at Caveman specifically, and mostly clears LeanCTX, Graphify, and Serena for that run. Running the full test would show whether that holds up, or whether it was a one-off.The haiku-weakness prediction above is easy to test. Run the same pilot on Sonnet or Opus and see if the correctness problem gets smaller. That's a real experiment you can run — not just "try a bigger model and see."
Agile
Career Development
Methodologies
Team Management
Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
October 5, 2026
by Akhil Madineni
CORE
Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases
October 2, 2026
by John Vester
CORE
AI on Top of a Dysfunctional System
October 2, 2026
by Stefan Wolpers
CORE
AI/ML
Big Data
Databases
IoT
Classification Never Left. It Just Got a New Home in LLMs.
October 7, 2026
by Vidyasagar (Sarath Chandra) Machupalli FBCS
CORE
Building High-Performance Time-Series Applications With Java and QuestDB
October 7, 2026
by Otavio Santana
CORE
Apache Phoenix: Global Secondary Indexes With Tunable Consistency
October 6, 2026 by Viraj Jasani
Frameworks
Java
JavaScript
Languages
Tools
Building High-Performance Time-Series Applications With Java and QuestDB
October 7, 2026
by Otavio Santana
CORE
Building and Serving a Custom Model With Azure ML, Then Wiring It Into a Foundry Agent
October 6, 2026
by Jubin Soni, FBCS
CORE
Building IoT Time-Series Applications With Java and Apache IoTDB
October 6, 2026
by Otavio Santana
CORE
AI/ML
Java
JavaScript
Open Source
Classification Never Left. It Just Got a New Home in LLMs.
October 7, 2026
by Vidyasagar (Sarath Chandra) Machupalli FBCS
CORE
Building High-Performance Time-Series Applications With Java and QuestDB
October 7, 2026
by Otavio Santana
CORE