DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

DZone Spotlight

Tuesday, October 6 View All Articles »
Building Enterprise File-Heavy AI Workflows: From Secure Uploads to Governed Document Intelligence

Building Enterprise File-Heavy AI Workflows: From Secure Uploads to Governed Document Intelligence

By Dr Gopala Krishna Behara DZone Core CORE
Enterprise AI is increasingly moving beyond clean, structured datasets and into the much larger world of documents, files, images, forms, emails, reports, contracts, claims, clinical records, and other unstructured content. These files often contain the information enterprises need most, but they are also among the hardest data assets to process reliably at scale. A document AI production system is much more than an LLM, a vector database, or an RAG application. Before a document can become useful AI context, it may need to be securely uploaded, validated, scanned, classified, parsed, OCR-processed, enriched with metadata, protected from sensitive-data exposure, and transformed into retrievable knowledge. The resulting content must then remain traceable to the source through retrieval, reasoning, human review, and downstream execution. That means the architecture must account for much more than the model: Ingestion → Validation → Processing → Enrichment → Indexing → Retrieval → Reasoning → Human Review → Action → Retention → Deletion This is especially important in file-heavy industries such as healthcare, financial services, insurance, and life sciences, where data can be large, heterogeneous, sensitive, and subject to long retention periods. Healthcare provides a useful example. A single enterprise may simultaneously manage clinical documents, medical images, pathology slides, claims data, genomic datasets, audio/video recordings, and medical-device telemetry. This article presents a reference architecture for building secure, scalable, observable, and governable AI workflows around these workloads. The focus is not on a particular LLM or cloud provider. Instead, the article examines the engineering capabilities required before, around, and after the LLM. Enterprise Document AI is not primarily an LLM problem. It is an end-to-end data, security, processing, retrieval, and workflow-engineering problem with AI embedded into it. Why File-Heavy AI Is Different Traditional application architectures often assume that inputs are relatively small and well-structured. An API may receive a JSON request containing a few kilobytes of data. The application validates it, executes business logic, and returns a response. Document-heavy AI workloads are fundamentally different. A single request may contain: A scanned PDF A multi-gigabyte medical image A pathology slide A spreadsheet containing thousands of records An email with multiple attachments A contract containing tables and embedded images A clinical document containing sensitive information A genomic data file An audio or video recording The system therefore must solve several problems simultaneously: Reliable ingestion Large-file transfer Security and malware protection Document understanding OCR and extraction Metadata management PII/PHI detection and protection Knowledge preparation Retrieval AI reasoning Human review Business workflow execution Observability Retention and deletion Lineage and provenance The important architectural implication is that the document pipeline cannot simply be treated as an extension of an LLM API. It needs to be engineered as a first-class enterprise data-processing platform. Healthcare as a File-Heavy Enterprise Healthcare organizations manage everything from electronic health records and claims to high-resolution imaging, pathology slides, genomic data, clinical recordings, and device-generated telemetry. Unlike traditional enterprise documents measured primarily in kilobytes or megabytes, healthcare workloads can contain files ranging from several gigabytes to hundreds of gigabytes. At enterprise scale, repositories can grow into petabytes and must remain accessible, secure, and governed for many years. The challenge therefore extends beyond storage. Healthcare platforms must simultaneously support: Secure ingestion Large-file transfer PHI/PII protection Document understanding Rapid retrieval AI processing Lineage Governance Retention Archival Deletion Major Healthcare Data Challenges challengedescriptionenterprise impact Security & Compliance PHI/PII must be protected throughout ingestion, processing, storage, and access Breach exposure, compliance risk, regulatory consequences, and loss of trust Large File Uploads Imaging, pathology, and genomic datasets can be multi-gigabyte or larger Failed transfers, delays, repeated uploads, and network utilization AI Pipeline Processing Files may require scanning, OCR, classification, extraction, redaction, summarization, and AI analysis Higher compute cost, latency, and processing complexity Massive File Sizes Imaging, pathology, genomics, and clinical video can produce very large objects Storage, transfer, retrieval, and processing challenges Storage Growth Long retention periods continuously increase repository size Infrastructure cost, replication requirements, and operational complexity Upload / Data Transfer Files originate from hospitals, clinics, laboratories, devices, and other locations Network latency, transfer failures, and processing delays AI Readiness & Governance Data must be scanned, validated, enriched, indexed, governed, and traceable before AI consumption Longer preparation time, higher costs, and compliance complexity The architecture must also account for the fact that "a document" is not a single data type. data typetypical scalekey challenge Medical Imaging Files MB to multiple GB DICOM, CT, MRI, PET, mammography, ultrasound files which are complex in nature Digital Pathology Files 2 GB to 100+ GB Large storage requirements, viewing performance, AI processing, replication Clinical Documents KBs to GBs EHR, CCD, FHIR bundles, metadata, PHI protection, compliance Claims & Encounter Data Hundreds of MBs to large, structured datasets Validation, processing at scale, reporting Forms & Patient Documents KBs to 50 MB OCR accuracy, poor image quality, extraction errors Genomics and Sequencing Data Tens to hundreds of GB FASTQ, BAM, CRAM, VCF processing and storage Audio & Video Clinical Content 1–100 GB per recording Transcription, retrieval, secure access, long retention Medical Device Data KBs to hundreds of MB per patient/device per day High-velocity ingestion and real-time analytics These workloads illustrate why one ingestion mechanism or one processing strategy is unlikely to be sufficient for enterprise Document AI. Modern healthcare platforms require specialized architecture incorporating cloud object storage, resumable uploads, event-driven processing, intelligent lifecycle management, and compliance-focused security controls. Characteristics of Enterprise Document Processing Using AI Reliable enterprise AI document workflows begin before a file reaches an LLM. Enterprises need a controlled document-processing pipeline that can securely accept large, varied files; validate and normalize them; extract usable content and metadata; identify and protect sensitive information; prepare content for downstream retrieval or inference; and enforce retention policies throughout the file lifecycle. Document Processing Pipeline for AI Raw enterprise documents should not be sent directly to an LLM. Build an event-driven, layered pipeline with separate stages for ingestion, validation, OCR/parsing, enrichment, indexing, retrieval, LLM reasoning, human review, and final workflow execution. The document-processing pipeline converts untrusted enterprise content into governed AI-ready knowledge. Large Uploading of Data Files Using AI Land uploads in secure object storage first, not directly into the app or LLM path. For very large files, use an asynchronous ingestion flow that can split/burst files into smaller units before downstream processing. PII in AI Document Workflow Treat PII/PHI as a policy-controlled data class. Minimize it, classify it, redact/tokenize where possible, and only store/process it in approved environments. Add access control, encryption, logging, monitoring, and output filtering. Requires approved production environments with encryption in transit/at rest, access controls, logging, and 24/7 monitoring. Sanitized inputs, structured prompts, and output moderation guardrails should be enforced. Document Chunking for AI and RAG Use layout-aware and semantic chunking (by paragraph, heading, table, or section) rather than naive fixed-character splitting. Maintain document metadata, parent-child section context, and overlap to preserve full context across chunk boundaries. Preserve source lineage, section headers, and page numbers in chunk metadata for accurate citation, auditability, and precise retrieval quality. OCR and Metadata Extraction in AI Pipeline OCR, layout parsing, and metadata extraction should happen in the asynchronous document-processing layer. CPU- and GPU-intensive workloads can be executed through scalable worker pools and queues rather than blocking interactive requests. Metadata is particularly important because it becomes an input to downstream retrieval and policy decisions. Document Archival Implement strict lifecycle, retention, and legal hold policies in object storage and databases. Automate document purging or archiving based on business retention schedules and regulatory compliance (e.g., GDPR, HIPAA). Enforce immutability during required retention windows, implement soft deletes with audit logs, and ensure cryptographic erasure or full deletion across object storage, indexes, vector stores, and cache layers when retention expires. Data Lifecycle Management for Enterprise Document AI Enterprise document AI requires lifecycle management that extends beyond ingestion and processing to the complete journey of data and its derived artifacts. This includes creation, ingestion, validation, storage, processing, enrichment, retrieval, AI consumption, retention, archival, and eventual deletion. Each stage should maintain appropriate security, governance, lineage, and audit controls while ensuring that policies are consistently applied to originals and derived artifacts such as OCR output, metadata, chunks, embeddings, and indexes. The following table describes the life cycle of documents and AI data life cycle, life cycle stageactivtieskey controls Create / Source Documents originate from users, applications, scanners, email, EHR/ECM, APIs, devices, and other systems Source identity, ownership, metadata, classification Ingest Files enter through APIs, connectors, SFTP, events, or resumable uploads Authentication, authorization, checksum, upload session, encryption Validate & Secure File type, integrity, malware, and policy checks are performed Malware scanning, quarantine, validation, content policy Store Original document is placed in durable object storage Immutable originals, encryption, access control, versioning Process & Transform OCR, parsing, layout analysis, table extraction and normalization occur Processing isolation, lineage, versioning, confidence Enrich & classify Classification, entities, metadata, PII/PHI detection, redaction and segmentation Policy enforcement, provenance, sensitivity tags Index & Retrieve Content becomes searchable through keyword, vector, and hybrid indexes ACL/ABAC, tenant isolation, metadata filtering Use AI RAG, summarization, extraction, agents and downstream workflows consume the data Grounding, guardrails, tool authorization, audit Retain / Archive Data and derived artifacts are retained according to business/regulatory policy Retention schedules, legal hold, archival, immutability Delete / Dispose / Verify Expired data and derivatives are removed, and deletion is verified Deletion propagation, audit trail Enterprise Document Intake and AI Pipeline Document-heavy enterprises usually need to: Ingest large volumes of PDFs, scans, emails, forms, images, and Office filesExtract structured and unstructured knowledgeRoute documents to the right workflowAnswer questions or generate outputs with traceabilityKeep security, compliance, and auditability intactScale across many business units, document types, and tenants A good enterprise design separates the system into layers that can scale ingestion, understanding, retrieval, and workflow execution independently. The following diagram depicts the detailed view of the Enterprise Document knowledge pipeline and the steps involved in processing. Fig. 1: Document Knowledge Pipeline - Detailed View Intake and Safety Gate Every incoming file should pass through authentication, authorization, file-type and integrity validation, malware/security scanning, size and policy checks, and quarantine handling before it becomes available to downstream processing. The intake layer should be isolated from the AI execution path so that untrusted content cannot directly reach parsers, models, tools, or retrieval systems. Source and Ingestion Layer It brings documents into the platform reliably. Core functions are connector management, deduplication, incremental sync, metadata capture, versioning, and checksum/integrity checks. Typical sources are: ECM/DMS systemsSharePoint/OneDrive/file sharesemail inboxesline-of-business appsAPIs, SFTP, event streamsscanner/OCR pipelines for paper documents Document Processing Layer This layer normalizes documents into AI-ready assets. The most common steps are file type detection, text extraction, OCR for scanned content, layout parsing, table extraction, image understanding, language detection, redaction or PII detection, and document classification It generates the output in the following formats: Raw textStructured blocksPage/layout coordinatesDocument metadataConfidence scores This layer should be asynchronous and queue-based so large batches don't block the system. Enrichment and Understanding Layer It converts raw extracted content into usable knowledge. Typical enrichments are: classification by document type, entity extraction, key-value extraction, taxonomy tagging, topic detection, summary generation, document segmentation by section, clause, or paragraph, duplicate and near-duplicate detection. For enterprise use, this layer often combines: Deterministic rulesML modelsLLMs for semantic interpretation Indexing and Retrieval Layer It makes content searchable and retrievable at scale. Usually includes a keyword index, vector index, metadata store, and object store for original and processed artifacts. Retrieval patterns are: Keyword search for exact matchVector search for semantic matchHybrid retrieval for best qualityMetadata filters for tenants, region, business unit, document type, date, retention class For document-heavy systems, hybrid retrieval is usually the most reliable. Orchestration and Workflow Layer This layer manages end-to-end business processes by connecting AI output to real business actions. Examples are routing a contract to legal review, sending a document to an extraction pipeline, triggering exception handling if confidence is low, creating a case, task, or approval when needed, and invoking downstream systems through APIs. AI Reasoning/Application Layer It performs the user-facing or workflow-facing AI task. Common patterns include RAG Q&A over enterprise documents, document summarization, version comparison, clause extraction, classification and routing, form-filling and auto-population, and agentic workflows with tool use. Best practice is to keep the model layer behind the controls as part of enterprise settings. It covers: Policy checks Prompt templates Retrieval constraints Grounding to approved sourcesGuardrails against unsupported generation Human-in-the-Loop Layer It handles exceptions and maintains quality. This is needed for: low-confidence extraction, policy-sensitive decisions, regulated content, disputed outputs, and approval workflows. This supports: Review queueSide-by-side source viewConfidence indicatorsCorrection captureFeedback loop to improve models and rules Reference Architecture for Enterprise Document AI Platform The complete platform can be organized into six major architectural layers, Users and Experience AI Application and Agent Layer AI Control Plane Knowledge and Tools Plane Document Knowledge Pipeline Data Plane Cross-cutting capabilities span all six layers: Security and Governance AI Safety Observability and MLOps Reliability The document pipeline therefore sits inside a broader enterprise AI platform rather than operating as an isolated OCR or RAG subsystem. Figure 2: Enterprise Document AI Platform Reference Architecture Users and Experience Layer This is the layer users and upstream systems interact with. It typically covers web and mobile apps, partner and batch APIs, and chat or Copilot-style interfaces. This is a single point of entry for the user, where the user gets authenticated and the intent of what is being uploaded and why is captured. Once the request is captured in the structured form, it is handed over to the AI application layer without embedding any business logic. Examples include claims processing and prior authorization workflows. AI Application and Agent Layer This layer hosts the actual AI-facing capabilities covering RAG-based question answering and summarization, extraction of structured fields from unstructured text, classification of documents and requests, and increasingly, agents that carry out multi-step tasks. Examples are, For example, a claims-processing agent might: retrieve a denial rationale, locate supporting attachment pages, extract relevant information, draft a response, request human approval. A prior-authorization agent might: classify an inbound document, identify the requested medication or procedure, extract relevant codes, route the request to the appropriate workflow AI Control Plane It governs model and AI behavior across applications. The model gateway is the governed entry point to every model. Model routing sends the request to the model to handle it reliably. Prompt management version controls the prompts. Guardrails and Content Safety screen both inputs and outputs. Knowledge and Tools Planes It has two distinct capabilities covering the Knowledge and RAG Plane and Tools/Action Plane. This separation is useful because retrieving information and taking an action are different capabilities with different security requirements. Knowledge/RAG Plane: Query processing that turns a natural language question into a retrieval query. Tools/Action Plane: These tools help in calling other systems via MCP, APIs, or agent-to-agent (A2A) protocols, integrating with enterprise applications, pushing work into workflow or case-management systems, and executing business transactions. Document Knowledge Pipeline This is the pipeline intake and the pre-ingestion safety gate, validation and OCR, enrichment including explicit chunking. It also addresses retrieval quality, PHI/PII detection and redaction, storage and indexing, and retrieval-time access control before the document reaches the Knowledge Plane above it. Claims processing: Claim attachments are split, OCR'd, tagged with claim ID and line-of-business metadata, chunked by section (correspondence, EOB, clinical attachment), and redacted for member PHI Prior authorization: Faxed clinical packet is validated and quarantined if password-protected, OCR'd, member and diagnosis identifiers are detected and minimized, and the remaining clinical content is chunked. Data Plane An immutable object store holding original source files, a metadata/state database, keyword and vector indexes for hybrid search, a cache layer for frequently accessed results, and lineage records tying every derived artifact back to its source. Cross-Cutting Platform Controls These categories apply continuously across every layer. It covers: Security and governance: It makes the system enterprise-safe. It implements Identity and access management, RBAC/ABAC, zero trust, encryption, key management, data-loss prevention, tenant isolation, data residency, retention, and audit. AI safety: It implements Prompt injection defense, PII handling, guardrails, grounding that ensures answers are tied to retrieved evidence rather than unsupported generation, and output validation. Observability and MLOps: Keep the system healthy and measurable. It implements logs, metrics, traces, AI-quality monitoring, cost tracking, latency, and SLOs. It also covers Prompt and model registries, evaluation pipelines, CI/CD, regression testing, and model/index/embedding versioning. Reliability: Autoscaling, retry logic, dead-letter queues, idempotency, disaster recovery, backup, and recovery-point/recovery-time objectives (RPO/RTO) Best Practices for Document Processing The recommended architecture patterns for enterprise scale are to use a decoupled, event-driven pipeline. These are best suited as they help: Ingestion spikes won't break the AI layerOCR/extraction can scale separately from retrievalWorkflow steps can be retried independentlySwap models without redesigning the systemEasier governance and auditability A good reference pattern to be used is: API gateway for ingestion and user requestsMessage queue/event bus for async processingObject storage for original and processed filesMetadata database for document state and workflow statusSearch index for keyword retrievalVector database for semantic retrievalWorkflow engine for business process coordinationModel gateway to route to approved AI modelsReview UI for exceptions and approvalsPolicy engine for security and compliance Engineering Principles for Production Document AI Never send untrusted files directly to an LLM. Use durable object storage as the document system of record. Separate ingestion, processing, retrieval, reasoning, and execution. Treat security and authorization as runtime controls, not post-processing. Preserve provenance from source document to AI response. Make every processing stage asynchronous, re-triable, and re-playable. Evaluate document AI quality independently from model quality. Design retention and deletion across originals and every derived artifact. The following are the best practices that need to be followed for document processing, Separate document processing from user interaction: Batch processing and interactive Q&A have different latency and scaling needs.Manage the Full Data and Artifact Lifecycle: Treat every document and its derived artifacts as lifecycle-managed assets. Maintain lineage from the original file through OCR output, extracted metadata, chunks, embeddings, indexes, AI-generated outputs, and downstream records. Apply retention, legal hold, archival, versioning, and deletion policies consistently across the entire artifact chain. When a document expires or is deleted, deletion should propagate to all applicable derived artifacts and be auditable and verifiable.Store source of truth separately from derived artifacts: Keep originals immutable. Generate processed outputs as versioned artifacts.Use hybrid retrieval: Keyword + semantic retrieval beats either one alone for messy enterprise documents.Ground AI outputs in evidence: Require source passages or citations internally so answers can be audited.Route by confidence: High-confidence automation; low-confidence human review.Treat security as a first-class workflow step: Do not bolt it on after retrieval or generation.Design for document variability: Expect scans, tables, handwriting, multi-language content, and inconsistent templates.Make feedback operational: User corrections should feed back into taxonomy, prompts, rules, and extraction models.Design for replayability: Every processing stage should be independently replayable without re-uploading the original document. Preserve immutable originals, processing versions, model versions, prompt versions, extraction configuration, and lineage so that documents can be reprocessed when parsers, models, or policies change. Conclusion Enterprise document AI is not primarily an LLM problem. It is an end-to-end data, security, processing, retrieval, and workflow engineering problem. For file-heavy enterprise workloads, documents must first be securely ingested, validated, scanned, normalized, extracted, enriched, classified, and protected. Their derived representations must retain provenance and access controls as they move into search, RAG, AI reasoning, human review, and downstream business workflows. Production-grade architecture therefore needs more than a model gateway and a vector database. It requires a durable document intake layer, asynchronous processing, scalable OCR and extraction, metadata and lineage management, ACL-aware retrieval, AI safety controls, policy enforcement, observability, evaluation, human oversight, and deterministic lifecycle management. The architecture presented in this paper separates these concerns into independently scalable layers while connecting them through common enterprise controls for security, governance, AI safety, reliability, and observability. This separation allows organizations to evolve document processors, retrieval technologies, models, and agent capabilities without redesigning the entire platform. Practical recommendations for the enterprise document intake and AI pipeline implementation are: Use object storage as the system landing zone; keep the LLM out of the upload path.Separate originals, processed text, metadata, and indexes.Chunk after OCR/extraction, before RAG.Treat PII as a workflow control point, not just a redaction step.Make retention/deletion deterministic and auditable. For healthcare and other highly regulated environments, the principle is even more important. Sensitive data must remain protected throughout its lifecycle, while every transformation from the original document through OCR, chunks, embeddings, retrieved evidence, model output, and business action must remain traceable and governed. Disclaimer The views and opinions expressed in this article are solely those of the authors and do not necessarily reflect the position or policies of any organization with which they are affiliated. More
OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

By Maxim Muzafarov
OpenSearch is an open-source, distributed search and analytics suite derived as a fork of Elasticsearch and maintained under the Apache 2.0 license. When it comes to memory configuration, the guidance is often reduced to a few rules of thumb: swapoff -a, vm.swappiness=1, or bootstrap.memory_lock, and allocating 50% of available memory to the JVM heap while leaving the rest for Lucene and the filesystem page cache, OpenSearch off-heap caches, network buffers, and other system needs. These recommendations are repeated throughout documentation, blog posts, and operational guides, yet their origins and the mechanisms that justify these specific values are rarely examined. Undoubtedly, they provide a reasonable and safe starting point or a safe upper bound in most of the cases, but a safe default is not necessarily an optimal configuration. All of this raises even more questions. How do these defaults affect cluster performance? What is the optimal JVM heap ratio? Does memory given up by the JVM actually become filesystem page cache, and at what point does that trade-off stop paying off? How do read/write latency correlate with the heap ratio? These questions become particularly important in resource-constrained environments and in the cloud, where long-term contracts may make existing instances significantly cheaper, making horizontal or vertical scaling a difficult decision. In this article, we'll try to answer these questions through benchmarking. This is Act 1 of a two-act series. Act 1 focuses on identifying the cause of the latency problems we observed with the current defaults. Act 2 will explore what other heap-ratio values might look like for read/write loads. Along the way, I'll share the tools and commands used throughout the investigation, making this article a practical reference as well for you and for myself when I inevitably need to retrace the investigation months later. Knowledge Context The story also crosses several boundaries, such as the Kernel VM, the JVM, and Lucene. So, it’s important to outline the concepts mentioned in this part of the article beforehand, both for the context and, optionally, to enrich the AI context if you'd like to summarize everything. AreaWhere memory livesWhy it matters hereLinux page cacheFile-backed RAMLucene relies heavily on it for index data; under memory pressure, these pages can be reclaimed and read again later.Linux swapDisk-backed anonymous memoryAnonymous process memory can be swapped out under pressure. vm.swappiness influences this decision but does not prohibit it.Linux PSIKernel pressure signalShows time tasks spend stalled due to CPU, memory, or I/O pressure. We'll use I/O PSI while investigating latency.JVM heapAnonymous memoryControlled by Xms/Xmx; contains Java objects and several OpenSearch data structures.JVM native memoryAnonymous/file-backed memory outside XmxIncludes code cache, metaspace, stacks, direct buffers, and native allocations. Heap metrics do not account for all of it.OpenSearch cachesHeap/off-heap, depending on cacheTheir sizes may depend on heap size, which becomes important when we change jvm_heap_ratio.OpenSearch indexing bufferHeapIts size depends on heap and therefore becomes an important variable in Act 2.Lucene mmapFile-backed/page cacheLucene index files mapped into the process do not consume JVM heap; resident pages compete for physical RAM. Environment I used Aiven for OpenSearch on Azure, with cluster metrics exported to Thanos. The OpenSearch Benchmark metrics don't provide everything we need to answer our questions, particularly host-level metrics such as Linux PSI and swap activity. Exporting the cluster metrics to Thanos allows us to use PromQL queries later to retrieve the additional metrics needed for the investigation. The cluster consists of 3 nodes: CPUAMD EPYC 7763v (Milan)vCPU / RAM2 vCPU, 8 GiBDisk Size175 GiB per nodeAzure Regionazure-westeuropeAzure DiskPremiumV2_LRSAzure SKUStandard_D2as_v5 OpenSearch Version3.6.0JDKjava-21-openjdk-headlessGCG1GC Act 1. The Latency and an Extra GB The symptom: elevated query latency across the cluster, first reported by the customer after a kernel and Azure image upgrade. The load pattern on OpenSearch itself remained unchanged, as did the cluster configuration and settings. The monitoring panels give us the first clue. In the screenshots below, the green vertical line marks the moment of the upgrade. After that point, the page cache grows by roughly a gigabyte, while I/O PSI, previously close to zero, starts showing significant spikes. Nothing crashed, no alert fired, and from the JVM point of view everything looked normal. So where did that extra gigabyte of page cache come from? Nothing was actually freed. It moved. The interesting part isn't just that memory went to swap; it's which memory. When you have thousands of running clusters, there is always a small fraction of them operating close to the edge: relatively stable, yet sensitive enough that even a small change can noticeably affect performance. Like a star nearing the end of its lifetime, they may look stable right up until something disturbs the balance. The immediate cause of the page-cache change was identified fairly quickly: Azure applies tuning parameters that differ from the Linux kernel defaults, and those parameters were not applied by the older image. Once applied, the larger buffers and read_ahead increased the filesystem cache footprint, putting additional pressure on anonymous memory and eventually pushing some of it to swap. But rather than stopping there, let's use this incident as an opportunity to experiment with the heap ratio and make the behavior of OpenSearch instances explicit and less dependent on such environmental changes. Evidence It Is on Swap; None of It Locked First, find what is going on on a node itself: Shell PID=$(pgrep -f 'org.opensearch.bootstrap.OpenSearch') grep -E 'VmRSS|RssAnon|RssFile|RssShmem|VmSwap|VmLck' /proc/$PID/status Shell VmRSS: 4366408 kB # resident RssAnon: 3329896 kB # heap + anonymous native RssFile: 1036496 kB # resident mmap'd Lucene pages RssShmem: 16 kB VmSwap: 2679496 kB # on swap VmLck: 0 kB # bootstrap.memory_lock=fasle, none locked Shell grep -E 'MemFree|MemAvailable|Cached|SwapFree' /proc/meminfo Shell MemFree: 258216 kB MemAvailable: 3479264 kB Cached: 3403196 kB SwapCached: 854396 kB SwapFree: 4745444 kB The Swap Device Is dm-crypt Then confirm there is somewhere for it to go, and on what kind of device: Shell swapon --show Plain Text NAME TYPE SIZE USED PRIO /dev/dm-4 partition 8G 3.5G -1 The dm-* swap device is the interesting detail to catch and to keep in mind. This is an encrypted device, so once a page is requested it could drive more I/O -> more dm-crypt allocations -> more high-order pressure. A self-reinforcing loop and a good example of read amplification. Paging Is Live, Not Historical The next logical step is to check whether it is live paging or just a stale historical tail. The vmstat 1 5 the Linux Virtual Memory Statistics Tool should give us an exact answer for this, where non-zero swap blocks in and out (marked as si, so): Shell vmstat 1 5 Plain Text procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 1 0 3620928 159716 5864 3704212 113 80 3361 577 3147 12 8 7 84 1 0 0 0 0 3625240 189036 5860 3705260 0 5228 24268 5228 4841 4072 14 12 69 4 0 0 0 0 3625240 166196 5860 3710696 0 0 32 0 3572 2556 21 5 74 0 0 0 0 0 3625240 162668 5860 3716200 8 0 164 0 3729 2619 20 7 73 0 0 0 12 0 3625312 147692 5860 3724288 0 140 9207 524 3679 4024 22 14 63 1 0 0 Non-zero si/so in 4 of 5 samples show the live paging process, swpd also climbing across five seconds, which is good proof. The Swapped Pages Are Anonymous, Not File-Backed Let's also check swap memory consumption for each of the process's mappings, to confirm that swap is heap-related: Shell PID=$(pgrep -f org.opensearch.bootstrap.OpenSearch) awk '/^[0-9a-f]/{h=$0} /^Swap:/{if($2>0)print $2" kB "h}' /proc/$PID/smaps | sort -rn | head -20 Plain Text 1160380 kB 708400000-7ffe00000 rw-p 00000000 00:00 0 62296 kB 7f30c0000000-7f30c3f4b000 rw-p 00000000 00:00 0 61452 kB 7f1fc8000000-7f1fcbc03000 rw-p 00000000 00:00 0 61448 kB 7f1fa8000000-7f1fabc02000 rw-p 00000000 00:00 0 61444 kB 7f2f38000000-7f2f3bc01000 rw-p 00000000 00:00 0 61444 kB 7f24b4000000-7f24b7c01000 rw-p 00000000 00:00 0 61444 kB 7f22ac000000-7f22afc01000 rw-p 00000000 00:00 0 61444 kB 7f2238000000-7f223bc01000 rw-p 00000000 00:00 0 61444 kB 7f216c000000-7f216fc01000 rw-p 00000000 00:00 0 61444 kB 7f2168000000-7f216bc01000 rw-p 00000000 00:00 0 61444 kB 7f2148000000-7f214bc01000 rw-p 00000000 00:00 0 61444 kB 7f1fec000000-7f1fefc01000 rw-p 00000000 00:00 0 61444 kB 7f1fe8000000-7f1febc01000 rw-p 00000000 00:00 0 61444 kB 7f1fcc000000-7f1fcfc01000 rw-p 00000000 00:00 0 61444 kB 7f1fac000000-7f1fafc01000 rw-p 00000000 00:00 0 60772 kB 7f1fc0000000-7f1fc3c14000 rw-p 00000000 00:00 0 59952 kB 7f311e000000-7f3123f41000 rw-p 00000000 00:00 0 59340 kB 7f30b4000000-7f30b7c9b000 rw-p 00000000 00:00 0 57180 kB 7f3110000000-7f3113e3e000 rw-p 00000000 00:00 0 43084 kB 7f3130800000-7f3133d70000 rwxp 00000000 00:00 0 The Largest Swapped Region Is But why is what we are seeing above a heap-related area? There are a few clues for that. The region size 0x708400000 - 0x7ffe00000 is exactly 4,154,458,112 bytes = 3,962 MiB as we use -Xms == -Xmx and the whole thing is committed at startup, and nothing else in a JVM process is a single contiguous ~4 GB anonymous rw-p mapping. Second, It's the lowest mapping in the address space, smaps_rollup [rollup] line starts at exactly 708400000: Shell cat /proc/$PID/smaps_rollup Shell 708400000-7ffd3bb56000 ---p 00000000 00:00 0 [rollup] Private_Dirty: 3213788 kB Swap: 2680532 kB SwapPss: 2679412 kB Locked: 0 kB The JIT Code Cache Is Swapped Too Decoding the top swapped regions: 1160380 kB at 708400000 – is the JVM heap, the compressed‑oops heap base and matches the [rollup] start from smaps_rollup.The dozens of 61444 kB regions – these areas are probably related to native/off‑heap: Netty, JNI, Lucene native, etc.43084 kB marked rwxp – the JIT code cache, also swapped out, a bad sign. Together, these regions account for almost exactly the ~2.5 GB of swapped memory we observed earlier: cold heap regions, native/off-heap allocations, code cache, and possibly thread stacks. Practically, this means two consequences: GC can amplify swap latency. 1 GB of the JVM heap was swapped out. G1 does not necessarily touch all of those pages during a mixed collection, but any GC phase that accesses a swapped page incurs a major fault and has to bring it back through the dm-crypt device. Hence, short GC work can produce significantly longer pauses.A swapped page can also contain executable code. The next call into a swapped-out compiled method can trigger a major fault before the code can run. The resulting latency may land on an otherwise random request and be difficult to attribute directly to GC, index I/O, or the query itself. Evidence of Sustained Anon Memory Churn workingset_refault_anon counts anonymous memory refault events after reclaim; it does not count unique pages. Together with pswpin and pswpout, it shows how much anonymous memory paging has accumulated since boot. Shell grep -E 'workingset_(refault|activate)_anon|pswpin|pswpout' /proc/vmstat Plain Text workingset_refault_anon 89061837 workingset_activate_anon 1856069 pswpin 86690374 pswpout 61618603 These counters are cumulative, so fetching them at 10-minute intervals clearly shows that this wasn't a one-time eviction. There was sustained process: nearly 89 million anonymous pages were repeatedly swapped out and faulted back in. vm.swappiness = 1 and GC Amplification In this story vm.swappiness=1 was set since the cluster's inception. It does what Linux defines it to do, but it doesn't provide the protection we wanted. I suspect this matters particularly in the most resource-constrained deployments. How do we know that? The entire result above is a counterexample. swappiness biases the kernel's choice between reclaiming file-backed and anonymous pages. It does not prevent anonymous pages from being swapped out. Even at 1, this can still happen under sustained memory pressure. On a resource-constrained node whose index is several times larger than its RAM, pressure on the filesystem cache is not an exceptional condition; this is the normal operating state. This has two important consequences: It is not self-healing. Nothing proactively pages anonymous memory back in on a schedule. A swapped-out page returns to RAM only when it is accessed again, and cold memory, as it's defined, may remain untouched for a long time. As a result, the cold JVM memory can remain in swap indefinitely.It may remain invisible until something touches it. At steady state, a cold tail of the heap can remain in swap without producing obvious symptoms. The problem becomes visible when those pages are touched again, causing major page faults and potentially amplifying GC and request latency. Key Takeaways So, the root cause of the latency problems is an oversized heap combined with page cache pressure (the cold heap tail has been swapped out). vm.swappiness=1 did not protect the heap. It biases what gets reclaimed; it does not prevent anonymous memory from being swapped out.vm.swappiness=1 should not be relied on with the other defaults in production. An oversized heap can lead to GC amplification that is difficult to detect.Most of the swapped-out memory wasn't heap at all, but malloc arenas and the JIT code cache, none of it inside Xmx, so heap metrics didn't show it.Swap on dm-crypt exacerbates the issue, resulting in longer GC pauses and random request latency spikes. dm-crypt may be unavoidable in production due to security requirements. The fix isn't another swappiness tweak. It's two things: stop committing heap you don't use, and make the heap you do commit non-evictable. Follow Up In Act 2, we'll answer the remaining questions raised at the beginning of this article and look more closely at the trade-off introduced by bootstrap.memory_lock. This setting makes the heap resident and swap-immune, but it also turns jvm_heap_ratio from a soft default into a permanent memory commitment. The question then becomes: what heap ratio best suits different read and write workloads? There is one more complication to mention in advance: in a resource-constrained environment, merge storms can distort benchmark results, making an otherwise good heap ratio appear poor. See the screenshot below. More
Building IoT Time-Series Applications With Java and Apache IoTDB
Building IoT Time-Series Applications With Java and Apache IoTDB
By Otavio Santana DZone Core CORE
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices
Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices
By Rahul Tewari

Refcard #405

Agentic AI Threat Intelligence Essentials

By Alessandro Cannarella
Agentic AI Threat Intelligence Essentials

Refcard #404

Getting Started With Agentic AI for SecOps

By Graziano Casto DZone Core CORE
Getting Started With Agentic AI for SecOps

More Articles

A New Chapter for DZone Newsletters
A New Chapter for DZone Newsletters

Hello DZone community! We’re refreshing our newsletters to help you follow the topics that matter most to your work, explore new ideas, and stay connected with the developer community. Our Zone newsletters are becoming six focused newsletters, with new names and related topics that were developed with input from some of our fantastic community members, and we’re excited to bring them to your inbox: Beyond the Rows: Big data and databasesMind the Model: AI for developers and engineersShip & Scale: Cloud, DevOps, performance, and AgileThe Attack Surface: SecurityDistributed by Design: Microservices, integration, and IoTCode & Craft: Java and web development Each topical newsletter will arrive twice a month, with editions scheduled on Tuesdays and Thursdays. Meet DZone Digest We’re also bringing DZone Daily and DZone Weekly together into DZone Digest, arriving every Wednesday. It will be your weekly roundup of articles and insights from across DZone. What Else Is Changing? All newsletters got a refreshed design, making room for the content you care about and more opportunities to discover upcoming events. We’re also opening newsletter subscriptions to everyone, including developers who aren’t DZone members. Our goal is to make DZone’s newsletters more useful, with clearer topic choices and a regular cadence that helps you keep learning. We’d love to hear from you: Which newsletter are you most interested in, and what topics would you like us to cover? Share your thoughts in the comments. You can subscribe to them here. (If you’re subscribed to any of our Zone newsletters, you’ll begin receiving the updated newsletter covering your topics. If you’re subscribed to DZone Daily or DZone Weekly, you’ll now receive DZone Digest.)

By Dominique Pugh
Building an AI-Ready Data Layer Without Rebuilding the Enterprise
Building an AI-Ready Data Layer Without Rebuilding the Enterprise

Most enterprise AI programs do not fail because the model is too weak. They stall because the data underneath the model is fragmented, delayed, poorly documented, or too expensive to access repeatedly. The common response is to propose a complete platform replacement. That sounds clean on a diagram and becomes dangerous in production. Existing warehouses often support financial reporting, operational dashboards, manufacturing analytics, and regulatory processes that cannot pause while a new AI platform is assembled. The better strategy is to build an AI-ready data layer around stable business contracts. The organization modernizes how data is stored, governed, observed, and served without forcing every existing consumer to migrate at once. AI Readiness Is a Data Contract Problem An AI-ready platform needs more than raw data in inexpensive storage. It needs trusted definitions, reproducible history, fresh operational signals, discoverable lineage, and predictable query behavior. A model trained on an ambiguous customer identifier or an unstable product hierarchy will produce unstable results regardless of model quality. Start by identifying the business entities that must remain consistent across old and new systems. Typical examples include customer, product, supplier, order, invoice, material, and account. Define a canonical contract for each entity before choosing the final storage engine. The contract should specify field names, data types, ownership, accepted values, freshness expectations, and compatibility rules. It should also separate business meaning from physical implementation. A field can move from a legacy warehouse to an open table format without forcing downstream users to learn a new definition. Here is a simplified contract expressed as YAML: YAML entity: product_component owner: supply_chain_data primary_key: product_id, component_id, effective_from freshness: maximum_delay_minutes: 30 fields: product_id: {type: string, nullable: false} component_id: {type: string, nullable: false} quantity: {type: decimal, nullable: false} effective_from: {type: timestamp, nullable: false} effective_to: {type: timestamp, nullable: true} compatibility: additive_columns: allowed destructive_changes: require_new_version This small document does something important. It gives legacy reports, data pipelines, and AI applications the same definition to depend on. Build an Abstraction Layer Before Moving Consumers The riskiest migration pattern moves data and consumers at the same time. When a report changes after cutover, the team cannot easily tell whether the problem came from extraction, transformation, business logic, or presentation. Instead, create a stable semantic or compatibility layer between consumers and physical tables. Existing reports continue reading familiar columns while the implementation behind the view changes gradually. SQL CREATE VIEW analytics.product_component_current AS SELECT product_id, component_id, CAST(quantity AS DECIMAL(18, 4)) AS quantity, effective_from, effective_to FROM modern_layer.product_component WHERE is_current = TRUE; The view is intentionally boring. That is a strength. It preserves a contract while engineers replace ingestion, storage, and transformation components behind it. During transition, the same interface can point to the legacy source, the modern source, or a reconciled combination. Consumers migrate when the new path is proven, not when the infrastructure team finishes installing it. Use Layering to Separate Ingestion From Business Meaning A practical architecture separates raw ingestion, normalized data, and business-ready models. The names are less important than the boundaries. The ingestion layer preserves source fidelity and arrival metadata. The normalized layer resolves types, keys, duplicates, and schema differences. The business layer applies reusable definitions for reporting, features, and AI retrieval. This separation prevents source-system changes from leaking directly into AI applications. It also allows the same governed business model to support batch analytics, streaming decisions, feature engineering, and retrieval-augmented generation. Open table formats can help because they support schema evolution, snapshot history, and rollback. The Apache Iceberg documentation explains how column additions, renames, and partition changes can occur as metadata operations without rewriting every historical file. Those capabilities are useful, but they do not replace contracts. A technically valid schema change can still break business meaning. Run Both Paths and Reconcile Continuously Dual running is not wasted infrastructure. It is how teams prove that a modern data layer is safe. For a defined period, execute legacy and modern pipelines from the same source data. Compare row counts, key coverage, financial totals, null rates, duplicate rates, and business-specific invariants. Do not rely only on aggregate equality, because two incorrect datasets can produce the same total. Python def compare_snapshots(legacy, modern): checks = { "row_count": legacy.count() == modern.count(), "key_coverage": legacy.keys() == modern.keys(), "amount_total": abs(legacy.sum("amount") - modern.sum("amount")) < 0.01, "duplicate_keys": modern.duplicate_count() == 0, } failed = name for name, passed in checks.items() if not passed if failed: raise ValueError(f"Reconciliation failed: {failed}") Real implementations need tolerance rules, exception handling, and audit records, but the principle remains simple. A migration is complete only when correctness is demonstrated repeatedly across normal operations, period close, late-arriving data, and recovery scenarios. Make Lineage and Observability Part of the Product AI systems often combine data from many pipelines. When an answer changes, teams need to know which source, transformation, or model version caused it. Capture lineage at execution time rather than asking engineers to document it later. The OpenLineage specification defines interoperable metadata around datasets, jobs, and runs. Whether a team adopts that standard or another approach, the important point is to connect every published dataset to its inputs, code version, execution, owner, and quality results. Monitor the data layer with service-level objectives. Useful signals include freshness delay, failed contract checks, schema drift, incomplete partitions, reconciliation differences, query latency, and cost per workload. Infrastructure uptime alone is not enough. A pipeline can be running while delivering yesterday's data or silently dropping a critical field. Add Real-Time Access Only Where the Decision Requires It AI readiness is often confused with making everything real time. That creates unnecessary cost and operational complexity. Classify datasets by decision latency. Fraud detection or equipment monitoring may need event-level updates. Product recommendations may accept a few minutes of delay. Financial reporting may prioritize completeness and controlled closing over speed. When streaming is justified, design for replay, idempotency, and explicit processing guarantees. The Apache Kafka Streams documentation describes transactional and idempotent processing for exactly-once behavior within supported read-process-write flows. Teams still need to test external side effects and recovery paths rather than assuming one configuration solves end-to-end correctness. Control Cost Through Workload Isolation Legacy warehouses often mix ingestion, transformation, dashboards, experiments, and ad hoc queries in one shared resource pool. AI adds expensive feature generation, embedding creation, and large scans to that competition. Separate workloads by purpose and apply budgets, concurrency limits, caching, and retention policies independently. Store reusable features and business models once instead of recomputing them in every notebook. Track cost by dataset and workload so teams can see whether freshness or model accuracy justifies the additional compute. Predictable cost is part of the data contract. A dataset that is technically available but economically impractical to query is not AI-ready. Modernize by Proving One Business Slice Do not begin with the entire enterprise. Choose one domain with meaningful AI potential and stable business ownership. Product structures, customer identity, inventory, or service events are common candidates. Build the canonical contract, abstraction layer, modern pipeline, reconciliation suite, lineage, and cost controls for that slice. Keep existing reports working. Then connect one AI use case to the same governed layer. The result becomes a reusable migration pattern. Future domains inherit working templates for contracts, quality checks, dual runs, observability, and cutover. Modernization accelerates because the organization is no longer debating the architecture from scratch. An AI-ready data layer is not a separate platform waiting for the enterprise to catch up. It is a controlled evolution of the enterprise data system itself. The safest path keeps trusted analytics running while gradually replacing the foundations beneath them.

By Rajaganapathi Rangdale Srinivasa Rao
Grounding AI Agents in Governed Data
Grounding AI Agents in Governed Data

Today, every vendor offering BI solutions has incorporated a chat box. Whether you use Copilot or some other natural-language interface that connects you to a data warehouse, just ask a question in simple words, and it will generate SQL automatically. While this is conducive to productivity in other industries, in banking it presents an opportunity for a new attack. It is not enough to simply say that wrong SQL can be produced. It’s that an ungoverned text-to-SQL layer may join tables it shouldn’t, return columns that should have been masked. As a result of a lack of oversight, a marketing analyst could receive a query containing raw account numbers, since none of the components of the stack told the system to do otherwise. Prompt-level guardrails (“please don’t show PII”) are not a security control. They’re just a suggestion, and a model under adversarial pressure ot just a confusing prompt will ignore a suggestion. The issue isn't just about providing a more intelligent prompt; rather, it's about placing the artificial intelligence assistant on the same layer of information as human analysts, meaning an environment where the database manages the relevant security details as per row and column criteria instead of relying on technology. Consequently, if the analyst does not have access to the specific column, there is no justification for the AI assistant to have access to it as well. The diagram below (Figure 1) shows the steps taken to build that layer in BigQuery: a validated semantic layer that allows human dashboards and AI-generated queries to be connected to the same quality definitions. This would allow the assistant to leverage existing security rather than creating it. Figure 1. The Semantic Layer Resolver Step 1: Stop Letting Anyone (Human or AI) Query Raw Tables The initial phase is architectural, not related to AI: there is no query made by a person or a system involving the base tables. Instead, each of the metrics that are accessible to consumers is defined at least once in a BigQuery view, and its calculation logic is embedded in that view. SQL -- Certified metric: Risk-Weighted Assets, defined once, queried everywhere CREATE VIEW analytics.risk_weighted_assets AS SELECT exposure.customer_id, exposure.region, exposure.exposure_class, exposure.outstanding_balance, risk_weights.weight_pct, ROUND(exposure.outstanding_balance * risk_weights.weight_pct / 100, 2) AS rwa_amount, CURRENT_TIMESTAMP() AS calculated_at FROM finance.exposures AS exposure JOIN reference.basel_risk_weights AS risk_weights ON exposure.exposure_class = risk_weights.exposure_class WHERE exposure.status = 'ACTIVE'; The view of "risk-weighted assets" created through a dashboard, a notebook, and an LLM agent is identical. There is no alternative version in a researcher’s spreadsheet, nor can an AI agent "helpfully" recreate the calculation based on exposure tables but use incorrect risk weightings. Step 2: Enforce Security at the Data Layer, Not the Application Layer It is important to ensure that BigQuery includes row-level and column-level security and associates it with the table. This means the principle will work irrespective of the entity making the query. SQL -- Row-level security: a regional analyst only ever sees their region's rows CREATE ROW ACCESS POLICY regional_filter ON analytics.risk_weighted_assets GRANT TO ('group:[email protected]') FILTER USING (region = 'EMEA'); Column masking works the same way, through policy tags rather than per-report logic: YAML # Dataplex policy tag: applied once, enforced everywhere the column is queried taxonomy: financial-pii policyTags: - displayName: "customer-account-number" description: "Masked for all roles except fraud-investigation" - displayName: "customer-ssn" description: "Masked for all roles except compliance-audit" Once a policy tag is applied to a column, a user, or any AI agent acting under that user's identity, who doesn’t have the appropriate fine-grained reader role, will receive either a null value or a hashed value. There is no mistake that the model has to circumvent; it’s simply a different result. This is what makes querying with AI safe, since whatever the query for the AI is, it cannot reveal anything that the column policy prohibits. Step 3: Give the Grounding Layer Metadata to Query Against It is impossible for an LLM to adhere to rules of governance it knows nothing about. Accordingly, a metadata directory is necessary for the semantic layer that contains a description of each certified metric with enough detail for the agent to turn an inquiry posed in natural language into the correct interpretation and filtering process, not simply provide it with a raw schema dump. JSON { "metric_id": "risk_weighted_assets", "display_name": "Risk-Weighted Assets", "view": "analytics.risk_weighted_assets", "owner": "[email protected]", "sensitivity": "internal", "allowed_dimensions": ["region", "exposure_class", "customer_id"], "definition": "Balance times Basel risk weight, summed by class.", "lineage": ["finance.exposures", "reference.basel_risk_weights"], "last_certified": "2026-06-01" } This record is the thing the AI agent actually reads. The document specifies which view will be interrogated, lists the dimensions available for filtering results, and identifies who to contact if something goes wrong. It is worth mentioning that in this record there is no schema given for the finance exposes table, which leaves the model nothing to "discover" about. Step 4: Route Natural-Language Requests Through the Semantic Layer, Not the Warehouse When the certified metrics with their metadata have been obtained, the resolution process consists of transforming the user's natural-language question into a query that uses an allowed view rather than directly referring to the underlying schema. Python class SemanticLayerResolver: def __init__(self, metric_catalog, bq_client): self.catalog = metric_catalog # metric_id -> metadata, from Step 3 self.bq_client = bq_client def resolve(self, nl_request: str, user_identity: str) -> QueryResult: # 1. Map the request to a certified metric, never to a raw table. # A constrained classifier over self.catalog.keys() works better # here than open-ended text-to-SQL against the full warehouse. metric = self.match_metric(nl_request) if metric is None: return QueryResult.refuse("No certified metric found.") # 2. Extract filters, restricted to the metric's allowed_dimensions. filters = self.extract_filters( nl_request, metric["allowed_dimensions"] ) # 3. Build SQL against the certified view only. sql = self.build_query(metric["view"], filters) # 4. Execute as the requesting user, so BigQuery's row/column # security applies exactly as it would for a human query. result = self.bq_client.query(sql, user=user_identity) # 5. Attach lineage and certification metadata to the answer, # so "what the AI said" is auditable like any report. return QueryResult( data=result, metric_id=metric["metric_id"], lineage=metric["lineage"], certified_at=metric["last_certified"], ) The important line is step 4: the query is executed under the requesting user instead of using a shared service account. Hence, all the downstream access control mechanisms are automatically applied. There is no need for a separate permission system for the resolver because it has no access rights that exceed the rights of the requesting user. Step 5: Audit Every AI-Generated Query Like You Would a Human's The governance teams will not agree on a system based on the suggestion of " having faith in the model." What they approve is proof in every case where a resolver has provided information, just as is done when an individual writes a report. Python def log_ai_query(user_identity, nl_request, result: QueryResult): audit_log.write({ "user": user_identity, "request": nl_request, "metric_id": result.metric_id, "lineage": result.lineage, "policy_version": result.certified_at, "row_count": result.row_count, "timestamp": now(), }) One financial institution successfully applied this approach. What used to be a lengthy project in which one would have to analyze whether an AI assistant could access customer information has been transformed into something evaluated right away: the assistant can perform the same functions as a human worker. The financial institution was also measuring the new trend of using a certified semantic layer, not only in regard to the AI being discussed. Conflicts over defining metrics across different business lines practically vanished when the organization no longer had to create a separate “AI-compliant” data model. The Real Insight: Governance Is What Makes AI Fast, Not What Slows It Down It’s easy to assume that the best approach to deal with the LLM and sensitive data combination is to include a review step in which a human sits in on every step of the process, or another model is deployed to analyze the first model’s outputs before they are used. This is not only unscalable, but it also misses the point. Another way is to make sure that the insecure path cannot be taken, rather than simply being shunned. If the data layer implements row-level security, column masking, and certified metric definitions, you can confirm that an AI agent querying the data cannot generate queries that reveal any data previously available to someone with the same role. As a result, there is no need to verify output against constraints, since they were already included in the model. This shift is suggested by this pattern. Governed self-service, the architecture that permits a business analyst to carry out data initiatives safely in the absence of ticket submission, also creates a secure basis for AI-enhanced analysis. But it wasn't the main purpose. It is just a coincidence that it has worked out this way.

By Jeevan reddy Geereddy
Reproducible WebRTC Failure Testing With Playwright and coturn
Reproducible WebRTC Failure Testing With Playwright and coturn

Testing how a WebRTC application behaves when the network fails is usually done by hand: open two browser windows, turn off Wi-Fi, count to ten, turn it back on, watch what happens. It works, after a fashion. It is also unrepeatable, untimed, impossible to run unattended, and useless for comparing two implementations — the interruption is never quite the same twice, and nothing records what actually occurred. I hit this while testing recovery behavior across browser engines. I needed the same interruption, applied at the same point in the connection lifecycle, repeated dozens of times, producing machine-readable output. Getting there took longer than expected, mostly because the obvious approach doesn't do what it appears to. This article describes the approach that worked, the failure modes that shaped it, and the parts worth reusing. The harness is open source; the code, raw trial records, and environment metadata are linked at the end. The Obvious Approach Is Misleading Playwright exposes BrowserContext.setOffline(true), which is the natural first reach: JavaScript await context.setOffline(true); It does something — just not the thing you want. setOffline operates at the network layer Playwright controls, so it reliably kills HTTP requests and WebSocket connections. Your signaling channel dies immediately, which looks convincing in the logs. What it does not reliably do is break an established peer connection whose media path runs over loopback or the local network. ICE keeps exchanging traffic and connectionState stays connected. If you are testing signaling recovery, this is a legitimate tool. If you are testing what an application does when the media path fails, you are testing nothing, and the dead signaling channel makes it look like you are. There is a second-order problem. Because setOffline blocks the page's WebSocket, any signaling that has to survive the outage — an ICE restart offer, for instance — cannot travel over the browser's own connection. In my harness, this forced signaling to be bridged through the Node process driving the test rather than the page, so that offers and answers could still cross between peers while one of them was offline from the browser's point of view. That is a workaround for a tool limitation, not a property of WebRTC, and it is worth knowing before you build around it. Make the Media Path Something You Control The alternative is to stop trying to break the network and instead force all traffic through a process you own. Setting iceTransportPolicy: "relay" in the RTCConfiguration restricts candidate gathering to relay candidates only — host and server-reflexive candidates are excluded from the pool entirely (MDN). Point that relays at a coturn instance on the local machine, and the entire media path now depends on one process you can signal. JavaScript const pc = new RTCPeerConnection({ iceTransportPolicy: "relay", iceServers: [{ urls: [ "turn:127.0.0.1:65050?transport=udp", "turn:127.0.0.1:65050?transport=tcp" ], username: "harness", credential: "harnesssecret" }] }); Interrupting the connection then becomes process control: Shell kill -STOP $TURN_PID # relay stops forwarding kill -CONT $TURN_PID # relay resumes SIGSTOP rather than SIGTERM matters here. The process is suspended, not terminated, so it stops relaying immediately while retaining its port bindings and allocation state. SIGCONT resumes it in place, with no restart and no re-binding race. The property that makes this useful for experiments is that the outage is *parameterized*. "Restore the relay three seconds after the connection reports failed" becomes a variable rather than a stopwatch and good intentions. Running the same interruption at one, three, and five seconds across two browsers is then just a loop. Choose the Method Once and Record It A harness that silently falls back between interruption methods produces data you cannot interpret. If trial 7 paused a relay and trial 8 called setOffline, the comparison is meaningless — and you will not know unless the selection is recorded. Select once at startup, log what was available alongside what was chosen, and write the selection into every trial record: JSON { "selected": "host_coturn_sigstop", "fallbackUsed": false, "attempted": [ { "available": true, "method": "host_coturn_stop_forward", "reason": "turnserver found at /opt/homebrew/opt/coturn/bin/turnserver" }, { "available": false, "method": "os_wifi_power", "reason": "ALLOW_WIFI_TOGGLE not set to 1" }, { "available": true, "method": "playwright_offline", "reason": "always available (may not break loopback WebRTC)" } ], "iceTransportPolicy": "relay" } The os_wifi_power entry deserves a comment. Toggling the actual interface via networksetup -setairportpower is closer to a real network event than pausing a relay, and is therefore a better test in principle. It is gated behind an explicit environment variable because a test suite that disables your machine's Wi-Fi without asking is a poor citizen, particularly in CI. Preflight, and Fail Loudly A long matrix run that breaks on trial 3 and produces 57 rows of garbage is worse than one that refuses to start. Before any measured trials, verify that each browser can reach every state the experiment depends on, and record what was observed: JSON { "engine": "chromium", "browserVersion": "151.0.7922.34", "steps": [ { "step": "ice_disconnected", "ok": true, "waitedMs": 5029 }, { "step": "ice_failed", "ok": true, "waitedMs": 9992 }, { "step": "full_reconnect", "ok": true, "elapsedMs": 145 } ] } The preflight also aborts hard on one specific error class: SDP negotiation failures, m-line ordering errors, and InvalidAccessError. These indicate the harness is broken rather than the connection under test, and allowing them through contaminates every downstream result. A run that does not finish with zero of these should be discarded. The preflight numbers are themselves a result. Under this interruption method, Chrome reached disconnected at 5.0 s and failed at 10.0 s; Firefox took 11.2 s and 19.9 s. Both engines are working within the consent-freshness bounds described in RFC 7675, which sets a 30-second consent expiry with checks roughly every five seconds — but the specific thresholds differ, and that difference is only visible because the interruption is byte-for-byte identical across engines. Manual Wi-Fi toggling cannot produce that comparison. Attribute Recovery to the Right Connection This is the detail that determines whether the results mean anything, and it is easy to omit. When testing whether a connection recovers, "the session is connected again" is not the same claim as "the original connection recovered." A freshly built RTCPeerConnection reaching connected is indistinguishable, in connectionState alone, from the original one returning. Counting both as recovery measures nothing. Tag every peer connection at construction and compare identity at the moment of recovery: JavaScript metrics.originalPeer = (pcInstanceId === iceRestartPcInstanceId); An epoch counter addresses the related hazard: callbacks from a torn-down connection fire after its replacement exists, and without a guard they write into the current trial's record. JavaScript if (epochAtStart !== pcEpoch) return; // stale callback, ignore Neither guard is exotic. Both are the difference between a dataset and a pile of numbers. Not Every Failure Test Needs a Network One of the two experiments in my harness tests whether a superseded recovery cycle can proceed into a destructive rebuild. That is a concurrency question, not a connectivity one, so it requires no interruption at all — an artificial delay forces the overlap window deterministically. The result runs identically anywhere, produces the same outcome every time, and completes in seconds. Before building network machinery, it is worth checking which of your failure scenarios are actually about the network. What the Setup Produces The stack is Node with Playwright driving two browser contexts, a small WebSocket signaling server, and coturn as the controllable relay. Each trial emits a JSON file with a timestamped event log covering connectionState, iceConnectionState, signalingState, restart and rebuild boundaries, cycle identifiers, and final outcome — plus a CSV summary row. Alongside those, the run captures an environment record: OS and architecture, Node and Playwright versions, both browser versions, ICE configuration, and the selected interruption method. That last file matters more than it sounds. The distance between "I tested this, and it worked" and a result someone else can check lies almost entirely in whether the environment was recorded next to the numbers. Two Things I Would Do Differently The environment record did not capture a git commit hash, so published results cannot be tied to an exact code revision. An obvious gap in hindsight, and the first thing I would add. More substantively: pausing a TURN relay is not a network interface change. It tests relay failure specifically. The engine timings above may not transfer to a genuine Wi-Fi-to-cellular handoff, where interface teardown, address changes, and gathering behavior all differ. The method buys reproducibility at the cost of realism, and that trade should be stated rather than discovered by a reader. Running It Yourself The harness is at github.com/jaynirmal15/webrtc-recovery-harness under an MIT license. The raw trial records, environment metadata, and preflight output from the runs described above are in the results/ directory. If you are testing recovery behavior in your own application, the interruption machinery is the part worth lifting.

By Jay Suresh Nirmal
The Context Window Trap: Why More Context Doesn’t Mean Better AI
The Context Window Trap: Why More Context Doesn’t Mean Better AI

When model providers announced 1-million-token and multi-million-token context windows, the software engineering world celebrated. The immediate narrative was simple and appealing: document chunking is dead, complex RAG pipelines are obsolete, and developers can now dump entire codebases, legal libraries, or multi-year enterprise datasets into a single model call. This breakthrough led engineering teams into a new trap: assuming that a model's capacity to accept context equals its ability to reason effectively over that context. In production, relying on massive context windows as a substitute for intelligent information retrieval leads to severe performance degradation, runaway infrastructure bills, and unpredictable hallucinations. Operating large-scale LLM architectures taught me a hard truth: A larger context window gives a model more surface area to get confused. Before you expand your context length, you must optimize your context quality. The Hidden Trap: "Needle in a Haystack" and Attention Degradation In traditional database systems, doubling the size of a query payload doesn't degrade the accuracy of the returned records. Relational algebra operates on exact logic. Large language models do not query data; they process spatial attention. As prompt lengths scale into hundreds of thousands of tokens, the self-attention mechanisms inside the transformer architecture begin to struggle. Key information buried deep in the middle of a massive prompt often suffers from "Lost in the Middle" syndrome, where the model strongly attends to tokens at the very beginning and very end of the prompt while ignoring crucial details placed in between. Diagram of Attention Degradation in Large Context Windows When an application fails under a massive context window, the model rarely throws an error. Instead, it silently ignores conflicting constraints, merges unrelated concepts, or produces plausible-sounding responses derived from irrelevant sections of your input. Step 1: Measure Prompt Noise-to-Signal Ratio In simple LLM integrations, teams measure throughput and token count. When scaling large-context applications, you must measure context density, the proportion of tokens directly relevant to the user's intent versus the filler text passed into the prompt. During a recent enterprise project, I built a utility to compute contextual relevance scores before submitting payloads to ultra-long context models: Python import logging from typing import List logger = logging.getLogger("ContextOptimizer") class ContextDensityAnalyzer: def __init__(self, key_terms: List[str]): self.key_terms = [term.lower() for term in key_terms] def analyze_density(self, document_text: str) -> dict: total_words = len(document_text.split()) if total_words == 0: return {"density_score": 0.0, "total_words": 0} # Calculate frequency of target domain terms in payload matched_terms = sum( document_text.lower().count(term) for term in self.key_terms ) density_score = round(matched_terms / total_words, 4) logger.info(f"Analyzed {total_words} words. Context Density: {density_score}") return { "density_score": density_score, "total_words": total_words, "status": "PASS" if density_score > 0.015 else "HIGH_NOISE" } # Example Usage analyzer = ContextDensityAnalyzer(key_terms=["quarterly revenue", "compliance", "EBITDA"]) payload = "..." # Large retrieved document payload metrics = analyzer.analyze_density(payload) By filtering out low-density documents prior to prompt assembly, I reduced token volume by 65% while simultaneously increasing precision on factual extraction tasks. Step 2: The Latency Penalty of Quadratic and Pre-Fill Processing In standard REST APIs, payload size marginally impacts network transport time. In transformer architectures, processing large context windows introduces a massive latency penalty during the pre-fill phase (time-to-first-token). While time-to-first-token (TTFT) for a 4K token prompt might take 300 milliseconds, pre-filling a 200K token prompt can take 8 to 15 seconds before the model generates a single word. Python import time def evaluate_prefill_latency(client, model: str, context_text: str, user_query: str): full_prompt = f"Context:\n{context_text}\n\nQuestion: {user_query}" start_time = time.time() # Measure time to first token response_stream = client.chat.completions.create( model=model, messages=[{"role": "user", "content": full_prompt}], stream=True ) ttft = None for chunk in response_stream: if chunk.choices[0].delta.content: ttft = time.time() - start_time break # Captured Time-To-First-Token logger.info(f"Model: {model} | TTFT: {ttft:.2f} seconds") return ttft If your application requires real-time user engagement (such as customer support or interactive co-pilots), high TTFT caused by bloated context windows will ruin the user experience long before the model finishes its output. Step 3: Compare Cost-to-Accuracy Across Context Tiers Passing massive amounts of unstructured text into a model for every query creates an exponential cost structure. What engineering teams miss is that the relationship between context length and accuracy is non-linear: doubling the context window doubles the cost but rarely doubles accuracy. Strategy Token Load Avg Latency (TTFT) Accuracy Rate Relative Cost Naive Context Dump 128,000+ tokens ~6.5 seconds 72% (Lost in Middle) 10x Baseline Hybrid RAG + Reranking 8,000 tokens ~0.8 seconds 91% (High Precision) 1x Baseline Hierarchical Summarization 16,000 tokens ~1.4 seconds 86% (Broad Context) 2x Baseline Table of Performance Comparison Across Context Optimization Strategies Optimizing for cost requires evaluating whether structured pre-retrieval (like semantic search paired with cross-encoder reranking) produces a better result at a fraction of the token cost. Step 4: Hybrid Architecture: Context Windows + Precision RAG The most resilient enterprise AI systems don't choose between large context windows and RAG; they use large context windows inside a structured retrieval architecture. Instead of dumping an entire database into the model, use vector retrieval and reranking to select the top relevant passages, and leverage the expanded context window exclusively to hold multi-turn conversation history and rich system instruction contracts. Python def assemble_intelligent_context(user_query: str, vector_db, reranker) -> str: # Step 1: Broad retrieval raw_docs = vector_db.similarity_search(user_query, k=25) # Step 2: Rerank to extract dense, high-signal passages ranked_docs = reranker.rank(query=user_query, documents=raw_docs, top_n=5) # Step 3: Assemble compact, structured context structured_context = "\n---\n".join([doc.page_content for doc in ranked_ranked_docs]) return f"RELEVANT CONTEXT:\n{structured_context}\n\nUSER QUERY: {user_query}" This hybrid approach ensures that the context window is populated only with dense, actionable information, preventing attention degradation and keeping latency low. Diagram of a Context Optimization Pipeline for LLM Inference Context Quality as the Engine for Enterprise AI Scalability The AI industry will continue pushing context limits from millions to tens of millions of tokens. But raw capacity is an infrastructure feature, not an architecture strategy. Relying on massive context windows as a crutch for poor data pipeline design is the modern equivalent of storing an entire relational database in server memory because you don't want to build indexes. Before you scale up your context window size, invest in context quality, intelligent chunking, and strict relevance filtering. When you feed your models high-density, low-noise prompts, you don't just get cheaper and faster applications; you build a system that executes predictably at scale. Conclusion Bigger context windows are an impressive engineering feat, but they are not a silver bullet for enterprise AI systems. As context size expands, the trade-offs in attention accuracy, latency, and operational cost become impossible to ignore. Real enterprise performance isn't achieved by seeing how much data a model can swallow in a single request; it is achieved by engineering precise, high-density context pipelines that deliver the exact right information at the right time. True intelligence in production starts with discipline, not volume.

By Chidiebere Njoku
I Tried Building a
I Tried Building a "Token Optimization Stack" for Coding Agents. Here's Why I Killed It.

How It Started This started from a plain problem: I kept hitting token limits at work. I was using Claude Code for real engineering work, and I was burning through budget faster than I wanted. The obvious question was: can I cut that down without hurting the quality of what the agent produces? Not "just use a cheaper model and hope." Something more deliberate — a set of tools that each attack a different part of the token bill. How much context gets read. How much gets re-read. How verbose the agent's own output is. How it finds its way around a codebase in the first place. That question turned into a side project: token-optimization-stack, a public repo with setup docs for tools that reduce token spend. And token-stack-benchmarks, a benchmark harness to actually test whether any of it worked. I'm writing this up because the project ended, a few weeks in, in a place I didn't expect. Not with a working stack and a savings number. Instead, with proof that the token-savings numbers I was looking at were actively misleading — and a cost problem that made the whole thing stop making sense before I could even publish a result. I think both of those are more useful to write about than a clean win would have been. What the Stack Looked Like Early On The first version of the stack had five tools in it: Graphify – turns a codebase into a queryable knowledge graph. The agent can ask "what calls this" instead of reading files to find out.Serena – lets the agent navigate and edit code by symbol, instead of raw file reads and text edits.Headroom – advertised as transparent context compression. Its docs said it needed "no behavioral changes" once installed.LiteLLM – a routing layer. The idea: send easy subtasks to a cheap model and hard ones to an expensive model.Caveman – compresses the agent's own output. Terser replies, compressed subagent output, less back-and-forth. Two of those five didn't survive contact with a real benchmark. Headroom didn't do what its own docs implied. The only way to register it without wrapping the whole claude command in a separate launcher is headroom init claude. That just adds an on-demand MCP tool — something the agent can call, not something that compresses context automatically. Running headroom doctor confirmed nothing was actually being routed through it unless you also ran a separate proxy process with an ANTHROPIC_BASE_URL override. That's a much heavier setup than "no behavioral changes" suggested. On top of that, its mcp serve command crashed against a current MCP SDK — it needed an old, pinned mcp<2 dependency just to start. LiteLLM had a different problem. Its usage-based routing is a load-balancing strategy across provider endpoints, not the complexity-based, per-task routing I actually wanted. And mechanically, Claude Code sends one fixed model for an entire session — there's no way to swap models mid-task based on how hard a step is. The tool I wanted didn't exist yet, at least not in this shape. I removed both rather than keep them in as unverified claims. What was left — Graphify, Serena, a compression/caching layer called LeanCTX, and Caveman — became the actual stack I tested. Why I Had to Stop Not because the idea was wrong. That came later. The experiment itself stopped making financial sense. The rigorous version of this test — real tasks from SWE-bench Verified and Multi-SWE-bench, sixteen repos, five versions of the stack, three repeats each — works out to about 4,800 agent runs. I never got close to that. Instead I ran a much cheaper pilot: 31 tasks, 2 versions of the stack, one repeat, on the cheapest model I had (claude-haiku-4-5, medium effort). Even that only partly finished — 11 of 31 task pairs — and it already cost about $5.60 in raw API spend. That's before EC2 costs, Docker builds, or the multi-day slog of getting this running cleanly on both EC2 and an Apple Silicon Mac. Scale that same per-run cost up to the full test and you pass $1,200 in API spend — on the cheapest model available, before a single result is even trustworthy. Sonnet costs 2x what Haiku does, on both input and output tokens ($2/$10 per million tokens vs. Haiku's $1/$5). So switching to it to get a trustworthy result would push the same test past $2,400. And Haiku wasn't trustworthy: it got zero correct fixes on the Java tasks, and it broke two of the three Python tasks the plain baseline had already solved. Is this savings number — or this correctness failure — actually about the stack? Or is it about the fact that I'm running everything on the cheapest model I could afford to run 4,800 times of? I didn't have a good answer. That's where I stopped. What I Actually Found, for What It's Worth Even the partial pilot data was worth sharing, because it directly contradicts what a token-savings-only view would have told me. On the three Python tasks the plain baseline agent solved correctly, adding the full stack did this: On one task, the agent treated a clear, self-contained bug report as if it were ambiguous. It asked a one-line clarifying question on its very first turn, then just stopped — num_turns: 1, no error, nothing left to score. The plain baseline took 30 turns on the exact same prompt and fixed the bug. Read only off the token dashboard, this run showed 97% fewer tokens used — the single best-looking number in the whole pilot, and it came from the one run that did no work at all.On a second task, the stack produced a real patch. It applied cleanly. The target test still failed.The third task stayed correct — but used more tokens than the baseline, not fewer. I'm not treating "2 of 3" as a rate. Three tasks are too small a sample to turn into a percentage. But something else holds, even at this size: the token numbers and the correctness numbers pointed in opposite directions, and the worst result in the batch produced the best-looking number. That doesn't need a bigger sample to be true — it happened, on a real task, and it's exactly the kind of failure a token-savings-only report can't catch. I'd also expect this to get worse on a weak model, not better. Haiku has less room to recover once a terser style takes away its ability to push back or think through whether a task is really ambiguous. A stronger model might ask the same question but keep working anyway — or not need to ask at all. I didn't get to test that. It's a specific, checkable prediction for whoever picks this up next, not just a guess. Token and cost savings numbers, without a real correctness check against the actual test suite, aren't just incomplete — they can point in exactly the wrong direction. And the biggest, flashiest savings number is a plausible place for that to happen, not an unlikely one. The fix is simple: report cost per solved task, not cost per task. Under that measure, the 97%-savings run isn't a win with an asterisk. It's infinitely expensive, because it solved zero tasks. That one change closes the trap — a dashboard built around it can't turn a silent failure into a headline number. Pair that with something even cheaper to check: turn count. A run that takes 1 turn when the baseline took 30 is a giant red flag, one that no token dashboard shows on its own. And unlike correctness scoring, checking it costs nothing — no test suite, no scoring setup, no Docker. It's already sitting in the same log that produced the token numbers. If You Want to Pick This Up I'm not going to keep running this. Not because I think the question is answered — I just can't afford to answer it properly right now. If you want to take it further, both repos are public: token-optimization-stack – the stack itself: setup scripts and docs for Graphify, Serena, LeanCTX, and Caveman, plus the benchmark methodology this pilot followed.token-stack-benchmarks – the test harness: a Dockerized runner for each version of the stack, task sampling, and the SWE-bench/Multi-SWE-bench scoring scripts. Also everything I ran into getting a Linux-shaped harness to work on both EC2 and an Apple Silicon Mac — case-sensitivity bugs, CPU architecture mismatches, a new Python version breaking a scoring dependency, and more. Contributions and forks are welcome. So are "here's why your pilot was wrong" pull requests. A few concrete places to start: The broken-patch case is still a mystery. The empty-patch failure has a clear cause now (see above). This one doesn't. On the second Python task, a real patch applied cleanly and still didn't fix the bug. Nothing in the logs explains why the stack produced a wrong-but-plausible answer instead of a right one. That's the harder failure mode, and nobody's looked into it yet.Running the full 5-arm test would confirm a real suspect, not just a guess. The 97% run failed because of behavior, not because context got lost. That points at Caveman specifically, and mostly clears LeanCTX, Graphify, and Serena for that run. Running the full test would show whether that holds up, or whether it was a one-off.The haiku-weakness prediction above is easy to test. Run the same pilot on Sonnet or Opus and see if the correctness problem gets smaller. That's a real experiment you can run — not just "try a bigger model and see."

By Shreyash Thakare
Nobody Designs an RBAC Mess; Everyone Ends Up With One
Nobody Designs an RBAC Mess; Everyone Ends Up With One

You're about to rename a column. Five minutes of work. Someone asks the obvious question first: who does this break? In a healthy Snowflake account, that's a query. In most accounts, it isn't. Someone runs SHOW GRANTS ON TABLE, gets a list of roles, and hits the real problem: those roles nest inside other roles, which nest inside more roles, and nobody can say with confidence where the chain ends. So the change waits. Or it ships anyway, because the maintenance window doesn't care about your archaeology project. That gap, between wanting to know who this impacts and actually being able to find out, is your role-based access control (RBAC) debt. It usually stays invisible until the worst possible moment. I've owned Snowflake governance at two different financial services organizations over the last several years, which means I've inherited this exact problem more than once, and cleaned it up more than once too. Here's how it happens, why it's so hard to undo, and the decisions that actually prevent it. Nobody Designs a Mess Nobody sits down and architects an unmanageable role hierarchy on purpose. It arrives in four stages, and every stage feels reasonable at the time. Stage 1 — Innocent beginnings. The account is new. A handful of people need access, so someone grants a privilege directly to a user "just for this sprint." The first role gets named whatever felt right that morning, ANALYTICS_ROLE maybe, because there's only one analytics team so far. Stage 2 — Growth by copy-paste. A second team needs access. The fastest path is cloning the first team's role and adjusting it, quirks included. There was never a naming convention, so there's nothing to copy correctly or incorrectly — every team just does what seemed sensible in isolation. Multiply this by a dozen teams, and you have a dozen roles with a dozen different mental models baked in. Stage 3 — The workarounds. New tables don't automatically show up in existing grants, because nobody set up future grants. People patch this with one-off grants on individual objects. An incident requires emergency access; the access gets granted, the incident resolves, and the grant quietly outlives its purpose because revoking it wasn't anyone's job. The most common version of this I've seen: non-prod needs realistic data to be useful for testing, and properly refreshing it on a schedule is expensive to build. So someone grants a non-prod role a "temporary" bridge into production raw data instead of fixing the actual refresh pipeline. The stale-data problem never gets solved, so the bridge never gets removed. Now your production PII exposure isn't just a function of who holds production roles. It's also a function of who holds non-production roles, because one of them quietly has a door into production. Stage 4 — Entanglement. Nobody can draw the hierarchy from memory anymore. Revoking anything feels dangerous, because nobody knows what depends on it. The mess is now load-bearing — people are relying on access paths nobody remembers granting, and everyone is afraid to touch it. Why You Can't Untangle It Later Reconstructing the role graph itself isn't actually that hard, technically. Snowflake's role hierarchy is fully queryable, and I've built automation before that keeps a daily-refreshed map of it straight from INFORMATION_SCHEMA.APPLICABLE_ROLES. The data is there. What you don't have is intent. Nobody recorded why a given grant exists, so every revocation becomes a gamble instead of a decision. Was this access load-bearing, or a leftover from an incident eighteen months ago? The only way to find out is to revoke it and see who complains, which is exactly the kind of change management nobody wants to sign off on. This is also where role chaining turns from convenient into dangerous. Snowflake won't let you build a literal cycle. The platform enforces the role graph as a DAG, so ROLE_A can never chain back to itself through ROLE_B. But it won't stop you from nesting roles in ways that quietly fan a privileged grant out to people who have no idea they inherited it. Chain an access role into another access role because it's faster than doing it properly, or nest a "shared utility" role into an unrelated hierarchy to save five minutes, and you've created inheritance nobody can see happening in real time. The difference between disciplined, one-directional nesting and an entangled graph isn't subtle once you draw it out. There's a second attribution problem, specific to Snowflake, that compounds the first. When secondary roles are active in a session, a user can exercise privileges from every role they hold, not just their primary role. So when you're reconstructing "how did this person access this table" from query history, the primary role on the query isn't necessarily the role whose grant actually authorized the access. It could have come from any secondary role active in that session. "Who can touch this" and "which specific grant let this exact query succeed" turn out to be two different questions, and both are harder to answer than they should be. Put together, the cost isn't slower audits. It's blast-radius analysis becoming unanswerable during an actual incident, right when you need the answer fastest. The Decisions That Prevent It None of this requires exotic tooling. It requires making a small number of decisions before the first grant, not after the hundredth. Separate access roles from functional roles, and nest in one direction only. Access roles describe what can be touched. Functional roles describe who someone is. Access roles nest into functional roles. Functional roles get granted to users. Never the reverse, and never access-role-to-access-role as a shortcut. This is the single rule that prevents the silent-inheritance problem above. If chaining only ever flows one direction, tracing who has this stays a bounded, predictable operation instead of an open-ended one. A minimal example of what that looks like as code: Plain Text # ACCESS ROLE — describes WHAT can be touched resource "snowflake_account_role" "access_select_customer_pii" { name = "ACCESS_SELECT_CUSTOMER_PII_MASKED" comment = "SELECT on masked customer PII views. Owner: data-governance-team." } resource "snowflake_grant_privileges_to_account_role" "select_customer_pii" { account_role_name = snowflake_account_role.access_select_customer_pii.name privileges = ["SELECT"] on_schema_object { future { object_type_plural = "VIEWS" in_schema = snowflake_schema.customer_masked.fully_qualified_name } } } # FUNCTIONAL ROLE — describes WHO someone is. # Composed only of access roles. Never nests another functional role. resource "snowflake_account_role" "functional_data_analyst" { name = "FUNCTIONAL_DATA_ANALYST" comment = "Standard analyst role." } # One-directional nesting: access role -> functional role. Never the reverse. resource "snowflake_grant_account_role" "analyst_gets_customer_access" { role_name = snowflake_account_role.access_select_customer_pii.name parent_role_name = snowflake_account_role.functional_data_analyst.name } # Functional role -> user is handled by SCIM/IdP sync in practice, # not a static grant like this — shown here only for illustration. resource "snowflake_grant_account_role" "assign_to_user" { role_name = snowflake_account_role.functional_data_analyst.name user_name = "jsmith" } #Resource names here are from the official snowflakedb/snowflake provider, v2.x. On the older Snowflake-Labs provider these were snowflake_role and snowflake_grant_role. The specific Terraform syntax doesn't matter much. What matters: every grant now has an owner, a comment explaining why it exists, and a diff sitting in a pull request. The why behind a grant lives in version control instead of nowhere, which is the actual fix for the lost-intent problem from earlier. Naming conventions before the first grant. Decide the pattern, role type, domain, environment, whatever fits your org, while there's exactly one role to name. Retrofitting a convention onto fifty existing roles is a project. Applying one from role one is free. Never grant directly to a user. Every exception to this rule is the seed of a future audit finding. It's tempting exactly when it matters most ("just for this sprint," during an incident, for someone leaving in two weeks), which is precisely why it needs to be a rule without exceptions, not a judgment call made under pressure. If non-prod needs production-realistic data, solve it with governed data sharing or masked replication, not a standing grant that blurs the line between environments. The bridge from Stage 3 is always framed as temporary. It never is. Identity-driven provisioning. Tying role assignment to your identity provider means access follows someone's actual employment status. Joiners get provisioned automatically; leavers get deprovisioned automatically. Manual provisioning is exactly where "I'll revoke this later" grants come from. Scheduled privilege audits, done as routine hygiene. Not incident response. A recurring, boring, calendar-driven review of who has what and whether it's still needed. The goal isn't to catch one dramatic violation. It's to stop small drift from compounding into Stage 4. Untangling a Mess That Already Exists If you're already past Stage 3, the fix looks less like a redesign and more like a controlled, incremental remediation: Inventory – pull the full grants graph, however ugly it looks.Cross-reference against actual usage – query history tells you who's really using an access path, versus who technically still has it.Stage revocations as tracked, reversible changes – not a big-bang cutover. Small batches, each one logged, each one able to be rolled back if something breaks.Monitor and iterate – the first pass won't be perfect. Treat it as a cycle, not a one-time project. For the blast-radius question specifically — who can reach this role through any depth of nesting — the shape of the fix is a recursive walk of the grants graph, something like: SQL WITH RECURSIVE role_chain AS ( SELECT name AS role_name, 0 AS depth FROM snowflake.account_usage.roles WHERE name = 'ACCESS_SELECT_CUSTOMER_PII_MASKED' AND deleted_on IS NULL UNION ALL SELECT g.grantee_name, rc.depth + 1 FROM snowflake.account_usage.grants_to_roles g JOIN role_chain rc ON g.name = rc.role_name WHERE g.granted_to = 'ROLE' AND g.granted_on = 'ROLE' AND g.deleted_on IS NULL ) SELECT u.grantee_name AS impacted_user, MIN(rc.depth) AS min_depth FROM role_chain rc JOIN snowflake.account_usage.grants_to_users u ON u.role = rc.role_name AND u.deleted_on IS NULL GROUP BY u.grantee_name ORDER BY min_depth; I've run a version of this remediation myself: roughly two dozen individually tracked changes across production and non-production, instead of one sweeping cutover. Tedious, honestly. But it works, as long as you're willing to do it in small, auditable steps instead of betting everything on one big rewrite fixing it in a single shot. The Bill Always Comes Due RBAC debt compounds quietly, the same way any technical debt does, except the interest gets paid in security exposure you can't quantify and audit findings you can't explain, not in slower deploys. The cheapest day to get this right was day one. Today's the next best option. I'd rather write that sentence than be in the room when someone asks who this impacts during a live incident, and the honest answer is that nobody actually knows.

By Mayank Sethi
Building a Migration Readiness Engine for Analytics Workflows
Building a Migration Readiness Engine for Analytics Workflows

Platform migrations are not only technology problems. They're visibility problems as well. Organizations often know what platforms they operate and who uses them. What they typically do not know is what runs on those platforms. Which workflows are business critical?Which workflows can be migrated with minimal effort?Which workflows require redesign?Which capabilities have no equivalent in a target platform? Without answers to these questions, migration planning becomes guesswork. During a large enterprise analytics modernization initiative, I helped evaluate an established platform environment that had evolved across multiple teams and environments over time. The challenge was not merely collecting infrastructure information. It was understanding workflow behavior at scale. What followed was the development of a workflow introspection platform that analyzed workflow definitions, categorized capabilities, and evaluated migration readiness. The most interesting outcome wasn't the migration itself. It was the framework that made migration planning more structured, reviewable, and evidence-based. The Visibility Problem in Analytics Platforms Most analytics platforms expose administrative metadata. This metadata can typically answer questions such as: Which workflows exist?Who owns them?When were they last modified? What it cannot usually answer is: What does this workflow do?Which capabilities does it use?How difficult would it be to migrate?What dependencies exist across workflows? This information typically resides within workflow definitions themselves. The management layer knows the workflow exists. The workflow definition explains how it operates. The challenge is connecting those two layers. To answer meaningful migration questions, we needed to move beyond administrative metadata and inspect workflow logic directly, subject to authorization, access controls, confidentiality objections, and security requirements. Building a Workflow Discovery Pipeline The first challenge involved collecting workflow artifacts at scale. Authorized workflow definitions were gathered and standardized into a common format that could be analyzed consistently across environments. Collection and storage should minimize or exclude embedded credentials, personal information, confidential business logic, and other sensitive content not required for the assessment. As workflows were collected, differences in metadata structures, naming conventions, and ownership information became apparent. To address this, the discovery process incorporated validation, normalization, and quality checks designed to improve consistency while preserving source provenance and documenting unresolved data quality issues. Once the collection process stabilized, workflow artifacts could be prepared for deeper analysis. Turning Workflow Files into Structured Data The next challenge was understanding what those workflows contained. Most workflow definitions were stored as structured configuration files. A simplified structure resembled: XML <Workflow> <Nodes> <Node Tool="Input" /> <Node Tool="Join" /> <Node Tool="Output" /> </Nodes> </Workflow> While simple on the surface, real-world workflows contained significantly more complexity. Workflows included: Data preparation logicTransformationsAnalytical operationsBusiness rulesAutomation componentsPlatform-specific functionality The parser extracted only the metadata needed for the migration assessment and transformed it into a structured profile that could be analyzed systematically. Production implementations should apply data minimization, retention limits, access controls, and audit logging to these profiles. For each workflow, we generated: Java public class WorkflowProfile { private String workflowId; private int toolCount; private List<String> tools; private List<String> dependencies; private ComplexityScore complexity; } This transformed workflow definitions into structured datasets that can be queried. Instead of analyzing workflow files individually, we could evaluate usage patterns across the environment. The Tool Parity Matrix Extracting workflow information was useful. It still didn't answer the most important question: Can this workflow be migrated? To solve this problem, I built what became the core component of the migration readiness engine: the tool parity matrix. The concept was simple. Each source platform capability was mapped to its nearest apparent equivalent in a target platform for the defined use case. Functional mapping was a starting point, not a substitute for validating behavior, performance, security, privacy, licensing, resilience, support, and regulatory requirements. Each mapping received two attributes. Parity Level Java enum ParityLevel { FULL, PARTIAL, NONE } Migration Effort Java enum MigrationEffort { LOW, MEDIUM, HIGH } The goal was not perfect accuracy. The goal was consistent decision support, with sourcing assumptions documented, calibrated, and subject to expert review. Calculating Migration Readiness Once tool mappings existed, preliminary migration scoring became possible. The resulting score was a prioritization aid rather than a determination that a workflow could be migrated safely or successfully. Workflows composed primarily of supported capabilities received higher readiness scores. Workflows containing specialized functionality received lower readiness scores and required additional evaluation. Conceptually: Python def migration_readiness(workflow): total_score = 0 for tool in workflow.tools: total_score += parity_score(tool) return total_score / len(workflow.tools) Similarly, relative effort could be estimated using weighted migration effort values. Translating those values into time or cost requires calibrated weights, historical data, documented assumptions, and validation for the relevant environment. This enabled workflows to be grouped into preliminary migration waves for engineering and stakeholder review. For example: Readiness / effort profileOutcomeHigh readiness / low effortMigration wave 1Medium readiness / medium effortMigration wave 2Lower readiness / higher effortAdditional review required What previously required largely subjective evaluation could now be informed by consistent, documented criteria, while retaining expert review for exceptions and consequential decisions. Discovering Hidden Platform Dependencies One unexpected benefit of the readiness engine was dependency discovery. As workflows were analyzed collectively, patterns began to emerge. Certain capabilities appeared repeatedly. Some workflow categories were straightforward to migrate. Others consistently required additional review because they relied on specialized functionality. Most importantly, platform-wide analysis revealed concentrations of functionality that might otherwise remain invisible until migration execution began. By surfacing these patterns early, architectural decisions could be made proactively rather than reactively, which can improve planning accuracy and help identify migration risks before execution. Building the Consolidation Layer Large platform environments rarely operate with perfectly consistent metadata. Different environments often use: Different naming conventionsDifferent ownership modelsDifferent deployment practicesDifferent classification standards To address this, a consolidation pipeline normalized workflow metadata into a unified model. The process included: Discovery → Normalization → Deduplication → Classification → Migration scoring Once normalized, workflows could be analyzed collectively rather than as isolated artifacts. This supported environment-wide reporting and migration planning, subject to the permitted use of the underlying data. Connecting Technical Data to Organizational Data One lesson from large migrations is that technical readiness alone is insufficient. Organizations also need to understand workload ownership. The readiness engine therefore connected workflow metadata to authorized organizational ownership information. Access to ownership data should be limited to legitimate planning purposes and handled in accordance with applicable privacy, employment, and records-management requirements. This allowed migration plans to answer both technical and operational questions. Instead of saying "These workflows require additional review," we could identify which groups owned those workflows and engage stakeholders earlier in the planning process. Technical analysis became actionable. Lessons Learned 1. Migration Planning Is a Data Problem Most migration programs begin with meetings. They should begin with visibility. Without understanding platform usage, migration planning becomes speculation. 2. Metadata Is Not Enough Administrative information provides useful context. It rarely provides sufficient context. Meaningful migration planning requires understanding workflow behavior. 3. Scoring Enables Scale Humans can review a limited number of workflows. Large environments require consistent evaluation frameworks. Well-designed scoring systems can support prioritization and repeatability, but they require representative inputs, documented limitations, monitoring, and expert review. Final Thoughts Many organizations approach platform migrations as technology replacement exercises. They are discovery exercises. Before deciding where workloads should move, you first need to understand what those workloads do. The migration readiness engine addressed part of that problem by combining workflow discovery, workflow analysis, dependency identification, and migration scoring into a repeatable planning framework. The result was not simply a migration plan. It was a structured approach for understanding complex analytics ecosystems and supporting migration decisions with better evidence, documented assumptions, and appropriate review. Disclaimer: This article represents my personal views and is not written on behalf of, endorsed by, or intended to represent the views of my employer. The architecture, code, scoring methods, and examples are simplified and illustrative and should be validated for the relevant technical, security, privacy, legal, and regulatory environment. This article is provided for general informational purposes and does not constitute legal, technical, or other professional advice.

By Ravi Krishna Palivela
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI

Regularly, someone reopens the same argument. XML or JSON or YAML, as if one has to win and the others lose. It usually comes up in a context like data contracts, where a team has to pick a format and defend it. The framing is wrong. These formats were built for different jobs in different eras, and the more useful question is which one fits the job in front of you. So here is the history, the trade-offs, and where each one still wins, including what changes now that LLMs and agents read and write structured data too. XML, JSON, and YAML at a Glance These three formats are different ways to represent structured data. XML is verbose and rigorous. JSON is compact and universal. YAML is readable and config-friendly. None is strictly best. Each won a different era and a different job, and validation became its own layer, led today by JSON Schema. Key Takeaways XML led enterprise integration for two decades and now lives mostly in legacy systems; JSON won web APIs; YAML won cloud-native config.XML, JSON, and YAML serialize data. JSON Schema and XSD validate it. They are different layers, not competitors.YAML has no schema language of its own. It borrows JSON Schema, which is how Kubernetes and similar tools validate YAML.JSON Schema now underpins LLM tool calling and structured outputs, which puts it at the center of agentic AI.Authored in YAML, validated by a schema, enforced in the pipeline: data contracts are the clearest example of a wider pattern in data governance and orchestration tools. Serialization vs. Validation: Two Different Jobs XML, JSON, and YAML are serialization formats. You author data in them, and a parser reads them back. JSON Schema and XML Schema (XSD) are validation languages. They describe what valid data looks like, and a validator checks a document against that description. So comparing YAML with JSON Schema is not a fair fight. One is a format you write. The other is a contract you check against. Keep that split in mind. Most of the real story is about how the two layers interact. A Short History of XML, JSON, and YAML Each format rose with a shift in how we built systems. XML came first, standardized by the W3C in 1998 with roots in SGML. It became the backbone of enterprise integration. SOAP, WSDL, and the early ESB and SOA stacks all spoke XML. It was verbose but rigorous, and it shipped with a full schema system in XSD. The ecosystem also grew heavy. The sprawling WS-* stack of SOAP extensions became so complex that many engineers came to call it WS-* hell, which is part of why lighter approaches eventually took over. JSON came out of the JavaScript world in the early 2000s. Douglas Crockford formalized it from JavaScript object literals, and json.org went up in 2002. As REST and AJAX replaced SOAP for web APIs, JSON replaced XML as the default wire format. It was lighter, easier to read, and mapped directly to data structures in most languages. JSON Schema followed later as a separate community effort. YAML appeared in 2001 as "YAML Ain't Markup Language," designed to be human-friendly first. Since version 1.2 in 2009, it is a superset of JSON. It found its home in the cloud-native era. Kubernetes, Terraform, Ansible, CI/CD pipelines, GitOps. Anywhere humans hand-write configuration that lives in Git. One detail matters for later. Each format handled schema differently. XML built it in with XSD. JSON bolted it on with JSON Schema. YAML never built one and borrowed JSON Schema instead. Other Formats Worth Knowing: TOML, HCL, Protobuf, and Config Languages A comparison limited to three formats would feel a decade out of date. The landscape is wider now. TOML is simple and serves as the config format for Rust's Cargo and Python's Poetry. HCL is HashiCorp's language for Terraform. In the streaming world, Protobuf and Avro take a schema-first approach and serialize to compact binary, which is why they sit under Kafka and gRPC. There is also a newer category built to fix YAML's weaknesses: configuration languages. CUE, Pkl from Apple, and KCL from the CNCF add expressions, validation, and reuse on top of the data model, then render plain YAML or JSON as output. They are not serialization formats. They are programs that generate configuration. For teams drowning in thousands of lines of near-duplicate YAML, they are worth a look. The rest of this post stays on XML, JSON, and YAML, since they are still the three you choose between most days. How XML, JSON, and YAML Differ in Practice The differences show up the moment you write them by hand. XML wraps everything in opening and closing tags and supports attributes, namespaces, and comments. It is precise and self-describing. It is also heavy. A small payload turns into a wall of angle brackets. JSON uses braces, brackets, and quoted keys. It is compact and unambiguous, and every major language parses it natively. Its one notable omission is comments. The spec does not allow them, which is a real constraint for anything humans need to annotate. YAML uses indentation instead of brackets and braces. It supports comments, multi-line strings, anchors for reuse, and multiple documents in one file. It reads closer to how people think about nested data. The cost is that whitespace carries meaning, so structure is easy to break. Pros and Cons of XML, JSON, and YAML XML's strength is rigor: namespaces, mature validation with XSD, XPath for querying, and decades of tooling. Its weakness is weight and friction. Few people enjoy writing it by hand, and it feels dated for new web APIs. JSON's strength is ubiquity and simplicity. It is the lingua franca of web APIs; it maps cleanly to data structures, and it parses fast everywhere. Its weaknesses are the lack of comments and the absence of native validation in the base spec. YAML's strength is readability. It is friendly to engineers and non-engineers, it diffs cleanly in Git, and it supports comments and reuse. The trade-offs come from the same design. Significant whitespace makes it fragile, so one wrong indentation can break the file, and loose typing causes surprises, like the Norway problem where the country code NO once parsed as the boolean false. YAML 1.2 and StrictYAML help, but the tension stays. The same whitespace that makes YAML readable makes it fragile. One caveat on YAML's downsides. They mostly bite humans. When tools generate and validate the files, as Kubernetes operators and config languages like CUE or Pkl do, the fragility matters far less. In fact, the sweet spot is machines generating and editing while humans mainly read and review, which plays to YAML's strengths. XML vs. JSON vs. YAML: A Comparison Table The table below sums up how XML, JSON, and YAML compare across the dimensions that matter most in practice, from readability and verbosity to schema support and failure modes. DimensionXMLJSONYAMLHuman readabilityLowMediumHighVerbosityHighMediumLowCommentsYesNoYesNative schemaXSD, built inJSON Schema, add-onNone, borrows JSON SchemaValidation maturityVery matureMature, now dominantMature, via JSON SchemaIDE toolingMatureMatureMatureType safetyStrong with XSDBasic, strong with schemaWeak, coercion surprisesGit-diff friendlinessPoorGoodExcellentLearning curveSteepEasyEasy to start, subtle trapsTypical useSOAP, documents, enterpriseWeb APIs, data exchangeConfig, contracts, pipelinesMain failure modeBloat and complexityNo comments, no base validationIndentation and type coercion Does YAML Have a Schema? YAML has no schema language of its own. It uses JSON Schema. Because YAML 1.2 maps onto the same data model as JSON, a JSON Schema validator can validate a YAML document without modification. The format you author in and the language that validates it are decoupled, and the decoupling is a feature. The mechanics are simple. A tool publishes a JSON Schema describing its YAML structure. Your IDE applies that schema as you type, giving autocompletion, inline validation, and error checking. In VS Code this runs through the YAML language server. SchemaStore acts as a public registry of JSON Schemas that editors auto-apply to hundreds of known config files, from Kubernetes manifests to GitHub Actions workflows. Kubernetes is the largest example. Custom resources validate against OpenAPI structural schemas, which are a JSON Schema dialect generated from the underlying Go types. CI and workflow tools like CircleCI follow the same idea, publishing a JSON Schema for their YAML that editors use for validation and autocompletion. The pattern is consistent. The code is the source of truth; it emits the schema, and the schema validates the YAML. So XSD was XML's built-in answer. JSON Schema is JSON's bolt-on answer. YAML's answer is to reuse JSON Schema. The result is that JSON Schema became the shared validation layer for both JSON and YAML. JSON Schema and Agentic AI: Tool Calling, Structured Outputs, and MCP Structured data formats used to be a backend concern. Now they sit at the center of how AI systems work, and JSON Schema is the format doing the work. When an LLM calls a tool, the tool is defined by a JSON Schema. When you ask a model for structured output, you hand it a JSON Schema and the model fills it in. Several providers go further with constrained decoding, which restricts the model token by token so the output cannot violate the schema. OpenAI's Structured Outputs guarantees schema compliance this way, and Google's Gemini and others offer similar structured-output modes. The Model Context Protocol (MCP), the emerging standard for connecting models to tools, adopted JSON Schema 2020-12 as its default dialect for tool inputs and outputs in 2025. The shift underneath is the interesting part. Schemas have always enforced structure, since XSD already validated and rejected SOAP messages at runtime. What is new is where the enforcement sits. The same JSON Schema that validates a config file now also shapes a language model's output as it generates, token by token. The contract moved from checking data after the fact to steering how it is produced. Notice the division of labor. Humans author agent and workflow config in YAML, because it is readable. Machines exchange JSON, because it is precise. JSON Schema validates both. Data Contracts and Governance: Where Format Choice Matters This is where the choice stops being academic. Data contracts are where formats, schemas, and governance meet. A data contract is a formal agreement between the team that produces a dataset and the teams that consume it. It defines fields, types, allowed values, freshness, ownership, and quality rules. The shift it represents is governance moving out of documents and into code. A contract in a Confluence page is documentation. A contract in version-controlled YAML, checked in CI/CD, is an enforceable control. The tooling has converged on YAML for the same reasons that make YAML good for config. The Open Data Contract Standard, at version 3.1.0 under the Linux Foundation's Bitol project, defines contracts in YAML and ships a JSON Schema so editors can validate them. Soda's SodaCL expresses data quality checks in YAML and runs them in pipelines, with a cloud layer that adds stakeholder approval. dbt embeds model contracts and tests in YAML. Great Expectations validates against a declarative spec. Different tools, same pattern. Author the contract in YAML, validate it with a schema, enforce it in the pipeline. The pattern is a familiar one. Databases use schemas to keep bad data out of storage. Streaming platforms use schema registries to keep bad data out of event streams. Analytical datasets now use contracts to do the same thing, one layer up. Data contracts are the clearest case, but the same model runs through orchestration and governance tooling. Workflow engines like Argo and Kestra define pipelines in YAML that a JSON Schema validates. Policy-as-code tools like Open Policy Agent and Kyverno keep rules in version control and enforce them in CI/CD or at deploy time. Author in YAML, validate with a schema, enforce in the pipeline. The format choice and the validation layer are the same story at every level. Which Format Should You Use? A Quick Decision Guide XML fits when you need namespaces, document-centric markup, or integration with SOAP and legacy enterprise systems that already speak it. JSON is the default for web APIs and machine-to-machine exchange, where precision and universal parsing matter more than human authoring. Reach for YAML for configuration and contracts that humans write and review in Git, where readability and comments earn their keep. For validation, reach for JSON Schema in almost every modern case, including for your YAML. Reach for XSD when you are already in the XML world. And when your YAML starts repeating itself across hundreds of files, look at a configuration language like CUE, Pkl, or KCL before the duplication gets worse. The same logic shows up across the modern data stack. Integration pipelines and APIs move JSON. Process mining still reads XML-based event logs like XES while newer event streams carry JSON. Data contracts and platform config are authored in YAML and validated by JSON Schema. Different layer, different format, same principle. Stop Asking Which Format Is Best XML, JSON, and YAML were never really competing for the same job. XML won enterprise integration for two decades, then its weight and WS-* complexity pushed teams toward lighter options, so today it lives mostly in legacy systems. JSON won the web API era and became the default wire format. YAML won the cloud-native config era. Validation sits on its own layer: XSD for XML and JSON Schema for JSON and YAML, and JSON Schema now also underpins how agents and AI systems exchange structured data. The useful question is not which format is best. It is which layer you are working in. Author where humans read. Exchange where machines parse. Validate everywhere. Get those three right and the format debate mostly takes care of itself.

By Kai Wähner DZone Core CORE
Building Time-Series Applications With Java and InfluxDB
Building Time-Series Applications With Java and InfluxDB

InfluxDB is essential for applications that analyze continuously changing data. In IoT, this includes tracking temperature, pressure, energy use, or machine telemetry over time. Financial systems use similar models for market prices, exchange rates, trading activity, portfolio values, and risk metrics. The key requirement is the ability to ingest large volumes of timestamped data and efficiently query current, historical, and evolving values. This versatility makes InfluxDB valuable beyond traditional monitoring. Enterprises use it for observability, infrastructure metrics, logistics, industrial systems, customer activity, fraud detection, transaction trends, and business KPIs. When time is central to data storage and queries, a dedicated time-series database simplifies architecture and enables more intuitive queries. Understanding InfluxDB InfluxDB is designed to preserve data value by keeping its relationship to time. Instead of treating timestamps as standard columns, InfluxDB organizes data by time-based measurements, attributes, and values. Each record represents an observation at a specific moment, such as a sensor’s temperature, API latency, or asset price. This structure supports direct queries, including retrieving the latest value, assessing recent trends, or calculating averages over defined intervals. This structure is ideal for workloads that produce continuous data. IoT devices generate millions of measurements, infrastructure platforms track ongoing metrics, and financial applications monitor prices and transactions. In these cases, data is usually appended, not updated, and queries often focus on recent values, time ranges, aggregations, or trends. As a result, InfluxDB is well suited for modern architectures. Distributed applications, cloud platforms, microservices, connected devices, and real-time business systems all generate continuous temporal data. Beyond storage, applications have to efficiently query and aggregate this data to assess current conditions and understand system evolution. The key point is that InfluxDB is not only an “IoT database.” InfluxDB is not limited to IoT use cases. It can serve as the temporal component in a more extensive polyglot architecture. For example, a relational database may manage transactional data, a document database may handle flexible aggregates, and InfluxDB can store the history of measurements, signals, and operational changes. When business needs require insight into how something evolves over time, this time-based specialization provides a significant architectural advantage. Hands-On: Java With InfluxDB Starting with Eclipse JNoSQL 1.1.18, Java applications can work with time-series databases through a consistent programming model. Let’s see this in practice with InfluxDB 3. For this example, InfluxDB will run locally without authentication to simplify setup. This approach is suitable for development and demonstrations, but --without-auth must not be used in production. Start InfluxDB 3 Core with Docker: Shell docker run -d \ --name influxdb-instance \ -p 8181:8181 \ influxdb:3.11.0-core \ influxdb3 serve \ --node-id jnosql \ --object-store memory \ --without-auth Once the server is running, create the database used by the application: Java docker exec influxdb-instance \ influxdb3 create database \ --host http://localhost:8181 \ metrics Setting Up the Java Application Eclipse JNoSQL uses Jakarta technologies such as CDI and JSON-B, along with Eclipse MicroProfile Config for externalized configuration. These APIs are available in runtimes including Open Liberty, Payara, Helidon, and Quarkus. Beyond the standard JNoSQL dependencies, add the InfluxDB driver: XML <dependency> <groupId>org.eclipse.jnosql.databases</groupId> <artifactId>jnosql-influxdb</artifactId> <version>${jnosql.version}</version> </dependency> Then configure the connection through microprofile-config.properties: Properties files jnosql.timeseries.database=metrics jnosql.influxdb.url=http://localhost:8181 jnosql.influxdb.token=jnosql-influxdb-test-token Since authentication is disabled for this local instance, the token serves only as a required placeholder for the driver configuration. In production, these values should be externalized and overridden using MicroProfile Config, aligning with Twelve-Factor App principles. Modeling Time-Series Data In this example, account transactions are modeled over time: Java @Entity public class AccountTransaction { @Id private Instant id; @Column private String account; @Column private BigDecimal amount; @Column private String currency; @Column private TransactionStatus status; // constructors, getters, and setters } The transaction status can be represented with a simple enum: Java public enum TransactionStatus { APPROVED, DECLINED, PENDING } The key detail is the Instant identifier. Each transaction records a specific point in time, enabling queries for the latest state and historical navigation. Using TimeSeriesTemplate You can now insert transactions and query them using TimeSeriesTemplate: Java public class App { public static void main(String[] args) { var firstTransaction = new AccountTransaction( Instant.parse("2026-09-20T08:00:00Z"), "account-42", new BigDecimal("79.90"), "EUR", TransactionStatus.APPROVED ); var secondTransaction = new AccountTransaction( Instant.parse("2026-09-20T09:00:00Z"), "account-42", new BigDecimal("24.50"), "EUR", TransactionStatus.APPROVED ); var latestTransaction = new AccountTransaction( Instant.parse("2026-09-20T10:15:00Z"), "account-42", new BigDecimal("120.00"), "EUR", TransactionStatus.DECLINED ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { TimeSeriesTemplate template = container.select(TimeSeriesTemplate.class).get(); template.insert(firstTransaction); template.insert(secondTransaction); template.insert(latestTransaction); var currentStatus = template .select(AccountTransaction.class) .where("account") .eq("account-42") .orderBy("id") .desc() .limit(1) .singleResult(); System.out.println( "Current account status: " + currentStatus ); var history = template .select(AccountTransaction.class) .where("account") .eq("account-42") .orderBy("id") .desc() .skip(1) .limit(10) .result(); System.out.println("Transaction history:"); history.forEach(System.out::println); } } } These queries address two common time-series scenarios: retrieving the latest observation for the account and obtaining its recent history, excluding the current record. Using Jakarta Data The same model can also be exposed through a Jakarta Data repository: Java @Repository public interface AccountTransactionRepository extends BasicRepository<AccountTransaction, Instant> { List<AccountTransaction> findByAccountOrderByIdDesc( String account, Limit limit); } The application code is now repository-oriented: Java public class App2 { public static void main(String[] args) { var firstTransaction = new AccountTransaction( Instant.parse("2026-09-20T08:00:00Z"), "account-42", new BigDecimal("79.90"), "EUR", TransactionStatus.APPROVED ); var secondTransaction = new AccountTransaction( Instant.parse("2026-09-20T09:00:00Z"), "account-42", new BigDecimal("24.50"), "EUR", TransactionStatus.APPROVED ); var latestTransaction = new AccountTransaction( Instant.parse("2026-09-20T10:15:00Z"), "account-42", new BigDecimal("120.00"), "EUR", TransactionStatus.DECLINED ); try (SeContainer container = SeContainerInitializer.newInstance().initialize()) { AccountTransactionRepository repository = container.select(AccountTransactionRepository.class).get(); repository.save(firstTransaction); repository.save(secondTransaction); repository.save(latestTransaction); var currentStatus = repository .findByAccountOrderByIdDesc( "account-42", Limit.of(1) ) .stream() .findFirst(); System.out.println( "Current account status: " + currentStatus ); var history = repository .findByAccountOrderByIdDesc( "account-42", Limit.range(2, 10) ); System.out.println("Recent transaction history:"); history.forEach(System.out::println); } } } Notably, the application expresses time-oriented business queries without relying directly on the InfluxDB client API. Familiar Java mapping and repository concepts are retained, while InfluxDB delivers specialized storage and query capabilities. Where InfluxDB Fits in an Enterprise Architecture While InfluxDB is commonly associated with monitoring or IoT use cases, its role in enterprise architecture is broader. It works best as a specialized time-series store that supplements, rather than replaces, other databases. Relational databases can continue to serve as systems of record for customers, orders, and transactions, while messaging platforms such as Kafka distribute events. InfluxDB then stores time-based operational metrics, telemetry, prices, business indicators, or application signals from these systems. This separation simplifies the overall architecture. Transactional systems are designed for preserving consistency, relationships, and state changes, while time-series databases excel at handling continuous data, recent-state queries, historical ranges, and time-based aggregations. Combining both workloads in a single database may work initially, but as temporal data grows, it frequently leads to complex indexing, partitioning, and retention strategies. For example, an e-commerce platform may store orders and payments in a relational database, while using InfluxDB for checkout latency, payment approval rates, inventory changes, and orders-per-minute. A financial platform can keep transactional records in its core database and use InfluxDB for exchange-rate history, portfolio measurements, or operational risk indicators. Similarly, an industrial platform might store machine metadata in a relational system and use InfluxDB for temperature, pressure, and vibration readings. The architectural value lies in specialization. InfluxDB is most effective when applications need to answer questions like “what is happening now?”, “what changed during this period?”, or “how is this metric trending?” without placing all temporal responsibilities on the transactional database. Conclusion InfluxDB is a strong fit when applications need to work with continuously changing data, recent state, and historical context. With Eclipse JNoSQL 1.1.18, Java developers can use InfluxDB through familiar Jakarta APIs instead of depending directly on database-specific client code, making time-series workloads easier to integrate into modern enterprise applications.

By Otavio Santana DZone Core CORE

Culture and Methodologies

Agile

Career Development

Methodologies

Team Management

Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal

October 5, 2026 by Akhil Madineni DZone Core CORE

Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases

October 2, 2026 by John Vester DZone Core CORE

AI on Top of a Dysfunctional System

October 2, 2026 by Stefan Wolpers DZone Core CORE

Data Engineering

AI/ML

Big Data

Databases

IoT

Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal

October 6, 2026 by Akhil Madineni DZone Core CORE

Building IoT Time-Series Applications With Java and Apache IoTDB

October 6, 2026 by Otavio Santana DZone Core CORE

Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices

October 6, 2026 by Rahul Tewari

Software Design and Architecture

Cloud Architecture

Integration

Microservices

Performance

Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices

October 6, 2026 by Rahul Tewari

OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

October 6, 2026 by Maxim Muzafarov

Nobody Designs an RBAC Mess; Everyone Ends Up With One

October 5, 2026 by Mayank Sethi

Coding

Frameworks

Java

JavaScript

Languages

Tools

Building IoT Time-Series Applications With Java and Apache IoTDB

October 6, 2026 by Otavio Santana DZone Core CORE

Beyond @Transactional: Solving the Dual-Write Problem in Distributed Microservices

October 6, 2026 by Rahul Tewari

OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

October 6, 2026 by Maxim Muzafarov

Testing, Deployment, and Maintenance

Deployment

DevOps and CI/CD

Maintenance

Monitoring and Observability

Reproducible WebRTC Failure Testing With Playwright and coturn

October 5, 2026 by Jay Suresh Nirmal

I Tried Building a "Token Optimization Stack" for Coding Agents. Here's Why I Killed It.

October 5, 2026 by Shreyash Thakare

Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases

October 2, 2026 by John Vester DZone Core CORE

Popular

AI/ML

Java

JavaScript

Open Source

Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal

October 6, 2026 by Akhil Madineni DZone Core CORE

Building IoT Time-Series Applications With Java and Apache IoTDB

October 6, 2026 by Otavio Santana DZone Core CORE

OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule

October 6, 2026 by Maxim Muzafarov

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×