DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Over 2 million developers have joined DZone.
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

The Latest Data Topics

article thumbnail
Building an OCR Data Pipeline: From Unstructured Images to Structured Data
How to treat OCR text as just another data source — build a repeatable ingestion, transformation, and validation workflow for unstructured data.
January 28, 2026
by Punitha Ponnuraj
· 3,107 Views · 1 Like
article thumbnail
Generating Schema-Valid Synthetic ISO 20022 Messages for Privacy-Preserving Fraud Detection
Leverage a schema-aware federated approach to generate synthetic ISO 20022 payments data with strict personal information privacy and XSD compliance.
January 28, 2026
by Senthilnathan Dhanasekaran
· 1,197 Views
article thumbnail
Building an Internal Document Search Tool with Retrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation (RAG) is transforming enterprise AI by bridging the gap between general-purpose language models and organization-specific knowledge.
January 27, 2026
by Manish Adawadkar
· 3,211 Views
article thumbnail
Cost-Aware GenAI Architecture: Caching, Model Routing, and Token Budgets That Don’t Explode
Keep GenAI cheap and fast: cache aggressively, route models by confidence, cap tokens and tools, compress context, and monitor cost per successful outcome.
January 27, 2026
by Mohan Sankaran
· 3,092 Views · 5 Likes
article thumbnail
Building Fault-Tolerant Data Pipelines in GCP
This article provides a practical guide to building a fault-tolerant Google Cloud data pipeline architecture with Firestore, Pub/Sub, Dataflow, and BigQuery.
January 26, 2026
by Krishnam Raju Narsepalle
· 1,219 Views
article thumbnail
Vibe Coding Part 3 — Building a Data Quality Framework in Scala and PySpark
How GenAI-assisted “vibe coding” can simplify building PySpark wrappers over Scala Spark libraries when you know what you’re doing.
January 23, 2026
by Bipin Patwardhan
· 1,339 Views · 2 Likes
article thumbnail
Securing AI/ML Workloads in the Cloud: Integrating DevSecOps with MLOps
ML systems introduce security risks most teams aren’t prepared for. The piece explores emerging ML-specific threats and what effective MLSecOps looks like in practice.
January 23, 2026
by Igboanugo David Ugochukwu DZone Core CORE
· 2,728 Views · 1 Like
article thumbnail
Design and Implementation of Cloud-Native Microservice Architectures for Scalable Insurance Analytics Platforms
How cloud-native microservices transform insurance analytics by enabling scalability, real-time processing, and seamless modernization of legacy platforms.
January 23, 2026
by Afroz Mohammad
· 2,053 Views · 1 Like
article thumbnail
Efficient Sampling Approach for Large Datasets
In this article, we will learn about the central limit theorem and how it helps with random sampling in big-data-related problems.
January 22, 2026
by Rajesh Vakkalagadda
· 1,259 Views
article thumbnail
Automating Visual Brand Compliance: A Multimodal LLM Approach
Manual review of marketing assets for brand consistency is a bottleneck. Here is an architectural pattern for building a compliance tool using Multimodal LLMs and Python.
January 22, 2026
by Dippu Kumar Singh
· 2,334 Views
article thumbnail
Why Semantic Layers Matter in Analytics: A Deep Dive into RAG Design
Analytics assistants/chatbots should trust the semantic layer — not documents. Retrieve metric definitions, run governed SQL, and attach an audit bundle to every KPI.
January 22, 2026
by Anusha Kovi DZone Core CORE
· 1,453 Views · 1 Like
article thumbnail
Data Engineering: Strategies for Data Retrieval on Multi-Dimensional Data
A comparative look at partitioning, indexing, clustering, and ordering techniques to match retrieval strategies with real-world query needs.
January 22, 2026
by Avi Yehuda
· 1,356 Views
article thumbnail
The Future of Data Streaming with Apache Flink for Agentic AI
Flink and Kafka enable real-time agentic AI by streaming fresh data and model context via the MCP standard for intelligent actions at scale.
January 21, 2026
by Kai Wähner DZone Core CORE
· 2,087 Views
article thumbnail
Where AI Fits and Fails in Workday Integrations
AI enhances Workday integrations by improving mapping, testing, and monitoring, but it fails when used without human oversight, domain expertise, and strong governance.
January 21, 2026
by Suresh Kurapati
· 2,619 Views · 1 Like
article thumbnail
MERGE and Liquid Clustering: Common Performance Issues
A practical look at common pitfalls and performance challenges when using MERGE operations on liquid-clustered Delta tables, and how to avoid them.
January 21, 2026
by Avi Yehuda
· 1,722 Views
article thumbnail
Caching Issues With the Spring Expression Language
Spring Expression Language is a flexible way to evaluate expressions at runtime. However, in the context of caching, this flexibility can lead to errors.
January 20, 2026
by Constantin Kwiatkowski
· 2,427 Views
article thumbnail
Top 5 Payment Gateway APIs for Indian SaaS: A Developer’s Analysis
Compare 5 international payment APIs for Indian SaaS. Choose APIs that enable automated FIRC retrieval — and test the sandbox.
January 19, 2026
by Sarang S Babu
· 5,576 Views
article thumbnail
RAG at Scale: The Data Engineering Challenges
The data engineering challenges that go beyond the basic concept of RAG when running RAG systems at scale in production.
January 16, 2026
by Guru Hegde
· 2,708 Views · 1 Like
article thumbnail
Parallel S3 Writes for Massive Sparse DataFrames: How to Maintain Row Order Without Blowing Memory
Learn how to write massive sparse Pandas DataFrames to S3 without OOM errors by using Spark to parallelize index-based chunks while preserving row order.
January 16, 2026
by pooja chhabra
· 1,601 Views · 2 Likes
article thumbnail
RAG on Android Done Right: Local Vector Cache Plus Cloud Retrieval Architecture
Most mobile RAG fails on latency and flaky networks. A local vector cache + cloud retrieval architecture keeps responses fast, fresh, and grounded.
January 16, 2026
by Mohan Sankaran
· 2,019 Views · 5 Likes
  • Previous
  • ...
  • 11
  • 12
  • 13
  • 14
  • 15
  • 16
  • 17
  • 18
  • 19
  • 20
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×