DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Over 2 million developers have joined DZone.
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

The Latest Big Data Topics

article thumbnail
When Coalesce Is Slower Than Repartition: A Spark Performance Paradox
In this article, learn why repartition() can outperform coalesce() in Apache Spark — and how Catalyst optimizer pushdown can throttle your job’s parallelism.
October 30, 2025
by Janani Annur Thiruvengadam DZone Core CORE
· 3,229 Views · 3 Likes
article thumbnail
Debugging a Spark Driver Out of Memory (OOM) Issue With Large JSON Data Processing
This article draws on real-world debugging experience and aims to provide insights into Spark's memory management challenges.
October 29, 2025
by Raju Ansari
· 2,619 Views · 2 Likes
article thumbnail
Unlocking Scalable Data Lakes: Building With Apache Iceberg, AWS Glue, and S3
Apache Iceberg + AWS Glue + S3 bring ACID, schema evolution, and time travel to data lakes—fixing schema drift, small files, and cost sprawl at enterprise scale.
October 28, 2025
by Vivek Venkatesan
· 4,022 Views · 1 Like
article thumbnail
Set Up Spring Data Elasticsearch With Basic Authentication
Guide to configure SSL communication with Elasticsearch via Spring Data Elasticsearch. Additionally, the communication is secured with BASIC authentication.
October 27, 2025
by Arnošt Havelka DZone Core CORE
· 2,363 Views
article thumbnail
Enterprise-Grade Document Intelligence: Cloud Big Data AI With YOLOv9 and Spark on AWS
Automate document analysis with YOLOv9, Apache Spark, and AWS. Boost speed, accuracy, and fraud detection across finance, healthcare, insurance, and more.
October 27, 2025
by Ram Ghadiyaram DZone Core CORE
· 2,701 Views · 4 Likes
article thumbnail
The Dark Side of Apache Iceberg’s Data Time Travel Feature
This article talks about the hidden aspects of the Apache Iceberg Time Travel Query feature. It also highlights how to address those hidden negative aspects.
October 22, 2025
by Pravin Dwiwedi
· 3,345 Views · 1 Like
article thumbnail
Building Scalable CRM Systems: Architecture Patterns and Data Modeling Strategies
A hands-on guide to building scalable CRM systems with the right architecture, data models, and performance and security strategies.
October 22, 2025
by Chitrapradha Ganesan
· 5,019 Views · 2 Likes
article thumbnail
A Fresh Look at Optimizing Apache Spark Programs
Optimize Spark jobs by tuning configurations, writing efficient code (Data Frames, broadcast joins), using optimized storage, and monitoring the Spark UI and logs.
October 14, 2025
by Nataraj Mocherla
· 3,570 Views · 2 Likes
article thumbnail
How Developers Use Synthetic Data to Stress-Test Models in Noisy Markets
Synthetic data lets quants stress-test equity strategies beyond noisy markets, preserving volatility, and building resilience before risking real capital.
October 14, 2025
by Jay Mehta
· 1,617 Views · 1 Like
article thumbnail
Operationalizing Responsible AI: Turning Ethics Into Engineering
This article will provide a direction on how to build a reliable AI system in production by incorporating bias mitigation strategies.
October 13, 2025
by Jofia Jose Prakash
· 2,402 Views · 2 Likes
article thumbnail
Apache Iceberg REST Catalog: The Key to Vendor-Agnostic Data Interoperability
In this article, I have demonstrated how Iceberg Data can be accessed through the Iceberg REST Catalog from Data Mesh with a simple Python application.
October 13, 2025
by Pravin Dwiwedi
· 2,831 Views · 1 Like
article thumbnail
Introduction to Spring Data Elasticsearch 5.5
Getting started with the latest version of Spring Data Elasticsearch 5.5 and Elasticsearch 8.18 as a NoSQL database for our data storage.
October 10, 2025
by Arnošt Havelka DZone Core CORE
· 3,340 Views · 2 Likes
article thumbnail
8 Challenges in Multimodal Training Data Creation
Creating high-quality multimodal training data is essential yet complex, involving challenges in synchronization, scalability, context capture, and tooling.
October 8, 2025
by Chirag Shivalker
· 3,354 Views
article thumbnail
7 AWS Services Every Data Engineer Should Master
In 2025, S3, Glue, Lambda, Athena, Redshift, EMR, and Kinesis form the core AWS toolkit for building fast, reliable, and scalable data pipelines.
October 6, 2025
by Sai Mounika Yedlapalli
· 4,522 Views · 3 Likes
article thumbnail
From Big Data to Agents: My Decade Building Systems
How a simple scraper, a few dashboards, and a lot of curiosity turned into agentic systems that actually ship value. A builder’s path.
October 3, 2025
by Nacho Corcuera
· 2,904 Views · 2 Likes
article thumbnail
Building a Scalable and Reliable Marketing Data Stack on GCP
A resilient marketing data stack on GCP leverages BigQuery, Pub/Sub, and Dataflow to deliver real-time insights, handle schema drift, and scale analytics.
October 2, 2025
by Shafeeq Ur Rahaman
· 1,612 Views · 2 Likes
article thumbnail
Salesforce Data Cloud: Setting Up and Using the Ingestion API
In this guide, learn to use Salesforce Data Cloud Ingestion API for real-time and bulk data ingestion to deliver accurate, personalized customer experiences.
October 2, 2025
by Ramesh Bellamkonda
· 4,595 Views · 2 Likes
article thumbnail
Implementing Governance on Databricks Using Unity Catalog
We learn how to treat data as a product through governance in Unity Catalog, ensuring the right people, metadata about the datasets.
October 1, 2025
by Junaith Haja
· 3,097 Views · 3 Likes
article thumbnail
Master Advanced Error-Handling to Make PySpark Pipelines Production-Ready
PySpark jobs often fail because of bad data, network issues, or logic errors. Sometimes, after hours of processing. Learn how to make your Spark pipelines more reliable.
September 30, 2025
by Ram Ghadiyaram DZone Core CORE
· 5,622 Views · 6 Likes
article thumbnail
Complex Data Tasks Are Now One-Liners With AI in Databricks SQL
In this guide, learn how to simplify data tasks with AI in Databricks SQL — summarize, translate, analyze sentiment, and mask PII with one-liner queries.
September 26, 2025
by Junaith Haja
· 3,863 Views · 2 Likes
  • Previous
  • ...
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • 11
  • 12
  • 13
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×