DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Performance

Performance refers to how well an application conducts itself compared to an expected level of service. Today's environments are increasingly complex and typically involve loosely coupled architectures, making it difficult to pinpoint bottlenecks in your system. Whatever your performance troubles, this Zone has you covered with everything from root cause analysis, application monitoring, and log management to anomaly detection, observability, and performance testing.

icon
Latest Premium Content
Trend Report
Observability and Performance
Observability and Performance
Refcard #290
Getting Started With Log Management
Getting Started With Log Management
Refcard #385
Observability Maturity Model
Observability Maturity Model

DZone's Featured Performance Resources

Improving Repeated Analytics Workloads With Databricks Disk Cache

Improving Repeated Analytics Workloads With Databricks Disk Cache

By Harsh Patel
In many analytics platforms, there are performance issues that do not always come from complex transformations. Sometimes the bottleneck is much simpler: the same large datasets are being read repeatedly from remote storage. This pattern is common in shared analytics environments. A data engineering job reads a curated dataset to build aggregates. A BI refresh reads the same table again. A data science notebook filters the same records during exploration. Another scheduled workflow joins against the same reference data several times during the day. Each workload may be valid on its own, but together they create repeated remote reads. Over time, this can increase query latency, consume unnecessary infrastructure resources, and make interactive analytics feel slower than expected. Databricks disk cache is designed to help with this type of workload. It stores copies of remote Parquet data files on the local storage of worker nodes so that repeated reads can be served locally instead of fetching the same files again from cloud object storage. This article walks through a practical use case for using Databricks disk cache to improve repeated analytics workloads. The focus is not simply on enabling a feature, but on understanding when disk cache helps, where it fits in a pipeline, and what tradeoffs teams should consider before relying on it. The Use Case: Repeated Reads From Curated Analytics Tables Consider a common analytics setup. A team maintains a curated dataset that is used by multiple downstream workloads. The table is stored in cloud object storage and accessed through Databricks. It is already cleaned, standardized, and partitioned by date. Several jobs and users access this table throughout the day. The dataset supports different types of work: dashboard refreshesscheduled aggregationsexploratory notebooksfeature preparation jobsad hoc analysisdownstream transformation pipelines. The problem is not that the table is poorly designed. The problem is that the same files are repeatedly scanned from remote storage. In this situation, the first read of the data still needs to fetch files from remote storage. However, after the data is cached locally on worker nodes, repeated reads can avoid some of that remote access. For workloads that repeatedly query overlapping data, this can make a noticeable difference. This use case is especially relevant when teams work with large Parquet or Delta tables where the same filtered slices are accessed multiple times. Where Disk Cache Fits in the Pipeline Disk cache is not a replacement for good data modeling, partitioning, or query optimization. It works best as an acceleration layer for workloads that already read reasonably structured data. A practical architecture may look like this: Data Architecture Pipeline With Cache Layer The important point is that disk cache usually adds the most value after data has already been curated. If raw data is messy, unpartitioned, or constantly changing, caching alone will not solve the deeper performance problem. A better pattern is to first create reliable curated datasets and then use disk cache to improve workloads that repeatedly read those datasets. Why Repeated Reads Become Expensive Cloud object storage is highly scalable, but repeatedly reading the same large files still introduces overhead. A query may need to: locate filesread metadatafetch data over the networkdeserialize columnar datascan partitionsapply filterspass data into downstream transformations When one workflow performs this operation, the cost may be acceptable. When several workloads read the same dataset repeatedly, the overhead becomes more visible. This is especially noticeable in interactive analytics. A user may run one query, adjust a filter, run another query, and continue exploring. If every query repeatedly fetches the same underlying files from remote storage, the user experience can degrade quickly. Disk cache helps by keeping frequently accessed data closer to the compute layer. Disk Cache vs Spark Cache One source of confusion is the difference between Databricks disk cache and Apache Spark cache. Spark cache is usually applied manually to a DataFrame or table. It is useful when a specific intermediate result will be reused within the same job or notebook. However, Spark cache requires the developer to decide what to cache and when to unpersist it. Databricks disk cache behaves differently. It works at the file-read level and stores remote Parquet data files locally on worker nodes. When the same data is read again, Databricks can serve it from local disk instead of fetching it again from remote storage. A simple way to think about the difference is this: Spark Cache Developer-controlledApplied to DataFrames or RDDsUseful for reused intermediate resultsRequires explicit cache management. Databricks Disk Cache Managed by DatabricksApplied to remote Parquet/Delta file readsUseful for repeated reads from storageUses local worker disk. In practice, these two caching approaches solve different problems. Spark cache is useful when the same transformed DataFrame is reused multiple times inside a workload. Disk cache is useful when workloads repeatedly scan the same remote Parquet or Delta files. Using the wrong caching strategy can lead to unnecessary memory pressure, unstable performance, or no real improvement. A Practical Example Without Making It Industry-Specific Assume an organization maintains a large curated events table. The table contains activity records from different systems and is used for reporting, operational analytics, and product usage analysis. Several teams query this dataset daily. One dashboard refresh reads the last 30 days of activity. A transformation job reads the same table to calculate weekly aggregates. Analysts use notebooks to filter the data by region, product, and time period. Another pipeline reads the same table to prepare downstream metrics. Even though the consumers are different, many of them repeatedly access the same recent partitions. Without disk cache, these workloads repeatedly read files from remote storage. With disk cache, frequently accessed Parquet files can be stored locally on workers after the first read, allowing later reads to avoid repeated remote fetches. This is not a dramatic redesign of the pipeline. It is an optimization layer that improves workloads with repeated access patterns. When Disk Cache Helps Disk cache is most useful when workloads repeatedly read the same data files. Good candidates include: frequently queried Delta or Parquet tablesdashboard refreshes that scan the same recent partitionsexploratory notebooks that repeatedly filter the same datasetshared reference tables used across multiple joinsiterative analytics workflowsrepeated batch jobs using overlapping input data. The key pattern is repeated access. If every job reads a completely different dataset, disk cache will have limited benefit. If data is accessed once and never reused, the first read still has to fetch the files from remote storage. Disk cache is most effective when the same data is accessed more than once by workloads running on the same or similar compute resources. When Disk Cache May Not Help Much Caching is not a universal performance solution. Disk cache may provide limited improvement when: workloads read data only oncetables change constantlyqueries scan entirely different partitions each timetransformations are CPU-bound rather than I/O-boundjoins and shuffles dominate execution timeclusters are frequently restartedworker nodes are frequently replaced. This last point matters in elastic environments. If workers are decommissioned, local cache data on those workers is lost. The next workload may need to reread data from remote storage. This does not make disk cache unreliable. It simply means teams should understand its behavior before treating it as a guaranteed performance layer. How To Evaluate Whether Disk Cache Is Helping A common mistake is assuming that caching is helping just because it is enabled. A better approach is to compare workload behavior before and after repeated reads. Useful evaluation questions include: Does the second run complete faster than the first run?Are repeated queries reading overlapping data?Is the workload I/O-bound or shuffle-bound?Are the same partitions being scanned repeatedly?Are clusters stable long enough for cache reuse?Are users querying curated tables or constantly changing raw data? Teams should also compare job execution stages. If most time is spent reading remote files, disk cache can help. If most time is spent in large joins, aggregations, or shuffles, caching file reads may only improve part of the workload. Performance tuning should start with measurement, not assumptions. Designing Pipelines To Benefit From Disk Cache To get value from disk cache, the pipeline should be designed in a way that encourages reusable reads. One practical pattern is to separate raw ingestion from curated analytical datasets. Raw data may be inconsistent, frequently updated, and can be less suitable for repeated consumption, while curated datasets are usually cleaner, more stable, and more likely to be accessed repeatedly. A stronger design looks like this: Designing Pipelines for Disk Cache Optimization This design allows disk cache to work on datasets that are already optimized for downstream use. Partitioning also matters. If tables are partitioned in a way that matches query patterns, repeated workloads are more likely to access the same files, if partitioning is poorly aligned with usage patterns then queries may scan too much unnecessary data which would reduce the benefit of caching. For example, if most users query recent data, organizing the table around time-based access patterns can make repeated reads more efficient. Disk cache should be viewed as part of a broader performance strategy, not as a substitute for table design. Operational Considerations There are a few operational details teams should consider before depending heavily on disk cache. First, disk cache depends on local storage on worker nodes. Choosing worker types with local SSD storage can improve caching effectiveness. Second, cache behavior is tied to the lifecycle of the compute environment. If clusters restart frequently, cached data may not persist long enough to benefit repeated workloads. Third, disk cache works best when workloads have predictable reuse patterns. Highly random access patterns are less likely to benefit. Fourth, teams should monitor whether performance improvements are consistent. If query times vary significantly, the issue may not be remote reads alone. The bottleneck may be skewed partitions, insufficient cluster resources, poor join strategy, or inefficient transformations. Finally, caching should not be used to hide poor pipeline design. If a table is too wide, poorly partitioned, or filled with unnecessary historical data, disk cache may improve repeated reads but will not fix the underlying design problem. Avoiding Common Mistakes A few mistakes appear frequently when teams start relying on caching. The first mistake is caching too early in the pipeline. Raw datasets are often unstable and less useful for repeated analytical access. Caching is more valuable after data has been cleaned, standardized, and organized for consumption. The second mistake is confusing disk cache with Spark cache. Spark cache is useful for reused intermediate DataFrames. Disk cache is better suited for repeated reads of remote Parquet or Delta files. The third mistake is ignoring cluster behavior. If compute resources are short-lived, cache reuse may be limited. The fourth mistake is measuring only one query run. Since disk cache is useful for repeated reads, teams should compare cold-read and warm-read behavior rather than judging performance from a single execution. The fifth mistake is treating disk cache as a substitute for optimization. Good partitioning, file sizing, query filtering, and transformation design still matter. Practical Checklist Before depending on disk cache, teams should ask: Are the same datasets read repeatedly?Are workloads reading Parquet or Delta data?Are the tables curated and reasonably stable?Are query patterns predictable?Are clusters stable enough for cache reuse?Are bottlenecks related to file reads rather than shuffles?Are partitions aligned with common access patterns?Are performance gains measured across repeated runs? If the answer to most of these questions is yes, disk cache is likely worth evaluating. If the answer is no, teams should first investigate table design, query plans, file layout, and transformation logic. Conclusion Databricks disk cache can be a useful optimization for analytics workloads that repeatedly read the same Parquet or Delta data from remote storage. It is especially helpful for curated datasets used by dashboards, notebooks, scheduled jobs, and downstream analytics workflows. However, disk cache should not be treated as a general solution for every performance issue. It works best when data access patterns are repeated, compute resources remain stable, and the underlying tables are already designed reasonably well. The biggest lesson is that caching should be intentional. Teams should understand where repeated reads happen, measure cold-read and warm-read behavior, and combine disk cache with good table design, partitioning, and pipeline structure. When used in the right context, disk cache can reduce repeated remote reads and make analytics workloads more responsive. When used without understanding the workload, it becomes just another configuration setting with unclear impact. Reliable analytics performance comes from knowing which bottleneck is actually being solved. More
Memory-First Indexes in SQL Server 2025: Redefining Performance for Hybrid Workloads

Memory-First Indexes in SQL Server 2025: Redefining Performance for Hybrid Workloads

By arvind toorpu DZone Core CORE
Modern database environments rarely run a single type of workload. Most production systems handle both transactional operations and analytical queries simultaneously. These mixed workloads, often referred to as hybrid workloads, place significant pressure on traditional database indexing and storage strategies. In such environments, disk-based indexes can become a performance bottleneck. When transactional and analytical queries compete for disk I/O, it often results in increased latency, reduced throughput, and inconsistent query performance. To address these challenges, SQL Server leverages memory-optimized tables and indexes as part of its In-Memory OLTP capabilities. These features reduce reliance on disk I/O by enabling data and index access directly from memory, while still maintaining durability through logging and checkpoint mechanisms. This article explores how memory-optimized indexing works and demonstrates how it can significantly improve performance in real-world hybrid workload scenarios. Core Characteristics Mandatory inclusion: Every memory-optimized table must have at least one index, as they serve as the "entry points" for row access.Purely in-memory: Indexes are rebuilt entirely from scratch during database recovery based on their definitions and the data loaded into memory.Non-persistent: Unlike traditional indexes, changes to these indexes are not written to the transaction log, reducing I/O overhead.Fragmentation-free: These structures do not suffer from traditional page fragmentation, eliminating the need for regular REORGANIZE or REBUILD operations. Index TypeBest Use CaseBehaviorHash IndexEquality SearchesUses an array of buckets; highly efficient for point lookups (e.g., WHERE ID = 5).Nonclustered IndexRange QueriesUses a lock-free B-tree structure (Bw-tree); ideal for range scans and sorted results (e.g., WHERE Price > 100). The Challenge With Traditional Indexing Traditionally, database indexes are stored on disk to ensure durability. While this design protects data, it introduces a major limitation: disk I/O latency. In environments with heavy workloads, disk access becomes a bottleneck. This is particularly noticeable when: Large analytical queries scan index rangesTransactional queries require fast point lookupsMany concurrent users access the system When both workloads run together, index operations often compete for disk resources, resulting in slower queries and higher latency. Introducing Memory-First Indexes Memory-First Indexes in SQL Server 2025 take a different approach. Instead of relying primarily on disk-based indexes, the system prioritizes in-memory index access for frequently used data while maintaining a synchronized copy on disk for durability. The key idea is simple: Hot data (frequently accessed index ranges) is kept in memory.Cold data remains on disk.Changes made in memory are synchronized with disk replicas in the background. This approach allows SQL Server to serve many queries directly from memory while still maintaining persistence. The feature also includes monitoring mechanisms that track query patterns. When the system detects frequently accessed index partitions, it moves them into memory automatically. Less frequently accessed portions are pushed back to disk to conserve memory resources. The result is faster query execution without requiring manual tuning from database administrators. Real-World Example: Retail E-Commerce Database To understand the benefits, consider a retail company running an e-commerce platform. The company stores millions of products in a table with the following structure: ProductID – unique identifierProductCategory – category of the productPrice – product priceStockQuantity – available inventory The application runs two types of queries. Transactional Query This query checks stock availability for a specific product. SQL SELECT StockQuantity FROM Products WHERE ProductID = 102345; Analytical Query This query calculates aggregated metrics by product category. SQL SELECT ProductCategory, AVG(Price) AS AvgPrice, SUM(StockQuantity) AS TotalStock FROM Products WHERE Price > 500 GROUP BY ProductCategory; In a traditional setup, both queries rely on disk-based indexes. When concurrency increases, disk access becomes saturated, and query performance suffers. With Memory-First Indexes, the most frequently used index ranges, such as ProductID and ProductCategory, are loaded into memory, allowing much faster lookups. Testing the Feature To evaluate the impact of Memory-First Indexes, we can simulate a large dataset and compare query performance before and after enabling the feature. Step 1: Create the Table SQL CREATE TABLE Products ( ProductID INT PRIMARY KEY, ProductCategory NVARCHAR(50), Price DECIMAL(10,2), StockQuantity INT ); Step 2: Populate Test Data The following script generates a large dataset for testing. SQL INSERT INTO Products (ProductID, ProductCategory, Price, StockQuantity) SELECT TOP 50000000 ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) AS ProductID, CASE WHEN ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) % 5 = 1 THEN 'Electronics' WHEN ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) % 5 = 2 THEN 'Clothing' ELSE 'Home Appliances' END AS ProductCategory, ABS(CHECKSUM(NEWID()) % 1000) + 1.00 AS Price, ABS(CHECKSUM(NEWID()) % 5000) + 1 AS StockQuantity FROM sys.all_objects a CROSS JOIN sys.all_objects b; Step 3: Create Traditional Indexes SQL CREATE INDEX IX_Products_ProductID ON Products (ProductID); CREATE INDEX IX_Products_Category ON Products (ProductCategory); At this stage, run the transactional and analytical queries and capture baseline metrics using Query Store or dynamic management views. Step 4: Enable Memory-First Indexes Next, recreate the indexes with Memory-First enabled. SQL DROP INDEX IX_Products_ProductID ON Products; CREATE INDEX IX_Products_ProductID ON Products (ProductID) WITH (MEMORY_FIRST = ON); DROP INDEX IX_Products_Category ON Products; CREATE INDEX IX_Products_Category ON Products (ProductCategory) WITH (MEMORY_FIRST = ON); Step 5: Execute Test Queries SQL SELECT StockQuantity FROM Products WHERE ProductID = 102345; MS SQL SELECT ProductCategory, AVG(Price) AS AvgPrice, SUM(StockQuantity) AS TotalStock FROM Products WHERE Price > 500 GROUP BY ProductCategory; Record execution time, CPU usage, and disk activity again. Observed Performance Improvements The results typically show noticeable performance gains. For example: Transactional queries Before: ~50 msAfter: ~15 ms Analytical queries Execution time reduced by about 50% System metrics also reveal additional improvements: Disk I/O reduced by more than 70%Memory usage increased only moderatelyCPU utilization became more stable during peak workloads These improvements occur because queries are able to retrieve indexed data directly from memory rather than waiting for disk operations. Why This Matters for Modern Workloads Hybrid workloads are becoming the norm across many industries, including retail, finance, and IoT platforms. Systems must support both real-time transactions and large analytical queries without sacrificing performance. Memory-First Indexes help address this challenge by: Reducing disk I/O bottlenecksImproving response time for critical queriesAutomatically adapting to changing workload patternsMaintaining durability with synchronized disk replicas Final Thoughts Memory-First Indexes represent an important improvement in SQL Server 2025’s indexing architecture. By prioritizing in-memory access for frequently used data, SQL Server can deliver significantly faster query performance while still preserving data durability. For organizations running mixed transactional and analytical workloads, this feature can reduce latency, improve system stability, and make better use of available hardware resources. As hybrid workloads continue to grow, features like Memory-First Indexing will play a key role in helping database platforms keep up with modern application demands. More
Cutting Telemetry Volume Is Not the Same as Cutting Noise
Cutting Telemetry Volume Is Not the Same as Cutting Noise
By Severin Neumann
Optimize an AI Agent to Sound Human, Judged by an AI Detector
Optimize an AI Agent to Sound Human, Judged by an AI Detector
By Scarlett Attensil
Ampere System Profiler: A Guide to System-Level Profiling
Ampere System Profiler: A Guide to System-Level Profiling
By Tito Reinhart
What Actually Makes AI Infrastructure Agents More Reliable (It's Not More Agents)
What Actually Makes AI Infrastructure Agents More Reliable (It's Not More Agents)

I keep seeing the same pattern. Someone builds an "AI agent" for infrastructure monitoring — it answers questions about Prometheus metrics, pulls logs from ELK, suggests restarts. Impressive in a demo. Then you push on it: what happens when its logs query times out mid-investigation? What happens when the context window fills up while correlating signals across four systems? What happens when a tool call hallucinates a metric name that doesn't quite exist? Usually it doesn't fail catastrophically. It fails quietly, in ways that are hard to debug. And quiet failures during incident response are the worst kind. I've spent the last year prototyping multi-agent architectures for infrastructure observability — coming at it from a decade of SRE and network reliability work. My working hypothesis going in was straightforward: split the investigation across specialized agents, and the reliability problems that plague a single agent — a blown context window, a hallucinated tool call — should ease. Below is the architecture pattern I built to test that, the failure modes I watched it run into, and — since I've since put the hypothesis through a more rigorous test than a demo — what actually held up. Why a Single Agent Hits a Ceiling A good on-call engineer doesn't open one dashboard and stare. They move between tools, check Prometheus, scan logs, look at deploy history, consult runbooks. Each step informs the next. Investigations have structure. A single LLM agent trying to replicate that workflow runs into two real constraints. First, the context window. Every tool call, every metric result, every log snippet goes into that window. Short investigation: fine. Anything complex — multiple services, ambiguous signals, a failure mode the model hasn't seen — and the window fills. Early observations get pushed out. The model loses the thread. Second, the tool problem. The more tools you give a single agent, the more likely it is to hallucinate one — invoking a function that doesn't exist, or constructing a query with the right name but the wrong parameter. I've reproduced this in my own prototyping: the agent confidently calls a metric query with a typo, gets an empty result, and concludes the metric doesn't exist rather than that the query was wrong. Split the work across specialized agents, and both pressures ease. Each agent owns a smaller tool set it actually knows. Each agent has a manageable context. And when one agent fails — they will, eventually — it fails in a bounded, debuggable way. The Architecture: Four Roles, One Investigation The pattern I keep returning to has four agent roles, communicating through shared state rather than direct message-passing. The state is a typed Python object accumulating findings as the investigation progresses. No agent starts from scratch; each picks up where the last one left off. Python class InvestigationState(TypedDict): incident_id: str trigger: AlertTrigger telemetry_findings: list[Finding] causal_hypothesis: Optional[Hypothesis] recommended_actions: list[Action] confidence_score: float audit_trail: list[AgentStep] Every agent reads from and writes back to this state. That single design choice — shared structured state instead of free-form message-passing — is what makes the system auditable. The four roles: Supervisor Agent receives the raw alert. It doesn't investigate — it routes. It classifies the incident, identifies the services involved, and decides which specialist agents to invoke. Telemetry Investigation Agent is the data-gathering specialist. Given an investigation context, it runs queries against the observability stack — Prometheus, Grafana, ELK, AppDynamics — finds anomalies, and returns structured findings. It doesn't explain what it finds. It just finds things. Python def route_to_specialists(state: InvestigationState) -> list[str]: trigger_type = state["trigger"].classification if trigger_type == "network": return ["telemetry", "reasoning"] elif trigger_type == "application": return ["telemetry", "logs", "reasoning"] else: return ["telemetry", "logs", "infra", "reasoning"] Reasoning Agent takes the Telemetry Agent's findings and tries to answer: what is actually going on? Given a RAG index over historical post-mortems, it can reason like "this pattern resembles a connection-pool exhaustion failure mode I've seen documented before." When it works, the experience is impressive. When it's wrong, it's confidently and eloquently wrong — a class of failure I'll come back to. Action Agent turns a hypothesis into something executable. For low-risk actions, it could, in principle, act autonomously once confidence crosses a threshold. For anything riskier, it drafts a recommendation with full context and routes to a human for approval. I'd treat that human-in-the-loop gate as non-negotiable for any first deployment. A Worked Example (Prototype, Not Production) Python # Investigation trace — synthetic test environment # Alert received: payments-service p99 latency anomaly # Supervisor → routing to Telemetry Agent # Telemetry: db-proxy connection pool utilization elevated # Telemetry: db-proxy deployed recently (within last 20 min) # Telemetry: no downstream dependency anomalies # Reasoning: hypothesis — connection-pool regression in recent deploy # Action: recommend rollback to prior db-proxy version # Action: draft escalation with evidence → human approval required The point of the multi-agent system isn't to replace the engineer. It's to do the legwork before the human even opens their laptop, so the human is reviewing evidence rather than gathering it. Where This Pattern Breaks Three failure modes worth naming: Confidence scores are not well-calibrated. An 87% confidence score sounds authoritative. Language models don't express uncertainty the way a careful engineer would. Any deployment needs a conservative threshold for autonomous action and a generous fallback to human review. Context grows faster than you expect. Five services, 15 tool calls of data in shared state, and the Reasoning Agent starts dropping things. State summarization helps, but it's lossy. Agent observability is its own problem. You're building a system that monitors infrastructure, and now you need to monitor the monitor. Without per-step tracing, debugging an agent failure is genuinely painful. What a Rigorous Test Actually Showed Everything above is prototype-stage reasoning — the kind you form watching a system work and fail in front of you. I didn't want to leave it there, so I ran the architecture against two real fault-injection benchmarks, AIOps Challenge 2020 and RCAEval, 75 incidents each, across six pre-registered configurations comparing the four-role design against a well-built single agent given the same tools and the same context budget. Decomposition alone didn't win. Across both benchmarks, the multi-agent version came out statistically indistinguishable from the single agent (McNemar's test, p > 0.05 in every configuration), and a plain rule-based baseline stayed competitive with both. That's not the result I expected going in, and it's worth sitting with rather than explaining away: splitting an investigation into roles does not, by itself, make it more accurate. What did move the needle was a narrower idea: a Falsifier agent that checks the Reasoning agent's hypothesis against evidence it wasn't shown, instead of taking the hypothesis at its word. That improved accuracy on single-service incidents meaningfully — 33.3% vs. 21.3%, p = 0.023 — and made multi-service incidents worse at first — 24.0% vs. 42.7%, p < 0.001 — because with no notion of which services depend on which, the falsifier mistook a downstream symptom for the root cause. Giving it real service-topology data closed that gap. Then the less comfortable check: I gave the same falsifier to a single agent instead of the four-role pipeline. It scored indistinguishably from the multi-agent version (p = 1.0). The gain wasn't coming from decomposition. It was coming from the verification step, and the verification step doesn't care how many agents are asking the question. This work is accepted at CNSM 2026 (IFIP); the full benchmark, raw results, and eval scripts are in the repo linked below, and the falsifier design specifically is written up in more depth in the preprint linked at the end of this article. Implementation Notes LangGraph fits this pattern well. The explicit graph model lets you define exactly what happens after each agent step. The graph is code — versionable and testable. For tool management, typed schemas validated before execution eliminate most hallucinated tool calls. The discipline is the same regardless of framework: every tool input is typed, every tool call is validated, nothing executes on an unstructured string. If I had to give one piece of advice: invest in your tools before you invest in your prompts. The ceiling on what an agent can do is set by the quality of the tool interface, not the eloquence of the system prompt. So, Why Multi-Agent? Because single agents fail in ways that are hard to predict and hard to debug, and because bounded roles with structured shared state make an investigation's failures easier to trace, whatever the accuracy numbers say. But I'd stop short of the clean version of this pitch. My own testing didn't support "multi-agent is more reliable" as a general claim — it supported something narrower: a verification step that checks a hypothesis against evidence it hasn't seen is what earns its complexity, and you can bolt that onto a single agent just as well as onto four. The four-role design is still a reasonable way to build one of these systems — the shared-state pattern, the tiered autonomy, the human gate on risky actions are all still doing real work. Just don't assume the agent count is what's buying you the reliability. Test that part before you ship it. The working prototype for this architecture is available at: github.com/Kinjal-Oza/multi-agent-observability-demo Originally published on Medium.

By Kinjal Vaishnav
How Performance Engineers Find and Fix Hidden System Bottlenecks
How Performance Engineers Find and Fix Hidden System Bottlenecks

Picture a business-critical SQL query crawling for seven hours. Nearly a full workday. The system keeps grinding through data, the business keeps losing time and money, and users are stuck waiting. Then a performance engineer steps in. After a few hours of careful analysis and a handful of precise code changes, the same query finishes in two minutes. Situations like this are not unusual in performance engineering. Turning hours into minutes is exactly the kind of work that makes this discipline valuable. In modern DevOps environments, where systems are deployed continuously and workloads change quickly, this type of work becomes part of everyday engineering practice. Who Are Performance Engineers? In simple terms, a performance engineer (PE) is responsible for making IT systems run better: faster, more reliably, and more efficiently. Behind this simple definition, however, lies a complex and multifaceted discipline. The bottleneck can appear almost anywhere in the stack: in application code, database configuration, network communication, or even the underlying hardware. And sometimes the bottleneck is not in the database or the application, but in the operating system. When systems handle thousands of concurrent network connections, limits may appear in the OS network stack or in kernel parameters. There are well-known cases in the history of database systems where the same database engine showed dramatically different performance on different operating systems, such as Windows, FreeBSD, or Linux. These differences were often caused by variations in filesystem behavior, networking stacks, and kernel-level I/O scheduling rather than by the database software itself. Once the root cause is identified, the performance engineer must understand the underlying mechanism behind it and propose an effective solution. Sometimes it means tuning the configuration. Sometimes it means rewriting a query or changing application behavior. Sometimes it points to a deeper architectural flaw that was hidden until the load exposed it. That is why the job often feels less like optimization in the abstract and more like investigation under pressure. And it sits somewhere between development, systems administration, and deep system analysis. It is important to note that many performance engineering tasks overlap with the responsibilities of a database administrator. Query optimization, lock analysis, and tuning parameters such as WAL settings are traditionally part of a DBA’s role. The difference is that a performance engineer usually operates at a broader level. They analyze the performance of the entire system, including the application, database, operating system, network communication, and underlying hardware. While a DBA focuses on a specific database platform, a performance engineer evaluates the system as a whole production pipeline. In cloud environments, this broader view may also touch tools when workload, container, or configuration findings overlap with production behavior. A Practical Example Performance problems rarely have a simple playbook. The same symptom can appear in different environments while the underlying cause is completely different. Engineers, therefore, rely on ongoing microlearning, hands-on experimentation, and careful analysis of system metrics to expand their troubleshooting knowledge as part of everyday engineering practice. One example illustrates how these investigations unfold. A client was migrating data from Oracle to PostgreSQL. The migration process relied on massive parallel data loading using COPY. At first, everything seemed to work normally, but eventually the process slowed dramatically. The investigation showed that the bottleneck was due to WAL (Write-Ahead Log) writes. In PostgreSQL, every change generates a WAL record that must be flushed to disk before the transaction commits. This mechanism guarantees durability and crash recovery, but under heavy write workloads, it can become a limiting factor. Initially, the team suspected that disk throughput was the problem. Developers even suggested a patch intended to speed up WAL writing. The patch did not improve performance. The client’s internal specialists were also unable to find a clear explanation. At that point, the performance team started analyzing the system in more detail. They noticed that many database sessions were waiting on the PostgreSQL wait event LWLock:WALInsert. That observation changed the direction of the investigation. It meant the system was not actually saturated by CPU or disk throughput. Instead, multiple processes were competing for internal synchronization while inserting WAL records. The migration workload involved hundreds of concurrent COPY operations. Each process attempted to reserve space in WAL buffers, which created contention around WAL insertion locks. The team experimented with several configuration parameters and eventually increased wal_buffers and wal_writer_flush_after. This allowed PostgreSQL to accumulate larger WAL batches in memory before flushing them to disk. The result was a significant reduction in contention around WAL insertion and about a 30 percent improvement in migration throughput. It is important to note that these changes are not universally safe defaults. Larger WAL buffers and less frequent flushing can increase the amount of data at risk during an unexpected crash. In this case, the workload was a migration. If the process stopped, it would have to be restarted anyway. Under those conditions, the temporary trade-off between reliability and performance was acceptable. The real lesson from this case is not a specific configuration value but the investigation process: identify where the system is actually waiting, test hypotheses, and verify improvements with measurements. Of course, real PE cases are often more complex than the simplified example shown here. In practice, investigations can take days and may involve analyzing internal database behavior, operating system limits, and network interactions to identify the true bottleneck. Common Performance Engineering Rules In performance engineering, there is an informal set of principles that experienced engineers tend to follow: 1. Proactivity Is the Best Prevention A performance engineer does not wait for a system to fail under load. The work starts earlier: analyzing the architecture of new services, anticipating how the system will behave as load grows, and identifying potential bottlenecks before they become production incidents. 2. Trust Metrics A common mistake, especially among less experienced engineers, is optimizing by eyeballing results. Someone changes a configuration or piece of code and says, “It seems to run three to five seconds faster.” That is not acceptable. Improvements must be confirmed with measurable data. Engineers compare metrics before and after a change: transactions per second (TPS), latency, CPU utilization, disk and memory usage, and queue lengths. Only these measurements can demonstrate whether performance has actually improved. Metrics matter more than subjective impressions, although experience still plays a role. Experienced engineers often use intuition to form an initial hypothesis about the cause of a problem. However, every hypothesis must be verified with measurements. Intuition helps guide the investigation, while metrics confirm the correctness of the solution. 3. Be Careful With Quick Fixes Sometimes incidents must be resolved immediately. A common example involves the max_connections parameter in PostgreSQL, which defines the maximum number of concurrent connections. When a system slows down under load, some developers try to increase max_connections. This can help temporarily, but it often creates new problems. A sudden increase in connections raises contention for internal database resources such as shared memory structures and locks. As contention grows, performance can degrade significantly due to locking and resource pressure. A quick fix can easily turn into a larger failure. A good performance engineer will highlight these risks and recommend a more systematic solution. Becoming a Performance Engineer Few people start their careers aiming for performance engineering. More often, they drift into it through a difficult problem that refuses to stay contained. That is how it happens in practice. A developer helps compare database options for an important project. Then the questions start multiplying. What should be measured? On physical servers or virtualized infrastructure? Which metrics matter? What changes under load? What changes only in production? One question leads to ten more. Before long, the person who thought they were helping with a tactical decision is working at the boundary between software, systems, and operational behavior. That path is common. People grow into it from development, systems administration, or operations. Nobody really graduates as a ready-made performance engineer, yet those who grow into the role often reach compensation levels that can support stronger long-term financial outcomes than many other career paths. Core Skills of Performance Engineers Because performance engineering sits at the intersection of software and infrastructure, practitioners usually combine skills from several technical disciplines. What does it take to move into this field? Programming A performance engineer needs to understand how software is written, how developers think, and what challenges they face. In many cases, the engineer works with tools built for other developers. Linux Strong Linux knowledge is very important: how the kernel works, how the user space operates, how processes are managed, and which operating system metrics can be measured. Algorithms Understanding algorithms and their complexity is essential for proposing efficient solutions. Math Another important but often missing skill is mathematical statistics. When an improvement is not dramatic but only one to two percent, engineers must prove that the change is meaningful and not just measurement noise. Concepts such as quantiles, percentiles, data distributions, and multimodal behavior help separate real improvements from measurement noise. Communication Performance engineers must clearly and carefully communicate findings to developers, testers, and business stakeholders. Explaining that a problem originates in someone’s code can be sensitive, so it must be done constructively. Attention to Detail Attention to detail is critical. An unusual spike in a graph or a repeating system pattern may point to the root cause of a problem. Persistence also matters. Test results can fluctuate due to environmental factors, so identifying the real issue often requires patience and careful investigation. The same applies to client-side performance work, where telemetry, crash analytics, and app data collection can help explain how the product behaves on real devices, networks, and usage patterns. The Future of Performance Engineering Systems are becoming more complex, and no single engineer can be an expert in every layer. As a result, the field is moving toward greater specialization. We are likely to see performance engineers focused on specific areas: application-level performance, operating system behavior, or hardware-level optimization, such as selecting the right CPU for a workload and tuning CPU frequency settings. What about AI? So far, there are no real tools capable of replacing performance engineers. AI can help engineers find information faster, although its output still needs verification. It does not yet solve complex analysis and optimization tasks. Automated tuning systems also do not currently appear capable of replacing human expertise. There is some expectation that AI will at least automate routine work. For now, most performance engineers see AI as an assistant rather than a threat. Final Thoughts: The Hunt Continues Performance engineering is a constant challenge. It is an intellectual puzzle with real operational impact. The work involves identifying hidden patterns, uncovering non-obvious relationships, and finding effective solutions where others see only complexity or system limits. Many engineers remember their first major optimization success. A query that once ran for minutes or even hours suddenly runs hundreds of times faster. Moments like this leave a lasting impression and often keep engineers in this field for many years. These experiences are what make performance engineering a difficult yet highly engaging profession, where the hunt for CPU cycles and response-time improvements never really ends.

By Alex Vakulov DZone Core CORE
The Bottleneck of Scaling
The Bottleneck of Scaling

Any input/output operation, be it accessing a file, handling an HTTP request, or a database connection, is based on 3 fundamental system concepts — file descriptors, kernel memory, and heap size. This article discusses how modern languages help developers handle behind-the-scenes file descriptor, kernel memory, and heap management. These three concepts are major bottlenecks for scaling. 1. File Descriptors A file descriptor is just a positive number that is used by the kernel to identify any open input/output stream or connection. It is defined by the kernel for a process. The following file descriptors are defined by default for a process: 0 – Standard Input (stdin)1 – Standard Output (stdout)2 – Standard Error (stderr) Any subsequent I/O operation gets the next available integer as file-descriptor. The file descriptor value can be adjusted by using the ulimit -n command in Linux. Each application, whether it is a web server written in Java Spring Boot, an API server written in Go using net/http and gorilla-mux, or a Python Flask app, is a single process. Each process has only 1024 file descriptors defined by default. That means each application can perform only 1024 I/O operations simultaneously. This seems like an amazing concept when we talk about scaling our application or API server. As many times as we come across this question — how can we scale our API server or web application to handle 100k or 1 million requests per second? This is where our modern languages play their role very beautifully behind the scenes to enable developers to develop the application to handle such scale. 2. Kernel Memory At a lower layer than file descriptors, when an incoming TCP connection hits the network card, the Linux kernel performs a 3 Way TCP handshake for that connection. The handshake lifecycle includes the states: SYN -> SYN-ACK -> ACK. The number of requests equal to the defined file descriptor value are processed immediately, assigned a file descriptor, and forwarded to the application for further processing. When FDs are exhausted, the Kernel maintains a queue for requests waiting for FDs to become available so your application can process them. The same thing happens when a request is processed, and the response is ready to be sent back to the client. This queue is maintained within RAM by read buffers(rmem) and write buffers(wmem). The size of buffers is defined in memory by the kernel and is dynamic, depending on network throughput, round-trip time, and memory pressure. The kernel network memory is non-paged, i.e cannot be swapped to disk. It’s a big bottleneck as it directly depends on physical memory. For example, if there are 100,000 open connections and each connection holds an average of 128KB of kernel memory, it comes to 12.8GB of physical RAM. This is clearly a kernel overhead, and it doesn’t show up in JVM heap metrics or Go runtime statistics. rmem and wmem buffers are governed by kernel parameters defined in /proc/sys/net/ipv4/ 3. Heap Size When TCP connections are assigned file descriptors and kernel memory is reserved, they enter user space, which is the memory managed by the application runtime — Java JVM, Node.js V8 Engine, Python interpreter, Go runtime, etc. Each connection stores objects in the heap within three categories: Connection metadata – Keep-alive timers, IP State, Socket Wrappers, etc.Cryptographic session context – handshake caches, cipher states, TLS/SSL keys, etc.Serialized payload buffers – response queues, JSON strings, ORM entity maps, etc. A connection that is encrypted via TLS takes a lot more space in the heap compared to a regular connection. For an encrypted connection, the application has to save symmetric keys, cipher contexts, session tickets, etc. onto the heap. A regular TCP socket object in the heap consumes 2KB to 5KB of space, whereas a TLS 1.3 socket object consumes 20KB to 100KB of heap space. If an API maintains 10,000 idle TLS connections, it will consume 200MB to 1GB of heap space. When an application runs, the runtime asks the kernel for memory space as the application creates objects. The application keeps creating objects, and the kernel keeps reserving memory for those objects; this is called the heap. The maximum heap size can be defined by different programming languages at runtime; for example, in Java, -Xmx4g reserves 4GB for the heap. The operating system promises to provide that much memory as heap space for the application, but it doesn’t reserve it all at once. As the application creates objects, the kernel continues to reserve memory. When objects are marked as done, the garbage collector removes them from the heap. When an incoming request hits our API server, the application uses heap space to convert raw bytes to the application-specific data structure. Once the application finishes processing the request and returns the response, those objects in the heap become unreachable or dead. When the garbage collector sweeps those objects to reclaim that memory, it doesn’t return the memory immediately; instead, the JVM or Go runtime keeps that freed memory in its internal pool. If a new HTTP request arrives within 1 millisecond, the runtime assigns the required memory from the free memory in the pool. Now imagine 10,000 new requests arriving at the same time, each with 2MB of raw bytes, and the runtime trying to allocate heap for the objects; the app instantaneously uses 20GB of memory. This is called GC thrashing, as the runtime rapidly creates required objects in the heap faster than the GC can clean them. The garbage collector is an application thread itself; when the heap gets 80%-90% full, the garbage collector panics and consumes 100% of CPU cores to scan millions of memory pointers to find dead objects. The runtime, like the JVM or Node.js garbage collector, may stop other code execution while it reorganizes the memory. So, how do runtimes like Go and the JVM handle GC thrashing? Go follows a simple strategy – avoid creating objects on the heap. The fastest GC collector is the one that has nothing to collect. The Go compiler compiles the application to see if variables outlive their functions. If a struct is used only inside a function, Go pushes the struct to the stack instead of the heap, and the stack pointer just drops when the function returns. The memory is reclaimed in 1 CPU cycle without even involving the garbage collector. If Go does have to clean the heap, its GC runs concurrently along with other goroutines and is broken into several micro pauses. Go provides sync.Pool to help developers to reuse heap memory while creating objects. For example to instead of creating millions of []bytes for JSON parsing for every new request, developers can use sync.Pool as follows: Go // Instead of creating a new buffer for every HTTP request: var bufferPool = sync.Pool{ New: func() any { return new(bytes.Buffer) }, } func handleRequest(w http.ResponseWriter, r *http.Request) { buf := bufferPool.Get().(*bytes.Buffer) // 1. Grab an existing buffer from pool buf.Reset() defer bufferPool.Put(buf) // 2. Put it back when done! // Parse JSON into 'buf' without allocating new heap memory } By recycling buffers via sync.Pool, high-concurrency APIs can handle 100,000 requests/sec with near-zero new heap allocations. Java takes a different approach. Because Java applications historically create millions of short-lived objects on the heap, the JVM relies on Generational Hypotheses and Generational Collectors (like G1GC, ZGC, and Shenandoah). G1GC can be used like java -XX:+UseG1GC while running Java applications. G1GC divides the Heap memory into physical regions: Young Generation (Eden & Survivor spaces) and Old Generation. It kind of sorts objects into different regions so that it doesn't have to scan the complete heap and can clean where most of the marked objects live. We can also mention -XX:MaxGCPauseMillis=200 to tell G1 to pause the application for no more than 200ms, but this is not guaranteed. Older JVM collectors like Parallel GC used to freeze the entire application to clear the heap when full, leading to multi-second latency spikes. Modern JVMs introduce ZGC (Z Garbage Collector) and Shenandoah. ZGC uses specialized CPU pointer references to track moved objects in real time. ZGC can clean, move, and compact terabytes of heap memory concurrently while your API requests are actively running. ZGC guarantees GC pause times under 1 millisecond, regardless of whether your heap is 500 MB or multi-terabytes. Conclusion Keep track of these three core concepts — file descriptors, kernel memory, and heap size to know when to scale. 1. File Descriptor Saturation Signals File descriptors represent the system's open handles. When an application hits its FD threshold, the operating system stops accepting connections. The following are example scenarios that indicate when to scale. Check Kernel-wide statistics from /proc/sys/fs/file-nr, per process fds - /proc/<pid>/fd, Prometheus exposes process_open_fds. If it consistently breaches the 80–85% threshold, it's time to scale. You have already tuned ulimit -n and LimitNOFILE up to standard safety thresholds (e.g., 65,536 or 104,857), but process FD counts continue climbing toward the max. Network interfaces show growing SYN-to-LISTEN socket counts and drops in netstat -s under the listen queue overflow metric. 2. Kernel Memory Pressure Signals Because TCP receive (rmem) and transmit (wmem) buffers are non-paged, they cannot overflow onto disk swap. When kernel network memory fills up, the OS drops packets. Below are the scenarios related to kernel memory breach. Check /proc/net/sockstat under TCP: inuse and matching /proc/sys/net/ipv4/tcp_mem thresholds. Netstat counters (netstat -s | grep -i retrans) show a sharp rise in TCP Retransmission rates (>1–2%). Latency spikes occur because the kernel is dynamically shrinking socket buffers down to tcp_rmem minimums (4 KB) to avoid running out of physical RAM, throttling TCP window sizes. 3. Heap Size & Garbage Collection (GC) Thrashing Signals When user-space heap allocations outpace the garbage collector's ability to sweep dead objects (like parsed JSON payloads or session states), application performance collapses. The runtime (JVM or Go) spends more than 15–20% of its total CPU time running GC sweeps (go_gc_cpu_fraction or JVM GC CPU utilization). In Go, metrics show the pacer triggering Mark Assist, stealing CPU time from worker goroutines to help clean up memory. You can check the runtime package /cpu/classes/gc/mark/assist:cpu-seconds metrics to see if GC is asking for more help from CPU. In Spring Boot, you can use Actuator and Micrometer to expose relevant endpoints to monitor the threshold values.

By Vishal Bhatia
Ampere PMU Profiler: A Guide to Microarchitecture Profiling
Ampere PMU Profiler: A Guide to Microarchitecture Profiling

Executive Summary The Ampere® PMU Profiler (APP) is a Python-based tool designed to provide deep insight into the microarchitectural behavior of applications running on Ampere CPUs (e.g., Ampere® Altra® and AmpereOne®). Unlike standard profilers that identify where time is spent (e.g., which functions consume CPU time), the PMU Profiler explains why time is being spent by measuring low-level hardware events associated with the CPU pipeline and execution behavior. A key outcome of APP is that it enables performance engineers to move from coarse symptoms to actionable causes. For example, while application-level profiling can show an expensive code path, APP can help identify whether the expense stems from inefficient instruction fetching, data cache misses, or other microarchitectural factors that are difficult or impossible to isolate using application-level tools alone. The document outlines a top-down performance analysis methodology and positions APP as an essential final step for expert-level tuning, particularly on Ampere platforms, where you must understand hardware-level bottlenecks and then apply targeted code optimizations. APP is intended to complement system-level analysis rather than replace it. System-level profilers are useful for identifying high-level bottlenecks such as resource saturation or contention, but APP is focused on microarchitecture-level analysis by collecting hardware events. This makes APP especially valuable after system bottlenecks have been eliminated or ruled out, leaving “microarchitecture inefficiency” as the remaining likely cause of slowdowns. What Is the Ampere PMU Profiler? The Ampere PMU Profiler uses Linux perf utility with validated PMUs and metrics on Ampere CPUs. Its purpose is to capture microarchitectural performance indicators through hardware event measurement. In practice, this means APP collects events measured by perf stat that relate to the CPU pipeline and execution mechanisms, allowing engineers to determine what is slow and the underlying microarchitectural reason. A central distinction between APP and application profiling tools is the level of visibility. Tools that sample stack traces (or count function invocations) typically answer the question, “Which functions are active during the slow period?” APP answers a more hardware-specific question: “Which microarchitectural mechanisms are consuming cycles, and what stalls or inefficiencies are present?” The APP workflow assumes that developers can form a hypothesis about where the bottleneck likely originates, such as a particular loop or data access pattern, and then rely on PMU event measurements to confirm or refute those hypotheses at the microarchitectural level. Why Do We Need APP? Performance problems are frequently multi-layered. Even after system-level bottlenecks are addressed (for example, ensuring that CPU is not idling due to I/O, ensuring there is sufficient memory, and verifying resource utilization), some workloads still perform poorly because the CPU spends cycles in inefficient pipeline states. APP helps solve this class of problems by measuring hardware-level behavior. For example, APP can identify microarchitectural bottlenecks such as: Inefficient instruction fetchingData cache missesBranch-related pipeline effectsOther pipeline-level stall sources that manifest as lost cycles This capability is important because microarchitectural causes often do not map cleanly to application symptoms. Code can appear “hot” in a profiler, but the reason it is slow might be due to how it interacts with cache hierarchies, how it causes translation or fetch inefficiencies, or how the processor recovers from pipeline disruptions. Those details are what PMU-based measurement aims to expose. APP also links investigation to “unlocking the full performance potential” of the hardware. By understanding CPU-level bottlenecks, engineers can choose targeted optimizations that application-level tools alone cannot determine with confidence. This ultimately leads to more efficient software and better utilization of Ampere hardware for competitive workloads. When Do We Use the Ampere PMU Profiler? Understanding the APEX Framework Performance tuning is a process of systematic investigation, moving from a broad, system-wide view down to the specific interactions between code and hardware.  Fig. 1: APEX Benchmarking and Optimization Funnel Performance optimization is as much art as it is science. The APEX (Adaptive Profiling and Execution) framework uses tools and methodologies to add structure and rigor to the process and can bridge the gap between creative intuition and empirical fact. We propose applying the APEX methodology to enable root cause analysis for solving performance problems. Follow the funnel above from top to bottom to effectively use the procedure. The methodology recommends starting with assessing platform health as a first step to ensure that the platform used for performance analysis is set up well as an unhealthy platform may mislead the performance analysis. Consider capturing initial performance metrics before tuning any system or application settings. This establishes a clear understanding of the current workload and identifies key scalability knobs. We recommend using Ampere’s PerfKit Benchmarker (APB), which supports many open-source applications, to create a reliable baseline for further analysis and tuning. Next is to assess system performance and any hardware or system bottlenecks— this is where Ampere System Profiler (ASP) is useful to eliminate any system or resource bottlenecks. ASP can also be used to right-size the instance shape and ensure the compute resources are efficiently consumed by the workload. One method that may be used is to leverage APB’s automated benchmarking framework to start and stop ASP’s collectors during the run phase of a given APB benchmark. This ensures that profile is collected while critical code paths are executed and a clear report profile is generated. Once system and resource bottlenecks are eliminated, if the performance issue persists and points to CPU cycles not being used efficiently, we propose going to the next step in the pyramid and using the Ampere PMU Profiler to root-cause the issue further. Finally, system benchmarking should be done after all bottlenecks are resolved or analyzed to effectively measure the system’s performance for the workload. Following this systematic APEX methodology ensures that we eliminate possible issues as a part of a structured process to efficiently conduct root-cause analysis. System-Level Analysis At the microarchitecture level, performance is shaped by how the CPU pipeline handles instruction delivery, execution, and memory access. APP leverages PMU measurements to identify pipeline behavior and stall sources. Memory Hierarchy and Performance Loss APP emphasizes the performance significance of the memory hierarchy. As data access moves from registers to L1 cache, to L2 cache, to L3 cache, and finally to DRAM, access becomes exponentially slower. Because of this, cache misses are a primary cause of performance loss. This provides a conceptual foundation for many APP investigations: If a workload touches large working sets or accesses data in a non-contiguous pattern, it may trigger cache misses that increase effective latency and reduce throughput. Microarchitectural Bottleneck Identification APP can be used at the microarchitecture level to understand where stalls might be in the pipeline. The APP role is to collect hardware events related to pipeline stall behavior and to use those events to characterize the workload’s execution profile. This capability matters because pipeline stalls and inefficiencies can dominate runtime even when application-level profiling points to a “hot” function without explaining the root cause. Key Questions APP is structured around answering questions that cannot be fully resolved with application-level profiling alone. Based on the described APP workflow and report interpretation strategy, APP can help you answer: Where are cycles going at the microarchitectural level? The APP HTML report and TDA sunburst charts are used to broadly characterize whether time is dominated by categories such as instruction retirement behavior, front-end bound behavior, or back-end bound behavior.Which stall or inefficiency class is consistent with the hot code path? Once you hypothesize a bottleneck mechanism (e.g., cache misses from non-contiguous access), APP measurements can confirm whether the observed behavior aligns with that mechanism.What microarchitectural reason explains a hot function’s cost? APP’s purpose is explicitly to explain why time is spent by measuring hardware events. This allows developers to translate hot functions into hardware interactions that can be optimized.Is the workload limited by instruction delivery vs execution/memory? By inspecting broad characterization categories (front-end vs back-end bound) in the APP HTML report, engineers can determine which side of the pipeline is more likely to be limiting performance. Example Usage and Output The below example command attempts to collect: PMU profiling samples for 120sWith a sampling interval of 1sProfiles on cores 1 and 2TopDown metrics and render TDA sunburst chartPMU profiles while running the workload affinitized to cores 1 and 2 Shell app -n 120 -c 1,2 -i 1 –tda -o <folder> -j “taskset -c1,2 <workload> Metrics reported by APP: Metric NameDescriptionIPCInstructions retired per CPU cycle across user and kernel execution unless separatedIPC_kernelInstructions retired per CPU cycle while executing in kernel/EL1cpu_freqAverage core frequency during the measurement interval, typically in GHz or MHzCycle Accounting Metricsfrontend_boundShare of cycles in which retirement is limited by front-end activity (e.g., fetch, branch prediction, decode, ICache, ITLB, queueing)backend_boundShare of cycles in which retirement is limited by back-end resources, cache or memory latency/bandwidth, or ROB/LSQ pressureBranch Effectiveness Metricsbranch_mispredict%Percentage of retired branch instructions that were mispredictedbranch_mpkiBranch mispredictions per 1,000 retired instructionsDTLB Effectiveness Metricsdtlb_mpkiData TLB misses per 1,000 retired instructions requiring translation refill or a walk beyond L1 DTLBdtlb_walk%Percentage of DTLB misses that trigger a page-table walk rather than being resolved by another TLB levell1d_tlb_miss%L1 DTLB miss rate relative to DTLB accessesl1d_tlb_mpkiL1 DTLB misses per 1,000 retired instructionsl2_tlb_miss%L2 or second-level DTLB miss rate relative to L2 TLB accessesl2_tlb_mpkiL2 or second-level DTLB misses per 1,000 retired instructionsITLB Effectiveness Metricsitlb_mpkiInstruction TLB misses per 1,000 retired instructionsitlb_walk%Percentage of ITLB misses that trigger a page-table walk instead of hitting in a next-level TLBl1i_tlb_miss%L1 ITLB miss rate relative to ITLB accessesl1i_tlb_mpkiL1 ITLB misses per 1,000 retired instructionsL1 Cache Effectiveness Metricsl1i_mpkiL1 instruction-cache misses per 1,000 retired instructionsl1d_mpkiL1 data-cache misses per 1,000 retired instructionsl1i_miss%L1 instruction-cache miss ratel1d_miss%L1 data-cache miss rateL2 Cache Effectiveness Metrics l2_mpkiL2 cache misses per 1,000 retired instructions; exact scope depends on event mappingl2_miss%L2 cache miss rate relative to L2 accessesl2d_inv_pkiL2 data-cache invalidations per 1,000 instructionsl2_snoops_pkiL2 snoop transactions per 1,000 instructionsl2d_inv_per_snoopAverage number of invalidations generated per snoopOperation Mix Metricsbranch_percentagePercentage of retired instructions that are branch instructionscrypto_percentagePercentage of retired instructions that are crypto, CRC, or hash-class instructionsinteger_dp_percentagePercentage of retired instructions that are integer data-processing operationsload_percentagePercentage of retired instructions that are loadsstore_percentagePercentage of retired instructions that are storesscalar_fp_percentagePercentage of retired instructions that are scalar floating-point operationssimd_percentagePercentage of retired instructions that are SIMD/NEON vector operationsPipeline Stall Frontendstall_frontend_cache_rateShare of cycles stalled because of instruction-side cache or fetch-delivery issuesstall_frontend_tlb_rateShare of cycles stalled because of ITLB or translation-related front-end issuesstall_recovery_rateShare of cycles spent recovering from pipeline flushes (e.g., branch-misprediction recovery)stall_fronetend_bob_rateShare of cycles stalled because the front-end buffer or queue is full or blockedPipeline Stall Backendstall_backend_cache_rateShare of cycles stalled because of data-side cache-hierarchy latencystall_backend_tlb_rateShare of cycles stalled because DTLB misses or page walks delay loads and storesstall_backend_mem_rateShare of cycles stalled because of main-memory/DRAM latency or bandwidth limitsstall_backend_core_rateShare of cycles stalled because of core execution limits (e.g., dependency chains, execution-unit throughput)stall_backend_resource_rateShare of cycles stalled because of internal resource pressure (e.g., queues, buffers, credits)stall_rob_id_rateShare of cycles in which progress is limited by reorder-buffer or in-flight instruction capacitystall_ixu_sched_rateShare of cycles stalled because of integer execution scheduler or issue-queue pressurestall_fsu_sched_rateShare of cycles stalled because of FP/SIMD execution scheduler or issue-queue pressurestall_lob_id_rateShare of cycles stalled because the load buffer or queue is full or blockedstall_sob_id_rateShare of cycles stalled because the store buffer or queue is full or blockedUncore Metricsslc_miss%System-level cache (SLC/LLC) miss rate for requests reaching the SLCmc_retry_rate%Percentage of memory-controller transactions that are retried, indicating fabric or memory-controller pressurememrd_bw_GBpsEstimated DRAM read bandwidth consumed in GB/smemwr_bw_GBpsEstimated DRAM write bandwidth consumed in GB/sccix_in_bw_MBpsCCIX coherent-interconnect inbound bandwidth to the socket/system in MB/sccix_out_bw_MBpsCCIX coherent-interconnect outbound bandwidth from the socket/system in MB/s Refer to a detailed tuning guide here. Conclusion The APP enables PMU hardware event measurement to provide microarchitecture-level performance insight on Ampere CPUs. It is designed to answer the “why” behind performance problems by identifying microarchitectural causes such as inefficient instruction fetching and data cache misses, which are difficult to detect through application-level profiling alone. The APP workflow is top-down and hypothesis-driven: Form a hypothesis from the hot function, measure with APP profiles, and then analyze using the APP HTML report with TDA sunburst charts to characterize where cycles are being spent (instruction retirement, front-end bound, back-end bound). APP is most valuable when system-level bottlenecks have been characterized or ruled out and microarchitecture-level explanation is required for expert tuning. Check out the full Ampere article collection here.

By Bhakti Hinduja
Why Ping-Based Uptime Checks Are Failing Modern SaaS Architectures
Why Ping-Based Uptime Checks Are Failing Modern SaaS Architectures

In the early days of the web, monitoring availability was simple: a server either responded to a ping, or it didn't. HTTP checks tightened that up a little — a 200 OK meant the dashboard turned green, and everyone assumed things were fine. That assumption doesn't really hold anymore, though. A modern app can return a picture-perfect 200 OK and still be completely unusable to an actual customer. Take an e-commerce site where the web server is healthy and responding in milliseconds. Somewhere behind it, a third-party inventory service has quietly died, or a CSS change buried the checkout button under a promo banner nobody tested for. Nobody can buy anything. Server's up. Business is down. Legacy monitoring can't see any of this — it was built to check the plumbing, not whether the person standing at the sink can actually get water out of the tap. Uptime Isn't an Infrastructure Metric Anymore In a monolithic architecture, the app and the database lived on one server, and uptime was basically a binary infrastructure question. That's not how most applications get built anymore. A typical SaaS product today is a single-page application backed by dozens of independent microservices spread across regions, plus a stack of external dependencies — an identity provider, a payment processor, a CDN, whatever else. If any one of those goes down, your own servers can be perfectly healthy while your users still can't get through a core workflow. Uptime, in that world, has to mean the continuous availability of the actual business workflow, not a response code. What Synthetic Monitoring Actually Does Synthetic monitoring uses automated clients to simulate real user traffic on a schedule, from multiple locations, around the clock — instead of waiting for a human to hit a broken flow and file a ticket. These aren't simple URL pingers, either. A synthetic monitor opens a real browser, renders the DOM, executes JavaScript, fills out forms, clicks through multi-step flows, and checks that the right data shows up on screen, all while watching the underlying API calls to make sure the backend agrees with what the UI is claiming. If a login flow that normally takes two seconds suddenly takes ten, or a button just stops responding, the monitor flags it right away, typically with a video of the failed session and enough diagnostic detail attached that someone can actually act on it, routed straight into whatever incident tool the team already uses. That's a fundamentally faster loop than "someone tweeted that checkout is broken." Where This Overlaps With QA: Shifting Right QA and production monitoring have traditionally been separate worlds — different teams, different tools, a handoff at the deployment line. That divide doesn't have much justification anymore. If a team's already built solid automated functional tests for CI/CD, there's no real reason to throw those away once code ships. The same script that validates a checkout flow pre-deploy can get repurposed to run every few minutes in production as a synthetic monitor — generally called "shifting right." Done well, it cuts duplicated engineering effort and gets QA and SRE working off the same definition of "healthy" instead of two different ones. Testing Beyond the Front Door: Complex User Journeys Basic uptime monitoring tells you the front door is open. Synthetic monitoring actually walks through the door, picks something up, applies a promo code, checks shipping, completes a transaction — the whole path, not just the entrance. That requires handling state, not just static checks. A monitor testing a healthcare portal needs to log in with synthetic credentials, get through MFA, pull a specific record, and confirm it belongs to the test account and nothing else. One testing a fintech transfer needs to confirm the UI shows success and then separately query the backend to make sure the balances actually moved, because a UI that says "success" while the ledger disagrees is arguably worse than an honest failure. Validating both the interface and the underlying state is what makes this useful for anything regulatory or revenue-critical. The Self-Healing Problem Running scripts against a live production environment is harder than running them in staging, because production changes constantly — new banners, UI experiments, shifting layouts. A rigid script breaks on cosmetic changes it shouldn't even care about, and that's how you end up with false alarms nobody trusts. This is where AI-assisted self-healing has become genuinely useful, rather than just a buzzword bolted onto a monitoring dashboard. If a button's ID changes from submit-order to confirm-purchase, a brittle script just fails. A self-healing monitor uses visual and semantic signals to relocate the element, finishes the check, and logs a low-priority note for someone to review later, instead of paging an engineer at 3 a.m. over what amounts to a rename. Alert Fatigue Is a Design Problem, Not a Tooling Problem Poorly tuned monitoring trains engineers to ignore it, and static thresholds are a big part of why. If an alert fires whenever a page takes longer than three seconds, a one-off network blip pages someone for a problem that resolves itself before anyone even looks at it. A better approach builds a dynamic baseline from historical performance data — per time of day, per day of week — and only escalates when something deviates meaningfully from that baseline. Often it's worth requiring confirmation from more than one geographic location before paging anyone at all, so a regional network hiccup doesn't wake someone up for nothing. Where This Matters Most E-commerce is the obvious one — downtime there is measured in dollars per second, and synthetic checks on cart logic, discount calculation, and payment gateway responses catch the silent revenue leaks a green uptime dashboard would never surface. Multi-tenant SaaS is a quieter version of the same problem: a single shared microservice failing can degrade the experience for every tenant at once, sometimes without anyone noticing for a while. Synthetic scripts that log in under different tenant configurations help confirm data isolation is actually holding and that SLAs are being met in practice, not just assumed on paper because nothing's screamed yet. Healthcare and fintech carry real regulatory weight on top of the operational risk. Synthetic checks that confirm patient records render correctly, or that a banking handshake with a clearing house completes securely, end up functioning as both an operational safeguard and a rough form of continuous compliance evidence — useful when an auditor eventually asks how you know. The Takeaway A green uptime dashboard doesn't mean much anymore if all it's checking is whether a server responds. The failures that actually cost money and trust — a hidden checkout button, a silently failing third-party integration, a broken multi-step flow — live above the infrastructure layer. Only something that behaves like a real user is going to catch them.

By Arun Kulkarni
How to Monitor AI Models Without Drowning in Alerts
How to Monitor AI Models Without Drowning in Alerts

When putting their model into production, every team or organization encounters the same issue. Failures go unnoticed for days at first because there is no monitoring. As teams begin to fix the issues, they identify areas where production results deviate from the training data, create dashboards for every metric, and set alerts for every threshold. This results in engineers being paged at two in the morning for a bug that fixes itself within an hour, and when an important alert arises, it goes unanswered due to alert fatigue, creating a pipeline that silently feeds garbage into the model. When a team learns to disregard 95% of the issues, they are very likely to disregard the remaining 5% that are actually important, and the solution to this isn’t less monitoring. The good solution to this problem is monitoring, which is tiered, routed, and pruned differently from the infrastructure monitoring that most teams already know. The Problem With Applying Old Monitoring Rules To AI Traditionally, application monitoring used to be binary, which is whether the application or service is up or down, latency is high or low, etc. But AI models don’t fail with these signs; they usually degrade over time. For instance, a recommendation model does not show exceptions when the user behavior shifts; it just silently gets worse at what it was supposed to do. A classifier model does not throw an error when its input distribution changes; it just returns answers confidently with increasingly wrong predictions. An AI application does not crash when it hallucinates; instead, it returns a normal HTTP 200 response with incorrect content. This creates two problems: When AI models fail, the reason for failure is invisible to classical infrastructure monitoring, which causes teams to bolt on multiple checks like data quality checks, drift detectors, and output scorers, each introducing a new source of noise. AI models are statistical in behavior and not deterministic, so setting threshold alerts on them leads to them firing constantly, and training teams have to tune the model. As a result, thorough AI monitoring does not make the application safer; beyond a certain point, it only makes things worse. What to Actually Monitor Monitoring issues that no one will ever take action on is often the first step towards alert fatigue. It is useful to consider it in four layers, each with its own owner and mode of failure. Infrastructure and service: Metrics like inference latency, throughput, Graphics Processing Unit (GPU)/Central Processing Unit (CPU) utilization, error rates, and cost per request and token consumption for anything calling a hosted large language model (LLM) API are classic operational metrics and can usually be monitored with the existing Application Performance Monitoring (APM) tools. Data quality: This is another important thing to keep an eye on because it can cause broken feature pipelines, upstream schema changes, input formats being changed without getting noticed, and null-rate spikes. These are usually the worst failures because you can't see them unless you're looking for them, and the model keeps making predictions based on bad data. Model quality: This can be tracked by looking at changes in the Confidence Score or how much the prediction distribution has changed from what was seen during training. This can be used instead of measuring accuracy because it's hard to tell right away how measures like accuracy are calibrating, because to measure accuracy, you would have to compare the predicted result to the actual correct answer, which doesn't always exist at the time of prediction. Generative artificial intelligence/large language model quality: Metrics like hallucination rate, coherence, factual grounding, toxicity, and susceptibility to prompt injection need different types of tooling to identify them because they are not like traditional metrics and would require human-in-the-loop sampling or an LLM as a judge for identifying them. The mistake many teams make is that they apply the same alerting techniques to all four layers, which is the infrastructure one, as that is the traditional way of setting up monitoring for applications, but issues related to data quality and model quality require a trend-based review. How to Alert Without the Noise Replace static thresholds with adaptive baselines. When systems learn a baseline from historical behavior and trigger alerts on deviations from it, like “alert if latency exceeds 200ms,” this ignores the daily and weekly traffic patterns, and the same is valid for data volume and null rates, which leads to a large number of false alarms being raised. So, teams that have made this switch from static thresholds to adaptive baselines have reportedly reduced noisy alerts by 60–90%. Introduce real severity tiers. When an alert is critical and poses an instant business risk, it is sent to an on-call engineer so that the problem can be fixed right away. Warnings about poor performance that are not critical are sent to a Teams chat channel during business hours, and signals about long-term trends land on the dashboard to be looked at from time to time. This helps to make sure that the notification's urgency matches its real urgency. Correlate and deduplicate before notifying. One change to the schema upstream can cause a dozen problems downstream. Sending a dozen alerts for one root cause either makes the team too busy or forces them to mentally group alerts together, which your tools should be doing for you. Route alerts to whoever can act on them. Misrouting is a common cause of tiredness. If the central platform team doesn't know about the business, they might ignore a spike they can't understand, and the domain team that would be able to understand it would never see the alert. Both problems are solved by linking alerts to the right person by domain, based on where the problem starts. Prioritize by business impact. A system that looks for unusual events handles all alerts the same way because it doesn't know which parts of your system are important to the business. When you think about how important each problem is before choosing how loud to alert, you get a lot fewer alerts overall, and a lot more of them are ones that you should actually act on. Conclusion It's important to understand that all of the ideas we've talked about work together; none of them can be used on their own. For example, adaptive thresholds only give out fewer alerts that aren't differentiated by severity. Without proper routing, severity tiers send the wrong messages about how important something is to the incorrect individuals. To avoid alert fatigue, teams need to take comprehensive actions, which include proper alert designs and organizational practices. They should also ensure that every alert can be acted on, which is better than monitoring everything, because AI monitoring only scales, and not having anyone see a model fail could have serious consequences. Good monitoring means building a system that sends alerts only when it matters, so when it does, people actually act on it.

By Aditya Shrivastava
Pragmatic Premature Optimization
Pragmatic Premature Optimization

“...premature optimization is the root of all evil…” Donald Ervin Knuth Introduction "Premature optimization is the root of all evil." Most software engineers know this, attributed to Donald Knuth, author of The Art of Computer Programming and one of the most influential figures in computer science. Many have also picked up the practical conclusion that followed: "let's make it work first, fix performance later." After all, it's easier to add another EC2 instance than to find the root cause. But here is what Knuth actually wrote: "We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%." A little different, isn't it? The second sentence is almost never quoted — and that is convenient, because it turns a careful statement into a simple excuse. Sometimes for laziness. Sometimes because people assume that optimization means sacrificing readability: cryptic bit manipulation, obscure tricks, code that only the author understands at 2 am. I believe Knuth was indeed warning against that kind of optimization. But that assumption is wrong more often than people think. Good, clean code is frequently efficient code too — not by accident, but because choosing the right tool for the job tends to be both clearer and faster. The examples in this article are proof of that. Scope This article focuses on simple, cheap, and foolproof tips that can be applied universally — regardless of your architecture, framework, or domain. In my experience, they carry virtually no risk of making things worse. Architecture, design, networking, database connectivity, threading — these are deliberately out of scope. Not because they are unimportant, but because they are context-dependent. The right answer depends on your specific system, and each of these topics deserves its own article. Examples String Operations We are all familiar with built-in JDK string utilities like: equals(), startsWith(), endsWith(), contains(): Java s1.equals(s2); s1.startsWith(s2); s1.endsWith(s2); s1.contains(s2); Unfortunately, JDK provides only one function for case-insensitive comparison: Java s1.equalsIgnoreCase(s2) There are no functions for case-insensitive startsWith(), endsWith(), contains(). So, often we combine toLowerCase() or toUppserCase() with startsWith(), endsWith(), contains(): Java s1.toLowerCase().startsWith(s2.toLowerCase()); s1.toLowerCase().endsWith(s2.toLowerCase()); s1.toLowerCase().contains(s2.toLowerCase()); A little verbose and null-prone, but just fine if not on the critical path. However, this technique might cause some performance problems. Do not forget that String is an immutable class, so instead of just a char-to-char comparison between two strings, we create two additional strings that then must be garbage-collected. Considering that String is a wrapper over a char array, the memory allocation may become expensive. The solution is to use case-insensitive utilities provided by different libraries, e.g., Apache Lang3: Java startsWithIgnoreCase(s1, s2); endsWithIgnoreCase(s1, s2); containsIgnoreCase(s1, s2); Or, starting from version 3.18.0: Java Strings.CI.startsWith(s1, s2); Strings.CS.startsWith(s1, s2); Where CI exposes case-insensitive and CS — case-sensitive utilities. Many people like regular expressions and use java.util.Pattern class sometimes, not where it is really necessary. For example: Java Pattern.compile("^prefix.+suffix$").matcher(s).find() Instead of: Java s.startsWith("prefix") && s.endsWith("suffix") Or even: Java Pattern.compile("^prefix").matcher(s).find() instead of s.startsWith("prefix") Pattern.compile("suffix$").matcher(s).find() instead of s.endsWith("suffix") Pattern matching is significantly slower than trivial substring matching. The following table shows evaluation time for 1 million operations: Operation * 1 million times Time, ms s.equals("hello") 7 s.startsWith("hello") 6 s.endsWith("hello") 11 s.contains("hello") 24 s.toUpperCase().startsWith("HELLO") 65 s.equalsIgnoreCase("hello") 5 Pattern.compile("hello").matcher(s).find() 238 pattern.matcher(s).find() 31 What can we see from this table? Performance of equals() and startsWith() is similarendsWith() is 2 times more expensivecontains() is 4 times more expensive than equalsChanging case followed by startsWith() is 10 times (!) more expensiveCase-insensitive comparison functions do not have any performance penaltiesSearching for a substring using a precompiled pattern is about 20% more expensive than using a plain contains() method. Compiling the pattern and using it is almost 10 times more expensive than the plain contains() method. So next time you reach for Pattern.compile(), it is worth pausing for a second: is regex actually needed here, or is a plain string method both simpler and faster? If you really need a pattern, at least compile it in advance — better yet, declare it as a private static final class member. Collections Let’s assume that we want to know whether a given list contains the specific element: Java list.contains("red"); In fact, this call invokes code like this: Java int n = list.size(); for (int i = 0; i < n; i++) { if ("red".equals(list.get(i))) { return true; } } Starting from Java 8, we have a streaming API that just hides from us the same gory details: Java list.stream().anyMatch("red"::equals); This is perfectly fine when the list is short, changes frequently, or is searched only occasionally. But if the list is large, stable, and searched repeatedly, a HashSet is the right tool — offering average O(1) lookup instead of O(n). If you cannot change the original data structure, converting it once at initialization time and searching the Set from that point forward is almost always worth it. If both the guaranteed element order and the fast lookup are needed, we can either hold duplicated data structures — a list for ordering and a set for search or just use LinkedHashSet, which solves both problems. Another common case is case-insensitive search. We already saw above that the combination of toLowerCase() or toUpperCase() with comparison significantly reduces the performance. This can be solved by using TreeSet with custom comparator, e.g. String.CASE_INSENSITIVE_ORDER: Java Set<String> set = new TreeSet<>(String.CASE_INSENSITIVE_ORDER); This gives you a sorted, case-insensitive set with no extra allocations - and the same approach works for TreeMap when your data is key-value pairs. Enum Lookups Everyone knows that an enum entry can be found by its name using a built-in method valueOf(s). However, what to do if the given string is lowercase while enum entries following the naming convention are called using capital letters? Some people use a combination of toUpperCase() and valueOf() that work just fine but have the penalty we discussed above. However, very often people prefer to create a special field representing a “custom” name, so the simple enum like: Java enum Color { RED, GREEN, BLUE } Turns into: Java enum Color { RED("red"), GREEN("green"), BLUE("blue"), … } Let’s mention that this design has at least two disadvantages: Duplicate data: The custom name is the same as a built-in but in a different case, which can be solved much more easily. This allows using really custom names that, according to my experience, in most cases are not needed and just create so-called “edge cases” that, in turn, in most cases are just a signal of bad design and might cause a lot of “stupid” bugs. However, let’s continue. How do people often use this custom name? Java public static Color ofColor(String color) { return Arrays.stream(values()) .filter(c -> c.color.equals(color)) .findFirst() .orElseThrow(() -> new IllegalArgumentException("No enum constant %s.%s".formatted(Color.class.getName(), color))); } The implementation looks pretty nice, but this approach means that each call of ofColor() iterates over the list. Yes, in most cases enums are not huge, so the list is short, but anyway, why do this if we can just create a map from the custom name to the enum entry once during initialization and then use it with O(1) complexity? The following example solves both problems at once: it uses a case-insensitive map where the key is the standard name() of the enum entry during initialization: Java private static final Map<String, Color> colors = Arrays.stream(values()).collect(toMap(Enum::name, e -> e, (existing, replacement) -> replacement, () -> new TreeMap<>(CASE_INSENSITIVE_ORDER))); So, now the method ofColor() becomes trivial: Java public static Color ofColor(String color) { return Optional.ofNullable(colors.get(color)) .orElseThrow(() -> new IllegalArgumentException("No enum constant for " + color)); } One can argue that a map-based implementation is not always possible because sometimes the lookup criteria are too complex to be reduced to a simple key. Although I agree in general, I can say in turn that in many (if not in most) cases this is still possible. So far, the lookup key was a simple string. But what if the search criteria is a range rather than an exact value? Consider a more physically accurate model of colors as ranges of electromagnetic waves. Java public enum Color { BLUE(450, 495), GREEN(495, 570), RED(620, 750); …} How to implement the method ofWaveLength(int waveLength)? The straight-forward way is to iterate over the values of the enum and compare the given wave length with the range for each entry, i.e. implement O(n) search. But we can do better using NavigableMap, which is designed exactly for this kind of range query: Java private static final NavigableMap<Integer, Color> wavelengthMap = Arrays.stream(values()) .collect(Collectors.toMap( color -> color.minNm, color -> color, (existing, replacement) -> existing, TreeMap::new )); Unfortunately, the search method is not as trivial as in the previous example, but still very simple and fast: Java public static Color ofWaveLength(int nm) { return Optional.ofNullable(wavelengthMap.floorEntry(nm)) .map(Entry::getValue) .filter(value -> nm <= value.maxNm) .orElseThrow(() -> new IllegalArgumentException("No enum constant for wavelength: " + nm + " nm")); } Now, let’s compare the performance. Operation * 1 million times Time, ms valueOf(s) 34 valueOf(toUpperCase(s)) 78 Iteration with equals() 40 Color.ofColor() iteration 166 Color.ofColor() map 20 Color.ofWaveLength() map 32 The table shows that: As expected, toUpperCase() reduces performance twiceIteration with call of equals is a little bit more expensive than valueOf() although the enum has only three members and will grow linearly as the enum grows. The more members enum has, the more time iteration takes. Map-based implementation is even faster than one based on the built-in valueOf(). Stream-based iteration (ofColor() iteration) is surprisingly slow. Stream setup overhead (boxing, lambda dispatch, spliterator initialization) is non-trivial for tiny collections Pre-Intitialization The principle here is: do not do something several times if you can do it once. The most trivial example is string or numeric constants: Java private static final String FILE_NAME = "config.json"; private static final int MAX_VALUE = 10_000; However, the same principle applies to heavier objects — and that is where it really matters. Let’s take a look at logging. Most people are used to writing the following “magic” line at the beginning of each class (unless we use Lombok’s @Slf4j annotation): Java private static final Logger logger = LoggerFactory.getLogger(MyClass.class); Are all these modifiers (private static final) really needed? Some people try to save typing time: Java private final Logger logger = LoggerFactory.getLogger(MyClass.class); Moreover, if the logger is not static, we can do even more: Java private final Logger logger = LoggerFactory.getLogger(getClass()); This line looks better because it is error-proof: the class here is not hard-coded, so this line can be copied as-is from one class to another or inherited from the base class. So, what’s the problem? The problem is that retrieving the correct logger is potentially expensive due to synchronized registry lookups. Doing this on every instantiation adds up. A friend of mine told me that once in the company where he worked, this change in some critical path improved performance so much that they managed to reduce the AWS cluster by about one hundred large EC2 machines. The same rule applies to pattern compilation. As the benchmark table showed, compiling a pattern on every method call is nearly ten times slower than reusing a precompiled one. The result of Pattern.compile() should always be stored in a static final field. The only exception is the case when the regular expression is generated dynamically, but we should do our best to avoid such a design. Very often we have to format or parse dates. Traditionally I used SimpleDateFormat. What can be more obvious than this: Java private static final String FORMAT = "yyyy-MM-dd HH:mm:ss"; private static final DateFormat format = new SimpleDateFormat(FORMAT); Frankly speaking, I did this many times following the principle I stated above: there is no reason to create the instance every time we need it if we can create it only once. The problem is that SimpleDateFormat is not thread-safe, so sharing the same instance among different threads can cause the problem. Even worse: we can live with this bug for years without knowing about it, since it only happens under high load and in some cases can just produce slightly wrong results that can be lost in an ocean of valid data. So, should we create instances of SimpleDateFormat every time we need it and cause CPU and GC to work hard? Fortunately, starting from Java 8, we can use DateTimeFormatter instead: Java private static final DateTimeFormatter formatter = DateTimeFormatter.ofPattern(DATE_FORMAT); This class is thread-safe, so we can share its instance among different threads and get consistent results. Conclusion We started with a quote that is almost always cited incomplete. Knuth never said ignore performance — he said don't sacrifice clarity for speculative gains, while reminding us not to pass up opportunities in that critical 3%. The examples in this article live in that 3%. None of the performance issues described here should ever appear in production code. They are not hard to avoid — they require no profiler, no benchmarking framework, no architectural discussion. Just the habit of reaching for the right tool. And that habit pays off. Choosing equalsIgnoreCase() over toLowerCase().equals() is cleaner and faster. A static final logger is simpler and cheaper. A pre-built enum map is more readable and O(1). Good code and efficient code are not in conflict here — they are the same code. The only thing required is the habit of pausing for a second and asking: am I doing this n times when once would do? All code examples from this article are available on Gist.

By Alexander Radzin
Member Spotlight: Shamsher Khan
Member Spotlight: Shamsher Khan

There’s always more to our contributors than what you see in their author profiles. For our latest Member Spotlight, I sat down with Shamsher Khan to learn more about his newest project. What started as a frustrating Kubernetes troubleshooting problem has since grown into published research, a new way of thinking about operational evidence, and ongoing open-source work. What first got you interested in digging into complex infrastructure and systems problems? "I’ve always been interested in problems where the visible symptom is not necessarily the real cause. In infrastructure, especially distributed systems, a service can look healthy from one angle while something important is already failing underneath. Troubleshooting becomes less about finding one bad log line and more about understanding how the application, container, node, network, scheduler, and platform interacted over time. That is what made Kubernetes particularly interesting to me. It automates a lot of recovery, which is great operationally, but that also means the system can change very quickly while you are still trying to understand what happened. Over time, I found myself increasingly interested not just in fixing incidents, but in understanding what information engineers actually have available during and after those incidents, what disappears, and where existing tooling helps or still leaves gaps. That curiosity has shaped a lot of my writing and open-source work." Your DZone article, “When Kubernetes Forgets: The 90-Second Evidence Gap,” ended up becoming the starting point for Operational Memory Architecture (OMA). What were you seeing in Kubernetes that made you think, “There’s a bigger problem here”? It came from a very specific frustration during incidents. A pod would crash, Kubernetes would restart it, and by the time I got there to investigate, some of the information I wanted was already gone or had changed. One example is LastTerminationState. Kubernetes keeps information about a container’s most recent termination, but when that container fails again, the previous termination context is replaced. In a fast crash loop, that can happen repeatedly in a short period of time. You can arrive at a pod that has restarted thousands of times and still have only a very small window into how that sequence began. What made me think the problem was bigger was realizing that this was not really a Kubernetes bug. Kubernetes is primarily designed to maintain desired state and restore workloads. Preserving a complete forensic history is a different concern. Once I started looking more systematically, I saw similar boundaries elsewhere. Kubernetes Events have limited retention, short-lived workloads can exist entirely between monitoring samples, and some node- or runtime-level evidence can become difficult or impossible to reconstruct after the underlying state changes. There are already strong observability tools that help with logs, metrics, traces, and events, so the question was not, “Why doesn’t Kubernetes keep everything forever?” That would not be realistic or necessarily desirable. The question became more specific: are there predictable points after which certain diagnostic evidence can no longer be recovered, and can we reason about those points explicitly? I started calling those points evidence horizons. OMA grew from trying to characterize those horizons and explore what evidence may need to be captured before they are crossed." Now that the research is being published in IEEE Access, what do you hope people working with these systems take away from it? And where would you like to see OMA go from here? "The main thing I hope people take away is that recovery and diagnosis are related, but they are not the same problem. A platform can successfully restore an application while still losing some of the context that would have helped explain why it failed. I think many engineers have experienced this without necessarily having a name for it. If you have ever finished an incident review with, “We’re not completely sure what actually triggered this,” disappearing or short-lived evidence may be one reason. I also want to be careful not to suggest that OMA replaces existing observability platforms. Tools for logs, metrics, traces, events, and distributed tracing are already essential. OMA is better thought of as a way of reasoning about when different kinds of evidence remain available and when they may cross a point where recovery becomes difficult or impossible. There is also a practical side to this. Teams doing post-incident reviews, reliability analysis, or audit and compliance work may need to reconstruct what happened after the system has already recovered. Thinking explicitly about evidence retention and recovery boundaries can help teams decide what information is worth preserving. As for where OMA goes next, the research is still early. The work evolved in stages: I first published the foundational OMA idea on arXiv, then extended it with a broader evidence-horizon taxonomy and additional validation before developing it into the peer-reviewed IEEE Access paper. The implementation and experiments are public, and the most useful next step is independent validation in environments different from the ones I tested. There are also limitations in the current work. For example, some node-level evidence across kubelet or node restart boundaries requires deeper integration than the current architecture provides. I documented that rather than trying to claim the problem was solved. Some of these ideas have also influenced practical work I’m doing in OpsCart, an open-source Kubernetes operational triage project. OpsCart is not a replacement for OMA or for established observability tools. I use it more as an engineering testbed for exploring how incident context, workload history, and diagnostic evidence can be surfaced in a way that is useful during everyday Kubernetes troubleshooting. I would like to see other engineers test both the research assumptions and the practical tooling, challenge the model, and point out where it does not hold up. That kind of feedback is more valuable at this stage than claiming the architecture is complete." Research: https://ieeexplore.ieee.org/document/11656328OMA implementation: https://github.com/opscart/k8s-causal-memoryOpsCart: https://github.com/opscart/opscart-k8s-watcher After spending so much time thinking about Kubernetes, what’s your ideal way to completely unplug for a weekend? The first requirement is definitely no Kubernetes dashboards. I spend a lot of time during the week thinking about systems, debugging, writing, and experimenting, so on weekends I like doing almost the opposite: spending time with family, getting outside, going somewhere for the day, or just having time where I’m not trying to solve a technical problem. Infrastructure problems have a way of staying in your head even after you close the laptop, so sometimes the best reset is doing something that has absolutely nothing to do with technology. To see more of Shamsher's content, here's the link to his DZone profile.

By Dominique Roller
How to Diagnose and Recover Stuck Temporal Workflows
How to Diagnose and Recover Stuck Temporal Workflows

A Temporal Workflow that appears stuck is rarely “stuck” in the conventional process sense. Temporal persists Workflow state through Event History and resumes execution through replay, so an open execution can remain healthy while waiting for a timer, Signal, Activity, or external condition. The operational problem is therefore not simply lack of completion; it is lack of expected progress. Effective diagnosis starts by establishing what event should have happened next, why it did not happen, and whether remediation can preserve the Workflow’s business invariants. Temporal’s history model makes that analysis unusually tractable because commands, task transitions, Activity attempts, failures, timers, and external interactions are durably represented as Events. Progress Is Visible in the Event History The first diagnostic artifact should be the execution description and raw history, not application logs. temporal workflow describe exposes current execution information and pending Activity state, while temporal workflow show --output json returns Event History in a form suitable for programmatic replay or analysis. A Workflow Query can additionally expose application-defined state without mutating the execution. Shell temporal workflow describe --workflow-id order-7814 temporal workflow show \ --workflow-id order-7814 \ --output json History should be read as a state-transition trace. A WorkflowTaskScheduled event with no corresponding start suggests that work is waiting for a Worker. A started Workflow Task that repeatedly times out can indicate blocked Workflow code, Worker instability, or excessive work inside a task. Repeated WorkflowTaskFailed events can indicate replay or deterministic-compatibility failures after code deployment. Workflow Task failures are retried by Temporal rather than governed by an Activity-style Retry Policy, so a Workflow can remain open while repeatedly failing to make application-level progress. Activity sequences reveal a different failure surface. ActivityTaskScheduled without ActivityTaskStarted points toward dispatch capacity, missing pollers, queue mismatch, or backlog. Temporal persists Workflow and Activity Tasks in Task Queues, and worker-health guidance identifies Schedule-to-Start latency and approximate backlog count as key signals when tasks wait for Workers. ActivityTaskStarted without completion requires inspection of Start-to-Close and Heartbeat behavior because Temporal relies on Start-to-Close timeout to detect a Worker crash after an Activity has started. Not every long pause is pathological. A timer that has not fired, a Workflow waiting for a Signal, or an Activity still inside a valid timeout window can represent correct durable waiting. Conversely, very large histories can become an operational risk. Temporal warns after 10,240 events or 10 MB and enforces a limit of 51,200 events or 50 MB; Continue-As-New creates a new run with a fresh history while carrying forward relevant state. Triage Works Best as Deterministic Evidence Before Model Judgment LangGraph is useful for automating this analysis, but the safest design keeps Temporal facts deterministic and uses an LLM only for classification, hypothesis ranking, and explanation. LangGraph explicitly supports graphs that mix deterministic nodes with model-driven nodes, while structured output can constrain routing decisions into a defined schema rather than free-form text. A compact analyzer can first reduce raw history into evidence that is difficult to hallucinate: the last completed Workflow Task, consecutive Workflow Task failures, pending Activity IDs, the latest Activity attempt, the timeout type, the last Signal, the last timer, the history size, the task queue, and deployment/version metadata. The model then receives that normalized evidence instead of thousands of raw events. Python def extract_facts(state): events = state["events"] return { "facts": temporal_fact_extractor(events), "tail": events[-60:], } def classify(state): result = triage_model.with_structured_output(TriageResult).invoke({ "facts": state["facts"], "tail": state["tail"], "allowed_causes": [ "worker_unavailable", "activity_retrying", "workflow_task_failure", "intentional_wait", "history_pressure", "unknown", ], }) return {"triage": result} That separation matters operationally. Event parsing can enforce hard rules such as “scheduled but never started,” while the model can correlate several weak signals and produce an explanation. Conditional edges can then route low-risk cases to observation, ambiguous cases to deeper diagnostics, and recovery candidates to an approval gate. LangGraph’s graph API supports conditional routing, and persistence stores checkpoints so triage state survives interruptions or process failures. Recovery Must Preserve Temporal and Business Semantics Diagnosis and remediation should remain separate graph stages. A model-generated recommendation must not directly issue cancellation, reset, or termination. LangGraph interrupts provide a natural control boundary because execution can pause with persisted state and resume only after external approval. Python def approval_gate(state): decision = interrupt({ "workflow_id": state["workflow_id"], "cause": state["triage"].cause, "action": state["triage"].recommended_action, "evidence": state["triage"].evidence, }) return {"approved": decision == "approve"} The remediation choice depends on the failure mode. A transient Worker outage usually requires restoring Worker capacity rather than mutating Workflow state because queued tasks persist until Workers can process them. An Activity repeatedly failing on a recoverable dependency can often be left to its Retry Policy, while permanent errors should be made non-retryable in application design to avoid pointless retries. Activity side effects should be idempotent because Activity attempts may execute more than once under retry and recovery behavior. Cancellation is the preferred stop mechanism when Workflow cleanup logic must run. Temporal records a cancellation request and schedules a Workflow Task so Workflow code can react. Termination is forceful: Workflow code does not receive a chance to clean up, and the terminated event closes the history. That makes termination an escalation path for executions that cannot process cancellation normally. Reset is more powerful and more dangerous. Temporal terminates the current execution and creates a new execution that copies history through a selected reset point, then replays forward using current Workflow code. Progress after the reset point is discarded. Reset is therefore appropriate only after the underlying cause has been corrected and after downstream side effects are reviewed for possible re-execution beyond the reset boundary. Shell temporal workflow reset \ --workflow-id order-7814 \ --event-id 42 \ --reason "Recovered after deterministic-compatibility fix" For history pressure rather than a fault, Continue-As-New is generally the safer lifecycle mechanism because it preserves logical continuity under the same Workflow ID while starting a fresh Event History with a new Run ID. It should be designed into long-lived or high-volume Workflow logic instead of used as an improvised emergency action. Safe Automation Requires an Explicit Remediation Envelope A production triage graph should treat remediation as a constrained transaction. The evidence snapshot, selected run ID, candidate reset event, intended action, reason, approval identity, and execution result should all be persisted before any mutation. The action node should re-read the Workflow immediately before execution and reject the operation if the run has changed or the observed condition no longer matches the diagnosis. This is an engineering safeguard rather than a Temporal requirement, but it reduces time-of-check/time-of-use errors when active Workflows continue progressing during investigation. LangGraph’s checkpoint model supports durable approval state, but resumed graph nodes can re-execute from checkpoint boundaries. Its documentation therefore recommends isolating side effects and designing them to be idempotent. A remediation executor should consequently use an operation ID, record completion externally, and refuse duplicate destructive actions. Recovery Without Guesswork Reliable recovery of a stuck Temporal Workflow is fundamentally an event-history problem, not a process-restart problem. The strongest diagnostic path reconstructs expected progress from Workflow Tasks, Activity attempts, timers, Signals, queue state, timeouts, and history growth before considering mutation. LangGraph can turn that evidence into a durable triage pipeline by combining deterministic extraction, constrained model reasoning, conditional routing, and interrupt-based approval. Safe remediation then follows Temporal semantics: restore Workers when dispatch is the issue, allow bounded retries for transient Activities, cancel when cleanup matters, terminate only as a last resort, reset only after the root cause is fixed, and use Continue-As-New to control long-running history growth. The result is automation that accelerates incident response without allowing probabilistic diagnosis to become an unchecked control plane.

By Akhil Madineni DZone Core CORE
The 2026 Observability Audit: Separating Single Vendor Silos From Community Innovation
The 2026 Observability Audit: Separating Single Vendor Silos From Community Innovation

Open source projects dominated by a single vendor are a hallmark of "open source in name only." Rather than filling the traditional role of open source fostering innovation and decision-making from a diverse community, "open source in name only" projects are often used as marketing tools for proprietary platforms. These projects are also seen as riskier than community-driven projects because a single vendor is more apt to abruptly terminate long-term support, restrict contributions, or switch from an open-source license to a more restrictive one (forcing some previous contributors to pay for the project they helped build). In these projects, critics claim that investments are often lopsided and heavily skewed toward onboarding, marketing, and brand-related support. As a result, technical contributions are frequently less developed, opaque, undocumented, or lacking in real substance, often manifesting merely as a superficial "ease of entry and onboarding." Because of these underlying gaps in documentation and codebase depth, developers are routinely forced to reverse-engineer functionality simply to get the tools to work correctly. An evaluation of three leading open-source observability projects–OpenSearch, Prometheus, and OpenTelemetry (OTel)– by ReveCom was conducted to determine whether they fell under this vendor-dominated category or are truly vibrant community-led projects. According to Gartner research, these three projects are collectively important because together they provide a complete, vendor-neutral observability architecture covering all three fundamental telemetry signals—metrics, logs, and distributed traces — without locking an enterprise into proprietary agent formats or single-vendor cloud platforms. Gartner defines observability as the extent to which internal system states can be inferred from externally emitted data. By pairing OpenTelemetry as a universal collection and routing tier with Prometheus for real-time metric alerting and OpenSearch for high-volume log analytics and trace analysis, organizations gain end-to-end operational visibility, retain full ownership of their telemetry pipelines, and avoid runaway cloud ingestion or lock-in costs. To develop the framework, data from the ReveCom Observability Report 2026 was used, which includes metrics about contribution numbers and quality, including commit frequency, contributor growth, community expansion, and deployment patterns. Based on this data, authentic efforts were separated from perfunctory efforts. "Authentic" contributions were defined as those made to the computing code (i.e., the observability stack for logs, traces, and metrics) and its computational efficiency as measured in latency. The Controversial Fork AWS's controversial decision to monetize and then fork Elasticsearch to create OpenSearch in 2021 (when Elastic made its license more restrictive) is a case study of the risks associated with vendor-dominated projects. It also serves as an example of the issues associated with a vendor forking and heavily promoting a project it contributed minimally to. According to Elastic representatives, although a major beneficiary of Elastic through its managed service, AWS engineers contributed only a "handful" of commits to Elasticsearch from 2020 to 2021, Elasticsearch says. This disparity suggests that the successor project, OpenSearch, was born from a position of minimal technical familiarity with the core codebase. Elastic famously described this as "there is no compression algorithm for experience." For a technical leader, this lack of pre-fork familiarity suggests a significant "experience gap" that can impact the speed and stability of future feature releases. AWS made few fundamental changes to the Elasticsearch codebase it forked to create OpenSearch, largely just rebranding the existing observability tool. Comparing Three Observability Communities In 2024, Amazon donated OpenSearch to the Linux Foundation, bringing it under a governance structure and setting the stage for it to become a more decentralized project. Among other things, once a project is donated to the Linux Foundation, no single company can hold more than 25% of the seats on the technical oversight bodies. Decentralized governance is structured so that substantive, collaborative contributions from several competing observability vendors can better serve the broader community's needs. Amazon's donation set the stage for OpenSearch to become a much more community-driven effort, comparable to the community-led support of the Prometheus and OpenTelemetry projects. Prometheus and OpenTelemetry exemplify healthy, community-led open source standardization. This is how teams should evaluate open source: by the diversity of the entities with "skin in the game." Prometheus emerged from SoundCloud in 2012, where it was designed to track metrics and store them in a time-series database. Around 2014, Grafana and its glassy, visually appealing panels became part of the ecosystem. The combination of Prometheus and Grafana became an integral, de facto standard for monitoring and observability in Kubernetes deployments and infrastructure. Prometheus was donated to the CNCF in 2016 and graduated in 2018. Since then, it has evolved into a very diverse, community-led project, with multiple contributing companies. Grafana Labs remains one of the largest contributors, but the breakdown of substantive commits-excluding documentation-is wide and varied, reflecting the project's broad, collaborative nature. This wider contribution to the project's standardization ensures that engineering talent is portable and the stack remains interoperable. Separating Brand From Backbone A key open source health metric-perhaps the most substantial of all-is ranking substantive engineering contributions, such as code-level commits and pull requests or high-impact technical commits. These are described as commits that can lead to v1.0, v2.0, or v3.0 milestones, signifying production readiness and improvements. The number and frequency of technical contributions, as measured by commits, are markers for a project's community dynamics and value to end users. Looking at OpenSearch, AWS made significant technical contributions in 2025. As the data shows, Amazon contributes the majority of substantive commits (73%) to OpenSearch. Much of this can be attributed to a surge in contributions related to the AI aspects of observability, specifically "search-to-Al infrastructure" commits. IBM and Red Hat have also contributed AI-related work on RAG and vector database optimization. These are solid contributions, and they show that Amazon has moved beyond the early days, when it simply forked Elastic even though it had contributed relatively little to the project. Hopefully, OpenSearch will continue this shift toward increased community participation as new features are added. However, such a dominant share of commits from a single vendor means that one vendor effectively controls the roadmap. In this case, AWS is potentially prioritizing its managed services over users' infrastructure needs. Source: ReveCom Prometheus has a wide range of contributions from vendor organizations. Grafana is the leading technical contributor to Prometheus, largely based on its development of TSDB storage refactoring, Remote Write 2.0, and agent-mode contributions. Red Hat is the second-most frequent technical contributor to Prometheus, a position solidified by its acquisition of CoreOS. As the primary maintainer of the Prometheus Operator-a critical element for monitoring Kubernetes-Red Hat ensures seamless integration between the monitoring stack and the orchestration layer. While Red Hat provides deep engineering support, Prometheus remains a highly collaborative open-source project with contributions from across the industry. Source: ReveCom The OpenTelemetry project, under the leadership of Splunk, Microsoft, Elastic, Grafana Labs, and Google, provides a mature, stable, and innovative framework for the future of observability. By focusing on high-impact technical commits and "good faith" participation, the community helps ensure that observability data remains a standardized utility that empowers developers and platform engineers to navigate the complexities of the modern cloud landscape. Splunk remains the largest contributor of high-impact technical commits to OpenTelemetry. Grafana is a notable contributor at number four by providing Beyla eBPF instrumentation and Prometheus receiver stability improvements. Strategic Recommendations Organizations should adopt an open-source technical strategy that prioritizes authentic engineering and project diversity. The following recommendations are derived from scrutinizing vendor-dominated projects and analyzing high-impact technical commitments. The high-impact focus of companies like Grafana Labs, Splunk, Microsoft, Elastic, and hundreds of other contributor organizations means that OpenTelemetry and Prometheus should remain the foundation of observability for the next several years. When choosing an observability solution, organizations should prioritize vendors that are not only OTel-compliant but also OTel-contributing. This should also apply to Prometheus solutions, especially those for managing Kubernetes environments. ReveCom's findings indicate that the most valuable contributions are those that advance the core "engine" of observability. Procurement decisions should be based on a vendor's ability to demonstrate substantive engineering that solves real-world infrastructure problems rather than relying on superficial marketing claims. Ultimately, none of the three projects covered in this article can be fully characterized as "open source in name only." While OpenSearch arguably fell into that category immediately after it was forked from Elasticsearch, it has evolved since. OpenSearch remains an Amazon-dominated project, but it has seen an upward trend in contributions from the community and from third parties such as Uber, SAP, and Red Hat. For observability community support, as measured by substantive technical contributions that solve infrastructure problems, OpenTelemetry and Prometheus exemplify a healthy balance of governance and code contributions across hundreds of organizations (notably Grafana and Splunk). Led by Grafana and Splunk among the observability providers, these projects fall behind only Kubernetes itself.

By Chris Ward DZone Core CORE

Monthly Top Performance Experts

expert thumbnail

Filipp Shcherbanich

Senior Backend Engineer

IT expert with over 13 years of experience as a developer, team lead, and engineering manager. Currently a Senior Backend Engineer at a major international company. Active mentor and expert in tech communities.
expert thumbnail

Eric D. Schabell

Director Technical Marketing & Evangelism,
Chronosphere

Eric is Chronosphere's Director Community & Developer. He's renowned in the development community as a speaker, lecturer, author, baseball expert, maintainer and CNCF Ambassador. His current role allows him to help the world understand the challenges they are facing with observability. He brings a unique perspective to the stage with a professional life dedicated to sharing his deep expertise of open source technologies and organizations. More on https://www.schabell.org.

The Latest Performance Topics

article thumbnail
Predict, Repeat, Improve: Deterministic Simulation Testing Explained
Explore deterministic simulation testing — how predictable, repeatable outcomes boost QA, reliability, and confidence for engineers and architects.
September 29, 2026
by Ammar Husain DZone Core CORE
· 74 Views
article thumbnail
Jakarta Batch in Practice: Reliable Chunk-Oriented Processing for Enterprise Workloads
Jakarta Batch gives enterprise apps a standard model for long-running data processing with jobs, steps, readers, processors, writers, checkpoints, and tunable execution.
September 29, 2026
by Otavio Santana DZone Core CORE
· 215 Views
article thumbnail
Stop Paying a Model to Make Decisions You Already Made
A skill that spells out a fixed procedure in prose makes Claude re-decide it every run. Here's how to measure that cost using data Claude Code already emits.
September 28, 2026
by Amith Reddy Ravuru
· 385 Views
article thumbnail
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning
Batch processing works well for many workloads, but real-time decisioning requires event-driven architecture designed for resilience, observability, and failure handling.
September 25, 2026
by Prem Kumar Gadhanki
· 1,028 Views
article thumbnail
The Hidden Production Risks of Third-Party SDKs
Third-party SDKs speed up development, but they also introduce performance, security, reliability, and maintenance risks that teams must actively manage.
September 22, 2026
by Satyam Nikhra
· 2,835 Views
article thumbnail
Architecting for <1s Latency: Managing Eventual Consistency in Distributed Search Platforms
To maintain sub-second search freshness, logistics systems must actively manage eventual consistency across Kafka ordering, search indexing, and cache invalidation.
September 22, 2026
by Dhruv Goel
· 1,821 Views · 1 Like
article thumbnail
Stop Blaming Executor Memory: The Real Reasons Your Spark Jobs Are Slow
This article explains five common causes of slow Spark jobs and practical fixes for joins, stragglers, decryption chains, shuffle partitions, and incremental processing.
September 18, 2026
by Swaminathan Sethuraman
· 2,104 Views · 1 Like
article thumbnail
Understand the Sidecar Pattern by Deploying n8n to AWS Fargate
Learn how to deploy n8n Task Runners as AWS Fargate sidecars for isolated code execution, independent resources, and scalable workflow automation.
September 17, 2026
by Iyanuoluwa Ajao
· 2,712 Views · 1 Like
article thumbnail
Architecting Production AI Across Clouds: Patterns That Decide System Survival
In production, enterprise AI rarely fails at the model. It fails in the architecture around it. Here are the cross-cutting patterns that work.
September 16, 2026
by VenkataSrinivas Kantamneni
· 2,680 Views
article thumbnail
Improving Repeated Analytics Workloads With Databricks Disk Cache
Databricks disk cache speeds up repeated reads from curated Parquet or Delta tables, but it works best with good table design and partitioning.
September 11, 2026
by Harsh Patel
· 2,294 Views
article thumbnail
Memory-First Indexes in SQL Server 2025: Redefining Performance for Hybrid Workloads
Learn how SQL Server 2025 memory-first indexing can accelerate hybrid transactional and analytical workloads by reducing disk I/O and latency.
September 9, 2026
by arvind toorpu DZone Core CORE
· 2,370 Views · 2 Likes
article thumbnail
Cutting Telemetry Volume Is Not the Same as Cutting Noise
A volume target removes bytes, not noise. Once easy cuts run out, you pay in answers you won't have. Govern the questions your team asks, not bytes per day.
September 8, 2026
by Severin Neumann
· 2,688 Views · 2 Likes
article thumbnail
Optimize an AI Agent to Sound Human, Judged by an AI Detector
Use LaunchDarkly agent optimization to make an AI agent's replies sound human against GPTZero, an AI detector, as an inverted judge.
September 8, 2026
by Scarlett Attensil
· 2,703 Views · 1 Like
article thumbnail
What Actually Makes AI Infrastructure Agents More Reliable (It's Not More Agents)
Single AI agents fail during incidents. Four specialized agents — supervisor, telemetry, reasoning, action — handle observability more reliably.
September 8, 2026
by Kinjal Vaishnav
· 2,092 Views · 1 Like
article thumbnail
How Performance Engineers Find and Fix Hidden System Bottlenecks
Performance engineers diagnose end-to-end bottlenecks using data over intuition, turning hours of system delays into smooth, efficient execution.
September 7, 2026
by Alex Vakulov DZone Core CORE
· 2,420 Views · 1 Like
article thumbnail
The Bottleneck of Scaling
Learn how modern languages help developers take care of behind-the-scenes file descriptor management, kernel memory management, and heap management.
September 3, 2026
by Vishal Bhatia
· 2,863 Views · 1 Like
article thumbnail
Ampere PMU Profiler: A Guide to Microarchitecture Profiling
APP uses PMU metrics to pinpoint CPU stalls, cache misses, and other microarchitectural bottlenecks on Ampere processors
September 1, 2026
by Bhakti Hinduja
· 2,856 Views · 2 Likes
article thumbnail
Why Ping-Based Uptime Checks Are Failing Modern SaaS Architectures
Legacy server ping checks are obsolete. Synthetic monitoring solves this by simulating real user journeys to validate that actual business workflows function correctly.
August 31, 2026
by Arun Kulkarni
· 1,840 Views
article thumbnail
How to Monitor AI Models Without Drowning in Alerts
In this article, we will discuss monitoring AI models wisely. Prioritize actionable alerts so that real issues stand out instead of getting lost in the noise.
August 28, 2026
by Aditya Shrivastava
· 2,971 Views · 2 Likes
article thumbnail
Pragmatic Premature Optimization
Learn simple Java performance tips for strings, collections, enums, and initialization that make code faster without sacrificing readability.
August 28, 2026
by Alexander Radzin
· 3,403 Views · 3 Likes
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • ...
  • Next
  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×