Java is an object-oriented programming language that allows engineers to produce software for multiple platforms. Our resources in this Zone are designed to help engineers with Java program development, Java SDKs, compilers, interpreters, documentation generators, and other tools used to produce a complete application.
Wasm Inside Neo4j: Building the Example That Didn't Exist
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
Artificial intelligence is transforming software testing by enabling faster test creation, smarter execution, and more efficient quality assurance processes. With AI agents, we can generate test cases and scripts, run tests, and produce detailed reports with minimal manual effort. AI agents are software systems that leverage artificial intelligence to achieve goals and perform tasks on behalf of users. They can think through problems, plan actions, and remember things, while also making decisions on their own and improving over time. In this tutorial, we will build an AI-powered agent that does exactly this. It takes a .txt file with a simple use case written in plain English, processes it using an AI model such as OpenAI or Ollama, and generates Selenium test automation scripts in Java. How to Build an AI Agent for Test Automation We will be building an AI agent for test automation that will perform the following tasks: Reads a test scenario from a .txt fileSends the scenario with the right prompt to OpenAI/Ollama, based on the provider selected in the config fileGenerates the Selenium WebDriver Java test automation scripts bifurcated into: Page Object ClassesTest ClassTestng.xmlReadme.md with notes and steps to run the test Prerequisites The following prerequisites are required to start building the AI agent for test automation: Python 3.14 and aboveCode Editor (VS Code is preferred)Basic understanding of Python and the TerminalTest scenario written in plain English text LLM Setup OpenAI API KeyOllama setup on the local machine Understanding the Selenium AI Test Generation Workflow Before we deep dive into building the AI agent, let’s get a high-level overview of how this AI Agent works for generating Selenium Java test automation scripts: The process is explained step by step below: User input: The user provides a plain English test scenario in a .txt file.Processing layer: The Python main script reads the test scenario file and forwards it to the LLM model (OpenAI or Ollama), depending on the configuration.AI engine: Before sending requests to the LLM, the AI agent converts the test scenario from plain English to a structured prompt and adds clear instructions like: Generate Selenium WebDriver code in JavaUse the TestNG frameworkFollow the Page Object ModelGenerate Multiple Page Object classes(if needed)Rules for generating the codeOutput format for files The LLM understands the test description, steps, and expected behavior, and converts them into the respective files.File generation: The AI agent creates a new timestamp-based folder and stores the generated code in separate files, bifurcating the page object classes, test class, testng.xml, and README.md. Markdown . ├── config.py ├── tools.py ├── main.py ├── requirements.txt ├── .env ├── test_cases/ │ ├── input/ │ │ └── sample_test_case.txt │ └── output/ Let’s start building the AI agent using the following step-by-step guide: Step 1: Creating and Activating a Python Virtual Environment Run the following command in the terminal to create a Python virtual environment: Plain Text python -m venv venv Next, run the following command to activate the virtual environment: On MacOS/Linux: Plain Text source venv/bin/activate On Windows: Plain Text venv\Scripts\activate Step 2: Installing Dependencies Let’s create a requirements.txt as it serves as a blueprint for recreating the exact environment needed to run the project consistently across different environments. Plain Text #requirements.txt openai==2.28.0 requests==2.32.5 python-dotenv==1.2.2 The following command should be run from the terminal to install the dependencies: Plain Text pip install -r requirements.txt Step 3: Writing the Configuration File The configuration file holds details related to OpenAI, Ollama, input and output files, and settings for using the desired LLM provider to generate test automation scripts. Python #config.py from dataclasses import dataclass, field from pathlib import Path from dotenv import load_dotenv from dotenv import load_dotenv from openai import OpenAI import os load_dotenv() @dataclass class OpenAIConfig: model_name: str = os.getenv("OPENAI_MODEL_NAME", "gpt-5") max_tokens: int = int(os.getenv("OPENAI_MAX_TOKENS", "4096")) temperature: float = float(os.getenv("OPENAI_TEMPERATURE", "0.3")) api_key: str = os.getenv("OPENAI_API_KEY", "") @property def client(self): if not self.api_key: raise ValueError( "OPENAI_API_KEY is required when using OpenAI provider" ) return OpenAI(api_key=self.api_key) @dataclass class OllamaConfig: model: str = os.getenv("OLLAMA_MODEL_NAME", "llama3") prompt: str = "You are a Selenium WebDriver test automation expert. Generate Selenium WebDriver Java test scripts. STRICTLY follow the format. Any deviation is not acceptable." stream: bool = False temperature: float = float(os.getenv("OLLAMA_TEMPERATURE", "0.3")) ollama_baseurl:str = os.getenv("OLLAMA_BASEURL", "http://localhost:11434") ollama_endpoint: str = os.getenv("OLLAMA_ENDPOINT", "http://localhost:11434/api/generate") @dataclass class FileConfig: input_file: Path = Path("test_cases/input/sample_test_case.txt") output_file_path: Path = Path("test_cases/output/") @dataclass class AppConfig: provider:str = "ollama" #openai or ollama openai: OpenAIConfig = field(default_factory=OpenAIConfig) ollama: OllamaConfig = field(default_factory=OllamaConfig) files: FileConfig = field(default_factory=FileConfig) config = AppConfig() The configuration file serves as a centralized configuration system for switching between LLM providers and managing file settings. Let’s break it down to understand it further: Dataclasses: The @dataclass automatically generates the __init__() method, making it easier to create and manage objects without writing boilerplate code.Separate Configs: Separate config classes are defined for OpenAI, Ollama, and input/output file handling, each with default values. The input file name is currently configured as “sample_test_case.txt.”AppConfig: The AppConfig class serves as a central wrapper, enabling switching between providers (OpenAI or Ollama) with a single flag.field(default_factory=…): It ensures each config gets its own instance, avoiding shared state issues. Step 4: Implementing Utility Functions for Reading and Generating Code In this step, we’ll create a new Python file named tools.py and implement the following utility functions within it: load_test_case_from_file()build_prompt()generate_with_openai()generate_with_ollama()generate_selenium_test_script()create_timestamped_output_dir()split_and_save_files() The following packages should be imported in the tools.py file: Python import requests from config import config from typing import Optional from pathlib import Path from datetime import datetime from logger import logger import requests from config import config from logger import logger Let’s learn the utility functions one by one. load_test_case_from_file() function: Python def load_test_case_from_file(file_path: str | Path) -> str: logger.info(f"Loading test case from: {file_path}") try: with open(file_path, "r", encoding="utf-8") as file: content = file.read() logger.info(f"Test case loaded successfully ({len(content)} characters)") return content except FileNotFoundError: logger.error(f"Test case file not found: {file_path}") raise FileNotFoundError(f"Test case file not found: {file_path}") except Exception as e: logger.exception(f"Error reading test case file: {e}") raise RuntimeError(f"Error reading test case file: {e}") The load_test_case_from_file() function takes the test case file path as a parameter, reads its contents, and returns them as a string. It safely handles errors by raising a clear message if the file is not found and wraps any other unexpected issues in a RuntimeError. build_prompt() function: Python def build_prompt(use_case_text: str) -> str: logger.info("Building AI prompt") prompt = f""" You are a test automation expert specializing in Selenium WebDriver with Java. Generate Selenium automation test script using the following instructions and skills: - Use Java 17 to write the code - Use latest Selenium WebDriver Java dependency version to write code - Do not write code statement "System.setProperty()" to add chromedriver path, in the tests - Follow Page Object Model (POM) - Use latest version of TestNG dependency - Apply best coding practices for writing Java code - Add comments explaining each step - Add assertions using TestNG assertion - Do not add random assertion statements in the code - Do not mention any text in README that says that ChromeDriver path should be added to the Path - Never use brittle XPATH and CSS Selectors selectors such as .btn-primary, .container > div:nth-child(2), #content div span, or auto-generated classes. IMPORTANT: You MUST follow the exact output format below. Rules: - ALWAYS start each file with ===FILE: filename=== - Use class names based on the web page(e.g. HomPage.java, LoginPage.java, etc. These names are for instructions only, use class name specific to the web page) - Do not add "Page" to the test class name - Do NOT add explanations outside file blocks - DO NOT skip this format - If you do not follow this format, the output will be rejected - DO NOT use markdown (no **, no ``` blocks) - DO NOT add file names outside ===FILE: markers - ONLY use ===FILE: filename=== format - OUTPUT FORMAT FOR FILES(STRICT): ===FILE: filename=== file content - The following files MUST only be generated in the same order(STRICT). No deviation is acceptable: - Multiple Page Object classes(if needed) (Strictly Page object class, no WebDriver instantiation in these classes, Do not create duplicate page object classes) - Test class(WebDriver should be instantiated in the Test class, Do not use WebDriverManager to instantiate WebDriver, Use TestNG's @BeforeMethod annotation and define a method to instantiate the WebDriver, Use TestNG's @AfterMethod to quit the WebDriver) - Add assertions using TestNG assertion - Do not add random assertion statements in the code - testng.xml(Follow correct structure as per TestNG guidelines) - README.md (Include notes and steps to run the test using testng.xml file) Use Case: {use_case_text} """ logger.info(f"Prompt created ({len(prompt)} characters)") return prompt The build_prompt() function is a core part of the application because it constructs the prompt sent to the LLM to generate Selenium Java test automation code. It takes the use case text as input and embeds it into a detailed instruction template, guiding the model to generate Java-based Selenium tests using best practices such as POM and TestNG. The prompt also enforces file-formatting rules like (===FILE: filename===) to ensure the output is structured, consistent, and ready for file generation without manual cleanup. It also enforces rules to generate the POM, test class, testng.xml, and README in strict order to keep the output consistent. Using Effective Prompts Effective prompts are essential for generating clean, reliable Selenium test scripts with AI. By clearly defining the requirements, the model can be guided to use best coding practices, create clean test scripts, and use the Page Object Model (POM) to generate well-structured code. Prompt Design for Better Selenium Java Tests Clear, specific prompts help the AI generate cleaner Selenium test scripts. Mention details such as using Java for Selenium, using the latest versions of Selenium and TestNG, and following coding best practices so the AI can better meet the requirements and produce well-designed output. Customizing Prompts for Page Object Model (POM) The following prompts can be used to train the AI to generate the page object files we need : Use best practices to generate the Page Object classes.Generate Multiple Page Object classes (if needed).Do not instantiate the WebDriver in the Page Object class.Do not create duplicate Page Object classes.Additionally, we can also mention using best locator strategies, like preferring IDs and CSS selectors over complex XPaths, to make the tests more stable, readable, and easy to maintain. generate_with_openai() function: Python def generate_with_openai(prompt: str) -> Optional[str]: response = config.openai.client.chat.completions.create( model=config.openai.model_name, temperature=config.openai.temperature, max_tokens=config.openai.max_tokens, messages=[ {"role": "system", "content": "You are a Selenium WebDriver test automation expert. Generate Selenium WebDriver Java test scripts. STRICTLY follow the format. Any deviation is not acceptable."}, {"role": "user", "content": prompt}, ], ) return response.choices[0].message.content The generate_with_openai() function sends a prompt to the OpenAI API to generate Selenium WebDriver test scripts in Java. The OpenAI client is initialized using the API key from an environment variable, allowing the code to connect and interact with OpenAI services securely. It uses the model, temperature, and token limits from the configuration. It generates the request with a message that defines the AI’s role and user prompt (the build_prompt() function will be supplied here). The API returns multiple choices, and the function extracts the generated content from the first response. Finally, it returns the generated test script as a string. generate_with_ollama() function: Python def generate_with_ollama(prompt: str) -> str: logger.info( f"Generating test code with Ollama model: {config.ollama.model}" ) response = requests.post( config.ollama.ollama_endpoint, json={ "model": config.ollama.model, "prompt": config.ollama.prompt + "\n" + prompt, "stream": config.ollama.stream, }, ) response.raise_for_status() data = response.json() generated_text = data.get("response", "") logger.info( f"Ollama generation completed ({len(generated_text)} characters)" ) return generated_text The generate_with_ollama() function generates text by sending a POST request to the configured Ollama API endpoint. It combines a predefined base prompt with the user-provided prompt and sends it along with the model name and streaming configuration. After receiving the response, raise_for_status() checks for HTTP errors, and response.json() parses the response into a Python dictionary. The generated text is extracted from the response field, with an empty string used if the field is missing. Finally, the function logs the response length and returns the generated text. generate_selenium_test_script() function: Python def generate_selenium_test_script(test_case_text: str) -> Optional[str]: logger.info("Starting Selenium test generation") prompt = build_prompt(test_case_text) provider = config.provider.lower() logger.info(f"Using AI provider: {provider}") if provider == "openai": logger.info("Sending prompt to OpenAI") return generate_with_openai(prompt) elif provider == "ollama": logger.info("Sending prompt to Ollama") return generate_with_ollama(prompt) else: logger.error(f"Unsupported provider: {provider}") raise ValueError(f"Unsupported provider: {provider}") The generate_selenium_test_script() function acts as a wrapper to generate a Selenium WebDriver test script based on the input test case. It first builds a prompt using the build_prompt (test_case_text) function, which prepares the input for the LLM. Next, based on the configured provider, i.e., OpenAI or Ollama, it dynamically calls the respective function to generate the script. If an unsupported provider is specified, it raises a ValueError to prevent unexpected behavior. create_timestamped_output_dir() function: Python def create_timestamped_output_dir(base_output_path: Path) -> Path: timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S") output_dir = base_output_path / timestamp output_dir.mkdir(parents=True, exist_ok=True) logger.info(f"Created output directory: {output_dir}") return output_dir The create_timestamped_output_dir() function creates a new folder named with the current date and time, so each run gets its own unique output directory. It ensures the folder exists (creating it if needed) and returns its path for saving files. split_and_save_files() function: Python def split_and_save_files(generated_text: str, base_output_path: Path) -> None: sections = generated_text.split("===FILE:") if len(sections)<=1: logger.error("No structured files found in AI response") raise ValueError ("No Structured files found in AI response!") logger.info(f"Found {len(sections) - 1} generated files") pageobject_dir = base_output_path/"pageobjects" pageobject_dir.mkdir(parents=True, exist_ok=True) for section in sections[1:]: section = section.strip() parts = section.split("\n",1) raw_filename = parts[0].strip() filename = raw_filename.replace("===", "").strip().split()[0] content = parts[1].strip() if len(parts)>1 else "" content = content.split("===FILE:")[0].strip() if filename.endswith("Page.java"): file_path = pageobject_dir / filename else: file_path = base_output_path / filename logger.info(f"Writing file: {file_path}") with open(file_path, "w", encoding="utf-8") as f: f.write(content) logger.info(f"Created: {file_path} ({len(content)} characters)") The split_and_save_files() function takes the AI-generated response and splits it into multiple files based on a marker (===FILE:). It extracts each filename and its content, then saves them into the appropriate folders. The Page Object files go into a pageobjects directory, while the test class, testng.xml, and README.md files go into the main output/<timestamped> folder. If the expected structure is missing, it throws an error to avoid saving incorrect output. Step 5: Writing the Main Script to Glue Everything In this step, we’ll create a main.py file to orchestrate the complete workflow of generating Selenium test scripts using AI: Python import re from tools import load_test_case_from_file, generate_selenium_test_script,split_and_save_files,create_timestamped_output_dir, check_llm_connection from config import config from pathlib import Path from logger import logger def main() -> None: try: logger.info("Starting Selenium AI Test Generator") if not check_llm_connection(): logger.error( "❌ LLM connection check failed. " "Stopping test generation." ) return input_file = config.files.input_file output_file_path = config.files.output_file_path run_output_dir = create_timestamped_output_dir(output_file_path) use_case_text = load_test_case_from_file(input_file) if not use_case_text: logger.error("Test case file is empty.") raise ValueError("Test case file is empty.") generated_output = generate_selenium_test_script(use_case_text) if not generated_output: logger.error("Failed to generate Selenium WebDriver Java test automation scripts.") raise RuntimeError("Failed to generate Selenium WebDriver Java test automation scripts.") split_and_save_files(generated_output,run_output_dir) logger.info(f"✅ All generated files saved successfully to {run_output_dir}") except Exception as e: logger.error(f"❌ Error occurred while generating output files: {e}") if __name__ == "__main__": main() The main.py file is where the program starts and manages the whole process of generating Selenium test scripts from start to finish. It reads the input test case file, sends the test case to the AI to generate automation scripts, and generates the timestamped output folder. Once the scripts are generated, it splits them into multiple files and saves them in the appropriate directories. In case of any exception, it catches the error and prints the message “Error occurred while generating output files” with the exception details. Generating First Test Scripts Let’s create a new text file, ”sample_test_case.txt,” and place it in the input/ folder with the following test scenario to generate the Selenium test scripts using the AI agent we created: Plain Text Title: Application Login scenario Precondition: User is registered in the application. Steps: 1. Open Chrome browser 2. Navigate to https://ecommerce-playground.lambdatest.io/index.php?route=account/login 3. Enter "[email protected]" in the E-Mail Address field 4. Enter "Password@321" in the Password field 5. Click on the Login Button 5. Add an assert statement to check that "My Account" page is displayed. Configuration and Setup Using OpenAI Create a .env file and add your OpenAI API key in it. This file should always remain on your local machine and should not be committed to the remote repository. Plain Text OPENAI_MODEL_NAME=<model name> OPENAI_MAX_TOKENS=<max tokens> OPENAI_TEMPERATURE=<temperature value> OPENAI_API_KEY=<Your OpenAI API Key> Using Ollama Download and install Ollama.Start Ollama by running the command ollama serve in the terminal.Pull the Qwen3:8b model by running the command — ollama pull qwen3:8b. Create a .env file and update the following details in it: Plain Text OLLAMA_MODEL_NAME=qwen3:8b OLLAMA_TEMPERATURE=0.3 OLLAMA_ENDPOINT=http://localhost:11434/api/generate OLLAMA_BASEURL=http://localhost:11434 Before running the model, ensure that Ollama is running in the background. Open Terminal and run the command ollama serve Note: Running the model explicitly is not required. Running the AI Agent Open the terminal and run the following command to generate the test scripts: Plain Text python main.py The following log should be printed in the console after the Agent is run successfully: The test scripts should be generated in a new timestamped(YYYY-MM-DD_hh-mm-ss) folder that is generated inside the output/ folder: Reviewing the Generated Java Selenium Code Let’s check each file generated in the output/ folder and review the code to verify if it fits the test scenario we defined and follows the expected structure and best practices. Let’s check each file generated in the output/ folder and review the code to verify if it fits the test scenario we defined and follows the expected structure and best practices. The following page object class is generated: LoginPage.java Java public class LoginPage { private WebDriver driver; public LoginPage(WebDriver driver) { this.driver = driver; } public void navigateToLoginPage() { driver.get("https://ecommerce-playground.lambdatest.io/index.php?route=account/login"); } public void enterEmail(String email) { driver.findElement(By.name("email")).sendKeys(email); } public void enterPassword(String password) { driver.findElement(By.name("password")).sendKeys(password); } public void clickLoginButton() { driver.findElement(By.xpath("//button[@type='submit']")).click(); } } The page object file is generated correctly using the name and XPath locator strategies, and appropriate methods are created to interact with the respective WebElements on the page. However, the import statements are missing, which should be added to avoid errors. Additionally, locators should be checked, and explicit waits could be added in this class while locating the elements to reduce flakiness during test execution. The following test class is generated: LoginTest.java Java import org.openqa.selenium.By; import org.openqa.selenium.WebDriver; import org.openqa.selenium.WebElement; import org.openqa.selenium.chrome.ChromeDriver; import org.testng.Assert; import org.testng.annotations.BeforeMethod; import org.testng.annotations.Test; public class LoginTest { private WebDriver driver; @BeforeMethod public void setup() { System.setProperty("webdriver.chrome.driver", "path_to_chrome_driver"); driver = new ChromeDriver(); } @AfterMethod public void teardown() { driver.quit(); } @Test public void testLoginScenario() { LoginPage loginPage = new LoginPage(driver); loginPage.navigateToLoginPage(); loginPage.enterEmail("[email protected]"); loginPage.enterPassword("Password@321"); loginPage.clickLoginButton(); WebElement myAccountPageTitle = driver.findElement(By.xpath("//h1[@class='page-title']")); Assert.assertTrue(myAccountPageTitle.isDisplayed(), "My Account page is not displayed"); } } The test class is generated per the provided prompt and includes the @BeforeMethod and @AfterMethod annotations for setup and teardown. However, it includes System.setProperty(), which can be removed because the latest version of Selenium doesn't require it to set the ChromeDriver path. The import statement for the @AfterMethod annotation is missing, which should be added. Additionally, the myAccountPageTitle WebElement can be moved to a separate MyAccount Page Object class, and its XPath should also be verified to ensure it's pointing to the correct element on the page. Testng.xml XML <?xml version="1.0" encoding="UTF-8"?> <!DOCTYPE suite SYSTEM "http://testng.org/testng-1.0.dtd"> <suite name="LoginTestSuite"> <test name="LoginTest"> <classes> <class name="LoginTest"/> </classes> </test> </suite> The qualified class name should be updated once the code is moved to the project. README.md Markdown **How to run the test** 1. Install TestNG framework and Selenium WebDriver. 2. Place the Selenium ChromeDriver executable in a directory, update the path_to_chrome_driver variable in `LoginTest.java` with this directory. 3. Run the test using the testng.xml file by executing the command: `java -cp .;test-classes org.testng.TestNG testng.xml` Note: Update the path_to_chrome_driver variable with your actual ChromeDriver executable path. The steps for the ChromeDriver executable and the command to run testng.xml can be removed, as they are not accurate. Instead, the step to run the testng.xml by right-clicking on it can be updated here. Overall, the AI agent generated the test code quickly, saving manual effort and providing a good starting point. However, review, validate, and refine the code before using it in a project. Best Practices While Using AI Agents The following best practices should be considered while using AI Agents to generate automation code: Improving prompt quality for reliable tests: Clear and detailed prompts help the AI generate stable test scripts. Mention requirements like framework names, naming conventions, and coding standards to improve structure and avoid errors.Adding assertions and validations: Always ensure the generated tests include proper assertions to validate expected outcomes. This helps confirm that the application behaves correctly instead of just performing actions.Review and refactor generated code before execution: Always review AI-generated code before running it. Refactoring helps remove redundancies, improve readability, and align the code with the project standards.Handling invalid/empty inputs: Validate inputs like test case files before processing them. Ensure you add proper steps and verification checks for empty or invalid data to prevent unexpected failures during execution.Integrating OpenAI API error handling: Proper error handling for API calls helps manage issues like timeouts, rate limits, or failed responses. Using try-catch blocks and logging error messages improves error readability. Exception handling helps interpret error messages in simple language and makes debugging easier.Logging and debugging tips: Adding logs at key steps helps track execution flow and quickly identify issues. Clear logging messages make debugging easier, especially when working with AI-generated outputs and API calls. Watch the step-by-step YouTube tutorial on How to Build an AI Agent to Generate Selenium WebDriver Tests in Java. Limitations and Challenges The following limitations and challenges can be faced while working with AI Agents for test script generation: AI-generated scripts may fail: AI-generated test scripts can be a starting point, but they may fail when the input prompt is unclear or missing important details. They can also break because of dynamic elements, incorrect UI assumptions, or environment-specific issues.Manual cleanup/selector accuracy: AI may generate locators that are not always reliable or optimal. Always review and refine locators in AI-generated scripts, preferring stable options like IDs or CSS selectors over complex XPaths.Test maintenance and flakiness considerations: AI-generated tests may be flaky if they don’t handle waits, dynamic content, or synchronization properly. Adding proper waits and improving test design can reduce flakiness. Regular maintenance is required to keep tests stable as the application evolves.Reviewing hallucinations: AI can sometimes generate incorrect or non-existent methods, elements, or logic. Review the code carefully to catch such hallucinations before execution. Validating against actual application behavior ensures the tests are accurate and usable. Conclusion To design an AI agent to generate test scripts, create a comprehensive prompt with clear rules, a structured output format, and strict guidelines to ensure consistent results. It’s also important to validate and refine the generated code and handle edge cases and failures gracefully. It all depends on the team’s decision; since AI is evolving, designing an AI agent that generates test automation scripts can also be a viable and flexible approach, especially when customization and control are important.
Enterprise applications commonly face multiple data challenges. Some data requires transactional integrity and relationships, while other data prioritizes fast, predictable access. Sessions, counters, rate limits, temporary state, often-accessed objects, and coordination data may not benefit from the complexity of a relational model. In these cases, a key-value database's simplicity becomes an architectural advantage. This simplicity is especially valuable in distributed and cloud-native systems, where latency, throughput, plus scalability directly shape user experience and infrastructure costs. A key-value database offers a focused approach: identify data by a key and retrieve or update it efficiently. The challenge is selecting a technology that delivers this performance while meeting the operational maturity, ecosystem support, and governance standards required for enterprise applications. Valkey meets these needs successfully. Originating from the Redis OSS lineage and developed as a vendor-neutral open-source project under the Linux Foundation, Valkey delivers a high-performance key-value platform suitable for caching, application state, messaging, and primary data storage. Beyond being another database option, it lets organizations explore how key-value persistence fits into modern enterprise architecture and lets Java applications use its benefits without tightly coupling to a specific datastore. Why Key-Value Databases Matter Key-value databases use a simple data model in which each unique key identifies a value. This simplicity is effective when applications can directly locate the required data. By enabling direct reads and writes, key-value databases typically deliver low latency, high throughput, and a horizontally scalable operational model. In enterprise systems, this model suits scenarios such as distributed sessions, caching, counters, rate limiting, feature flags, shopping carts, temporary workflow state, idempotency keys, leaderboards, and frequently accessed application data. These workloads prioritize fast access by identifier over joins, ad hoc queries, or complex relational constraints. The main architectural advantage of key-value databases is their specialization for specific access patterns, rather than universal speed or simplicity. When the primary requirement is to retrieve the current value for a given key, adding a more complex persistence model can introduce unnecessary overhead. As part of a polyglot persistence strategy, key-value stores enable architects to align the database model with the workload, rather than forcing all workloads into a single database. Putting Valkey Into Practice With Jakarta NoSQL A key advantage of using Valkey in enterprise Java is that it does not require a new programming model. With Jakarta NoSQL and Eclipse JNoSQL, Valkey serves as another key-value implementation behind a consistent API and mapping model. Domain annotations remain unchanged, so switching between key-value databases usually involves only updating the driver and its configuration, not rewriting the application. This abstraction is valuable architecturally. The application relies on the Jakarta NoSQL contract, while Eclipse JNoSQL manages integration with the database. Although database-specific features may introduce some coupling, applications that use the portable API can switch key-value implementations with minimal impact. For this article, we will use a simple Java SE example. This persistence layer can later support a REST API, messaging consumer, scheduled process, or other enterprise architecture without altering the core database interaction. Starting Valkey The first step is to make a Valkey instance available. Docker provides a convenient way to start one locally: Shell docker run --name valkey-instance \ -p 6379:6379 \ -d valkey/valkey:latest With Valkey running, add the Eclipse JNoSQL Valkey driver to the Jakarta NoSQL infrastructure, which includes CDI, Eclipse MicroProfile Config, and Jakarta JSON Processing. XML <dependency> <groupId>org.eclipse.jnosql.databases</groupId> <artifactId>jnosql-valkey</artifactId> <version>${jnosql.version}</version> </dependency> Configure the connection externally: Properties files jnosql.keyvalue.database=developers jnosql.valkey.port=6379 jnosql.valkey.host=localhost Since Eclipse JNoSQL integrates with Eclipse MicroProfile Config, you do not need to hard-code these values. They can be provided through configuration sources such as environment variables, in line with the Twelve-Factor App methodology. Mapping an Entity The mapping model for a key-value database is intentionally simple. Identify the class as an entity and specify the field that represents its key: Java @Entity public class User { @Id private String userName; private String name; private List<String> phones; // constructors, getters, setters... } Importantly, @Entity and @Id are part of the mapping abstraction, not Valkey itself. The domain model does not require Valkey-specific annotations. Using Jakarta NoSQL Eclipse JNoSQL provides KeyValueTemplate, a specialization of the Jakarta NoSQL Template API for key-value databases. This allows direct persistence and retrieval of entities: Java User user = User.builder() .phones(Arrays.asList("234", "432")) .username("username") .name("Name") .build(); KeyValueTemplate template = container.select(KeyValueTemplate.class).get(); User userSaved = template.put(user); System.out.println("User saved: " + userSaved); Optional<User> userFound = template.get("username", User.class); System.out.println("Entity found: " + userFound); For applications that prefer a repository abstraction, Eclipse JNoSQL integrates with Jakarta Data: Java @Repository public interface UserRepository extends CrudRepository<User, String> { } This approach allows the application code to focus more directly on domain operations: Java User user = User.builder() .phones(Arrays.asList("234", "432")) .username("username") .name("Name") .build(); UserRepository repository = container .select( UserRepository.class, DatabaseQualifier.ofKeyValue() ) .get(); repository.save(user); Optional<User> userFound = repository.findById("username"); System.out.println("User found: " + userFound); Notably, this code includes no Valkey-specific API in the entity or repository. Valkey is an infrastructure choice, while Jakarta NoSQL and Jakarta Data remain the application-facing abstractions. This separation guarantees the architecture remains reusable if the underlying key-value technology changes. Conclusion Key-value databases are highly effective for workloads that require direct access, low latency, and high throughput, rather than complex queries or relational navigation. This article examined how this model fits within enterprise architecture and how Valkey can integrate via Eclipse JNoSQL, allowing applications to avoid direct dependencies on vendor-specific APIs. By maintaining consistent entity mapping and using Jakarta NoSQL or Jakarta Data abstractions, switching key-value implementations becomes mainly a matter of infrastructure and configuration. This shift reflects a broader evolution in enterprise Java, as the platform expands its persistence capabilities beyond traditional relational databases. With Jakarta Persistence, Jakarta Data, Jakarta NoSQL, and tools like Eclipse JNoSQL, architects can choose the best data model for each workload while keeping familiar programming abstractions. Valkey enhances this ecosystem by providing a robust key-value option, making polyglot persistence both feasible and practical.
Data engineers managing batch SQL pipelines on Snowflake, BigQuery, and increasingly Databricks, and streaming pipelines on Apache Flink face a familiar problem: two toolchains, two skill sets, two CI/CD pipelines.dbt is now extending into stream processing. This post explains what that means in practice, why it matters for data engineering teams, and what a concrete implementation looks like with Apache Flink on Confluent Cloud. Data Streaming Meets the Lakehouse Data lakes promised to solve the enterprise data problem. The reality has been messier. Batch pipelines produce stale information, and analytical workloads run hours after the business event occurred. By the time a query runs, the window for action is often already closed. The lakehouse pattern has improved matters. Apache Iceberg has become the dominant open table format, supported across Snowflake, Databricks, BigQuery, and a growing number of query engines. Teams can run SQL analytics directly on data in object storage without duplicating it into a proprietary warehouse. But the lakehouse alone does not solve the real-time problem. Data still arrives as a batch, minutes or hours after the source event. That gap reflects a deeper architectural split. Data streaming with Apache Kafka and Flink is the operational layer: it handles critical SLAs, powers event-driven applications, and keeps business systems running in real time. The lakehouse is the analytical layer: it stores historical data for reporting, ML, and near real-time or batch analytics. These are two distinct workloads with different requirements regarding uptime, data loss, latency, and throughput. They need to coexist without forcing engineers to build and maintain two separate pipelines. How Kafka, Flink, and Iceberg Work Together That is what the combination of Apache Kafka, Apache Flink, and Apache Iceberg addresses. Kafka captures every event at the source and serves as the operational backbone for real-time systems. Flink processes and enriches data in motion, supporting both immediate operational decisions and the preparation of data for downstream analytics. Iceberg stores the result as a governed, queryable table for any analytical engine, whether that is Snowflake, BigQuery, or Databricks. A full treatment of this architecture, including schema evolution, compaction, and catalog integration, is covered here: Data Streaming Meets Lakehouse: Apache Iceberg for Unified Real-Time and Batch Analytics. The question is no longer whether streaming and lakehouse architectures can coexist. They already do. The question is how data engineering teams can work across both without maintaining separate toolchains. That is where dbt enters the picture. What Is dbt? dbt, the data build tool, is an open-source framework for SQL-based data transformation. A dbt model is a SQL SELECT statement saved as a file. dbt infers execution order from how models reference each other using ref(). The standard commands cover the full engineering workflow: dbt run executes the SQL against the target platform, dbt test validates data quality, and dbt docs generate produces a browsable documentation catalog. What made dbt successful is the discipline it brings to SQL work. Before dbt, transformation logic lived in scattered scripts and proprietary ETL tools. dbt replaced that with a code-first, version-controlled workflow with built-in lineage, testing, and documentation. Snowflake and BigQuery are where most dbt adoption lives today. Both are SQL-native and optimized for the ELT pattern dbt was built around. Redshift is a strong third platform in AWS environments. Databricks has seen growing dbt adoption more recently, driven by investments in serverless SQL Warehousing, but its roots are in Spark and Python, making it a newer entrant in the dbt ecosystem. dbt Labs crossed $100 million in ARR in early 2025, with over 5,000 paying customers. Around 90,000 dbt projects are running in production today. The Fivetran and dbt Labs merger, announced in October 2025, created a combined data infrastructure company with nearly $600 million in annual revenue — a clear signal that dbt has moved well beyond a popular open-source tool and into foundational enterprise data infrastructure. dbt Meets Apache Flink: One Workflow for Data Engineers Data engineering teams managing both batch and streaming today operate in two separate realities. Snowflake or BigQuery on one side: dbt models, version-controlled SQL, automated tests, generated docs. Apache Flink on the other: Terraform scripts, custom deployment code, or the Flink console. Skills and practices do not transfer between the two. That separation has a real cost. Streaming pipelines are harder to test, harder to document, and harder to hand over. Many teams compensate by keeping streaming logic minimal and pushing transformation work downstream into the warehouse, which reintroduces latency and undermines the point of streaming. The vision is straightforward: one SQL workflow for both. The engineer who builds dbt models on Snowflake or BigQuery should be able to apply the same approach to an Apache Flink streaming pipeline, without switching tools or rebuilding CI/CD from scratch. Two toolchains mean two testing strategies, two documentation systems, and two skill sets to hire and retain. Governance enforcement becomes inconsistent across the two environments. SQL is the shared foundation that makes this realistic. Flink SQL is mature and production-proven. Snowflake and BigQuery are SQL-native. Apache Iceberg tables are queryable via SQL across multiple engines. dbt wraps SQL with engineering discipline. The model files look the same. The ref() dependency resolution works the same way. Tests and documentation generation work through the same commands. Organizations do not need to hire separate Flink infrastructure specialists. The existing data engineering team can own both sides. Apache Iceberg connects the two worlds at the storage layer. A Flink pipeline writes structured, governed events into an Iceberg table in the organization's own S3 bucket. That same table is immediately readable by Snowflake, BigQuery, or Databricks without any additional ETL step. dbt can model data across the full pipeline: shaping it as it streams through Flink, and transforming it again when it lands in the warehouse for analytics. This is also a direct enabler of the Shift Left Architecture 2.0. The Shift Left approach moves data integration logic closer to the source, applying quality checks, enrichment, and governance in the streaming layer before data lands in the lakehouse. Until now, that required streaming-specific skills that most dbt-native teams did not have. dbt for Flink lowers that barrier considerably. The full architectural detail is covered here: The Shift Left Architecture 2.0: Operational, Analytical and AI Interfaces for Real-Time Data Products. Concrete Example: dbt on Confluent Cloud with Apache Flink The most concrete implementation available today is the dbt-confluent adapter, released by Confluent alongside the confluent-sql Python driver. Both are open source and available on PyPI and GitHub. Data engineers define streaming pipelines as dbt models and deploy them to Flink compute pools using the standard dbt run command. Getting started is a single step: pip install dbt-confluent Three materializations are supported: view for a virtual Flink SQL view over a Kafka topic, streaming_table for a continuous always-current result set, and streaming_source for defining a Kafka topic as a dbt source. Testing is deterministic, using Confluent Cloud's snapshot query capability to return bounded point-in-time results rather than silently passing on timeout. Documentation generation works through INFORMATION_SCHEMA integration, producing the same browsable catalog that Snowflake and BigQuery projects generate. The underlying confluent-sql driver is DB-API v2 compliant, meaning any compatible tool can connect directly to Confluent Cloud Flink: Airflow and Dagster for orchestration, Pandas for snapshot queries, Streamlit for live dashboards, and LangChain for AI agent workflows. For data engineers already working in dbt, this means the skills and practices built around Snowflake or BigQuery transfer directly to the streaming side of the architecture. The Data Engineer Owns Batch and Streaming with dbt The separation between batch and streaming engineering has always been more organizational than technical. Both worlds use SQL. Both require testing, documentation, and reliable deployment. The tools just never bridged the gap, so organizations staffed and operated two distinct engineering disciplines. dbt extending to Apache Flink changes that equation. The data engineer who runs dbt on Snowflake or BigQuery today can apply the same mental model, commands, and CI/CD pipeline to Flink streaming pipelines. No Flink infrastructure specialization required. They write SQL models, define tests, generate documentation, and deploy, exactly as they do for batch. The implication is straightforward. The investment in dbt skills and tooling now extends further into the architecture. Streaming can be adopted incrementally by the same data engineering teams already trusted for batch. One team, one tool, one governance standard, across both operational and analytical workloads. The Flink adapter for dbt is earlier in maturity compared to dbt on Snowflake or BigQuery, and teams should expect to work with an evolving ecosystem. But the foundation is solid, the direction is clear, and the core architectural components are already running in production at scale across multiple industries. The demand from data engineering teams is real and growing.
Running Apache Flink on a mainframe sounds odd at first. A modern stream processing engine on a platform most people call legacy? But take a closer look. It is not only possible. It might be a smart move for some of the largest financial institutions in the world. This post explores why some enterprises want Apache Flink on the mainframe, how it could work, and whether it is a brilliant innovation or a technical detour. Disclaimer: The views and opinions expressed in this blog are strictly my own and do not necessarily reflect the official policy or position of my employer. Mainframes Will Still Matter in 203X! A few months ago, I wrote about integrating Apache Kafka with mainframe systems. The blog covered various real-world examples across industries. The key message: Mainframes are still in use. In many organizations, they are not going away. They remain a central part of IT strategy, especially in banking, insurance, and the public sector. But they are not just legacy systems. Modern mainframes such as the IBM z17 offer the latest Telum II processor and support up to 64 terabytes of system memory. The z17 enables very large in‑memory workloads and faster processing for analytics and real‑time use cases. These systems also integrate on‑chip AI acceleration and optional AI‑focused hardware to support machine learning and real‑time decisions directly where mission‑critical data resides, while running modern Linux environments and container platforms. Some companies are still on the mainframe because they cannot easily migrate. But many others do not want to move away. Instead, they modernize around the mainframe. Apache Kafka and Flink play a key role in this journey. They enable a real-time data foundation that connects core systems with modern applications across environments. In future hybrid cloud strategies, this becomes even more critical. Kafka acts as the central nervous system, delivering the right data and context at the right time between on-prem mainframes and cloud-based AI services, including agentic AI and large language models. An event-driven architecture with hybrid streaming replication ensures business-critical decisions are made on fresh, reliable, and contextual information. Mainframe Migration Has Not Happened Ask any architect or CTO in banking. Mainframe migration has been on the roadmap for over two decades. Full replacement of core systems is still rare. However, it is important to distinguish between migration and offloading. Mainframe migration means shutting down mainframe workloads entirely and moving all applications and data to a new platform. There are many reasons: Risk is too highOrganizational resistance is strongMainframe skills are still needed but hard to findSystems are complex and deeply integratedThese applications run reliably and perform well Mainframe offloading, on the other hand, is much more common. It means moving selected workloads, queries, or processing tasks off the mainframe to more flexible and scalable platforms. This reduces load and cost on the mainframe while enabling innovation elsewhere. I have shared several real-world examples of offloading in action, using Kafka, IBM MQ, and Change Data Capture (CDC) tools like IBM IIDR or Precisely to synchronize and replicate data between mainframe systems and the cloud or distributed infrastructure in real time: Mainframe Offloading and Integration Examples. Because of this, many firms choose mainframe integration and a slow lift-and-shift leveraging the Strangler Fig design pattern over migration. Kafka is already helping. Flink is the next step. Apache Flink Meets the Mainframe: Unlikely Combo, Real Potential At first glance, Apache Flink and the mainframe seem like technologies from two different worlds. But combining them can unlock surprising value. What Is Apache Flink? Apache Flink is the leading open-source stream processing engine. It is designed to process high volumes of data continuously and in real time, rather than in batches. Flink is widely used to support use cases like fraud detection, customer personalization, operational monitoring, and data transformation at scale. Many of the largest tech companies and digital natives rely on Flink to process billions of events per day with low latency and high throughput. It supports both event streaming and batch workloads, but its true strength lies in real-time use cases. Here is an example of continuous stream processing leveraging Apache Flink together with OpenAI for Generative AI in real-time: Flink is built for modern environments. It runs natively on Kubernetes, integrates with Apache Kafka for real-time data ingestion, and is commonly deployed in public cloud, private cloud, or hybrid architectures. This makes it an ideal fit for enterprises looking to build fast, intelligent applications on fresh and contextual data. How to Run Apache Flink on the Mainframe? Yes, Apache Flink can run on the mainframe. In fact, it already does. I have already seen this deployed in a real-world environment. A large global financial institution is preparing to invest massively to expand its use of Apache Flink. Running Flink on IBM LinuxONE is a central part of that strategy. This is NOT a lab experiment, but a production-focused initiative. This bank already uses Kafka and Flink in production. Now they want to move Flink compute workloads onto the mainframe. The reason is simple. They already have unused compute on LinuxONE. Running Flink there is cheaper and easier to scale (for some companies) than scaling out other systems. The architecture is modern. IBM LinuxONE runs OpenShift. IBM LinuxONE is a high-performance, enterprise-grade server built on IBM Z architecture. It is designed to run Linux workloads with extreme reliability, scalability, and security. Unlike traditional mainframes focused on COBOL and legacy apps, LinuxONE is optimized for modern Linux applications. Flink is deployed in containers inside OpenShift's Kubernetes infrastructure, just like in any other cloud or data center. From a technical perspective, you need to build Docker images for the s390x architecture to run Apache Flink on IBM LinuxONE. In addition, components like RocksDB, which is used as a state backend in Flink, must be compiled for s390x to ensure full functionality. Why Put Apache Flink on the IBM Mainframe? This is a very valid question! Nobody would buy a mainframe just to run Flink on it. However, this approach offers several benefits for organizations that already own and operate mainframe infrastructure: Available compute resources on the mainframe.Consume data directly from mainframe sources (such as IBM MQ or other integration interfaces) and process it directly on the mainframe; or consume data from external sources such as a Kafka cluster running on x86 infrastructure, enabling flexible integration across hybrid environments.Lower total cost of ownership (TCO) regarding hardware and license costs compared to adding new external x86 servers and bi-directional integration pipelines.Simplified operations within a single, familiar environment.Benefit from IBM actively promoting LinuxONE and driving more workloads onto the platform. Mainframes are not only still alive. They are growing. IBM’s infrastructure business, which includes the mainframe, is doing very well. In Q3 2025, IBM reported 3.6 billion dollars in revenue for the infrastructure segment. That is 17 percent growth. IBM Z alone grew 61 percent. In Q2 2025, infrastructure revenue was 4.14 billion dollars, beating expectations by a wide margin. This is not legacy tech in decline. It is a platform in transformation. A New Chapter for Stream Processing and Mainframes Apache Flink running on the mainframe may sound unusual at first, but it reflects a broader shift in how enterprises think about modernization. The mainframe is not just a legacy system to replace. IBM Mainframe can fit hybrid cloud strategies, especially in highly regulated industries like banking and insurance. Apache Flink brings real-time intelligence. The mainframe brings performance, reliability, and unmatched security. Together, they offer a powerful combination for building fast, contextual, and mission-critical applications, without abandoning existing infrastructure. With Kafka as the backbone and Flink as the engine for real-time processing, organizations can connect mainframe systems with cloud innovation, including advanced AI workloads. This is not just about preserving the past. It is about extending and reusing trusted systems to meet the demands of the future. Enterprises that embrace this model can reduce risk, increase agility, and unlock new value from the heart of their operations.
If you ask Java developers about the concept of ‘Sneaky Throws,’ I am almost sure there will be a couple of opinions that are quite differently expressed, but similar in their meaning. Some will sum it up as being able to throw checked exceptions without declaring them explicitly; others will amend that it means writing functional-style code (lambdas) and being allowed to call methods that throw checked exceptions. Most probably, it will be surely mentioned that there’s a Lombok annotation called exactly @SneakyThrows that solves the problem immediately when put on a method. Last but not least, to outline it in a more pragmatic manner, the concept allows tricking the Java compiler into treating checked exceptions as runtime exceptions. All of these are valid points of view, and to clarify the concept, this article aims to provide a straightforward yet useful approach to handling methods that throw checked exceptions. Let’s jump right in and imagine the following situation. The team is requested to enhance the currently delivered application and implement new functionalities. This obviously happens on a ‘sprint-ly’ basis. Nevertheless, the project has been successfully developed for quite a while now; it also deals with legacy code, and moreover, developers are interacting with other parts of code that were written, let’s say, in a less fortunate manner. Such an example is the class below. Java public class TwoDigitsInteger { private final Integer value; public TwoDigitsInteger(Integer value) { this.value = value; } public boolean isValid() throws NotSetException { if (value == null) { throw new NotSetException("Number value not set."); } return value >= 10 && value <= 99; } public Integer getValue() throws NotSetException { if (value == null) { throw new NotSetException("Number value not set."); } return value; } } Just as its name suggests, it models a two-digit integer number. Instances of this class are immutable; the value is set upon construction, and it declares two methods, one for reading the value — getValue() — and another one for validating it — isValid(). We’re not going to further elaborate on the quality of the code, as it helps in the experiment done. The main issue here, the plot of this article, is the fact that both methods declare a NotSetException as they might throw it under certain circumstances, and even that might be fine unless this Exception hadn’t been a checked one. Java public class NotSetException extends Exception { public NotSetException(String message) { super(message); } } One option (and definitely the one worth taking into account) is to profit and consider the moment a good opportunity to refactor this ‘legacy’ code and at least make the Exception a runtime one. A few unit tests can be written (in case these are missing), then the implementation improved, and focus can be moved on the newly requested features. Nevertheless, for the sake of the experiment in this article, it’s assumed the TwoDigitsInteger class is kept as it currently is and the Exception remains checked. Exception Function Let’s consider a very simple scenario: there is a collection of TwoDigitsIntegers and the intent is to create a string expression that outlines the sum of the numbers. Java List<TwoDigitsInteger> numbers = List.of(new TwoDigitsInteger(10), new TwoDigitsInteger(25), new TwoDigitsInteger(37)); If writing the code as in the test below, Java @Test void sumExpression() { String result = numbers.stream() .map(TwoDigitsInteger::getValue) .map(String::valueOf) .collect(Collectors.joining("+")); Assertions.assertEquals("10+25+37", result); } the Java compiler will complain, saying — Unhandled exception: com.hcd.utilities.NotSetException – as the getValue() method declares a checked Exception and obviously it cannot be used inside a stream. To solve the issue, a try-catch is needed, which makes the code quite difficult to read (and ugly). Not to mention that we’re modifying the state of the joiner as we loop the collection. Java @Test void sumExpression1() { StringJoiner joiner = new StringJoiner("+"); for (TwoDigitsInteger number : numbers) { try { joiner.add(String.valueOf(number.getValue())); } catch (NotSetException e) { throw new RuntimeException(e); } } String result = joiner.toString(); Assertions.assertEquals("10+25+37", result); } In order to overcome this and allow having a fluent API even in situations where checked Exceptions are present, the following ExceptionFunction interface is created. Java @FunctionalInterface public interface ExceptionFunction<T, R, E extends Exception> { R apply(T t) throws E; } It is general enough; it represents a function that accepts one argument (of type T), produces a result (of type R) and when applied, an Exception subclass (of type E) might be thrown. Implementers shall define a single method, which effectively applies the function. Additionally, the following class is defined. Java public final class ExceptionWrapper { public static <T, R, E extends Exception> Function<T, R> apply(ExceptionFunction<T, R, E> function) { return t -> { try { return function.apply(t); } catch (Exception e) { throw new RuntimeException(e); } }; } ExceptionWrapper() { throw new UnsupportedOperationException("No need to be called."); } } When the ExceptionWrapper#apply() method is called, in case an Exception is thrown, it is wrapped into a RuntimeException one and thrown further irrespective of the type of the initial one (the checked Exception case is obviously covered as well, so we’re good). The ExceptionFunction passed as a parameter represents the initial call that is wrapped to overcome the problem. The previously discussed test is modified to use the ExceptionWrapper#apply() method. Not only does it now compile and run successfully, but the code readability is definitely improved. Java @Test void sumExpression() { String result = numbers.stream() .map(ExceptionWrapper.apply(TwoDigitsInteger::getValue)) .map(String::valueOf) .collect(Collectors.joining("+")); Assertions.assertEquals("10+25+37", result); } Exception Predicate Let’s now consider another straightforward scenario, one in which we want to count only the valid two-digit integers that are found in a designated range. Also, for the sake of this experiment, it’s assumed the previous TwoDigitsInteger class is used. As in the previous case, the following piece of code that would do the job doesn’t compile because of the same reason – Unhandled exception: com.hcd.utilities.NotSetException — as the isValid() method declares a checked exception, and it cannot be used inside a stream. Java long count = IntStream.range(0, 150) .mapToObj(TwoDigitsInteger::new) .filter(TwoDigitsInteger::isValid) .count(); Again, assuming the TwoDigitsInteger is needed, one would have to loop through the numbers, check them in a try-catch for checked NotSetExceptions as isValid() declares it, then pack the Exception as a RuntimeException one and throw it further, finally count the valid number. This is already way too complicated even when only enumerating the steps in natural language. To be able to keep the API fluid and use streams when performing checks that declare checked Exception, the next interface is declared. Java @FunctionalInterface public interface ExceptionPredicate<T, E extends Exception> { boolean test(T t) throws E; } It represents a predicate (a boolean-valued function) of one argument that might throw an Exception subclass. The method evaluates the predicate on the given argument and returns true if the input argument matches, or false otherwise. In addition, the following method is added to the ExceptionWrapper class, very similar to the apply() one. Java public static <T, E extends Exception> Predicate<T> test(ExceptionPredicate<T, E> predicate) { return t -> { try { return predicate.test(t); } catch (Exception e) { throw new RuntimeException(e); } }; } When called, it effectively applies the provided predicate. In case an Exception is thrown, it is wrapped into a RuntimeException one and thrown further. The initial code can now be rewritten as below and successfully compiled and executed. Java @Test void count() { long count = IntStream.range(0, 150) .mapToObj(TwoDigitsInteger::new) .filter(ExceptionWrapper.test(TwoDigitsInteger::isValid)) .count(); Assertions.assertEquals(90, count); } Takeaways Although simple and to-the-point, the presented solution comes in very handy, especially when dealing with functions that declare checked Exceptions and are further used in the code that we produce. For sure, other ready-to-use alternatives already exist, an example being the Lombok @SneakyThrows annotation. Personally, I have very rarely included the Lombok library in any of my projects and as Java introduced the records, this becomes even more unlikely to happen in the future. That being said, the structures described in this article are very helpful, lightweight, and easy to understand and use when needed. ExceptionWrapper, ExceptionFunction and ExceptionPredicate source code is part of the asentinel-orm open-source project. To use it, one may either declare the Maven dependency in their pom.xml file (version 1.72.2 is the latest at the moment of this writing) XML <dependency> <groupId>com.asentinel.common</groupId> <artifactId>asentinel-common</artifactId> <version>1.72.2</version> </dependency> or use it directly if considering there’s too much overhead to include the whole library. Resources [1] – asentinel-orm open-source ORM project is here [2] – the picture was taken at ‘Harry Potter Warner Bros. Studios’, near London
In a previous article, Working with Spreadsheets in Java: A Practical Overview, we walked through the common scenarios where Java applications need to interact with spreadsheets and the categories of tools available for the job. One of the factors mentioned there was support for modern Excel formulas — a topic that deserves more space than a single bullet point. Java applications interact with Excel more often than most teams plan for: file uploads from finance, calculation logic authored in a workbook, reporting exports back to business users. The files these users produce today are not the same as the files they produced five years ago. Excel 365 and Excel 2021 introduced a new formula model, and workbooks authored in those versions routinely use it. Depending on which library you use, those formulas may evaluate correctly, fail silently with stale cached values, or throw exceptions at recalculation time. This article goes deeper on that topic: what dynamic arrays and spill behavior are, what the new function set looks like, why supporting them is technically difficult, and what Java developers should look for when evaluating whether a library handles them correctly. Why You're Seeing Them in Real-World Workbooks Dynamic arrays spread rapidly because they eliminate many of the helper columns, copied formulas, and Ctrl+Shift+Enter array formulas that older Excel workbooks depended on. Workbooks become shorter, easier to audit, and easier to maintain. As organizations migrate to Microsoft 365, these newer formulas increasingly appear in spreadsheets exchanged with Java applications, even when the application itself hasn't changed. What Changed: One Formula, Many Values Before dynamic arrays, formulas that returned multiple values generally required a pre-sized array range and legacy array-formula syntax. Dynamic arrays changed this by allowing a single formula to return a variable-sized array and automatically spill into neighboring cells. For example: =UNIQUE(A1:A6) Entered in one cell, this returns the full list of distinct values from A1:A6. The result fills as many cells as there are distinct values. If the source data changes and the number of unique values changes, the spill range automatically grows or shrinks. This is not just a new function. It is a change in the evaluation model itself. The Mechanics of Spill A spilled formula produces a region of cells with specific roles: The anchor cell is the one cell that contains the formula. It "owns" the result.The spilled cells are the neighboring cells that display the additional values. They do not contain formulas of their own; they mirror slices of the anchor's result. You can reference the entire spilled range from another formula using the # operator. The # reference is not a fixed cell range such as A1:A6; it refers to whatever range the anchor currently spills into. If A1 contains =UNIQUE(...) and the result spills into A1:A6, then =COUNTA(A1#) counts the values in the entire spilled range. If the spill range grows or shrinks, the # reference adjusts automatically. #SPILL! errors. If the range a formula needs to spill into is blocked by an existing value, a merged region, or an Excel Table, the formula cannot spill, and the anchor cell shows #SPILL! instead of a result. Clearing the obstruction allows the formula to complete. Implicit intersection with @. Older Excel silently reduced arrays to single values in many contexts. Modern Excel returns the full array unless the formula uses the @ prefix. For example, =A1:A10 entered in a cell in modern Excel spills the values from A1:A10, while =@A1:A10 applies implicit intersection and returns the value corresponding to the formula's row. Files migrated from older Excel versions often contain automatically inserted @ prefixes to preserve their original behavior. The New Function Family The modern functions commonly associated with Excel's dynamic-array model can be grouped into three broad categories. It is worth understanding the grouping because the groups behave differently. Group 1: Language Features These are not really functions in the traditional sense. They add expression-level constructs to Excel's formula language. LET binds names to intermediate values inside a formula, so you can write =LET(total, SUM(B2:B100), tax, total*0.1, total+tax) instead of repeating SUM(B2:B100) three times.LAMBDA defines a reusable function inside a workbook. Combined with named ranges, LAMBDA effectively adds user-defined functions without VBA.ISOMITTED is used inside LAMBDA to detect whether an optional argument was supplied. These functions do not inherently produce a spilled array. LET returns the result of its calculation, which may itself be an array. Group 2: Dynamic Array Functions These are the functions people usually mean when they talk about "the new Excel functions." These functions are designed to return arrays, and when their results contain multiple values, Excel can spill those results into neighboring cells. UNIQUE returns distinct values from a range.SORT and SORTBY return sorted arrays.FILTER returns rows that match a condition.SEQUENCE generates a sequence of numbers.RANDARRAY generates an array of random numbers. Array-shaping functions form a subset of this group. They take arrays as input and return reshaped arrays: CHOOSECOLS, CHOOSEROWS, DROP, EXPAND, HSTACK, VSTACK, TAKE, TOCOL, TOROW, WRAPCOLS, WRAPROWS. TEXTSPLIT also fits here — it splits a string into an array. Other functions, including BYROW, BYCOL, MAP, REDUCE, and SCAN, build on the same dynamic array model. Group 3: Scalar Functions Added in the Same Era Dynamic arrays are primarily an evaluation model; modern functions are a collection of functions that take advantage of, or coexist with, that model. Group 3 functions were introduced as part of the broader set of modern Excel functions, but they are not themselves primarily array-producing functions. XLOOKUP and XMATCH are modern replacements for VLOOKUP and MATCH. They normally return a single value, though they can return an array when passed an array of lookup values.TEXTAFTER and TEXTBEFORE return substrings.VALUETOTEXT and ARRAYTOTEXT convert values to text (ARRAYTOTEXT takes an array as input but returns a single string). These are often lumped in with dynamic array functions because they arrived together, but their evaluation model is closer to VLOOKUP than to UNIQUE. Why This Is Hard for a Formula Engine Supporting these features is not just a matter of adding new function names to a list. The dynamic array model requires substantial changes to the evaluation engine itself. A traditional one-cell-at-a-time formula model is not sufficient to implement dynamic arrays. An engine must be able to represent a formula whose result has a variable shape and propagate that result across multiple cells. A modern engine has to handle four additional concerns: Array-shaped results. A formula's return value may be a 2D array whose dimensions depend on the input data. =UNIQUE(A1:A100) returns a different number of rows depending on how many unique values the range contains. The engine must determine the result shape at evaluation time, not at parse time. Spill range tracking. The engine must reserve the cells the formula spills into and prevent other content from occupying them. When something occupies a spill target, the anchor must return #SPILL! rather than overwrite the obstruction. The reserved region must also update when the shape of the result changes. Downstream references. Expressions like A1# refer to the entire spilled range. When the shape of the anchor formula changes, every downstream reference must be re-evaluated with the new dimensions. This makes the dependency graph more dynamic than in a one-value-per-cell model. Implicit intersection compatibility. Older Excel silently collapsed arrays to single values in many contexts. Modern Excel returns the whole array. When files authored in older Excel are opened in modern Excel, @ prefixes are inserted automatically to preserve original behavior. An engine that reads modern .xlsx files needs to honor the @ operator, or the imported formulas will produce different results. Adding these behaviors to an engine designed around the one-formula-one-value model is a substantial rewrite, not an incremental feature addition. This is part of why support across the Java ecosystem has been uneven. What Java Developers Should Check Support for these capabilities varies significantly across Java spreadsheet libraries. Some engines were originally designed around traditional one-cell-one-result evaluation and only implement subsets of the modern Excel model. Others have extended or redesigned their evaluators to support dynamic arrays. Rather than relying on feature lists, it is worth validating behavior against workbooks representative of your own application. If your application needs to evaluate modern Excel formulas, the following checks are worth running before committing to a library. Test with a file containing a spilled formula. Create a small .xlsx with =UNIQUE(A1:A100) or =SORT(A1:A100) in a cell. Load it in your candidate library and try to recalculate the anchor cell. A library that supports dynamic arrays will return the array; one that does not will typically throw an exception or return only the first value. Check for the # spill operator. In the same file, add another cell containing =COUNTA(A1#) where A1 is the anchor. This tests whether the library understands spilled range references, which is a separate capability from evaluating the anchor formula itself. Test the @ operator. Add =@A1:A10 in a cell and check whether the library correctly returns the value at the current row rather than the full array. Files migrated from older Excel routinely contain @ prefixes; a library that doesn't handle them will produce different results than Excel. Test with LET and LAMBDA. Write a formula like =LET(total, SUM(A1:A100), total * 1.1) and check both evaluation and .xlsx round-trip. Test LET and LAMBDA independently. Parsing, preserving, and evaluating these functions are separate capabilities, so a library that can read or write the formula text may not necessarily be able to evaluate it correctly. Test round-trip. Save the workbook, reopen it in Excel, and check that the formulas still produce correct results. Some engines strip modern constructs on save. Check what happens on failure. When a library encounters a function it does not implement, does it raise an exception, return an error value, or silently fall back to the cached value from the file? Silent fallback is the most dangerous behavior because it masks the problem during development and only fails in production when the data changes. Conclusion Excel's formula language has changed more in the last few years than in the two decades before it. Dynamic arrays, spill behavior, and the new function set are not experimental — they are standard in Excel 365 and Excel 2021, and they show up in workbooks that Java applications routinely have to process. For Java developers, the practical implication is that "Excel formula support" is no longer a single property that a library either has or doesn't have. There are several distinct capabilities involved, and libraries vary widely on each. As covered in the previous article, the Java spreadsheet landscape spans open source libraries such as Apache POI, commercial headless engines, and embedded spreadsheet components like Keikai. Whichever category fits your use case, the checks above are a reasonable way to verify that a candidate library handles modern Excel behavior against the workbooks your real users produce.
A pull request arrives. A few hundred lines of Java implementing the new discount rule: tiered thresholds, a regional exception, something about loyalty tiers that nobody can quite explain. It compiles. The tests pass. An LLM wrote it in about forty seconds. Now: who reviews it? The person who owns that rule is in commercial operations. She knows exactly which customers should get the discount and why the regional exception exists, and she cannot read Java. The person who can read Java has no idea whether the thresholds are right. He will check that the code looks reasonable, because that is the only thing he is equipped to check. So the review that happens is not the review that matters. That is the problem I keep coming back to, and it has nothing to do with how good the model is. This Is Not an Argument About Whether the Model Is Good Enough Most objections to generated code are about competence. The model hallucinates an API. It gets an edge case wrong. It writes something that works on the happy path and falls over in production. I find these arguments unconvincing because they expire. Models get better. Any position resting on today's error rate is a position with a shelf life, and people who staked one out three years ago have mostly had to retreat from it. The durable question is different. It is not how well the model writes. It is what the thing it writes is permitted to say. A model that never makes a mistake, handed Java, can still emit Runtime.getRuntime().exec(...). Not because it is malicious or confused — because that sentence is available in the language it was asked to write. Competence and authority are separate axes, and improving the first does nothing to the second. "Write it in Java" Is a Much Bigger Grant Than Anyone Means Consider what you actually authorize when you ask for a discount rule in Java. You authorize file system access. Network sockets. Reflection. Thread creation. Process execution. Every class on the classpath, including the ones that talk to your database, your payment provider, and your secrets manager. You authorize the loading of new code at runtime. Nobody intends to grant any of this. It arrives free with the language, the way a house key also opens the shed. The task needed perhaps six operations — look up an order, total it, check a customer's tier, apply a discount, log the decision, approve or refuse — and the language you handed over contains everything Java contains. That gap, between the authority the task requires and the authority the language confers, is the whole of it. It exists whether or not the model is trustworthy. It exists whether or not anyone acts on it. It is just very large, and it is not visible in the pull request. The Usual Guardrails Are Denial Lists The standard responses all share a shape. Tell the model in the prompt not to touch the file system. Review the generated code. Run static analysis and flag dangerous calls. Run it in a sandbox with a restricted security policy. Every one of these asks you to enumerate what must not happen, over a space of things that can happen which is effectively unbounded. You are writing a deny-list against a general-purpose language. You have to think of exec. Then of reflection reaching exec. Then of the dependency that shells out on your behalf. Then of the next one. We learned this lesson in security a long time ago and reached a settled answer: allow-lists beat deny-lists, because the allow-list is finite and you wrote it. Somehow, when the subject is generated code, we reach for the deny-list again. Shrink the Language, Not the Model The alternative is to stop constraining a powerful language and instead supply a small one. Give the model a vocabulary that contains exactly the operations the domain has — the six from earlier, say — and nothing else. Not a restricted Java. A different, much smaller language, whose entire vocabulary is a list your team wrote in advance, in Java, on purpose. Generated business logic then looks like this: Python PROGRAM ApproveOrder(orderId INTEGER, limit DECIMAL) RETURNS BOOLEAN DECLARE purchase Order DECLARE total DECIMAL purchase = LOAD_ORDER(orderId) total = ORDER_TOTAL(purchase) IF total > limit THEN REJECT purchase, "over limit" RETURN FALSE END IF APPROVE purchase RETURN TRUE END. LOAD_ORDER, ORDER_TOTAL, REJECT and APPROVE are not part of the language. They are Java classes somebody decided to expose. Order is a Java object the program can hold and pass and never look inside — there is no purchase.customer.account.balance here, only the operations the domain chose to have. Two things change, and the second matters more than the first. The obvious one: dangerous programs are no longer forbidden, they are inexpressible. If the model emits DELETE_ALL_ORDERS, nothing rejects it on policy grounds. The name means nothing. The program does not compile, for the same reason a typo does not compile. There is no deny-list because there is nothing to deny. The less obvious one: the commercial operations manager can read the program above. She can tell you whether the threshold is right, whether the rejection reason is the one the contract requires, whether an approval should have been logged. The review moves to the person who owns the rule. That is the review that was missing at the start of this article, and no amount of static analysis over generated Java produces it. A small language buys something else, quietly. With no data structures, one global scope, no null, and a compiler that refuses to run a program that reads a variable before it is set, entire families of subtle wrongness have nowhere to live. Not caught — absent. What It Costs, and What It Does Not Buy I would not trust this argument from someone who only listed the advantages, so here are the bills. You have to design the vocabulary. Somebody sits down and decides that the domain has ORDER_TOTAL and CUSTOMER_RISK and not forty other things. That is real work, done before the first generated line, by someone who understands the domain. And if nobody on your team can write that list, this approach will not help you. It will only show you that the list does not exist. That is worth finding out, but it is not a pleasant morning. Complex algorithms stay in Java. Business rules are algorithms too, and they belong in the small language; that is the point. But route optimization, a scoring model, anything with real computational substance belongs behind a function the small language calls. The signal is usually that you want to build up a data structure, or that you want a helper you can call from three places. Both mean you have wandered out of business logic and should walk back. The boundary bounds naming, not doing. This is the limit people miss, and overstating it is how the idea gets dismissed. A function you expose can do anything Java can do. RUN_SHELL_COMMAND is a perfectly registrable operation. The vocabulary is only as narrow as the operations you chose, and choosing them badly gets you exactly the exposure you were avoiding. There are no resource limits yet. A generated program can still loop forever. This one is a gap rather than a decision: the interpreter walks the program one statement at a time, so a step budget or a deadline is a small addition rather than a redesign, and it will go in when somebody needs it. Until then, untrusted input needs the same containment any untrusted workload needs. What you get is narrower than "safe" and more useful than it sounds: the set of things a generated program can name is finite, written down, and reviewable by a human before anything is generated at all. When I Would Still Write Java If the thing is genuinely computational, write Java. If it is a one-off that will be deleted next week, use whatever is nearest — Java, Python, a shell script — and let the model write it; do not build a vocabulary for something with a life expectancy of days. If the rules change so fast that the vocabulary would be obsolete before it settled, the overhead will not pay for itself. And if your business logic is already reviewed by people who can read it, understand it, and are accountable for it being right — you may not have the problem this solves. Plenty of teams do not. But if you are about to let a model write business rules in Java, ask the question I started with, because the answer is usually uncomfortable. Somebody is going to approve that pull request. Are they the person who knows whether the rule is correct? If not, the language is too big. I have been building a small language along these lines: BUBAS, an orchestration language for subject-matter experts, embedded in Java. The example above is real BUBAS. The idea does not require my implementation, though — the argument is about the size of the language you hand over, and you can shrink yours however you like.
Every Java developer who runs services on Kubernetes has watched this scene play out. Traffic spikes, the autoscaler adds a pod, and then everyone waits. The container is running in two seconds. The application is not ready for another twelve seconds. During those ten seconds, your existing pods absorb the extra load, latency climbs, and if things are bad enough, the autoscaler panics and adds even more pods that are also not ready. I spent years treating Spring Boot startup time as a fact of life, the way you treat weather. Then I found out the JVM has had a fix for a big chunk of it since Java 12; it works beautifully inside Docker, and almost nobody bakes it into their images. It is called Class Data Sharing, CDS for short, and this article shows you how to make your Docker build do the work Where Those Twelve Seconds Actually Go When a Spring Boot application starts, the JVM is not mostly running your code. It is loading classes. A plain REST service with Spring Web, Spring Data, and a driver or two loads somewhere between ten and twenty thousand classes before it serves its first request. For every single one of those classes, the JVM does the same ritual. Find the class file inside a jar, read the bytes, parse them, verify the bytecode is legal, and build the internal metadata structures it needs at runtime. Thousands of times. Every startup. In every pod. Here is the part that should bother you. Your container image never changes after you build it. The same jar, the same classes, the same parsing work, repeated identically in every pod that ever starts from that image. The JVM is solving the same puzzle again and again and throwing away the answer each time. CDS is the JVM saying: let me solve it once, write the answer to a file, and just memory map that file next time. What a CDS Archive Is A CDS archive is a file, usually ending in .jsa, that contains classes already parsed and verified, stored in the exact internal format the JVM uses in memory. On startup, the JVM maps this file straight into memory. No finding, no parsing, no verifying. The work was done ahead of time. You have been using CDS without knowing it. Modern JDKs ship with a default archive covering the core JDK classes, which is why java -version is fast. The step almost everyone skips is creating an archive for your application classes, all fifteen thousand of them. That is where the real win lives. The mechanism has one rule that matters for us. The archive must be created with the same JVM and the same classpath that will use it. That rule sounds annoying until you realize a Docker image is the one place in your entire infrastructure where JVM and classpath are frozen forever. Docker is not just compatible with CDS. It is the perfect home for it. The Training Run Creating the archive takes two steps. First you do a training run, where the JVM starts your application, watches which classes get loaded, and writes the list down. Then you exit, and the JVM turns that list into the archive. Since Java 13, this is pleasantly simple: Shell java -XX:ArchiveClassesAtExit=app.jsa -jar app.jar Run the app, let it come up, stop it, and app.jsa appears. From then on you start the app like this: Shell java -XX:SharedArchiveFile=app.jsa -jar app.jar There is an obvious question here. The training run wants to actually start the application, and inside docker build there is no database, no message broker, nothing to connect to. A Spring Boot app that cannot reach Postgres will crash during training. Spring Boot 3.3 solved this neatly. Setting one property makes the application run through its entire startup sequence, create all bean definitions, and then exit just before touching the outside world: Shell java -Dspring.context.exit=onRefresh -XX:ArchiveClassesAtExit=app.jsa -jar app.jar The application loads nearly everything it will ever load, writes the archive, and exits cleanly with no infrastructure needed. This is exactly what a Docker build stage can do. The Dockerfile Here is the complete picture: a multi-stage build where the image trains itself: Shell FROM eclipse-temurin:21-jdk-alpine AS build WORKDIR /build COPY . . RUN ./mvnw -B package -DskipTests # Explode the jar so the classpath is stable RUN java -Djarmode=tools -jar target/app.jar extract --destination /app FROM eclipse-temurin:21-jre-alpine AS runtime WORKDIR /app COPY --from=build /app /app # Training run: start the context, record classes, exit RUN java -Dspring.context.exit=onRefresh \ -XX:ArchiveClassesAtExit=/app/app.jsa \ -jar /app/app.jar ENV JAVA_TOOL_OPTIONS="-XX:SharedArchiveFile=/app/app.jsa" ENTRYPOINT ["java", "-jar", "/app/app.jar"] Two details in there deserve a closer look. The extract step unpacks the fat jar into a folder with the dependencies laid out as plain files. CDS is picky about the classpath being identical between training and real runs, and a fat jar with nested jars inside it makes that fragile. The exploded layout keeps the classpath boring and stable, which is exactly what CDS wants. On Spring Boot 3.2 and older, the same idea works through the layertools jarmode instead. The training run happens as a RUN instruction, which means it executes once at build time on your CI server. Every container that ever starts from this image inherits the archive for free. You did the class loading homework once, in the build, and ten thousand pod starts copy the answer. What You Get Numbers vary with how heavy your application is, but the pattern is consistent. A typical Spring Boot 3 web service that started in 10 to 12 seconds lands somewhere between 5 and 7. The JVM portion of startup shrinks dramatically, and as a bonus, the archive is memory-mapped and shared, so if you run several JVMs on one node, they share those pages and total memory drops too. You can verify the archive is actually being used, which I recommend, because CDS fails silently by design. If something mismatches, it just quietly falls back to normal class loading: Shell docker run --rm my-service -Xlog:class+load=info | head -5 Classes loaded from the archive say source: shared objects file. If you see jar paths instead, the archive is being ignored, and the log will usually tell you why. The usual culprit is a classpath that differs from training, even by one entry. One honest caveat. The training run exercises startup, not your traffic. Classes that only load when a specific endpoint gets hit for the first time are not in the archive, so those first requests still do normal loading. The archive covers the framework and wiring, which is most of the cost, but it is not a magic warm-up for everything. Why This Beats the Alternatives You Have Heard Of Whenever container startup time comes up, someone mentions GraalVM native images, and native images are impressive. Millisecond startup is real. But they come with a price list: long build times, a closed-world assumption that fights with reflection, some libraries that simply do not work, and a different runtime profile you have to learn to debug. CDS costs you five lines of Dockerfile. Your application is still a completely normal JVM application. Same debugging, same profilers, same libraries, same behavior, just faster out of the gate. For most teams, that trade-off is not even close. It also stacks with what is coming. Project Leyden's AOT cache in Java 24 and beyond is essentially this same idea grown up, caching not just parsed classes but resolved linkage and compiled code. The Dockerfile pattern you build today, a training run at build time producing a cache file shipped in the image, is exactly the shape Leyden uses. Learning it now means the future is a flag change. The Takeaway Your Docker image is immutable. Your JVM does expensive, perfectly repeatable work on every startup. Those two facts fit together like puzzle pieces, and a training run inside docker build is where they connect. One extra build step, and every pod your autoscaler ever creates comes up in half the time. The next time you watch a rollout crawl because pods take forever to go ready, remember that the answer was hiding inside the build all along.
Any input/output operation, be it accessing a file, handling an HTTP request, or a database connection, is based on 3 fundamental system concepts — file descriptors, kernel memory, and heap size. This article discusses how modern languages help developers handle behind-the-scenes file descriptor, kernel memory, and heap management. These three concepts are major bottlenecks for scaling. 1. File Descriptors A file descriptor is just a positive number that is used by the kernel to identify any open input/output stream or connection. It is defined by the kernel for a process. The following file descriptors are defined by default for a process: 0 – Standard Input (stdin)1 – Standard Output (stdout)2 – Standard Error (stderr) Any subsequent I/O operation gets the next available integer as file-descriptor. The file descriptor value can be adjusted by using the ulimit -n command in Linux. Each application, whether it is a web server written in Java Spring Boot, an API server written in Go using net/http and gorilla-mux, or a Python Flask app, is a single process. Each process has only 1024 file descriptors defined by default. That means each application can perform only 1024 I/O operations simultaneously. This seems like an amazing concept when we talk about scaling our application or API server. As many times as we come across this question — how can we scale our API server or web application to handle 100k or 1 million requests per second? This is where our modern languages play their role very beautifully behind the scenes to enable developers to develop the application to handle such scale. 2. Kernel Memory At a lower layer than file descriptors, when an incoming TCP connection hits the network card, the Linux kernel performs a 3 Way TCP handshake for that connection. The handshake lifecycle includes the states: SYN -> SYN-ACK -> ACK. The number of requests equal to the defined file descriptor value are processed immediately, assigned a file descriptor, and forwarded to the application for further processing. When FDs are exhausted, the Kernel maintains a queue for requests waiting for FDs to become available so your application can process them. The same thing happens when a request is processed, and the response is ready to be sent back to the client. This queue is maintained within RAM by read buffers(rmem) and write buffers(wmem). The size of buffers is defined in memory by the kernel and is dynamic, depending on network throughput, round-trip time, and memory pressure. The kernel network memory is non-paged, i.e cannot be swapped to disk. It’s a big bottleneck as it directly depends on physical memory. For example, if there are 100,000 open connections and each connection holds an average of 128KB of kernel memory, it comes to 12.8GB of physical RAM. This is clearly a kernel overhead, and it doesn’t show up in JVM heap metrics or Go runtime statistics. rmem and wmem buffers are governed by kernel parameters defined in /proc/sys/net/ipv4/ 3. Heap Size When TCP connections are assigned file descriptors and kernel memory is reserved, they enter user space, which is the memory managed by the application runtime — Java JVM, Node.js V8 Engine, Python interpreter, Go runtime, etc. Each connection stores objects in the heap within three categories: Connection metadata – Keep-alive timers, IP State, Socket Wrappers, etc.Cryptographic session context – handshake caches, cipher states, TLS/SSL keys, etc.Serialized payload buffers – response queues, JSON strings, ORM entity maps, etc. A connection that is encrypted via TLS takes a lot more space in the heap compared to a regular connection. For an encrypted connection, the application has to save symmetric keys, cipher contexts, session tickets, etc. onto the heap. A regular TCP socket object in the heap consumes 2KB to 5KB of space, whereas a TLS 1.3 socket object consumes 20KB to 100KB of heap space. If an API maintains 10,000 idle TLS connections, it will consume 200MB to 1GB of heap space. When an application runs, the runtime asks the kernel for memory space as the application creates objects. The application keeps creating objects, and the kernel keeps reserving memory for those objects; this is called the heap. The maximum heap size can be defined by different programming languages at runtime; for example, in Java, -Xmx4g reserves 4GB for the heap. The operating system promises to provide that much memory as heap space for the application, but it doesn’t reserve it all at once. As the application creates objects, the kernel continues to reserve memory. When objects are marked as done, the garbage collector removes them from the heap. When an incoming request hits our API server, the application uses heap space to convert raw bytes to the application-specific data structure. Once the application finishes processing the request and returns the response, those objects in the heap become unreachable or dead. When the garbage collector sweeps those objects to reclaim that memory, it doesn’t return the memory immediately; instead, the JVM or Go runtime keeps that freed memory in its internal pool. If a new HTTP request arrives within 1 millisecond, the runtime assigns the required memory from the free memory in the pool. Now imagine 10,000 new requests arriving at the same time, each with 2MB of raw bytes, and the runtime trying to allocate heap for the objects; the app instantaneously uses 20GB of memory. This is called GC thrashing, as the runtime rapidly creates required objects in the heap faster than the GC can clean them. The garbage collector is an application thread itself; when the heap gets 80%-90% full, the garbage collector panics and consumes 100% of CPU cores to scan millions of memory pointers to find dead objects. The runtime, like the JVM or Node.js garbage collector, may stop other code execution while it reorganizes the memory. So, how do runtimes like Go and the JVM handle GC thrashing? Go follows a simple strategy – avoid creating objects on the heap. The fastest GC collector is the one that has nothing to collect. The Go compiler compiles the application to see if variables outlive their functions. If a struct is used only inside a function, Go pushes the struct to the stack instead of the heap, and the stack pointer just drops when the function returns. The memory is reclaimed in 1 CPU cycle without even involving the garbage collector. If Go does have to clean the heap, its GC runs concurrently along with other goroutines and is broken into several micro pauses. Go provides sync.Pool to help developers to reuse heap memory while creating objects. For example to instead of creating millions of []bytes for JSON parsing for every new request, developers can use sync.Pool as follows: Go // Instead of creating a new buffer for every HTTP request: var bufferPool = sync.Pool{ New: func() any { return new(bytes.Buffer) }, } func handleRequest(w http.ResponseWriter, r *http.Request) { buf := bufferPool.Get().(*bytes.Buffer) // 1. Grab an existing buffer from pool buf.Reset() defer bufferPool.Put(buf) // 2. Put it back when done! // Parse JSON into 'buf' without allocating new heap memory } By recycling buffers via sync.Pool, high-concurrency APIs can handle 100,000 requests/sec with near-zero new heap allocations. Java takes a different approach. Because Java applications historically create millions of short-lived objects on the heap, the JVM relies on Generational Hypotheses and Generational Collectors (like G1GC, ZGC, and Shenandoah). G1GC can be used like java -XX:+UseG1GC while running Java applications. G1GC divides the Heap memory into physical regions: Young Generation (Eden & Survivor spaces) and Old Generation. It kind of sorts objects into different regions so that it doesn't have to scan the complete heap and can clean where most of the marked objects live. We can also mention -XX:MaxGCPauseMillis=200 to tell G1 to pause the application for no more than 200ms, but this is not guaranteed. Older JVM collectors like Parallel GC used to freeze the entire application to clear the heap when full, leading to multi-second latency spikes. Modern JVMs introduce ZGC (Z Garbage Collector) and Shenandoah. ZGC uses specialized CPU pointer references to track moved objects in real time. ZGC can clean, move, and compact terabytes of heap memory concurrently while your API requests are actively running. ZGC guarantees GC pause times under 1 millisecond, regardless of whether your heap is 500 MB or multi-terabytes. Conclusion Keep track of these three core concepts — file descriptors, kernel memory, and heap size to know when to scale. 1. File Descriptor Saturation Signals File descriptors represent the system's open handles. When an application hits its FD threshold, the operating system stops accepting connections. The following are example scenarios that indicate when to scale. Check Kernel-wide statistics from /proc/sys/fs/file-nr, per process fds - /proc/<pid>/fd, Prometheus exposes process_open_fds. If it consistently breaches the 80–85% threshold, it's time to scale. You have already tuned ulimit -n and LimitNOFILE up to standard safety thresholds (e.g., 65,536 or 104,857), but process FD counts continue climbing toward the max. Network interfaces show growing SYN-to-LISTEN socket counts and drops in netstat -s under the listen queue overflow metric. 2. Kernel Memory Pressure Signals Because TCP receive (rmem) and transmit (wmem) buffers are non-paged, they cannot overflow onto disk swap. When kernel network memory fills up, the OS drops packets. Below are the scenarios related to kernel memory breach. Check /proc/net/sockstat under TCP: inuse and matching /proc/sys/net/ipv4/tcp_mem thresholds. Netstat counters (netstat -s | grep -i retrans) show a sharp rise in TCP Retransmission rates (>1–2%). Latency spikes occur because the kernel is dynamically shrinking socket buffers down to tcp_rmem minimums (4 KB) to avoid running out of physical RAM, throttling TCP window sizes. 3. Heap Size & Garbage Collection (GC) Thrashing Signals When user-space heap allocations outpace the garbage collector's ability to sweep dead objects (like parsed JSON payloads or session states), application performance collapses. The runtime (JVM or Go) spends more than 15–20% of its total CPU time running GC sweeps (go_gc_cpu_fraction or JVM GC CPU utilization). In Go, metrics show the pacer triggering Mark Assist, stealing CPU time from worker goroutines to help clean up memory. You can check the runtime package /cpu/classes/gc/mark/assist:cpu-seconds metrics to see if GC is asking for more help from CPU. In Spring Boot, you can use Actuator and Micrometer to expose relevant endpoints to monitor the threshold values.
“...premature optimization is the root of all evil…” Donald Ervin Knuth Introduction "Premature optimization is the root of all evil." Most software engineers know this, attributed to Donald Knuth, author of The Art of Computer Programming and one of the most influential figures in computer science. Many have also picked up the practical conclusion that followed: "let's make it work first, fix performance later." After all, it's easier to add another EC2 instance than to find the root cause. But here is what Knuth actually wrote: "We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%." A little different, isn't it? The second sentence is almost never quoted — and that is convenient, because it turns a careful statement into a simple excuse. Sometimes for laziness. Sometimes because people assume that optimization means sacrificing readability: cryptic bit manipulation, obscure tricks, code that only the author understands at 2 am. I believe Knuth was indeed warning against that kind of optimization. But that assumption is wrong more often than people think. Good, clean code is frequently efficient code too — not by accident, but because choosing the right tool for the job tends to be both clearer and faster. The examples in this article are proof of that. Scope This article focuses on simple, cheap, and foolproof tips that can be applied universally — regardless of your architecture, framework, or domain. In my experience, they carry virtually no risk of making things worse. Architecture, design, networking, database connectivity, threading — these are deliberately out of scope. Not because they are unimportant, but because they are context-dependent. The right answer depends on your specific system, and each of these topics deserves its own article. Examples String Operations We are all familiar with built-in JDK string utilities like: equals(), startsWith(), endsWith(), contains(): Java s1.equals(s2); s1.startsWith(s2); s1.endsWith(s2); s1.contains(s2); Unfortunately, JDK provides only one function for case-insensitive comparison: Java s1.equalsIgnoreCase(s2) There are no functions for case-insensitive startsWith(), endsWith(), contains(). So, often we combine toLowerCase() or toUppserCase() with startsWith(), endsWith(), contains(): Java s1.toLowerCase().startsWith(s2.toLowerCase()); s1.toLowerCase().endsWith(s2.toLowerCase()); s1.toLowerCase().contains(s2.toLowerCase()); A little verbose and null-prone, but just fine if not on the critical path. However, this technique might cause some performance problems. Do not forget that String is an immutable class, so instead of just a char-to-char comparison between two strings, we create two additional strings that then must be garbage-collected. Considering that String is a wrapper over a char array, the memory allocation may become expensive. The solution is to use case-insensitive utilities provided by different libraries, e.g., Apache Lang3: Java startsWithIgnoreCase(s1, s2); endsWithIgnoreCase(s1, s2); containsIgnoreCase(s1, s2); Or, starting from version 3.18.0: Java Strings.CI.startsWith(s1, s2); Strings.CS.startsWith(s1, s2); Where CI exposes case-insensitive and CS — case-sensitive utilities. Many people like regular expressions and use java.util.Pattern class sometimes, not where it is really necessary. For example: Java Pattern.compile("^prefix.+suffix$").matcher(s).find() Instead of: Java s.startsWith("prefix") && s.endsWith("suffix") Or even: Java Pattern.compile("^prefix").matcher(s).find() instead of s.startsWith("prefix") Pattern.compile("suffix$").matcher(s).find() instead of s.endsWith("suffix") Pattern matching is significantly slower than trivial substring matching. The following table shows evaluation time for 1 million operations: Operation * 1 million times Time, ms s.equals("hello") 7 s.startsWith("hello") 6 s.endsWith("hello") 11 s.contains("hello") 24 s.toUpperCase().startsWith("HELLO") 65 s.equalsIgnoreCase("hello") 5 Pattern.compile("hello").matcher(s).find() 238 pattern.matcher(s).find() 31 What can we see from this table? Performance of equals() and startsWith() is similarendsWith() is 2 times more expensivecontains() is 4 times more expensive than equalsChanging case followed by startsWith() is 10 times (!) more expensiveCase-insensitive comparison functions do not have any performance penaltiesSearching for a substring using a precompiled pattern is about 20% more expensive than using a plain contains() method. Compiling the pattern and using it is almost 10 times more expensive than the plain contains() method. So next time you reach for Pattern.compile(), it is worth pausing for a second: is regex actually needed here, or is a plain string method both simpler and faster? If you really need a pattern, at least compile it in advance — better yet, declare it as a private static final class member. Collections Let’s assume that we want to know whether a given list contains the specific element: Java list.contains("red"); In fact, this call invokes code like this: Java int n = list.size(); for (int i = 0; i < n; i++) { if ("red".equals(list.get(i))) { return true; } } Starting from Java 8, we have a streaming API that just hides from us the same gory details: Java list.stream().anyMatch("red"::equals); This is perfectly fine when the list is short, changes frequently, or is searched only occasionally. But if the list is large, stable, and searched repeatedly, a HashSet is the right tool — offering average O(1) lookup instead of O(n). If you cannot change the original data structure, converting it once at initialization time and searching the Set from that point forward is almost always worth it. If both the guaranteed element order and the fast lookup are needed, we can either hold duplicated data structures — a list for ordering and a set for search or just use LinkedHashSet, which solves both problems. Another common case is case-insensitive search. We already saw above that the combination of toLowerCase() or toUpperCase() with comparison significantly reduces the performance. This can be solved by using TreeSet with custom comparator, e.g. String.CASE_INSENSITIVE_ORDER: Java Set<String> set = new TreeSet<>(String.CASE_INSENSITIVE_ORDER); This gives you a sorted, case-insensitive set with no extra allocations - and the same approach works for TreeMap when your data is key-value pairs. Enum Lookups Everyone knows that an enum entry can be found by its name using a built-in method valueOf(s). However, what to do if the given string is lowercase while enum entries following the naming convention are called using capital letters? Some people use a combination of toUpperCase() and valueOf() that work just fine but have the penalty we discussed above. However, very often people prefer to create a special field representing a “custom” name, so the simple enum like: Java enum Color { RED, GREEN, BLUE } Turns into: Java enum Color { RED("red"), GREEN("green"), BLUE("blue"), … } Let’s mention that this design has at least two disadvantages: Duplicate data: The custom name is the same as a built-in but in a different case, which can be solved much more easily. This allows using really custom names that, according to my experience, in most cases are not needed and just create so-called “edge cases” that, in turn, in most cases are just a signal of bad design and might cause a lot of “stupid” bugs. However, let’s continue. How do people often use this custom name? Java public static Color ofColor(String color) { return Arrays.stream(values()) .filter(c -> c.color.equals(color)) .findFirst() .orElseThrow(() -> new IllegalArgumentException("No enum constant %s.%s".formatted(Color.class.getName(), color))); } The implementation looks pretty nice, but this approach means that each call of ofColor() iterates over the list. Yes, in most cases enums are not huge, so the list is short, but anyway, why do this if we can just create a map from the custom name to the enum entry once during initialization and then use it with O(1) complexity? The following example solves both problems at once: it uses a case-insensitive map where the key is the standard name() of the enum entry during initialization: Java private static final Map<String, Color> colors = Arrays.stream(values()).collect(toMap(Enum::name, e -> e, (existing, replacement) -> replacement, () -> new TreeMap<>(CASE_INSENSITIVE_ORDER))); So, now the method ofColor() becomes trivial: Java public static Color ofColor(String color) { return Optional.ofNullable(colors.get(color)) .orElseThrow(() -> new IllegalArgumentException("No enum constant for " + color)); } One can argue that a map-based implementation is not always possible because sometimes the lookup criteria are too complex to be reduced to a simple key. Although I agree in general, I can say in turn that in many (if not in most) cases this is still possible. So far, the lookup key was a simple string. But what if the search criteria is a range rather than an exact value? Consider a more physically accurate model of colors as ranges of electromagnetic waves. Java public enum Color { BLUE(450, 495), GREEN(495, 570), RED(620, 750); …} How to implement the method ofWaveLength(int waveLength)? The straight-forward way is to iterate over the values of the enum and compare the given wave length with the range for each entry, i.e. implement O(n) search. But we can do better using NavigableMap, which is designed exactly for this kind of range query: Java private static final NavigableMap<Integer, Color> wavelengthMap = Arrays.stream(values()) .collect(Collectors.toMap( color -> color.minNm, color -> color, (existing, replacement) -> existing, TreeMap::new )); Unfortunately, the search method is not as trivial as in the previous example, but still very simple and fast: Java public static Color ofWaveLength(int nm) { return Optional.ofNullable(wavelengthMap.floorEntry(nm)) .map(Entry::getValue) .filter(value -> nm <= value.maxNm) .orElseThrow(() -> new IllegalArgumentException("No enum constant for wavelength: " + nm + " nm")); } Now, let’s compare the performance. Operation * 1 million times Time, ms valueOf(s) 34 valueOf(toUpperCase(s)) 78 Iteration with equals() 40 Color.ofColor() iteration 166 Color.ofColor() map 20 Color.ofWaveLength() map 32 The table shows that: As expected, toUpperCase() reduces performance twiceIteration with call of equals is a little bit more expensive than valueOf() although the enum has only three members and will grow linearly as the enum grows. The more members enum has, the more time iteration takes. Map-based implementation is even faster than one based on the built-in valueOf(). Stream-based iteration (ofColor() iteration) is surprisingly slow. Stream setup overhead (boxing, lambda dispatch, spliterator initialization) is non-trivial for tiny collections Pre-Intitialization The principle here is: do not do something several times if you can do it once. The most trivial example is string or numeric constants: Java private static final String FILE_NAME = "config.json"; private static final int MAX_VALUE = 10_000; However, the same principle applies to heavier objects — and that is where it really matters. Let’s take a look at logging. Most people are used to writing the following “magic” line at the beginning of each class (unless we use Lombok’s @Slf4j annotation): Java private static final Logger logger = LoggerFactory.getLogger(MyClass.class); Are all these modifiers (private static final) really needed? Some people try to save typing time: Java private final Logger logger = LoggerFactory.getLogger(MyClass.class); Moreover, if the logger is not static, we can do even more: Java private final Logger logger = LoggerFactory.getLogger(getClass()); This line looks better because it is error-proof: the class here is not hard-coded, so this line can be copied as-is from one class to another or inherited from the base class. So, what’s the problem? The problem is that retrieving the correct logger is potentially expensive due to synchronized registry lookups. Doing this on every instantiation adds up. A friend of mine told me that once in the company where he worked, this change in some critical path improved performance so much that they managed to reduce the AWS cluster by about one hundred large EC2 machines. The same rule applies to pattern compilation. As the benchmark table showed, compiling a pattern on every method call is nearly ten times slower than reusing a precompiled one. The result of Pattern.compile() should always be stored in a static final field. The only exception is the case when the regular expression is generated dynamically, but we should do our best to avoid such a design. Very often we have to format or parse dates. Traditionally I used SimpleDateFormat. What can be more obvious than this: Java private static final String FORMAT = "yyyy-MM-dd HH:mm:ss"; private static final DateFormat format = new SimpleDateFormat(FORMAT); Frankly speaking, I did this many times following the principle I stated above: there is no reason to create the instance every time we need it if we can create it only once. The problem is that SimpleDateFormat is not thread-safe, so sharing the same instance among different threads can cause the problem. Even worse: we can live with this bug for years without knowing about it, since it only happens under high load and in some cases can just produce slightly wrong results that can be lost in an ocean of valid data. So, should we create instances of SimpleDateFormat every time we need it and cause CPU and GC to work hard? Fortunately, starting from Java 8, we can use DateTimeFormatter instead: Java private static final DateTimeFormatter formatter = DateTimeFormatter.ofPattern(DATE_FORMAT); This class is thread-safe, so we can share its instance among different threads and get consistent results. Conclusion We started with a quote that is almost always cited incomplete. Knuth never said ignore performance — he said don't sacrifice clarity for speculative gains, while reminding us not to pass up opportunities in that critical 3%. The examples in this article live in that 3%. None of the performance issues described here should ever appear in production code. They are not hard to avoid — they require no profiler, no benchmarking framework, no architectural discussion. Just the habit of reaching for the right tool. And that habit pays off. Choosing equalsIgnoreCase() over toLowerCase().equals() is cleaner and faster. A static final logger is simpler and cheaper. A pre-built enum map is more readable and O(1). Good code and efficient code are not in conflict here — they are the same code. The only thing required is the habit of pausing for a second and asking: am I doing this n times when once would do? All code examples from this article are available on Gist.
Technical Manager,
Baptist Health South Florida
Software Engineer Expert,
UPMC
Award-winning Software Engineer and Architect,
OS Expert
Senior Principal Developer Advocate,
IBM