A database is a collection of structured data that is stored in a computer system, and it can be hosted on-premises or in the cloud. As databases are designed to enable easy access to data, our resources are compiled here for smooth browsing of everything you need to know from database management systems to database languages.
Wasm Inside Neo4j: Building the Example That Didn't Exist
Why Time Series Databases Matter for Modern Enterprise Applications
One of my favorite things about F1 racing is the data behind it. F1 cars are the most complex and advanced in any racing series. They collect huge amounts of telemetry data. The tracks also gather data during events, and race engineers analyze it week after week. They study everything from weather and tire temperatures to corner exit speeds. Data drives the sport forward in a major way. While learning about graph databases and Neo4j, I realized it was the perfect tool for answering a question I was curious about. We've all seen or heard of "Six Degrees of Kevin Bacon," where nearly any actor can be traced back to the Footloose star. I wondered: could this work for F1 drivers? By comparison, it's a much smaller dataset than famous actors. Photos via Wikimedia Commons, licensed under CC BY‑SA 4.0. Is Max Verstappen connected to Juan Manuel Fangio? Could a driver who retired in 1958, decades before Max was born, connect to him through a chain of teammates? And if so, how many links does it take? This post is how I answered that, and it doubles as a gentle introduction to Neo4j and graph databases. By the end, you'll have built a real graph of every F1 driver since 1950 on your own machine, and you'll run a query that answers my Verstappen-to-Fangio question in a single line. No prior graph experience needed. Let's get into it. Why This Is a Graph Problem Before we start: My question isn't really about drivers. It's about the connections between drivers. If all I wanted was a list of drivers, or each driver's win count, or how many races happened at Monza, a plain old relational table handles that beautifully. Even a spreadsheet can do it. Spreadsheets are wonderful at facts about things. Where they start to sweat is questions about relationships and chains of relationships. Specifically, long relationship chains. Think about what "is Verstappen connected to Fangio?" actually requires. You don't know in advance whether the answer is three hops or nine. So in SQL you'd be writing a recursive common table expression that joins a results table to itself, over and over, to a depth you can't predict, while trying not to drown in duplicate paths. I tried to do this very thing and locked up the application trying. It's possible to do queries like this, but they rarely run fast, if they run at all. Relational databases weren't designed for things like this. A graph database flips the whole thing around. Instead of storing drivers in one table and hoping to reconstruct their connections later with joins, it stores the connections themselves as useful entities. That's the one idea underneath everything else in this post: In a graph database, the relationships are first class data. They're not something you compute at query time. They're something you store, traverse, and count directly. That single design choice is what turns this tough question into a one-liner. Let me show you the model before we build it. The Property Graph Model, in Four Pieces Neo4j uses the Labeled Property Graph model. It sounds fancy; it's just four building blocks. I'll introduce each one using our F1 data. Nodes are the things in your domain. The entities. For us, that's drivers and teams. In a diagram, you draw them as circles. Ayrton Senna is a node. McLaren is a node. Labels are the type of a node, written with a colon: :Driver, :Constructor. (Constructor is just F1's official word for "team".) Labels are how Neo4j knows a Senna node is a driver and a McLaren node is a team. By convention, they're written in PascalCase. Relationships are the connections between nodes, and this is where graphs shine. Every relationship has a type in SCREAMING_SNAKE_CASE, a direction, and a start and end node. Senna DROVE_FOR McLaren is a relationship. Crucially, that connection is stored in the database. Neo4j keeps a pointer from one node to the next. This is why hopping across relationships stays fast even when your graph gets huge. The cost of following one relationship remains small, whether your database has a thousand nodes or a billion. Properties are key-value pairs you can hang on either a node or a relationship. A :Driver node has forename: 'Ayrton', surname: 'Senna', nationality: 'Brazilian'. Here's something that might be surprising if you come from a table world: a relationship can carry properties too. Our DROVE_FOR relationship will carry season: 1988, because which season someone drove for a team is a fact about the connection, not about the driver or the team on their own. That last point is worth thinking about, because it was an "aha" moment for relational-to-graph thinking. Senna drove for McLaren, but when he did, it doesn't belong to Senna and doesn't belong to McLaren. It belongs to the link between them. Put it on the relationship, and a whole category of modeling headaches evaporates. Here's our entire starting model: Two circles, one labeled arrow between them. If you want to make that concrete right now, open arrows.neo4jlabs.com (a free browser diagramming tool) and draw it. Click to make a node, give it a label and some properties, drag from its edge to a second node to create the relationship. It's helpful to sketch your schema there before writing any code. Design for the Question You Want to Ask Before building anything, I did the single most useful thing you can do when modeling a graph: I wrote down the question I actually wanted to answer first, and let it drive every decision afterward. My question: "How are two drivers connected through shared teammates?" This tells me exactly what my graph needs. It needs drivers. It needs some notion of two drivers being teammates. Everything else is optional scaffolding. But notice the dataset doesn't hand me "teammate" directly. It gives me who drove for which team in which season. Two drivers are teammates when they drove for the same team in the same season. So my plan has a nice shape to it: Load Driver and Constructor nodes.Connect them with DROVE_FOR relationships (one per driver, per team, per season).Derive a brand-new TEAMMATE_OF relationship between any two drivers who share a team and season.Walk the TEAMMATE_OF web to answer my question. In step three, we are creating relationships that weren't in the raw data, by reasoning about the graph you already have. This is one of my favorite things about working in Neo4j, and you'll see why shortly. Let's build. What You'll Need Neo4j Desktop – free, from neo4j.com/download. I'm on the current version (Desktop 2.x) on a Mac; Windows and Linux are the same journey.The dataset – the "Formula 1 World Championship (1950–2024)" dataset by Rohan Rao on Kaggle. Free Kaggle account, one download, clean CSVs.An afternoon. Realistically, a couple of hours, most of it spent going "oh that's cool" at the results.No prior Cypher. I'll explain every query as we go. Cypher is Neo4j's query language, and it's genuinely readable. If you already know SQL, you'll be nodding along within minutes. A quick note on the data: For about a decade, the go-to source for F1 data was the Ergast API. It shut down at the end of 2024. The Kaggle dataset we're using preserves Ergast's exact structure as downloadable CSVs, and if you later want live current-season data, the community-run Jolpica-F1 API (api.jolpi.ca/ergast/f1/) serves the same schema. Build against the CSVs today; top up from Jolpica whenever you like. Nothing in this tutorial changes. Once Neo4j Desktop is installed, under local instances, click create instance to create a new local instance. Give it a name and a password you'll remember, and start it. Then open the Query tool (in Desktop 2.x this is where you run Cypher — it's the modern replacement for what older tutorials call "Neo4j Browser"). That's our workbench. We need four files from the Kaggle download: drivers.csvconstructors.csvraces.csvresults.csv We can stage these files for import by placing them in our imports folder. Select your local instance, and look for the ... button. Select Open then Instance folder: This is the folder for your Neo4j instance. Next, select the import folder. This is where you want to place the files. macOS: Plain Text /Users/[your username]/neo4j-community-2026.x/import Windows: Plain Text C:\Neo4j\import Linux: Plain Text /var/lib/neo4j/import Now that the files are in the import folder, we can access them with Cypher later. Step 1: Constraints First Before loading a single row, I created uniqueness constraints. A constraint guarantees you'll never accidentally create two copies of the same driver. Neo4j automatically builds an index behind each one, which makes all the lookups during import dramatically faster. In Neo4j Desktop, you can run queries by selecting Query from the left-hand panel and entering your queries in the window in the upper right. Here's the query to create the constraints: Cypher CREATE CONSTRAINT driver_id IF NOT EXISTS FOR (d:Driver) REQUIRE d.driverId IS UNIQUE; CREATE CONSTRAINT constructor_id IF NOT EXISTS FOR (c:Constructor) REQUIRE c.constructorId IS UNIQUE; CREATE CONSTRAINT race_id IF NOT EXISTS FOR (r:Race) REQUIRE r.raceId IS UNIQUE; Run SHOW CONSTRAINTS to confirm all three landed. That's your data-integrity seatbelt fastened. Step 2: Load the Nodes Now we bring in the entities. LOAD CSV reads a file row by row; MERGE is Cypher's "create this if it doesn't already exist, otherwise match the existing one" command. Drivers: Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/drivers.csv' AS row MERGE (d:Driver {driverId: toInteger(row.driverId)}) SET d.forename = row.forename, d.surname = row.surname, d.fullName = row.forename + ' ' + row.surname, d.nationality = row.nationality, d.dob = CASE WHEN row.dob <> '\\N' THEN date(row.dob) END; That CASE WHEN row.dob <> '\\N' is guarding against a quirk you'll hit constantly with this dataset: missing values are stored as the literal text \N. If you don't filter them out, you'll end up with drivers whose birthday is the string "backslash-N", which is exactly as useful as it sounds. Consider that your first real-world data-cleaning lesson, delivered by Formula 1. Constructors: Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/constructors.csv' AS row MERGE (c:Constructor {constructorId: toInteger(row.constructorId)}) SET c.name = row.name, c.nationality = row.nationality; Races (we mostly need these to know which season a result belongs to, but full Race nodes cost nothing and set you up for future projects): Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/races.csv' AS row MERGE (r:Race {raceId: toInteger(row.raceId)}) SET r.year = toInteger(row.year), r.round = toInteger(row.round), r.name = row.name, r.date = date(row.date); Quick check: you should see something in the ballpark of 861 drivers, 212 constructors, and 1,100-plus races: Cypher MATCH (n) RETURN labels(n)[0] AS label, count(*) ORDER BY label; If those numbers look right, you've just loaded three-quarters of a century of motorsport into a graph. We haven't done anything clever yet, but we're about to. Step 3: Connect Drivers to Teams The results.csv file has one row per driver, per race — around 26,000 rows. I don't want 26,000 relationships cluttering my graph. I want one clean fact per driver, per team, per season: this person drove for this team that year. So I aggregate as I load. Cypher :auto LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/results.csv' AS row CALL (row) { MATCH (d:Driver {driverId: toInteger(row.driverId)}) MATCH (c:Constructor {constructorId: toInteger(row.constructorId)}) MATCH (r:Race {raceId: toInteger(row.raceId)}) MERGE (d)-[s:DROVE_FOR {season: r.year, constructorId: c.constructorId}]->(c) ON CREATE SET s.entries = 1, s.wins = CASE WHEN row.positionOrder = '1' THEN 1 ELSE 0 END ON MATCH SET s.entries = s.entries + 1, s.wins = s.wins + CASE WHEN row.positionOrder = '1' THEN 1 ELSE 0 END } IN TRANSACTIONS OF 2000 ROWS; There's a lot of learning packed into that one statement, so let's unpack it: MERGE (d)-[:DROVE_FOR {season, constructorId}]->(c) is the key move. The first time we see Senna-at-McLaren-in-1988, this creates the relationship. Every subsequent race that season just finds the existing one and bumps its counters. Twenty-six thousand rows collapse into roughly 3,600 clean driver-season-team facts. This is the graph-modeling principle in action: store relationships at the granularity you plan to query.ON CREATE / ON MATCH let you do one thing when the relationship is brand new and a different thing when it already exists — here, initialize the counters versus increment them.:auto and IN TRANSACTIONS OF 2000 ROWS tell Neo4j to commit the import in batches rather than one giant transaction, which keeps memory happy on a big file. Version note: CALL (row) { … } is the modern syntax (Neo4j 5.23 and up) for passing a variable into a subquery. If your Neo4j is older, you'll get a syntax error on that line — use the legacy form CALL { WITH row … } instead. And don't copy an abbreviated snippet with ... in the middle into the Query editor; Cypher will try to parse the dots. Use the full block above. Let's make sure it worked by looking at a career I know: Cypher MATCH (d:Driver {surname:'Senna', forename:'Ayrton'})-[s:DROVE_FOR]->(c:Constructor) RETURN c.name AS team, s.season AS season, s.wins AS wins ORDER BY season; Toleman in '84, Lotus '85 to '87, McLaren '88 to '93, Williams in '94. If that's what you see, your graph is alive and correct. Step 4: Derive the Teammate Network Everything so far was set up. This is the payoff of graph thinking. Nowhere in the data does it say "Senna and Prost were teammates." But we can derive that: two drivers are teammates if they each have a DROVE_FOR relationship to the same constructor with the same season. And in Cypher, describing that pattern is close to describing it in English: Cypher MATCH (d1:Driver)-[r1:DROVE_FOR]->(c:Constructor)<-[r2:DROVE_FOR]-(d2:Driver) WHERE r1.season = r2.season AND d1.driverId < d2.driverId MERGE (d1)-[t:TEAMMATE_OF {season: r1.season, team: c.name}]->(d2); Read that MATCH line like a picture: driver one points to a constructor, and driver two points to the same constructor from the other side. The WHERE says "same season." And then we MERGE a shiny new TEAMMATE_OF relationship between them. We just created around 10,000 relationships that didn't exist in the source data, purely by reasoning about the shape of the graph. Two small things worth understanding: d1.driverId < d2.driverId stops us creating each pairing twice (Senna→Prost and Prost→Senna). By only linking the lower ID to the higher one, each pair gets a single relationship. When we query it, we'll just ignore direction — because "teammate" goes both ways, and Cypher happily traverses a relationship in either direction when you leave the arrowhead off.A deliberately imperfect definition, and why I'm keeping it. "Same team, same season" isn't exactly "raced side by side." Midseason driver swaps mean, for example, that Senna and David Coulthard both count as 1994 Williams drivers. Coulthard was Senna's replacement, and they never actually raced as teammates. I could tighten this up by deriving teammate links per-race instead of per-season. I'm keeping the looser version on purpose, for two reasons. First, it makes the network richer and more connected across eras, which is the whole point. Second, every graph model is an argument about what a relationship means. There's no universally correct answer; there's only the definition that serves your question. Naming that trade-off out loud is important. Step 5: Answer the Question Here it is. The reason I built the whole thing. Two drivers separated by half a century, and one line of Cypher to connect them: Cypher MATCH (max:Driver {surname:'Verstappen', forename:'Max'}), (fangio:Driver {surname:'Fangio'}) MATCH p = shortestPath((max)-[:TEAMMATE_OF*]-(fangio)) RETURN p; That * after TEAMMATE_OF is the star of the show. It means "follow this relationship any number of times" — a variable-length path. shortestPath then finds the tightest chain of teammate links between the two drivers. This is the exact query that would've been a page of recursive SQL. In Cypher, it fits on a napkin. When you run it, the Query tool draws the answer as a chain of driver nodes, each link labeled with the team and season that connects them. This is how close Max really is to Fangio. Once that lands, you'll want to push further. Here are the three queries I couldn't stop running. Every driver's "Senna number" — like the Bacon number, but for F1. How many teammate-hops is each driver from Ayrton Senna? Cypher MATCH (senna:Driver {surname:'Senna', forename:'Ayrton'}) MATCH (d:Driver) WHERE d <> senna MATCH p = shortestPath((senna)-[:TEAMMATE_OF*..25]-(d)) RETURN length(p) AS sennaNumber, count(d) AS drivers ORDER BY sennaNumber; The shape of those results is the real insight: almost the entire history of the sport sits within a handful of hops of Senna. That's a "small-world network," demonstrated with race cars. The most connected drivers in history — the human bridges holding the whole web together: Cypher MATCH (d:Driver)-[:TEAMMATE_OF]-(other:Driver) RETURN d.fullName AS driver, count(DISTINCT other) AS teammates ORDER BY teammates DESC LIMIT 10; Watch who tops this list: long-career journeymen and team-hoppers, not necessarily the champions. Connectedness rewards longevity and movement, not podiums. I find that interesting. The people stitching F1's social fabric together are often not the ones holding the trophies. Is it really all one network? Are there isolated islands of drivers? Cypher MATCH (d:Driver) WHERE NOT (d)-[:TEAMMATE_OF]-() RETURN count(d) AS unconnectedDrivers; A small handful of true loners from F1's chaotic early days, and then one enormous connected web containing basically everyone else. Seventy-five years, one family. What You Just Learned (It Wasn't Really About F1) The property graph model – nodes, labels, relationships, and properties, including the quietly powerful idea that a relationship can carry properties of its own.Designing for questions, not entities – writing the question first and letting it shape the model.LOAD CSV, MERGE, and constraints – the everyday mechanics of getting real data into Neo4j cleanly.Deriving new relationships – creating structure that wasn't in your source data by reasoning about the graph you already have.Variable-length paths and shortestPath – the thing graphs do effortlessly and relational databases do through gritted teeth. And here's the part that matters beyond motorsport: swap the dataset and every one of these skills transfers directly. The teammate network is structurally identical to a fraud ring, a supply chain, a social graph, an org chart, or the knowledge graph behind an AI application. "Who is connected to whom, and how?" is one of the most valuable questions in software, and you now know how to ask it. Download the code here. Where to Go Next If this clicked for you, the best next move is to get the fundamentals properly, in order. That's exactly what GraphAcademy is for. It's Neo4j's free, hands-on learning platform. For future articles, I'm thinking: turn this same graph into a fair fight, deriving "who-beat-whom" links between teammates and running an algorithm called PageRank to settle the greatest-of-all-time argument without ever touching the points table.
Most AWS migration projects don't fail because of technical complexity. They fail because teams treat migration as a single activity rather than a set of distinct strategies applied to different workloads. AWS defines seven migration strategies — the 7Rs — that determine how each application moves to the cloud. The decision of which strategy applies to which workload has more impact on project cost, timeline, and outcome than any architectural choice you'll make after. Yet in practice, most teams default to "lift-and-shift everything" without evaluating whether that's appropriate. This article presents a practitioner's framework for classifying workloads into the 7Rs, based on delivering 50+ AWS migrations across fintech, SaaS, healthcare, and e-commerce. The 7R Strategies Retire Not every workload deserves migration. During discovery, you will invariably find applications that are redundant, unmaintained, or replaceable. In a typical enterprise portfolio of 20–40 applications, 10–20% qualify for retirement. Decision criteria: No active users, duplicate functionality already covered by another system, or maintenance cost exceeds business value. Common mistake: Teams skip this step because retiring applications requires stakeholder conversations. The result is migrating dead applications that consume compute budget indefinitely. Retain Some workloads shouldn't migrate in this wave. Applications with deep hardware dependencies, pending end-of-life within 12 months, or complex regulatory constraints that require legal review before cloud deployment are candidates for retention. Decision criteria: High migration complexity combined with low business urgency, or external constraints that prevent cloud deployment within the project timeline. Retain is not "never migrate." It's "not now." Document these workloads with a future migration path and trigger conditions. Rehost (Lift-and-Shift) Moving applications to EC2 or containers without code modifications. AWS Application Migration Service (MGN) automates this by continuously replicating servers and orchestrating cutover with minutes of downtime. Decision criteria: Application has a short remaining lifespan (1-2 years), speed of migration matters more than optimization, or the application is a black box with no available source code. Timeline: Days to weeks per workload. Trade-off: You gain cloud elasticity and pay-as-you-go pricing immediately, but you inherit all existing architectural inefficiencies. A poorly designed monolith on-premises becomes a poorly designed monolith on EC2. Relocate Hypervisor-level migration, primarily for VMware workloads moving to VMware Cloud on AWS. The OS, application, and configuration remain untouched. Decision criteria: Large VMware estate, tight data center exit deadline, and applications that cannot tolerate any configuration change. Replatform Migration with targeted adaptations to managed services. The application architecture stays intact, but you replace self-managed infrastructure components with AWS equivalents: Self-ManagedAWS ManagedOperational BenefitSelf-hosted PostgreSQLRDS for PostgreSQLAutomated backups, patching, failoverCron jobs on EC2EventBridge + LambdaNo server to maintain, pay-per-invocationSelf-managed RedisElastiCacheAutomatic failover, scalingNginx load balancerApplication Load BalancerManaged TLS termination, WAF integrationSelf-hosted ElasticsearchOpenSearch ServiceManaged cluster scaling, snapshots Decision criteria: Application is well-structured but operationally expensive. The team spends significant time on database maintenance, patching, backup verification, or scaling. Timeline: 2-4 weeks additional per workload compared to rehost. Trade-off: Moderate additional effort (schema compatibility testing, connection string changes) in exchange for a 40-60% reduction in ongoing operational cost. For most mid-complexity applications, replatforming represents the optimal balance between migration effort and long-term benefit. Refactor (Re-Architect) Rebuilding applications for cloud-native patterns: microservices decomposition, containerization (ECS/EKS), serverless (Lambda), event-driven architecture (EventBridge, SQS, SNS, Step Functions). Decision criteria: The application is a core business asset that needs capabilities the current architecture cannot deliver — true horizontal scaling, independent service deployments, multi-region active-active, or zero-downtime deployments. Timeline: Months. Budget accordingly. Trade-off: Highest upfront investment, but delivers the best long-term results in terms of deployment velocity, fault isolation, and scaling capability. Reserve this for 2-3 applications maximum in a migration portfolio. Repurchase Replacing custom-built software with a commercial SaaS product. The application doesn't move to AWS; it moves to a vendor. Decision criteria: The in-house application solves a problem that is not a core competency and commercially available alternatives have matured to cover your requirements. Common candidates: CRM, HR systems, monitoring, project management. The Decision Framework Classification should happen during the assessment phase, before any infrastructure work begins. For each workload, evaluate four dimensions: 1. Business Value How critical is this application to revenue generation or core operations? High: Core product, customer-facing, revenue-generatingMedium: Internal operations, supports core processesLow: Legacy, rarely used, or duplicate functionality 2. Technical Complexity How difficult is it to migrate given current architecture, dependencies, and state management? High: Stateful, tightly coupled, hardware dependencies, proprietary protocolsMedium: Standard web application with database, some external integrationsLow: Stateless, containerizable, standard protocols 3. Team Capacity Does your engineering team have the skills and bandwidth to support a complex migration approach? High capacity: Can support re-architecting alongside other workLimited capacity: Can handle replatforming with some external supportMinimal capacity: Rehost or retain is the only realistic option 4. Time Constraint How quickly must this workload be operational on AWS? Immediate (weeks): Data center exit, contract expiryStandard (1-3 months): Planned migration within a programFlexible (3-6 months): Can wait for deeper optimization Mapping Dimensions to Strategy Business ValueComplexityCapacityTimeRecommended StrategyLowAnyAnyAnyRetire or RepurchaseAnyHighLowImmediateRehost (with future replatform plan)MediumMediumMediumStandardReplatformHighMedium-HighHighFlexibleRefactorAnyAnyAnyBlockedRetain A Practical Example Consider a portfolio of 15 applications for a mid-size SaaS company: Plain Text ┌─────────────────────────────────────────────────────┐ │ RETIRE (3) │ │ - Legacy admin panel (replaced by new one 2024) │ │ - Internal wiki (moved to Confluence) │ │ - Prototype service (never went to production) │ ├─────────────────────────────────────────────────────┤ │ RETAIN (1) │ │ - Hardware security module integration │ │ (requires legal review for cloud deployment) │ ├─────────────────────────────────────────────────────┤ │ REHOST (4) │ │ - Backoffice tools (low traffic, stable) │ │ - Legacy reporting engine (EOL in 18 months) │ │ - Monitoring collector agents │ │ - Staging environment clone │ ├─────────────────────────────────────────────────────┤ │ REPLATFORM (5) │ │ - Main API (PostgreSQL → RDS, cron → Lambda) │ │ - Worker services (EC2 → ECS Fargate) │ │ - File processing pipeline (S3 + Lambda) │ │ - Authentication service (→ElastiCache for sessions│ │ - Notification service (→ SES + SQS) │ ├─────────────────────────────────────────────────────┤ │ REFACTOR (1) │ │ - Core product platform (monolith → microservices) │ ├─────────────────────────────────────────────────────┤ │ REPURCHASE (1) │ │ - Custom CRM (→ HubSpot) │ └─────────────────────────────────────────────────────┘ This distribution — 20% retire, 7% retain, 27% rehost, 33% replatform, 7% refactor, 7% repurchase — is representative of what I see in practice. The replatform bucket is almost always the largest. Migration Tooling Alignment Each strategy maps to specific AWS tooling: StrategyPrimary ToolsRehostAWS Application Migration Service (MGN), Migration HubReplatformDMS (databases), manual adaptation, Terraform/IaCRefactorECS/EKS, Lambda, Step Functions, custom developmentRelocateVMware Cloud on AWS AWS Migration Hub provides a unified tracking dashboard across all strategies. For database migrations specifically, AWS Database Migration Service (DMS) handles both homogeneous and heterogeneous migrations with continuous replication (CDC), enabling near-zero-downtime cutovers. Common Anti-Patterns "Rehost everything, optimize later." Teams that plan to rehost first and replatform in a second phase rarely execute phase two. The urgency disappears once applications are running, and the team moves to other priorities. If replatforming is the right strategy, do it during migration. "Refactor everything for cloud-native." The opposite extreme. Not every application needs microservices. A well-structured monolith running on ECS Fargate can serve thousands of requests per second with simpler operations than a distributed system. "One strategy for all workloads." Every application in the portfolio has different characteristics. The decision framework exists because one size does not fit all. Conclusion The 7R classification exercise takes 3-5 days for a typical portfolio. It requires involvement from engineering leads, product owners, and sometimes finance (for retire/repurchase decisions). The output - a workload-by-workload strategy map - becomes the foundation for accurate timeline estimates, resource planning, and budget allocation. Without it, you're building infrastructure for workloads that might not need to exist. For a comprehensive breakdown of migration costs, the full 6-phase delivery process, and AWS tooling details, see my complete AWS cloud migration guide.
A self-healing SQL pipeline should not mean autonomous SQL generation followed by privileged execution. In production, the safer interpretation is narrower, where a language model proposes a repair, while deterministic controls decide whether that repair is syntactically valid, semantically plausible, operationally safe, and eligible for execution. This distinction matters because the same mechanism that corrects a renamed column can also generate an unintended DELETE, widen a join, or scan an unexpectedly large dataset. Structured-output features can constrain an LLM response to a defined schema, but schema conformance is not equivalent to database correctness or authorization. OpenAI’s Structured Outputs is designed to make generated output conform to supplied JSON Schemas, and it does not validate SQL semantics or execution safety. Treat the Failure as Evidence, Not Merely a Prompt The repair loop should begin by classifying the failure before any model call. Parser errors, missing relations, unknown columns, type mismatches, permission failures, timeouts, cardinality explosions, and upstream freshness problems require different responses. A permission error should not trigger SQL rewriting, while an unknown-column error may justify metadata inspection. This gate keeps deterministic failure classes deterministic. Consider a pipeline that previously executed: SQL SELECT customer_id, customer_segment FROM analytics.customer_profile WHERE active = TRUE; An upstream migration renames customer_segment to segment_name. The database returns an unknown-column error. The repair service should collect the failed statement, SQL dialect, error code, schema version, referenced objects, and recent catalog changes. Metadata inspection can then establish that customer_profile still exists, customer_segment no longer exists, and segment_name appeared in the latest schema version. That evidence is stronger than asking a model to infer a replacement from exception text alone. Schema drift detection should therefore precede generation. Catalog snapshots can be hashed and compared between successful and failed runs. Candidate mappings can incorporate data type compatibility, nullability, lineage metadata, column comments, and migration records. The LLM then receives only evidence relevant to the suspected failure class. Candidate generation should return a structured proposal rather than free-form SQL. A production contract can require the proposed SQL, repair category, changed identifiers, evidence references, and assumptions: Python def generate_candidate(failure, metadata): return llm.generate( schema=RepairProposal, context={"failure": failure, "metadata": metadata}, constraints={"max_statements": 1, "allow_dml": False} ) The allow_dml flag is an application policy, not an instruction trusted merely because it appears in a prompt. Structured generation narrows output shape, while authorization remains outside the model. Anthropic’s evaluation guidance similarly emphasizes measurable success criteria and testable thresholds rather than treating model behavior as inherently reliable. Parse the Candidate Before the Database Sees It String matching is too weak for SQL safety. A rule such as "DELETE" not in sql.upper() can miss nested statements and dialect-specific constructs. The candidate should be parsed into an abstract syntax tree using the correct database dialect. SQLGlot parses SQL into expression trees and supports multiple dialects, enabling structural inspection before execution. A validator can reject statement types outside an allowlist and verify referenced tables and columns against current metadata: Python def validate_ast(sql, dialect, catalog): tree = parse_one(sql, dialect=dialect) if tree.find(Delete) or tree.find(Update) or tree.find(Insert): raise PolicyViolation("mutating statement rejected") for table in tree.find_all(Table): catalog.require_table(table.name) for column in tree.find_all(Column): catalog.require_column(column.table, column.name) return tree AST validation should also enforce tenant boundaries, prohibited schemas, mandatory predicates, join limits, and function restrictions. A generated query can be syntactically valid yet unsafe because an omitted filter changes a targeted lookup into a full table operation. Semantic checks therefore need context from the original successful query, expected output columns, and data quality assertions. The repaired query may be: SQL SELECT customer_id, segment_name AS customer_segment FROM analytics.customer_profile WHERE active = TRUE; Preserving the original output alias matters because downstream consumers may depend on customer_segment even though the physical source column changed. A repair that only substitutes the new identifier could restore execution while silently breaking the pipeline contract. Dry Runs Should Prove More Than Syntax A candidate that survives static validation still should not immediately reach production data. Database-native validation can catch failures that an AST cannot. BigQuery dry runs validate query structure and estimate bytes processed without executing the query, making them useful for rejecting unexpectedly expensive repairs. Google also documents that successful dry runs do not guarantee successful runtime execution and that multi-statement dry runs have special limitations. A guarded validation step can combine dry run results with policy limits: Python def dry_run(candidate): result = warehouse.validate(candidate.sql) if result.bytes_scanned > MAX_BYTES: raise PolicyViolation("scan budget exceeded") if result.output_schema != candidate.expected_schema: raise ContractViolation("output schema changed") return result For engines without native dry run support, a read-only transaction, isolated replica, sandbox database, or planner-only operation can provide a safer boundary. PostgreSQL supports read-only transaction modes that prevent changes to non-temporary tables, adding a database-enforced control rather than relying only on application logic. Operational safety also depends on continuous monitoring after a repair is deployed. Query latency, row counts, null rates, schema changes, and downstream data quality indicators can reveal subtle regressions that static validation may miss. Post execution monitoring therefore provides another deterministic checkpoint, allowing suspicious behavior to trigger rollback or human review immediately. Confidence should come from independently observable signals, not from an LLM declaring confidence in its own answer. A score can combine schema evidence, AST-policy results, dry run success, output-schema stability, and historical repair success: Python def repair_score(signals): return ( 0.30 * signals.schema_evidence + 0.25 * signals.ast_validation + 0.20 * signals.dry_run + 0.15 * signals.contract_match + 0.10 * signals.history ) Weights should be calibrated against labeled historical failures. LLM-based judging can contribute a secondary semantic signal, but it should not authorize execution. G-Eval shows that model-based evaluators can correlate with human judgments while also identifying evaluator bias as a concern. A useful inference from RAGAS is that component-level measurements are more diagnosable than one opaque score; the same principle fits SQL repair by keeping schema, syntax, execution, and contract evidence separately observable. Bound Autonomy and Make Every Repair Reversible Self-healing becomes dangerous when retries are unbounded. A failed candidate should feed only new deterministic evidence into the next attempt, such as a parser error or dry-run diagnostic, and the loop should stop after a small configured limit. Repeated failures, low confidence, ambiguous schema mappings, contract changes, or any proposed mutation should escalate to human review. LangSmith distinguishes offline evaluation from online production evaluation and describes a feedback loop in which production failures become future evaluation cases and the same operational pattern fits SQL repair systems. Every attempt should produce an immutable audit record containing the original SQL hash, failure evidence, metadata version, model and prompt version, candidate hash, validation outcomes, confidence signals, execution identity, and final disposition. That record supports debugging and regression testing of future repair policies. Production monitoring should also track changes in failure distributions and repair success rates, as Google’s MLOps guidance treats monitoring as a trigger for new experimentation when production quality degrades. Automatic mutation deserves a higher bar than automatic read-only repair. When writes are permitted, idempotency keys should prevent duplicate side effects across retries, and execution should remain transactional whenever supported. AWS reliability guidance recommends idempotency for database insert, update, and delete operations. Transaction savepoints and rollback mechanisms add another containment layer, as PostgreSQL savepoints allow effects after a savepoint to be selectively discarded without abandoning the entire transaction. A production-grade self-healing SQL pipeline is not an autonomous database administrator implemented with a prompt. It is a controlled repair system in which probabilistic generation is surrounded by deterministic evidence collection, AST inspection, schema validation, database-enforced dry runs, calibrated confidence thresholds, bounded retries, auditability, and explicit escalation. The safest design grants the LLM authority to propose change, not authority to approve or execute it. With that separation preserved, LLMs can reduce recovery time for routine SQL failures while the database, policy engine, and human review path retain control over production state.
Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq
Testing POST API requests is an important skill for modern QA and automation engineers working with backend services and microservices. In this article, we’ll explore how to test POST API requests with Playwright and TypeScript, focusing on sending a request body using different approaches. By the end, you will learn how to send POST API requests using the following approaches for adding a request body: JSON Object/Array(data)Stringified JSONJSON FileFaker library Application Under Test We will use the POST /addOrder API of the RESTful e-commerce demo application for this demo. The API schema is provided below: JSON { "user_id": "string", "product_id": "string", "product_name": "string", "product_amount": 0, "qty": 0, "tax_amt": 0, "total_amt": 0 } ] Testing POST API Requests With Playwright TypeScript Playwright provides a powerful request API that allows us to create and manage HTTP request contexts. Let’s walk through how to send POST API requests step-by-step using different approaches for passing the request body: Request Body as JSON Object/Array Let’s send a POST API request with a static JSON Array and verify that the response status code is 201. TypeScript test("POST order details API with static JSON Array", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data:[{ user_id: "1", product_id: "82", product_name: "Cadbury Bar", product_amount: 12, qty: 2, tax_amt: 1, total_amt: 25 }, { user_id: "2", product_id: "80", product_name: "MilkyBar", product_amount: 10, qty: 1, tax_amt: 1, total_amt: 11 } ], }); expect(response.status()).toBe(201); }); Code Walkthrough A Playwright test case for testing the POST API request is created using the request API, which is a built-in Playwright APIRequestContext fixture used to send HTTP requests. Sending the POST Request: The request.post() sends a POST request to the “http://localhost:3004/addOrder/” endpoint. The response received after sending a POST request is stored in the response variable.Passing the Request Body (Static JSON Array): The data property is used to send the request body, which accepts data in JSON Array and JSON object formats. Since the POST /addOrder API accepts an array of order details, allowing multiple orders to be submitted within a single JSON array, the test provides the data in JSON array format.Validating the Response: The response.status() retrieves the HTTP status code. The assertion ensures that the status code 201 is returned, indicating that the orders were successfully created. Request Body Using JSON.Stringify() JSON.stringify() can be used while sending raw payloads, testing malformed JSON, or sending custom-formatted JSON. However, the Content-Type header should be set to application/json so the server correctly interprets the request body as JSON data. Let’s write the same POST /addOrder test using JSON.stringify() and verify that a status code 201 is returned in the response. TypeScript test("POST order details API using JSON.Stringify", async ({ request }) => { const orderData = [{ user_id: "5", product_id: "64", product_name: "Cadbury Mini", product_amount: 5, qty: 3, tax_amt: 1, total_amt: 16 }]; const response = await request.post("http://localhost:3004/addOrder/", { data: JSON.stringify(orderData), headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); The orderData array is defined to contain one order object. Each field represents the order details such as user_id, product_id, product_name, and so on. This is a normal JavaScript object at this stage, not JSON yet. The JSON.stringify(orderData )converts the JavaScript object into a JSON string. As we manually stringify the request payload, we should explicitly set the Content-Type header to application/json. It tells the server to treat the request body as JSON data. Playwright sends the POST API request using the post() method and then validates that the response returns a 201 status code. Request Body as a JSON File Using a JSON file as a request body is a handy approach when testing POST API requests with Playwright TypeScript. It supports large payloads and allows the same file to be easily reused across multiple tests. The following orders.json file will be used as a request payload in the POST /addOrder API: JSON [ { "user_id": "1", "product_id": "79", "product_name": "5 star 10gm Chocobar", "product_amount": 5, "qty": 1, "tax_amt": 0.5, "total_amt": 5.5 }, { "user_id": "2", "product_id": "71", "product_name": "Lindt Milk Chocolate", "product_amount": 15, "qty": 3, "tax_amt": 2.5, "total_amt": 47.5 } ] The following configurations should be in place before we proceed to write the test using a JSON file as the request payload. tsconfig.json JSON { "compilerOptions": { "module": "NodeNext", "moduleResolution": "NodeNext", "resolveJsonModule": true } } We need to ensure that the JSON file is imported first and then attached to the data parameter while sending the POST request. TypeScript import orders from '../test_data/orders.json' with {type: 'json'}; test("POST order details API using JSON file", async ({ request }) => { const response = await request.post("http://localhost:3004/addOrder/", { data: orders, }); expect(response.status()).toBe(201); }); This test imports test data from an external orders.json file and uses it as the request body for a POST API call in Playwright. The request.post() method sends the imported JSON directly as the payload, and the test verifies that the API returns a 201 status code. Using a JSON file makes it easier to manage, reuse, and update request payloads without modifying the test logic, improving the maintainability and readability of API tests. Request Body With Faker Library Using the Faker library to create a request body for a POST API helps generate realistic and dynamic test data, reduces hardcoded values, and improves test coverage. It is especially useful for simulating real-world scenarios and avoiding duplicate data issues. However, it also has limitations. Since Faker generates random data, it may lead to inconsistent test results if not properly handled, and debugging failures could become more difficult without fixed or reproducible inputs. To use the Faker library, we need to install it first using the following command: Plain Text npm install --save-dev @faker-js/faker Next, let’s create two helper functions: one to create order objects with the required order details, and another to generate an array of orders based on the count provided by the user. TypeScript import { faker } from '@faker-js/faker'; export function createOrderDetails() { const productAmount:number = faker.number.int({ min: 1, max: 100 }); const qty:number = faker.number.int({ min: 1, max: 5 }); const taxAmt:number = faker.number.int({ min: 2, max: 10 }); const totalAmt:number = (productAmount*qty)+taxAmt; return { user_id: faker.number.int({ min: 1, max: 50 }), product_id: faker.number.int({ min: 1, max: 100 }), product_name: faker.commerce.productName(), product_amount: productAmount, qty:qty, tax_amt: taxAmt, total_amt: totalAmt }; } The createOrderDetails() function returns the order details with the required fields. It also calculates the total amount by summing the total product value and tax amount. The values in all fields are updated randomly using the appropriate methods provided by the Faker library. TypeScript import { faker } from '@faker-js/faker'; export function createRandomOrders(count:number) { return faker.helpers.multiple(createOrderDetails, {count}); } The createRandomOrders(count) method accepts a count parameter and returns randomly generated orders that can be directly used in the test. TypeScript return faker.helpers.multiple(createOrderDetails, {count}); This is where the work happens. multiple() is a helper function provided by Faker. Its job is to execute another function multiple times and collect all the results into an array. It takes two arguments: A function to execute repeatedly.An options object specifying how many times to execute it. It asks Faker to call createOrderDetails() exactly count times, collect all the generated orders into an array, and return that array to the caller. TypeScript test('POST order details API using Faker library', async({request}) => { const orderData = createRandomOrders(5); const response = await request.post("http://localhost:3004/addOrder/", { data: orderData, headers: { "Content-Type": "application/json", }, }); expect(response.status()).toBe(201); }); This test sends a POST request to the /addOrder endpoint API using dynamically generated test data from the Faker library. The createRandomOrders(5) function generates an array of 5 random order objects, which are passed as the request body using the data property. The test then verifies that the API responds with a 201 status code, confirming that the orders were successfully created. Summary In this tutorial, we explored multiple approaches to testing POST API requests using Playwright with TypeScript, including sending static JSON payloads, using JSON.stringify, importing data from external JSON files, and generating dynamic test data with the Faker library. We also covered how to handle headers correctly and validate API responses using status code assertions. In my experience, using external JSON files is a practical approach for testing POST API requests because it lets us send bulk data from a single file. Similarly, the Faker library can also be used to generate dynamic test data; however, in some cases, using a third-party library may not be permitted. Ultimately, the software team chooses the approach that best fits their testing strategy. Happy testing!
When many developers think about recommendation engines, they think of machine learning: collaborative filtering models, matrix factorization, embedding vectors, and training pipelines. What surprises many people is that you can build a genuinely useful recommendation system with nothing more than a graph database and several Cypher queries. No scikit-learn, no TensorFlow, no model training. Just the natural structure of the data doing the work. In this article, we'll build a product recommendation engine on top of Neo4j Aura using two Jupyter notebooks. The first generates a realistic synthetic dataset and loads it into Aura. The second runs four recommendation queries directly in Cypher and visualizes the results with Plotly. Everything runs locally in a Python virtual environment against a free cloud Neo4j instance. The full source code is available on GitHub. Why Graphs Are a Natural Fit for Recommendations The core intuition behind most recommendation approaches is relationship: this customer bought that product, those products appear together in the same order, this product shares attributes with that one. In a relational database, capturing these relationships means multiple self-joins across large tables. A query like "find products bought by customers who also bought what this customer bought" quickly becomes difficult to write and expensive to execute at scale. In a graph, that same question is a traversal. We follow edges from a customer to the products they purchased, hop across to other customers who share those products, and collect what else those customers bought. The query is short, the intent is clear, and the graph engine is optimized for exactly this kind of path-following work. Prerequisites AuraDB is Neo4j's fully managed cloud database. A free tier is available with no credit card required. Sign up at Get Started for Free.Create a new AuraDB Free instance.When the instance is created, download or note the credentials — the connection URI, username, and password.Once the instance is running, open the Query tab and connect to the instance.Confirm it's empty with MATCH (n) RETURN count(n) which should return 0 A virtual environment is highly recommended. For example: Shell python3 -m venv ~/recommendation-engine-env source ~/recommendation-engine-env/bin/activate Before starting Jupyter, export the connection details as environment variables in your shell: Shell export NEO4J_URI="neo4j+s://xxxx.databases.neo4j.io" export NEO4J_USERNAME="your_username_here" export NEO4J_PASSWORD="your_password_here" The Graph Model Before we write any code, let's define the graph. We have four node types and three relationship types. Nodes Customer – id, name, email, city, country.Product – id, name, description, price.Category – name (e.g., Electronics, Clothing, Books).Tag – name (e.g. "wireless", "eco-friendly", "premium"). Relationships (:Customer)-[:PURCHASED {order_id, quantity, order_date}]->(:Product) — order metadata lives on the relationship rather than a separate Order node, which keeps our Cypher clean.(:Product)-[:BELONGS_TO]->(:Category)(:Product)-[:TAGGED_WITH]->(:Tag) The decision to put order_id, quantity and order_date on the PURCHASED relationship is worth discussing. It means a single customer can have multiple PURCHASED relationships to the same product (each with a different order_id) and we can group by order_id to find products that appeared together in the same basket — which is exactly what our co-purchase query needs. Figure 1 illustrates exactly this point, as we have a customer, two products, and the same order_id. Figure 1. Shared order_id enables co-purchase queries Notebook 1: Data Generation and Loading Rather than sourcing an external dataset, we'll generate synthetic data using Faker. This keeps the notebook fully self-contained, and readers can run it as-is without downloading anything. We'll generate 2,000 customers, 500 products across 15 categories, and 20,000 orders. Each order is a basket of several products sharing the same order_id — this is the key design decision that makes the frequently-bought-together query work. With an average basket of 3 products, we end up with around 60,000 PURCHASED relationships in the graph. Realistic Product Names Faker's default catch_phrase() method produces output like "Proactive exuding encoding" — readable enough for a demo but not really useful in an article. Instead, we define a PRODUCT_VOCAB dictionary keyed by category, each containing lists of adjectives, nouns, use cases, and benefit statements. A product name is then a simple combination, as follows: Python def make_product_name(category): vocab = PRODUCT_VOCAB[category] adj = random.choice(vocab["adjectives"]) noun = random.choice(vocab["nouns"]) return f"{adj} {noun}" def make_product_description(category, name): vocab = PRODUCT_VOCAB[category] use_case = random.choice(vocab["use_cases"]) benefit = random.choice(vocab["benefits"]) return f"The {name} is designed for {use_case}. {benefit}." This gives us names like "Wireless Noise-Canceling Earbuds," "Organic Ground Coffee" and "Ergonomic Lumbar Support Cushion" — realistic enough to make the recommendation output meaningful. Basket-Based Order Generation Each order picks a random customer, generates a unique order_id, and samples several products into a basket. We then flatten the basket into individual order lines, each carrying the shared order_id: Python orders = [] for _ in range(NUM_ORDERS): order_id = str(uuid.uuid4()) customer = random.choice(customers) order_date = (start_date + timedelta(days=random.randint(0, 730))).strftime("%Y-%m-%d") basket = random.sample(products, k=random.randint(2, 4)) for product in basket: orders.append({ "order_id": order_id, "customer_id": customer["id"], "product_id": product["id"], "quantity": random.randint(1, 5), "order_date": order_date }) Loading Into Aura Data loading is in batches of 100 using MERGE statements. To show progress during the load, we'll use tqdm as ~60,000 order lines can take several minutes, and the progress bars make it easy to see what's happening: Python with driver.session() as session: customer_batches = range(0, len(customers), BATCH_SIZE) for i in tqdm(customer_batches, desc="Loading customers", unit="batch", colour="#1f77b4"): session.execute_write(load_customers, customers[i:i+BATCH_SIZE]) product_batches = range(0, len(products), BATCH_SIZE) for i in tqdm(product_batches, desc="Loading products ", unit="batch", colour="#1f77b4"): session.execute_write(load_products, products[i:i+BATCH_SIZE]) for product_id, tags in tqdm(product_tags.items(), desc="Loading tags ", unit="product", colour="#1f77b4"): session.execute_write(load_tags, product_id, tags) order_batches = range(0, len(orders), BATCH_SIZE) for i in tqdm(order_batches, desc="Loading orders ", unit="batch", colour="#1f77b4"): session.execute_write(load_orders, orders[i:i+BATCH_SIZE]) A verification query at the end confirms the counts. The Four Recommendation Queries Notebook 2 runs four Cypher queries against the loaded graph, each implementing a different recommendation strategy. Before running any query, we fetch a stable seed customer, product, and category: Python with driver.session() as session: customer = session.run(""" MATCH (c:Customer) RETURN c.id AS customer_id, c.name AS customer_name ORDER BY c.name ASC LIMIT 1 """).single() product = session.run(""" MATCH (p:Product)<-[r:PURCHASED]-() RETURN p.id AS product_id, p.name AS product_name, count(r) AS order_count ORDER BY order_count DESC LIMIT 1 """).single() top_cat = session.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category) RETURN cat.name AS category, count(p) AS total ORDER BY total DESC LIMIT 1 """).single() We pick the alphabetically first customer for consistency, the most-purchased product to ensure co-purchase data exists, and the category with the most products for the trending query. This makes the notebook reproducible across runs. Query 1: Collaborative Filtering The classic "customers who bought this also bought" approach. We find customers who share at least one purchased product with the seed customer, then collect what else those customers bought — excluding anything the seed customer already purchased. Python def collaborative_filtering(tx, customer_id, limit=5): result = tx.run(""" MATCH (target:Customer {id: $customer_id})-[:PURCHASED]->(p:Product) <-[:PURCHASED]-(other:Customer)-[:PURCHASED]->(rec:Product) WHERE NOT (target)-[:PURCHASED]->(rec) RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(other) AS score ORDER BY score DESC, id ASC LIMIT $limit """, customer_id=customer_id, limit=limit) return result.data() The score is the number of other customers whose purchasing overlap with our target customer also led them to buy the recommended product. A higher score means more customers in the overlap group bought it, making it a stronger signal. In Cypher, the traversal reads almost like the description: start at the target customer, follow PURCHASED edges to products, hop to other customers who bought the same products, then follow their PURCHASED edges to new products. Query 2: Frequently Bought Together This query finds products that appeared in the same order as the seed product. The key is matching on order_id across two PURCHASED relationships from the same customer: Python def frequently_bought_together(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})<-[r1:PURCHASED]-(c:Customer) -[r2:PURCHASED]->(other:Product) WHERE r1.order_id = r2.order_id AND other.id <> $product_id RETURN other.id AS id, other.name AS product, other.price AS price, count(c) AS frequency ORDER BY frequency DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The WHERE r1.order_id = r2.order_id clause is what makes this work. It constrains the traversal to only consider cases where both products were part of the same order, not just bought by the same customer at different times. frequency counts how many distinct customers placed an order containing both products together. Query 3: Content-Based Filtering Rather than looking at purchase behavior, this query finds products similar to the seed product based on shared tags. The more tags two products have in common, the more similar they are: Python def content_based(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})-[:TAGGED_WITH]->(t:Tag) <-[:TAGGED_WITH]-(rec:Product) WHERE rec.id <> $product_id RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(t) AS shared_tags ORDER BY shared_tags DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The traversal goes outward from the seed product through its tags, then back inward to any other product that shares those same tags. count(t) gives the number of shared tags, which serves as a simple but effective similarity score. This approach works without any purchase history, making it useful for recommending products to new customers or for newly listed products with no order data yet. Query 4: Trending in Category This query finds the most purchased products in the top category within a fixed date window. In our case, this is from 2024-10-01 onwards: Python def trending_in_category(tx, category_name, cutoff="2024-10-01", limit=5): result = tx.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category {name: $category_name}) MATCH (:Customer)-[r:PURCHASED]->(p) WHERE date(r.order_date) >= date($cutoff) RETURN p.id AS id, p.name AS product, p.price AS price, count(r) AS purchases ORDER BY purchases DESC, id ASC LIMIT $limit """, category_name=category_name, cutoff=cutoff, limit=limit) return result.data() We use date() conversion on the stored string order_date to enable date comparison. count(r) counts individual PURCHASED relationships rather than distinct customers, so a customer who bought the same product multiple times within the window is counted each time — reflecting genuine demand volume rather than unique buyer count. Notebook 2: Results Each query outputs a table followed by a Plotly horizontal bar chart. Here are the results for our seed data. Collaborative Filtering Figure 2 returns five products. The top recommendation is An Introduction to Public Speaking, driven by the number of customers whose purchasing overlap with Aaron Boyd also led them to buy it. Heavy-Duty Cable Management Box and Educational Coding Robot follow closely, showing that the overlap group bought broadly across categories rather than clustering in one area. Figure 2. Collaborative filtering Frequently Bought Together Figure 3 shows products co-purchased with the Durable Grooming Brush in the same order basket. The top results — Waterproof Hammock and Natural Body Lotion at frequency 4, followed by Adjustable Lumbar Support Cushion, Sugar-Free Collagen Powder and Slim-Fit Hiking Vest at frequency 3 — show which products most commonly appeared alongside the seed product in the same order. The cross-category spread here (Beauty, Outdoor, Clothing, Health, Office) is a feature of random synthetic data; in a real system, we'd expect more category clustering. Figure 3. Frequently bought together Content-Based Filtering Figure 4 finds products sharing the most tags with the seed product. All five results share 2 tags with the Durable Grooming Brush — Smart Mechanical Keyboard, Waterproof Toiletry Bag, Ergonomic Whiteboard, Cold-Pressed Hot Sauce, and Durable Dumbbell Pair. The cross-category reach (Sports, Food & Drink, Office, Travel, Electronics) illustrates the tag graph doing its job: shared attributes like "durable" or "waterproof" create similarity links that cross category boundaries, which is useful for surface-level discovery recommendations. Figure 4. Content-based filtering Trending in Category Figure 5 shows the top 5 products in Toys — the category with the most products in our graph — with purchase counts from 2024-10-01 onwards. Battery-Free Coding Robot leads, followed by Battery-Free Building Blocks Set, Interactive Remote Control Car, Wooden Magnetic Drawing Board, and Creative Puzzle Game. The scores are tight here, which makes sense because within a single category over a fixed time window, popular products tend to cluster around similar purchase volumes. Figure 5. Trending Summary We've built a working product recommendation engine using nothing but Neo4j, Cypher, and a few Python libraries. No ML framework, no training data, no model deployment. The four queries cover the most common recommendation patterns in production systems: Collaborative filteringCo-purchase analysisContent similarityTrending detection The graph model is the foundation that makes this possible. Storing orders as relationships with properties means co-purchase queries are a natural traversal rather than a complex join. Adding tags as nodes means similarity queries are just path-matching. Because everything lives in the same graph, we can also combine these approaches. For example, filtering collaborative filtering results by tag similarity using a single extended Cypher query. The full source code is available on GitHub.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
Combining Temporal and LangGraph creates a deceptively simple question: which runtime owns the state of the agent? Both preserve execution progress, but they preserve different kinds of progress. Temporal reconstructs Workflow state from Event History and reuses recorded Activity results during replay. LangGraph persists thread-scoped graph state as checkpoints and resumes from super-step boundaries. Treating those mechanisms as interchangeable creates ambiguous recovery semantics. Production integration therefore needs explicit authority for business progress, agent working state, and the handoff between them. Deployment language, storage backend, model provider, and hosting topology remain unspecified assumptions. The current Temporal LangGraph integration narrows the problem. Its public-preview Python plugin can run LangGraph nodes as Temporal Activities or deterministic Workflow code, while Temporal provides durability; the documentation recommends an in-memory LangGraph checkpointer rather than a separate PostgreSQL or Redis checkpointer. Continue-As-New can carry cached task results into the next Workflow Run. The harder two-runtime problem appears when LangGraph retains an independent persistent checkpointer while Temporal separately orchestrates the business lifecycle. That case needs an application-level consistency contract. Ownership and Handoff The cleanest ownership rule is semantic. Temporal should own business lifecycle state: whether an execution is open, waiting for approval, canceled, timed out, compensated, or complete. LangGraph should own agent working state: messages, retrieved evidence, tentative plans, tool proposals, and graph position. External systems should remain authoritative for effects in their own domains. A payment processor owns whether a charge occurred; a deployment service owns whether a release exists. Temporal and LangGraph may retain receipts, but neither should invent contradictory domain truth. This matches their native models: Event History records Workflow progress, while LangGraph checkpointers persist thread state. The handoff should be modeled as a command protocol. A business execution ID identifies the long-lived application process and can map to a stable Temporal Workflow ID. A graph thread ID identifies checkpoint lineage. A command ID identifies one requested graph advance, while a monotonically increasing application revision identifies the state version against which that command was accepted. Temporal Run ID should not become the business identifier because Continue-As-New preserves Workflow ID while creating a new Run ID. LangGraph uses thread_id to load checkpoint history. Stable application identities must therefore outlive either runtime’s individual run instance. A graph-advance Activity can enforce that contract without exposing persistence details to Workflow code: Python def advance_agent(cmd): receipt = receipts.get(cmd.command_id) if receipt and receipt.status == "completed": return receipt.result state = threads.acquire( cmd.thread_id, fencing_token=cmd.revision, ) if state.revision != cmd.expected_revision: raise StaleCommand(cmd.command_id) result = graph.invoke( cmd.input, {"configurable": {"thread_id": cmd.thread_id}, durability="sync", ) return receipts.complete(cmd, result) The key property is the admission rule. The same command ID must never become new input merely because an Activity retried. A command accepted at revision 17 remains command 17 across worker crashes and timeouts. A genuinely new turn receives a new command ID and expected revision. LangGraph’s synchronous durability mode persists each checkpoint before the next step starts, reducing checkpoint-loss risk, but it does not create a transaction with Temporal Event History. Failure and Re-Execution The critical failure window begins after LangGraph commits a checkpoint and before Temporal records Activity completion. Temporal documents the analogous edge directly: an Activity can finish, the worker can crash before reporting completion, and the Activity can then execute again. Completed Activities are not re-executed during Workflow replay, but an unrecorded completion is indistinguishable from unfinished work at the orchestration boundary. Blindly injecting the same graph input on retry can therefore advance the agent twice for one logical transition. A durable command receipt closes that ambiguity. It records command ID, thread ID, expected revision, resulting revision, checkpoint reference, status, and result reference. When checkpoint and receipt records share a database, a custom persistence adapter can commit command acceptance and checkpoint metadata in one local transaction. When stores cannot share a transaction, recovery can reconstruct a missing receipt from checkpoint metadata containing the command ID and resulting revision. That reconstruction is an application-level inference based on LangGraph checkpoint metadata and lookup capabilities. A separate receipt written only after graph completion leaves another crash window. External effects require a second deduplication boundary. Temporal recommends idempotent Activities because Activity execution may happen more than once, and idempotency keys must ultimately be enforced by the called service. The graph should therefore propose an effect before performing it. An interrupt can expose the proposal, allowing Temporal to own approval, deadlines, and cancellation while a dedicated Activity performs the mutation with a stable operation ID. LangGraph documents that an interrupted node restarts from its beginning when resumed, so code before interrupt() executes again and pre-interrupt effects must be safe to repeat. Python def await_effect(state): receipt = interrupt({ "operation_id": state["operation_id"], "proposal": state["proposal"], "revision": state["revision"], }) if receipt["operation_id"] != state["operation_id"]: raise ValueError("effect receipt mismatch") return {"effect_receipt": receipt} This arrangement also clarifies replay. Temporal replay re-executes Workflow code while matching Commands against Event History; recorded Activity results are reused. LangGraph replay from an older checkpoint instead re-executes nodes after that checkpoint, including LLM calls, API requests, and interrupts. A LangGraph fork is therefore a new computational branch, not restoration of external reality. Previously completed business effects remain attached to their original operation receipts, while newly proposed effects require fresh authorization. Operational Semantics Concurrency control must prevent overlapping Activity attempts from advancing one thread simultaneously. A lease alone is insufficient if an expired holder can still write. A monotonically increasing fencing token tied to the accepted application revision provides a stronger rule: persistence rejects writes from an older token after a newer command is admitted. This is an integration pattern rather than a built-in Temporal or LangGraph guarantee. Administrative retries, manual resumes, and human approvals should pass through the same admission path. Temporal’s documented retry model establishes the underlying reason for such protection: Activity execution can occur more than once even though successful completion is observed once by the Workflow. Retry policy should remain layered. Temporal should own Activity retry delivery, while LangGraph-level retries should remain narrowly scoped to graph operations whose repetition is safe. Independent retry loops at both layers can multiply attempts and obscure the failure budget. Cancellation also needs explicit semantics. Temporal delivers cancellation to heartbeat-enabled Activities through heartbeats, but cancellation of orchestration does not prove that an already accepted remote effect was reversed. Ambiguous operations therefore require reconciliation with the authoritative downstream system. Continue-As-New changes run identity but not business identity. Temporal starts a fresh Event History with the same Workflow ID and a different Run ID, carrying selected state forward. In a dual-runtime design, the business execution ID, graph thread ID, latest revision, outstanding commands, and unresolved effect receipts must cross that boundary. Temporal scopes Update-ID deduplication to a Workflow Run, so deduplication that must survive Continue-As-New cannot rely solely on server-side Update identity. The official LangGraph integration similarly carries serialized cached results across Continue-As-New. Operational tests should target boundaries rather than happy paths. Worker termination immediately after checkpoint commit, after an external service accepts an operation, and before Activity acknowledgment exposes duplicate-execution defects. Concurrent retries should prove that stale fencing tokens cannot write. Stale approvals should prove that proposal revision checks reject obsolete decisions. Continue-As-New tests should prove that logical identities and deduplication records survive the run transition. Temporal provides test hooks for exercising Continue-As-New, while LangGraph checkpoint history and replay make recovery assertions observable. Conclusion A Temporal-and-LangGraph agent becomes reliable only when durability is subordinate to ownership. Temporal should decide business progress, LangGraph should preserve agent working state, and external systems should remain authoritative for real-world effects. The handoff needs stable business, thread, command, and revision identities; durable receipts; idempotent effect execution; explicit replay semantics; fenced concurrency; and continuity across Continue-As-New. The official Temporal LangGraph plugin can collapse much of this complexity by making Temporal the durable execution substrate. When independent persistence remains on both sides, however, two durable runtimes do not become one consistent runtime automatically. A precise state-ownership protocol is what turns duplicated durability into controlled recovery.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. After a few years of operating a shared Kubernetes environment, the shift in the center of gravity becomes clear. Cluster provisioning, container scheduling, and upgrades become routine, yet releases still stall over ownership, access, telemetry, and cost allocation. Consider a hypothetical product team adding a stateful order-processing service to a shared platform. The service has an API, a worker, database migrations, and a data store backed by cluster-managed persistent storage, and it must run in staging and production. We will follow that service through its delivery path to examine where mature infrastructure stops helping, how local workflow differences compound, and which operating model decisions restore consistency without stripping teams of useful autonomy. When Infrastructure Maturity Stops Solving the Hard Part At first glance, onboarding the service should be routine. The cluster exists, the CI system can build an image, and infrastructure as code can create the namespace. Then the service reaches production and encounters a different StorageClass, quota profile, network policy, or service account configuration from staging. Each difference may be valid, but the delivery workflow didn’t surface the environment contract early enough. This is the practical limit of infrastructure maturity. Reliable clusters provide capable building blocks, while reliable delivery also requires a shared agreement about how teams use those blocks, what evidence a release produces, where exceptions go, and who owns the outcome. How Cloud Complexity Starts to Compound Follow the service through one release and the pattern becomes clearer: The same components use inconsistent service and environment identifiers across logs and traces.Ownership labels exist in one cluster but not the other.The team copies a pipeline because the shared template cannot sequence migrations.Production access and policy exceptions move through separate ticket queues. The operational cost shows up in the manual coordination required before each deployment. An engineer has to reconstruct which rules apply every time. During an incident, responders can’t move cleanly from an alert to the owning team, deployment record, runbook, and cost center. Finance sees shared-cluster spend that cannot be attributed reliably, while the security team receives evidence in different formats. OpenTelemetry semantic conventions and FinOps allocation practices rely on consistent service, environment, and allocation metadata, so local naming schemes undercut the value of the underlying tools. As the same pattern spreads across clusters and cloud accounts, small differences become a persistent operating burden. Operating Models Set the Terms of Scale The operating model decides who turns those building blocks into a usable delivery system. For our example, the product team owns the order domain, data model, migration safety, scaling behavior, SLOs, and on-call response. The platform team owns the interface through which the service receives a namespace, workload identity, baseline policy, deployment workflow, and telemetry defaults. Security, SRE, and FinOps teams contribute requirements and review the evidence that the workflow produces. This split keeps service-specific decisions close to the people who understand them while centralizing cross-cutting capabilities that every team would otherwise rebuild. CNCF’s platform guidance makes an important distinction here: A platform team is responsible for the interfaces and experience around shared capabilities, even when another team or provider operates the backing service. In practice, a mature platform offers versioned workflows, clear support boundaries, self-service for common requests, and feedback loops based on real usage. In this way, the platform team is an enabler of consistency rather than the operator of every component. Standardization, Autonomy, and Shared Operating Logic The defensible baseline is narrower than a universal application architecture. For this service, shared standards should cover: Service and environment identityWorkload identity and minimum network policyResource requests, quota expectations, and cost-allocation metadataRelease evidence, rollback behavior, and minimum telemetry These rules belong in the shared workflow because inconsistency affects other teams and complicates incident response, security, and cost allocation. The product team still chooses its schema, partitioning strategy, cache design, scaling thresholds, and release timing, and defines SLOs around the behavior users experience. The stateful workload then tests that boundary. A default pipeline built for stateless HTTP services may need a supported hook for migrations and worker rollout. An overly broad standard becomes an approval layer or bottleneck, while an overly narrow one leaves every team maintaining its own release and recovery logic. A bounded extension with an owner, tests, constraints, and review date preserves autonomy without creating an unsupported parallel system. Why Team Friction Turns Into a Scaling Tax Weaknesses in the operating model become most visible in the friction between teams. If a developer must request a namespace, ask another team for credentials, copy a pipeline, and find a production approver, the architecture may be automated while delivery remains ticket-driven. Each handoff adds queue time and loses context. Adding a portal without changing that path gives the developer one more place to check. Effective self-service completes the request, applies policy, records the change, and returns a clear support path. To see whether self-service is reducing friction, track metrics like request-to-environment time, time to first production deployment, exception rate, support demand, and failed-deployment recovery time. CNCF recommends tracking fulfillment and new-service delivery latency; DORA advises applying delivery metrics in the context of a specific service. Together, these measures show whether the workflow reduced coordination overhead or moved it to another queue. When Control Models Backfire The order-processing service example exposes two ways the control model can fail: Overly rigid standardization. A workflow designed only for stateless services forces the team to create a separate migration path, fragmenting release evidence.Unbounded local variation. Unrestricted cluster access allows identity, policy, and resource controls to drift between teams. The scalable approach pairs a narrow baseline, enforced through mechanisms like admission policies, with a documented extension path for legitimate workload-specific behavior, keeping the standard credible without turning each exception into a permanent fork. Operating Assumptions That Fail at Scale Old Assumption Why It Breaks Operating Model Replacement Healthy clusters make a workload portable Storage, identity, policy, and quota profiles differ by environment Versioned environment contract with a shared baseline One shared pipeline can serve every workload Stateful rollout and migration steps don’t fit the default sequence Core workflow with bounded, tested hooks A portal provides self-service Tickets and manual approvals remain behind the interface Workflow that provisions, enforces policy, and records evidence Local conventions remain harmless when teams own their services Metadata and controls drift across services Small enforced baseline with governed exceptions Making Cloud Complexity More Manageable To begin, you don’t need to redesign your entire platform. You can trace one representative delivery workflow and find where coordination breaks. For the order-processing service example, map the path from repository creation to production, including owners, queues, controls, evidence, and exceptions. Improvements should then be tested through adoption and outcomes such as lead time, failed-deployment recovery time, support demand, exception volume, and cost-attribution coverage. This sequence shows whether the platform is reducing operational variation for real workloads before the model expands to more teams and environments. References: Platforms for Cloud-Native Computing, CNCFResource Quotas, KubernetesStorage Classes, KubernetesAdmission Control in Kubernetes, KubernetesResource Semantic Conventions, OpenTelemetryAllocation FinOps Framework Capability, FinOps FoundationService Level Objectives, Google SRESoftware Delivery Performance Metrics, DORA This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
Before diving into solutions, it helps to understand the scale of the problem. Take a common production pattern: a customer support bot that processes 10,000 messages per day, each with a 2,000-token system prompt and a 200-token user message. Cost breakdown by the numbers: scenariomodeldaily costNo optimizationClaude Opus ($5/1M)$11.00Right model for taskClaude Haiku ($1/1M)$2.20Add prompt cachingHaiku + caching ($0.10/1M cached)$0.42Combined savings96% reduction That's not a benchmark; it's arithmetic. The techniques don't require magic; they require applying what providers already offer. Technique 1: Prompt Caching, Up to 90% Off Repeated Content The Problem Most LLM applications send the same system prompt on every request. If your system prompt is 2,000 tokens and you make 10,000 requests per day, you're paying for 20 million input tokens daily, even though the content never changes. How It Works Anthropic's prompt caching lets you mark content blocks with cache_control. The first request pays full price and writes to cache. Every subsequent request that hits the same cached content pays 10% of the normal input price. Cache entries last 5 minutes and reset on each hit. The key insight: cache the most stable content first. Your base instructions change rarely. Your few-shot examples change occasionally. Your per-request context changes every time. Structure your prompt from most stable to least stable. Python from llm_optimizer import OptimizedClient, build_cached_system_prompt import anthropic client = OptimizedClient(anthropic_client=anthropic.Anthropic()) # Build an optimally structured cached system prompt system = client.build_cached_system( base_instructions=""" You are an expert customer support agent for a SaaS company. You have deep knowledge of our product, billing, and technical issues. Always be empathetic, clear, and solution-focused. [... 1,500 more tokens of stable instructions ...] """, # ← cached after first request — 10% cost on all subsequent calls few_shot_examples=""" Example 1: Billing question → here's how to handle it Example 2: Technical issue → here's the escalation path [... 500 tokens of examples ...] """, # ← also cached separately ) # First call: pays full price, writes to cache response1 = client.complete(messages=[{"role": "user", "content": "How do I cancel?"}], system=system) # Second+ calls: system prompt served from cache at 10% cost response2 = client.complete(messages=[{"role": "user", "content": "Where's my invoice?"}], system=system) What the Library Does llm-optimizer automatically injects cache_control breakpoints at optimal positions, system prompt, few-shot examples, and long conversation history, respecting Anthropic's 4-breakpoint limit. You don't touch the API directly. Savings Calculation Shell 2,000 token system prompt × 10,000 requests/day = 20M tokens/day Without caching: 20M × $3.00/1M (Sonnet) = $60.00/day With caching: 2M × $3.00 + 18M × $0.30 = $11.40/day Savings: $48.60/day = $17,739/year Technique 2: Model Routing, 60% to 80% Off by Using the Right Model The Problem Routing every request to your best model is the most common and most expensive mistake. Claude Opus costs 5x more than Claude Haiku. For tasks that Haiku handles perfectly, such as classification, extraction, translation, and simple Q&A, you're paying a 500% premium for no benefit. The Naive Approach and Why It Fails The obvious solution is to route by keyword: if the prompt contains "classify," use Haiku; if it contains "analyze," use Sonnet. This works until it doesn't. A prompt like "Explain the constitutional implications of this clause" is 8 words. Short, simple-looking. A keyword router sees no complexity signals and routes it to Haiku. But the task requires expert-level legal reasoning. This class of error intent-heavy short prompts is the primary failure mode of heuristic routing. The Better Approach: Use a Classifier llm-optimizer solves this by optionally using Haiku itself to classify task complexity before routing. The cost is approximately 15 tokens, about $0.000015. If that classification prevents one wrong Opus call (2,000 tokens × $5/1M = $0.01), it pays for itself 666 times over. Python from llm_optimizer import OptimizedClient, Provider client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_llm_classifier=True, # uses Haiku to assess complexity — ~$0.000015/call preferred_provider=Provider.ANTHROPIC, ) # Short prompt, complex intent → correctly routed to Opus response = client.complete( messages=[{"role": "user", "content": "Explain the constitutional implications of this clause"}] ) # You can audit the routing decision from llm_optimizer import ModelRouter router = ModelRouter(enable_llm_classifier=False) print(router.explain("classify this email as spam or not")) # { # "detected_complexity": "simple", # "routed_model": "claude-haiku-4-5", # "keyword_signals_fired": {"simple": ["classify"]}, # "token_count": 8 # } Complexity Tiers A closer look at complexity: Tierexamplesdefault modelSimpleClassification, extraction, yes/no, translationClaude HaikuMediumSummarization, paraphrasing, short Q&AClaude HaikuComplexCode generation, analysis, evaluationClaude SonnetExpertLegal reasoning, research, system design, math proofsClaude Opus Shell 1,000 requests/day — mixed complexity Without routing: all → Opus ($5/1M input) 1,000 × 500 tokens = 500K tokens × $5 = $2.50/day With routing: 70% Haiku, 20% Sonnet, 10% Opus 700 × 500 × $1 + 200 × 500 × $3 + 100 × 500 × $5 = $1.05/day Savings: 58% reduction Technique 3: Prompt Optimization, 5% to 20% Off Token Count The Problem Prompts written by humans, especially in collaborative or enterprise settings, accumulate filler. Phrases like "please note that", "it is important to note that", "in order to", and "due to the fact that" add tokens without adding meaning. At scale, this is a measurable cost. What the Library Strips Python from llm_optimizer import PromptOptimizer opt = PromptOptimizer() result = opt.optimize(""" In order to complete this task, please note that you should carefully analyze the following text. It is important to note that accuracy matters. Please be aware that your response should be concise. """) print(result.optimized_text) # "To complete this task, carefully analyze the following text. # Accuracy matters. Your response should be concise." print(f"Saved {result.tokens_saved} tokens ({result.savings_pct}%)") # Saved 18 tokens (31%) What is never touched: Code blocks, factual content, user-specified phrasing. The optimizer is conservative by default. It only removes patterns with no semantic value. Conversation history trimming: In long conversations, the library keeps the last N turns and drops older context, preventing unbounded token grow. Technique 4: Batch Processing, 50% Off Non-Urgent Requests The Problem Not every LLM call needs an immediate response. Nightly report generation, document indexing, data enrichment pipelines, and offline classification jobs all of these run fine with a delay. But most teams send them as real-time requests anyway, paying full price. How Anthropic's Batch API Works Anthropic's Message Batch API processes up to 10,000 requests per batch at 50% of the normal price. Results are available within minutes to hours. The trade-off is explicit: cost for latency. Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), enable_batching=True, ) # Queue 1,000 document summaries throughout the day for doc in documents: client.queue( custom_id=doc["id"], messages=[{"role": "user", "content": f"Summarize: {doc['text']}"}], max_tokens=200, ) # Submit as one batch — 50% cheaper than 1,000 individual calls batch_id = client.submit_batch() # Poll when ready — minutes to hours depending on load results = client.poll_batch(batch_id, wait=True) for r in results: print(f"{r.custom_id}: {r.content}") When to Use It Nightly data processing pipelinesDocument indexing and enrichment Offline classification and tagging Report generation Technique 5: Document Compression to Reduce Context Before Sending The Honest Tradeoff This technique requires a direct warning: document compression is lossy. Removing content from a document to reduce token count means the model works with less information. For some tasks this is fine; for others it produces wrong answers. Use it only when: You've verified empirically that compression doesn't hurt your answer quality.You're doing rough extraction where completeness isn't required.You have re-ranking downstream (e.g., RAG pipelines). Do not use it for legal documents, compliance reviews, or any task where every sentence may be relevant. TF-IDF Extractive Compression When you do use compression, naive truncation (cutting from the end) is the worst strategy. llm-optimizer implements TF-IDF paragraph scoring where each paragraph is scored by its term overlap with your query, weighted by how unique those terms are across the document. The most relevant paragraphs fill the token budget; the rest are dropped. Python from llm_optimizer import DocumentCompressor # ⚠️ Read the accuracy warning before using in production comp = DocumentCompressor( max_tokens=4000, strategy="extractive", # TF-IDF scoring — best accuracy ) compressed, tokens_saved = comp.compress( document=long_contract, # 50,000 tokens query="payment terms and termination clauses" # focus compression here ) print(f"Compressed to {4000} tokens, saved {tokens_saved} tokens") # All compressed output includes a visible [⚠️ COMPRESSION WARNING] marker There are three strategies available: Strategy options: StrategyHow it worksbest forExtractiveTF-IDF scoring against queryWhen you have a specific querySmartKeeps first 60% + last 20%Structured documents with summariesTruncateHard cutoffWhen you need predictable behavior Technique 6: Cost Tracking That Observes Before You Optimize Why This Matters You can't optimize what you don't measure. Before applying any of the above techniques, you need to know: Which models you're actually usingWhere your token spend is goingWhether your optimizations are working Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), persist_tracking="usage.jsonl", # survives restarts ) # ... run your application ... client.print_summary() # ═══════════════════════════════════════════════════════ # LLM Cost Optimizer — Usage Summary # ═══════════════════════════════════════════════════════ # Total Requests : 1,247 # Total Cost : $0.8432 # Total Saved : $7.2180 (89.5% savings) # Cached Tokens : 8,432,000 # By Model : haiku: 891 reqs ($0.12) | sonnet: 312 ($0.58) # Optimizations : prompt_caching: 1247x | model_routing: 1247x # ═══════════════════════════════════════════════════════ The tracker records every request's tokens, cost, cached tokens, savings, latency, and which optimizations fired. Data persists to JSONL so you can analyze it across sessions or pipe it to your observability stack. Architecture: Why Not Just Use LiteLLM? The obvious question. LiteLLM is excellent and covers a lot of ground, including unified provider API, routing, cost tracking, batch processing. If you're not already using it, you should evaluate it. llm-optimizer does three things LiteLLM doesn't: Automatic cache_control injection: LiteLLM passes caching headers through but doesn't inject breakpoints at optimal positions automatically.Prompt filler stripping: LiteLLM has no token-level prompt optimization.TF-IDF document compression: LiteLLM has no query-aware document compression. The intended use is actually as a complement: llm-optimizer can sit on top of a LiteLLM setup, handling the prompt-level optimizations that LiteLLM doesn't touch. Streaming Support For user-facing applications, the library supports streaming: Python with client.stream( messages=[{"role": "user", "content": "Explain quantum entanglement"}], system="You are a physics tutor.", max_tokens=512, ) as stream: for chunk in stream: print(chunk, end="", flush=True) # Access token usage after stream completes usage = stream.usage() All optimizations, including caching, routing, prompt, and optimization, apply identically to streaming requests. Error Handling Production LLM applications need to handle rate limits and model overloads gracefully. The library handles this automatically: Python client = OptimizedClient( anthropic_client=anthropic.Anthropic(), max_retries=3, # retry on rate limit with exponential backoff retry_base_delay=1.0, # 1s, 2s, 4s ) # Rate limit (429) → retried with backoff # Model overloaded (529) → falls back to next capable model automatically # Non-retriable error → raises immediately Limitations Honest about what this doesn't do yet: Not org-scale validated: v0.4.0 is tested against ai-core Bedrock and Anthropic direct (14 live tests, all passing). Not yet run against production workloads at team scale. The pilot measures this.No async support: complete() and stream() are synchronous. Async support planned for a future release.OpenAI and Google partially tested: Anthropic and AWS Bedrock are the validated providers. OpenAI is implemented but not end-to-end tested in CI. Google Gemini streaming is not yet implemented.Token counting is approximate: The default estimator is within ~20% of the actual count. Install tiktoken for exact counts: pip install llm-optimizer[tiktoken].Pricing data can go stale: Stored in pricing.json with a version stamp. The library warns automatically if data is older than 30 days.Model allowlist: Only models listed in pricing.json can be routed to. Mythos and Fable 5 are not in the registry and cannot be called. Adding a new model requires a deliberate update to pricing.json. Python # Basic install pip install llm-optimizer # With exact token counting pip install llm-optimizer[tiktoken] # All providers pip install llm-optimizer[all] Python import anthropic from llm_optimizer import OptimizedClient client = OptimizedClient( anthropic_client=anthropic.Anthropic(), # All optimizations on by default except compression (lossy — opt-in) ) response = client.complete( messages=[{"role": "user", "content": "Your prompt here"}], system="Your system prompt here", ) client.print_summary() Links: PyPI: https://pypi.org/project/llm-optimizeGitHub: https://github.com/banerjeeso/llm-optimiz What's Next Async support (acomplete(), astream())LiteLLM adapterBudget guard: raise before a request exceeds a cost thresholdReal production benchmarks once I've run this against a live workload. Feedback welcome, especially from anyone who works with LLM APIs in production and can stress-test the routing logic or compression accuracy. Published on PyPI as llm-optimizer. MIT license. Contributions welcome.
Principal PM, Azure Cosmos DB,
Microsoft
Cloud Ops DBA,
Allspring
Award-winning Software Engineer and Architect,
OS Expert