In our Culture and Methodologies category, dive into Agile, career development, team management, and methodologies such as Waterfall, Lean, and Kanban. Whether you're looking for tips on how to integrate Scrum theory into your team's Agile practices or you need help prepping for your next interview, our resources can help set you up for success.
The Agile methodology is a project management approach that breaks larger projects into several phases. It is a process of planning, executing, and evaluating with stakeholders. Our resources provide information on processes and tools, documentation, customer collaboration, and adjustments to make when planning meetings.
There are several paths to starting a career in software development, including the more non-traditional routes that are now more accessible than ever. Whether you're interested in front-end, back-end, or full-stack development, we offer more than 10,000 resources that can help you grow your current career or *develop* a new one.
Agile, Waterfall, and Lean are just a few of the project-centric methodologies for software development that you'll find in this Zone. Whether your team is focused on goals like achieving greater speed, having well-defined project scopes, or using fewer resources, the approach you adopt will offer clear guidelines to help structure your team's work. In this Zone, you'll find resources on user stories, implementation examples, and more to help you decide which methodology is the best fit and apply it in your development practices.
Development team management involves a combination of technical leadership, project management, and the ability to grow and nurture a team. These skills have never been more important, especially with the rise of remote work both across industries and around the world. The ability to delegate decision-making is key to team engagement. Review our inventory of tutorials, interviews, and first-hand accounts of improving the team dynamic.
Can Your Team Name the Work It Already Runs With AI?
One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph
A mobile request can fail without the server-side work failing. An iOS app may time out, lose the response after a POST has reached the service, or retry after connectivity changes while the original execution is still progressing. Apple explicitly distinguishes safe retry behavior by HTTP method and notes that URLSession can retry requests in some connection-loss cases, waitsForConnectivity can also cause the system to continue a request when connectivity returns. The dangerous state is therefore not “request failed,” but “completion is unknown.” If that request starts an agent that charges an account, reserves inventory, sends a message, or invokes an MCP tool, a second submission can become a second side effect. The Retry Boundary Is the Real Transaction Boundary “Exactly once” is too strong for a workflow crossing an iPhone, HTTP, an agent runtime, an MCP server, Kafka, a database, and an external API. Kafka can provide exactly-once guarantees within defined Kafka processing boundaries, but those guarantees do not atomically include arbitrary remote tool effects. The practical target is effectively-once behavior, and retries are expected, but every effect is guarded by a stable operation identity and converges on one committed outcome. Kafka’s idempotent producer suppresses duplicate records caused by producer retries, while transactional producers can atomically publish across Kafka partitions; the producer documentation also limits idempotence guarantees to a producer session and requires read_committed consumers for end-to-end transactional visibility. The operation identity must exist before the first network attempt. An iOS client can create an operationId when an action becomes durable local intent, persist it, and reuse it across transport retries. Transport material such as a server challenge may change, but the business ID must not. The server treats (subjectId, operationId) as a uniqueness boundary and stores a canonical payload hash with it. PostgreSQL unique constraints enforce row uniqueness, while INSERT ... ON CONFLICT provides an atomic conflict path under concurrency. SQL INSERT INTO agent_operation(subject_id, operation_id, payload_hash, status) VALUES (:subject, :operationId, :payloadHash, 'ACCEPTED') ON CONFLICT (subject_id, operation_id) DO NOTHING; A conflict with the same payload hash returns the existing operation; a different hash rejects key reuse. The record should exist before agent execution starts, and the accepted response should expose the durable operation identity. Let LangGraph Resume Without Repeating Effects LangGraph persistence is useful precisely because durable execution can replay code. With a checkpointer, LangGraph saves state at super-step boundaries; if execution resumes after a failure, an affected node can run again from the beginning. Official guidance consequently requires idempotent node logic, and task results can be checkpointed so completed task work can be reused during resume instead of recomputed. Replaying from an earlier checkpoint can also re-trigger later LLM calls and API requests. A stable business operation should therefore map to a stable LangGraph thread, while every effectful tool boundary receives the same operation ID. Python config = {"configurable": {"thread_id": operation_id} result = graph.invoke( {"operation_id": operation_id, "command": command}, config ) Checkpointing reduces recomputation but does not replace downstream idempotency. A reservation can succeed before its task result is durably checkpointed. LangGraph’s functional API therefore recommends idempotent tasks because an incomplete task can execute again during resume. Python @task def reserve_inventory(operation_id, sku, quantity): return mcp.call_tool("reserve_inventory", { "operationId": operation_id, "sku": sku, "quantity": quantity }) The significant property in this snippet is not the decorator. The important part is that the business identity crosses the graph boundary and reaches the tool implementation. A downstream inventory service can then use that identity to return a previously committed reservation rather than creating another one. MCP Tasks Are Durable Handles, Not Deduplication Keys The current MCP Tasks design is especially relevant to long-running agent tools. In the July 28, 2026 protocol revision, Tasks moved into the io.modelcontextprotocol/tasks extension. A server can return a durable task handle, and the client can poll with tasks/get, provide input with tasks/update, or request cancellation with tasks/cancel. The task is durably created before its handle is returned, which allows polling after a disconnect. That durability solves result retrieval after task creation, but it does not by itself deduplicate the request that creates the task. The task ID is server-generated. If the server creates task A, the response disappears, and the original tools/call is sent again, a naïve implementation can create task B. Therefore, the business operationId must be part of the tool arguments or equivalent application metadata, and task creation must first look up an existing operation. This follows directly from MCP’s server-generated task-ID model combined with retry ambiguity at the HTTP boundary. The MCP server can return an existing task handle for the same authenticated subject, operation ID, and payload hash, and later return the stored terminal result. Cancellation should also be idempotent because MCP defines it as cooperative rather than a guarantee that underlying work stops immediately. Keep Kafka Guarantees Inside Kafka Kafka is most valuable after the operation has been claimed. A database transaction can persist operation state with an outbox row carrying the same ID. Kafka producer idempotence protects against duplicates caused by producer retries, while consumers can still use the operation ID for application-level deduplication. Kafka transactions can atomically cover Kafka writes, but they do not extend over an MCP server or payment API. The event contract should preserve causality rather than inventing a new identity at each hop. JSON { "operationId": "8E7B6D9E-...", "type": "AgentToolCompleted", "tool": "reserve_inventory", "status": "SUCCEEDED" } A consumer can enforce uniqueness on (consumerName, operationId, eventType) or make the state transition conditional. Kafka delivery guarantees and application idempotency then reinforce each other instead of being treated as interchangeable. Bind Retry Identity to App Attest Without Blocking Legitimate Retries App Attest addresses a different failure mode: whether a request comes from a legitimate app instance and whether signed request material has been replayed or altered. Apple’s current guidance uses a server-provided challenge for assertions and requires the server to validate a strictly increasing assertion counter; that counter is specifically an anti-replay signal. Assertions are generated locally on the device after key attestation. The App Attest assertion must not become the business idempotency token. A legitimate retry should obtain fresh challenge material and generate a fresh assertion while retaining the original operation ID. The data hashed for the assertion can bind the server challenge, operation ID, and canonical payload hash together. Swift let payloadHash = SHA256.hash(data: body) let clientData = challenge + operationID.data + Data(payloadHash) let clientDataHash = Data(SHA256.hash(data: clientData)) let assertion = try await service.generateAssertion( keyID, clientDataHash: clientDataHash ) Apple recommends server-controlled challenges, server-side validation, and assertion-counter tracking as assertions are generated on demand without a round trip to Apple’s servers. The server verifies App Attest, checks that the challenge binds the operation ID and payload, then performs the idempotency lookup. A fresh assertion can retry the same operation; a replayed assertion fails anti-replay validation; an altered payload fails the hash check. Effectively-Once Behavior Is a Composition Property Reliable agent execution does not come from asking iOS to retry less often or from labeling a Kafka pipeline “exactly once.” It comes from carrying one durable business identity across every retry and every boundary, claiming that identity atomically before execution, making LangGraph effects idempotent under resume, using MCP Tasks as durable result handles rather than creation-time deduplication keys, restricting Kafka’s exactly-once guarantees to Kafka’s transactional domain, and using App Attest to prove request integrity without confusing anti-replay state with business deduplication. When those boundaries align, a lost mobile response can cause another HTTP attempt, another graph invocation, or another poll, but it does not cause another business effect. That is the operational meaning of effectively once.
Most enterprise AI post-mortems do not blame the model. They blame the storage tier that starved the accelerators, the identity policy that over-granted access, the cost model that ignored egress, the forecast that leaked future data, or the region that failed and took a business process with it. The hard part of production AI was never intelligence. It was the engineering discipline around it. This article distills the architectural patterns that decide whether a cloud AI system is trustworthy at scale, spanning infrastructure, identity, cost, operations, the applied domains, low-code assembly, platform selection, and multi-cloud resilience. It is written for engineers who have to keep these systems running, not for a keynote. Infrastructure: The Interconnect Is the Bottleneck Distributed training is a systems problem before it is a machine learning problem. When a job spans many graphics processing units (GPUs), the fabric connecting them (e.g., NVLink within a node, InfiniBand, or a vendor fabric across nodes) frequently caps throughput more than raw compute does. Accelerators wired through an ordinary network idle while they wait to synchronize gradients. Storage is the symmetric constraint. If the file system cannot deliver data at the rate the accelerators consume it, utilization collapses. The pattern is a tiered design: Hot tier: parallel or block storage feeding active training at high input/output operations per second (IOPS).Warm tier: recent data staged for quick promotion.Durable lake: object storage providing petabyte-scale durability, partitioned and lifecycle-managed underneath. Two cost drivers hide from the pricing page: data egress (moving data across regions or out of a provider) and idle warm capacity. Optimizing only the advertised compute line item guarantees a surprise on the invoice. Identity Is the Perimeter In a service-to-service AI architecture, the network perimeter is gone; identity is the boundary. A zero-trust posture, where every request authenticates and receives least privilege, contains the blast radius when a component is compromised. Across providers, identity federation is the load-bearing pattern: a principal authenticates once and is recognized everywhere, so access is granted and revoked centrally instead of reconciled across three identity systems. Policy must travel with the workload; a rule enforced on one cloud and forgotten on another is not a policy. Model authorization is the emerging frontier. As models call tools and take actions, the question moves from who can query this model to what may this model do on a user's behalf. Least privilege applied to an autonomous agent is the boundary between useful and unbounded. Cost and Operations Are a Control Loop Cost management is not a spreadsheet; it is automation. Consistent resource tagging across every cloud is the prerequisite for attribution. On top sit budgets, alerts, and automated remediation that throttles runaway spend before it escalates. Site reliability engineering (SRE) supplies measurable targets. For AI workloads, the golden signals extend beyond latency and errors to accelerator utilization, queue depth, and prediction quality. A model can be fully available and quietly wrong, so define a service level objective (SLO) for output quality, not just uptime. Three techniques earn their complexity: Spot or preemptible capacity plus checkpointing cuts training cost sharply when jobs resume cleanly after reclamation.Predictive scaling anticipates load instead of reacting to it.LLM inference optimization becomes architectural: batch requests, cache frequent responses, route easy queries to smaller models, reserve the expensive model for queries that need it. The Applied Domains Share a Spine, Differ in Physics Vision is byte-heavy. High-resolution images and video streams make the data and network layers dominant. For real-time video, decouple frame capture from analysis and sample frames rather than processing every one. Critically, a business-rule layer, never the model alone, owns consequential decisions. Every extraction should carry a confidence score used as a routing gate: Python def route_extraction(field, threshold=0.90): if field["confidence"] >= threshold: return "auto_process" return "human_review" Language is byte-light but semantically treacherous, and because it replies directly to users, errors are visible. The defining risk of generative systems is hallucination. The strongest architectural defense is retrieval grounding, forcing answers from verified sources with citations: Python def answer(question, knowledge_base): passages = knowledge_base.search(question, top_k=3) context = "\n".join(p.text for p in passages) prompt = f"Answer using ONLY this context.\n{context}\n\nQ: {question}" return model.generate(prompt), [p.source for p in passages] Forecasting is defined by time order. You cannot shuffle a time series into random splits, and the most common failure is data leakage, using information unavailable at prediction time. Test on a fair, time-ordered holdout, and always emit a prediction interval; a point forecast that hides its uncertainty invites overconfident decisions. No-Code and Low-Code: Governed or Ungoverned No-code and low-code platforms collapse build cost from a scoped project to an afternoon, which is why adoption is exploding. The symmetric risk is sprawl: hundreds of ungoverned flows handling sensitive data, owned by no one. Govern with guardrails, not gates. Restrict which connectors and data sources are permitted, assign an owner and an SLO to every production flow, then let builders move freely inside the boundary. The goal is to make the safe path the easy path. Platform Selection Without Self-Deception Vendors all claim to be fastest, cheapest, and most reliable. Benchmark to replace claims with evidence: Latency: report percentiles (p95, p99), never averages that hide the slow tail.Quality: measure on your own representative data, not a public leaderboard.Cost: model total cost of ownership, including transfer, storage, idle capacity, operations, and migration, not the headline compute rate.Reliability: verify the platform meets your recovery time objective (RTO) and recovery point objective (RPO). Combine dimensions in a weighted scorecard whose weights are fixed before scores are seen. Adjusting weights afterward to crown a favorite converts analysis into rationalization. Multi-Cloud Resilience: Design for the Day a Cloud Fails For systems a business cannot lose, a single provider is a gamble. Multi-cloud resilience deliberately places critical workloads so no single provider failure takes the business down, applied only where the cost of failure exceeds the cost of prevention. Predict rather than react. Combine leading signals into a health score and fail over proactively: Python def health_score(latency_ms, error_rate, saturation): latency_factor = max(0, 1 - (latency_ms / 1000)) error_factor = max(0, 1 - (error_rate / 0.05)) saturation_factor = max(0, 1 - saturation) return round(0.4*latency_factor + 0.4*error_factor + 0.2*saturation_factor, 3) Kubernetes makes workloads portable; data replication (with the consistency-versus-availability trade-off decided per workload) keeps data ready on the other side; and a portable foundation of federated identity, uniform policy, and centralized monitoring makes failover routine rather than heroic. The discipline that separates real resilience from a slide deck is rehearsing failure on purpose. An untested failover path is a promise, not a capability. The Judgment Layer Across every layer, value came not from the most powerful component but from the judgment applied to it: matching effort to problem difficulty, keeping humans on consequential decisions, measuring before deciding, building governance in early, and designing for change. Tools will churn; foundation models will make today's designs look quaint. That is precisely why principles outlast product knowledge. The scarce resource in enterprise AI was never intelligence. It was judgment, and judgment does not ship from the cloud.
TL;DR: A Déjà-Vu? AI adoption seems to be scaling: 37% of respondents in McKinsey’s 2026 survey report an EBIT effect from AI, and Gartner finds that 22% of organizations have scaled it across business units. Now, Agile practitioners have seen this combination before, as AI transformations and Agile transformations rhyme. Five classic failure patterns from Agile transformation adventures are back under new names: mandates from above, licenses mistaken for training, greenfield showcases, parachuted consultants, and promised payroll savings dressed up as strategy. They share one condition: organizations make AI decisions at organizational scale without leaving inspectable evidence at the workflow level in the trenches. And for good measure, let us throw in ignoring culture and excluding most of the organization’s people in the process. History Does Not Repeat Itself, but AI Transformations and Agile Transformations Do Rhyme AI transformations in large organizations are scaling, individual productivity is up, leaders still plan to increase spending, and yet enterprise financial impact remains limited: McKinsey's 2026 State of AI survey (1,719 respondents, fieldwork May 4 to June 8, 2026) puts numbers on three of the four: 44 percent of respondents say AI is scaling across their enterprise, up from 38 percent a year earlier; 80 percent of those who use AI report improved individual productivity; and 37 percent attribute any EBIT impact to AI at all, with the "AI high performer" group flat at about 6 percent.Gartner's September 2026 survey of 1,303 respondents from organizations with at least $50 million in annual revenue supplies the spending picture: 85 percent of functional leaders plan to increase AI spending in 2026, 22 percent of organizations have scaled AI across multiple business units or adopted an AI-first approach, and 11 percent do not know what their function spent on AI in 2025. Something is happening, and something is also not translating. Agile practitioners have seen that combination before. "History does not repeat itself, but it rhymes," a line widely attributed to Mark Twain despite no evidence that he said it; the attribution to Twain dates back to 1970. The attribution is shaky; nevertheless, the observation holds. I wrote the Scrum Anti-Patterns Guide about what organizations do to Agile when they adopt it from the top down. The same organizations are now doing the same things to AI, with a new generation of leaders who consider the Agile years ancient history, and five rhymes stand out. The Five Rhymes of AI Transformations Rhyme 1: The Mandate From Above IBM's June 2026 study of 2,000 C-level technology executives found that 80% reported CEO-driven AI transformation mandates, and 77% said adoption is already outpacing their governance capabilities. The most public example is Shopify. In a late-March 2025 memo that he later posted on X, CEO Tobi Lütke told the company that "reflexive AI usage is now a baseline expectation at Shopify" and that, before asking for more headcount, teams "must demonstrate why they cannot get what they want done using AI," as Tom's Hardware and TechCrunch reported. Whether that works at Shopify, I cannot judge from the outside. What I can judge is the predictable risk when that kind of memo lands in an organization where governance is already falling behind: visible compliance and invisible workarounds. Agile practitioners remember the memo announcing "we are now an agile organization" and the Sprint Reviews that followed, which were ignored by everyone who could change a decision. Rhyme 2: The Belief That This Time Training Is Optional The Agile version bought a two-day certification class and called it a transformation. Often, AI transformations skip even that: buy Copilot or ChatGPT Enterprise licenses, send an email, done. The tool is "intuitive," so the reasoning goes; it is sold as the classic example of learning by applying. Lütke's own memo contradicts this, noting that "using AI well is a skill that needs to be carefully learned." The McKinsey gap between 80% reporting individual productivity gains and 37% reporting any EBIT impact shows why individual productivity is a poor proxy for organizational change. Individuals may get more productive, whatever that means in this context, which does not imply that the organization has changed at the same time. The same survey shows where the difference lies: nearly three-quarters of high performers report fundamentally redesigning workflows because of AI, against one-quarter of everyone else. Deloitte's June 2026 pulse check of nearly 3,700 professionals found that 48% were adding AI without redesigning workflows or roles, and only 12% were redesigning workflows or roles at scale. The divide runs between organizations that change the nature of work and those that bolt AI onto whatever structure they have. Rhyme 3: The Greenfield Showcase Every transformation needs a success story for the board, so a team with no dependencies on the legacy systems, no regulatory exposure, and no operational duty builds something impressive. The Agile version was the "pilot team" in the innovation lab with the fancy toys. The AI version is the internal chatbot that answers HR policy questions and was presented at the town hall as evidence of AI's great potential. Exploration detached from production constraints is useful. The anti-pattern is mistaking evidence that something can be built for evidence that the organization has created value. BCG's 2025 survey of 1,250 senior executives found that 70% of AI's potential value sits in core business functions such as sales and marketing, manufacturing, supply chain, and pricing, which is where the showcase usually never goes, due to the "unsexiness" of the use cases. Rhyme 4: The Consultancy That Sets It Up for You In come the slide decks, the "AI transformation office," and the currently fashionable forward-deployed engineers. The role name dates back to Palantir in the early 2010s; the practice is far older. Thomas Otter, who spent years at SAP, notes that "early chunks of SAP R/1 were built at ICI and John Deere," decades before anyone called the practice forward deployment. I do not consider the practice an anti-pattern. Engineers who join the organization, learn its culture, and stay long enough to hand over applications built on understanding are legitimate. The anti-pattern is the parachute version: the engineers arrive, do the tactical technical work, and leave, and the organization is now running systems it cannot explain. Agile had the consultancy-staffed transformation office that left when the budget line ended. Rhyme 5: The Cost Story Ask most leadership teams why the organization adopts AI, and you get a story about new business, better products, and faster learning. Ask what the business case they signed off actually contains, and you find payroll. Consultancies, in my observation, sell AI as they sold offshoring: a way to remove people who do repetitive work. Cost reduction, as such, is not the anti-pattern; however, turning it into the transformation objective is. About 80% of McKinsey's high performers, and everyone else, pursue efficiency, but most high performers also pursue growth or innovation, thereby distinguishing the two approaches. Klarna ran the other experiment in public. After claiming its AI assistant did the work of 700 customer service agents, CEO Sebastian Siemiatkowski told Bloomberg in May 2025, as CX Dive reported, that "cost unfortunately seems to have been a too predominant evaluation factor when organizing this; what you end up having is lower quality," and started hiring humans again. McKinsey's respondents have noticed which story their leadership actually believes: 39% now expect AI-related job cuts, up from 32% a year earlier. Agile had the same split. The board heard "faster and cheaper"; the teams heard "better products"; and when the two stories collided, the teams lost. What the Five Rhymes of AI Transformations Share Each AI transformation rhyme has a visibility problem. Leadership can see the headcount numbers perfectly well and still optimize them; a consultancy dependency is a capability-transfer problem, while a mandate is an authority and incentive problem. What the five have in common sits one level down. In each case, the organization makes its AI decisions at organizational scale (a mandate, a license contract, a showcase budget, a vendor engagement, a business case) without leaving inspectable evidence at the workflow scale. Too often, nobody can show, for a specific workflow, who decided that AI would do this work, on what terms, under what cost constraints, who checked it, and with what result. Visibility is the symptom, and missing evidence is the condition. Scrum already had low-tech answers to similar problems: an ordered Product Backlog, an explicit Definition of Done, and a recurring Retrospective. None of them needed a platform, and none of them made leadership act. What they did was let a team generate evidence about the system it worked inside. The A3 Delegation System borrows that design principle: make consequential decisions visible before buying another layer of tooling to manage them. Six stages (Decide, Route, Hand Over, Define Done, Inspect, Roll Up), seven artifacts, and no software beyond the AI the team already uses. It is an operating discipline for AI delegation, one workflow at a time, and the evidence is a byproduct of doing the work. Where Each Rhyme Meets a Countermeasure Let us come back to the five "rhymes" and how the A3 Delegation System can mitigate these AI transformation issues: The AI Workflow Inventory makes the license fallacy and the showcase inspectable: Before anything else, the team lists the workflows it already hands to AI, each with an owner. It takes an hour, and the assumption that "people will figure it out" collapses once the list shows what they figured out. You may discover personal AI habits that were never treated as organizational workflows at all, including some touching sensitive data. The inventory also refuses the greenfield showcase by construction. Only existing workflows with a named owner enter it. The A3 Framework decision and the Routing Policy put a countermeasure against the mandate: For each inventory entry, the team decides Assist (AI drafts, you decide), Automate (delegate execution, not responsibility), or Avoid (the cost of failure is trust). Then it routes the work to a model tier by stakes and cost. Leadership can set the boundaries: approved tools, prohibited data, risk limits, or spending constraints. It cannot make the workflow-specific delegation decision from a company-wide memo; the people who know the work can do so in minutes per entry, and the decision is then on paper for leadership to read. Routing is also where the token bill becomes a decision, and precision matters here: while the price per token keeps falling, the cost of operating AI keeps rising, because cheaper tokens invite longer, more autonomous workflows that consume far more of them. Gartner predicted in August 2026 that inference costs per agentic workflow will rise more than fivefold through 2028; its analyst, Will Sommer, said, "Product leaders cannot rely on more efficient token economics to rationalize AI costs." That is the economic problem the Routing Policy addresses at the workflow level: expensive intelligence is a deliberate choice, never a default. The A3 Handoff Canvas and the AI Definition of Done make the parachute inspectable: Six fields (task split, inputs, outputs, validation, failure response, records) and a one-page quality standard per task class. Here is the ownership test for anything a consultancy or a forward-deployed engineer built: can the team fill in these two documents for the system without calling the vendor? If yes, the team owns the delegation, whoever set it up. If no, the organization is renting understanding, and the rent comes due when the engineers leave. The Delegation Audit asks one question, and it is not the cost question: Monthly or every other Sprint, 45 to 60 minutes, four checks: output and source drift, model fit, reversibility, and category creep (Assist work that quietly became unreviewed Automate). Each finding gets an owner and a decision: change the A3 category, change the tier, update the AI Definition of Done, fix the stop rule, or retire the delegation. The Audit asks whether this delegation is still sound. It does not ask what the freed capacity produced. Roll Up, the last stage, compiles what the audits show for those who ask. Whether what they show is worth paying for is a leadership decision, and it sits outside the A3 Delegation System. The system produces evidence for the value conversation, but it does not own the value decision. The AI Working Agreement wraps the other six: It records the team's rules on data, disclosure, responsibility, and review, and it is the document the team hands upward when leadership asks what "AI adoption" looks like here. It is a page that beats a slide on every occasion. Where the A3 Delegation System Stops A skeptical reader, and my readers have watched frameworks overclaim for twenty years, will ask the obvious question: am I criticizing consultancies for selling transformation frameworks and then selling my own? A3 is not an AI-transformation methodology. It is an evidence-generating delegation discipline. It cannot make leadership respond to the evidence. It can make ignoring the evidence harder. It does not determine why your organization adopts AI, nor does it replace a strategy, a portfolio decision, or the conversation about what happens to the people whose repetitive work disappears. What it does is make the absence of those decisions visible within weeks, team by team, in writing, for the price of a few hours. There is a second limit: A3 can tell you whether AI should do a piece of work, which model, what it needs, what acceptable output means, whether the delegation has drifted, who owns it, and what it is allowed to cost. It does not tell you whether the workflow should exist. Suppose the system reduces a weekly reporting workflow from 4 hours to 40 minutes, with excellent output and impeccable governance. The question that remains is why the organization produces that report at all. The A3 Delegation system can prevent undisciplined delegation. It cannot prevent an organization from competently automating work that should have disappeared. The likely next development step of the A3 system is a single field on the AI Workflow Inventory, not another canvas: what changes if this workflow works? I have not added it yet. The team should expect the visibility A3 produces to be unwelcome. A team that runs the inventory, the decisions, and the Audit inside a mandate-driven transformation produces evidence the organization may refuse to absorb. I have watched organizations refuse the evidence their Retrospectives produced for years, and the refusal told the teams more about the transformation than any all-hands did. If your leadership will not read a one-page working agreement and a monthly audit log, you have learned what the AI transformation is for. Conclusion AI transformations may repeat many of the mistakes of Agile transformations. The A3 Delegation System does not prevent organizations from making them. However, it gives teams a simple way to make some of them visible before they become expensive: Count how many of the five rhymes are playing in your organization right now. Respect yourself and be honest while aggregating those. Then put a document against one of them.
For a while now, an idea has been gaining traction: with artificial intelligence, anyone can build an app without knowing how to code. The promise is incredibly seductive: with just a few prompts, we can generate code and instantly turn an idea into a product. It’s no coincidence that this vision took hold so quickly and gave rise to services like Lovable.dev, Blot.new, v0, and others. Every new technological evolution that narrows the gap between an idea and software tends to make developers' work look like an arcane ritual waiting to be dismantled by a simpler formula. There is something deeply familiar about all of this. Something that reminds me of a line from a song many of us grew up with, with its slightly childish enthusiasm: everybody wants to be a cat! Today, it seems like everybody wants to be a dev. The real question is whether everybody can be a dev. Joking aside, the attempt to make programming accessible to everyone is an old story, one that certainly didn't start with the advent of AI. A World Without Developers The idea that technological evolution can democratize programming is a recurring theme in the history of computer science. Every time a new abstraction emerges, someone proclaims that the job of writing software is about to become obsolete. Sometimes the promise is alluring; other times, it's just a clever way to sell a new tool. Yet, the core premise remains the same: if computers get closer and closer to understanding human language, then perhaps those seemingly indispensable technical skills are no longer needed. I’ve seen this pattern repeat itself multiple times. A demo takes half an hour to build, a prototype seems to work, and suddenly, the idea of building an app feels within anyone's reach. It’s fascinating, but the problem is that what you see at the beginning is often just the surface-level work: the interface, the screens, the user flow. What remains hidden is the hardest part, the work that determines whether the application will actually hold up when it goes live in production. Promises of the Past Looking back, the history of computing is full of waves that announced the end of developers. These waves didn't eliminate the profession; they transformed it. And that transformation should serve as a lesson to help us understand exactly what is happening today with AI. COBOL and Quasi-Natural Language In the 1950s and '60s, when programming meant working directly with hardware, assembly, and mathematical logic, COBOL was born. Its goal was clear: to bring programming closer to everyday language so that business managers could express rules more naturally. The idea was that a manager could describe a process's logic in English, and the computer would handle the rest. That promise didn't pan out the way people imagined. We didn't end up in a world where everyone wrote software the way they wrote letters. Instead, a vast ecosystem of specialists emerged who knew how to use that language rigorously, efficiently, and sustainably. In other words, the barrier to computer programming didn't disappear; it shifted. SQL and Fourth-Generation Languages In the 1970s and '80s, with the rise of databases, fourth-generation languages (4GLs) and SQL arrived. The concept was simple: instead of explaining every procedural step to the computer, you just had to declare what you wanted to achieve. In theory, a non-technical user could query a database and get a result. In practice, however, writing correct queries, managing complex schemas, and understanding how data connects requires a much deeper level of reasoning than it appears at first glance. As a result, the language became more accessible, but the need for expertise didn't vanish. If anything, it became more specialized. New roles, new professionals, and new problems to manage emerged. The computer kept doing its part, but the ability to think in a structured and precise way remained essential. HyperCard and the Dream of Democratic Programming In the 1980s, HyperCard truly felt like a revolution. With a simple card-based metaphor and a highly readable language, it promised to put software creation into the hands of anyone. Teachers, artists, students, everyday people: everyone could build interactive apps, games, or educational tools without a deep background in computer science. It was a captivating dream, and it partially worked. HyperCard became a massive tool for creativity, inspiring the evolution of the Web and early forms of digital collaboration. But when it came to building something truly robust, scalable, or professional, the system hit technical and organizational walls. The democratization of programming remained a promise that looked much better on paper than in industrial reality. CASE Tools and the Dream of Guided Software Between the 1980s and '90s, another promise attempted to make development more accessible: CASE (Computer-Aided Software Engineering) tools. The idea was simple: if a system could help map out an application's flow, generate pieces of code, and guide the design process, then even non-experts could build software in a more structured way. In practice, however, CASE tools didn't eliminate the need for expertise. They simplified certain steps, especially during the analysis and design phases, but they didn't replace the work of someone who could see the bigger picture. Visual Basic and the Drag-and-Drop Era In the 1990s, Visual Basic turned creating desktop applications into a near drag-and-drop affair. It was the modern version of the dream: just draw a window, drop a button, and tell the computer what to do when that button was clicked. To many, it felt like the moment the barrier between user and developer would dissolve once and for all. To an extent, it did. But it also opened up a different narrative. Many applications built this way were fast to construct but incredibly fragile without a solid architecture backing them. As a system grows, knowing how to make a window pop up isn't enough anymore. You need to know how to define architecture, manipulate state, handle errors, maintain code, and ensure quality. The initial simplicity didn't eliminate the need for technical skills; it just pushed the problem down the road to a later stage of the product lifecycle. No-Code and the Myth of the Citizen Developer In the 2010s, with the explosion of the web and APIs, no-code and low-code carried the torch of a new promise. Platforms like Bubble, Webflow, or Zapier suggested that even those who couldn't code could build personal tools, automations, or full-fledged applications. This birthed the idea of the "citizen developer", a business professional who creates their own solution without going through IT. Here too, reality proved more nuanced. These platforms are phenomenal for prototyping, automating minor processes, and creating straightforward experiences. But the moment a project requires complex integrations, security, scalability, or nontrivial logic, you hit a wall. People can navigate the system, but they don't always have full control over it. AI Is Not the End of Programming Today, AI is making it easier to build the first version of an application, but it isn't eliminating the developer's job. What's changing is how they work, as I mentioned in another article: less time spent writing code, and more time dedicated to understanding the problem, defining requirements, guiding the tools, and verifying the output. The skills that matter now aren't just about syntax; they are about choosing the right solution for the context, anticipating errors, and knowing if a system will truly hold up. Anyone using AI can churn out software faster, but they can't always tell if the result is correct, secure, or sustainable. And that's exactly where the difference lies. A beginner can make something simple work. An experienced developer also knows how to explain why the system holds together, where it might break, and how to prevent it. The history of computing has already taught us that while every new technology takes a step forward in making software development more accessible, it never eliminates the need for specific expertise. The type of skill required changes, but its importance never does.
Modern organizations operate under a persistent tension: they must both discover the future and deliver the present. These two modes of work — exploration and exploitation — are fundamentally different in goals, incentives, risk tolerance, and execution style. Yet both are essential for long-term success. The challenge is that most systems, teams, and incentives are not naturally designed to handle both well at the same time. Organizations that fail to balance these modes tend to collapse in predictable ways. Some become overly focused on optimization, refining existing products while missing shifts in technology or user behavior. Others become addicted to experimentation, constantly building new ideas without the discipline required to scale or sustain them. Sustainable companies learn to do both deliberately. What Exploration and Exploitation Really Mean Exploration is the process of discovering new opportunities. This includes new products, technologies, user behaviors, and markets. It is inherently uncertain. Success is measured not by stability or scale, but by learning. Exploration favors speed over perfection, and reversibility over permanence. It thrives in environments where failure is expected and inexpensive. Exploitation, on the other hand, is about scaling what already works. It is the phase where systems are hardened, performance is optimized, reliability is improved, and operational excellence becomes the focus. Exploitation favors predictability, consistency, and efficiency. It assumes that the underlying idea has already been validated and is worth investing in for long-term use. The key insight is that neither mode is superior. They are complementary, and the health of an organization depends on how well it can transition between them. Why Balance Is Difficult The difficulty arises because exploration and exploitation demand opposing behaviors. Exploration rewards experimentation, tolerance for ambiguity, and willingness to discard work. Exploitation rewards discipline, stability, and careful optimization. Teams often struggle because they try to apply the same engineering standards to both modes. If everything is treated as production-grade from day one, exploration slows down and innovation dies. If everything is treated as experimental, systems become unstable and difficult to maintain. Organizations that fail in this balance typically fall into one of two traps: Over-exploitation: Companies focus on improving existing systems until they become rigid and blind to change.Over-exploration: Companies generate many ideas but fail to turn them into reliable, scalable systems. The most successful organizations maintain what is often called organizational ambidexterity: the ability to explore and exploit simultaneously without letting one destroy the other. How Engineers Enable Exploration Engineers play a central role in making exploration safe and productive. During exploration, the goal is to maximize learning per unit of effort. This requires different design choices than those used in production systems. Key engineering principles for exploration include: 1. Optimize for Speed and Learning Early systems should prioritize rapid iteration. The goal is not correctness at scale, but fast validation of assumptions. 2. Keep Systems Lightweight and Reversible Exploration work should be easy to discard or rewrite. Heavy architecture decisions too early can slow learning and lock teams into premature constraints. 3. Use Isolation Mechanisms Feature flags, sandbox environments, and isolated services allow experimentation without risking core systems. 4. Limit Blast Radius Experimental work should be contained so failures do not cascade into production instability. 5. Treat Code as Temporary Exploration code should be written with the expectation that it may be replaced or removed entirely. The most important mindset shift is accepting that exploration is about learning, not longevity. How Engineers Enable Exploitation Once a direction is validated, the focus shifts from learning to scaling. This is where engineering discipline becomes critical. Exploitation requires different priorities: 1. Raise Quality Standards Reliability, performance, security, and maintainability become central concerns. Systems must now behave predictably under real-world conditions. 2. Simplify and Stabilize Complex experimental structures should be reduced or refactored into stable designs. What was once acceptable for speed may become unnecessary overhead. 3. Pay Down Technical Debt Shortcuts taken during exploration must be revisited. Debt that is ignored compounds and eventually slows down future progress. 4. Standardize and Automate As systems scale, consistency becomes critical. Automation, observability, and standardized patterns reduce operational burden. 5. Design for Longevity Exploitation systems should assume long-term operation. This means careful attention to interfaces, dependencies, and evolution paths. The transition from exploration to exploitation is one of the most important engineering inflection points. Many systems fail not because the idea was wrong, but because the transition was never properly completed. The Core Engineering Discipline At the center of this balance is a deceptively simple question: Are we exploring or exploiting right now? This question matters because it determines everything else — architecture, testing strategy, deployment rigor, and even communication style. When this intent is clear: Engineers can apply the right level of rigorTeams can consciously accept or reject technical debtTrade-offs become explicit instead of accidentalSystems evolve without losing coherence When this intent is unclear, teams often apply mismatched expectations. Experimental systems become over-engineered too early, or production systems remain under-documented and fragile. Clarity of intent is what enables disciplined flexibility. The Role of Product and Engineering Together The balance between exploration and exploitation cannot be managed by engineers alone. It requires close alignment with product thinking. Be Explicit About Intent Teams should clearly label work as exploratory or exploitative. This avoids confusion about expectations and quality standards. Define Success Appropriately Exploration should be evaluated based on learning outcomes: validated hypotheses, user insights, or technical feasibility. Exploitation should be evaluated based on reliability, efficiency, and scalability. Manage Technical Debt Intentionally Speed during exploration often introduces debt. The key is not to avoid it, but to make it visible and intentional, with a plan for when it will be addressed. Protect Capacity for Both Modes Healthy organizations allocate time for experimentation, operational improvement, and debt reduction. Without this balance, either innovation or reliability suffers. Make Transitions Explicit When an experiment proves successful, it should be consciously transitioned into a production system. Likewise, failed experiments should be retired decisively to avoid long-term clutter. The Bottom Line The long-term success of engineering organizations depends on their ability to explore new possibilities while reliably exploiting proven systems. This balance is not accidental — it must be designed. When exploration is clearly separated from exploitation, teams can move quickly without fear and scale confidently without chaos. Technical debt becomes a managed tool rather than an unintended burden. Systems evolve in a controlled way rather than accumulating uncontrolled complexity. Ultimately, the goal is not to choose between exploration and exploitation, but to build the discipline and systems that allow both to coexist. That is what enables continuous innovation while still delivering dependable value at scale.
Picture a business-critical SQL query crawling for seven hours. Nearly a full workday. The system keeps grinding through data, the business keeps losing time and money, and users are stuck waiting. Then a performance engineer steps in. After a few hours of careful analysis and a handful of precise code changes, the same query finishes in two minutes. Situations like this are not unusual in performance engineering. Turning hours into minutes is exactly the kind of work that makes this discipline valuable. In modern DevOps environments, where systems are deployed continuously and workloads change quickly, this type of work becomes part of everyday engineering practice. Who Are Performance Engineers? In simple terms, a performance engineer (PE) is responsible for making IT systems run better: faster, more reliably, and more efficiently. Behind this simple definition, however, lies a complex and multifaceted discipline. The bottleneck can appear almost anywhere in the stack: in application code, database configuration, network communication, or even the underlying hardware. And sometimes the bottleneck is not in the database or the application, but in the operating system. When systems handle thousands of concurrent network connections, limits may appear in the OS network stack or in kernel parameters. There are well-known cases in the history of database systems where the same database engine showed dramatically different performance on different operating systems, such as Windows, FreeBSD, or Linux. These differences were often caused by variations in filesystem behavior, networking stacks, and kernel-level I/O scheduling rather than by the database software itself. Once the root cause is identified, the performance engineer must understand the underlying mechanism behind it and propose an effective solution. Sometimes it means tuning the configuration. Sometimes it means rewriting a query or changing application behavior. Sometimes it points to a deeper architectural flaw that was hidden until the load exposed it. That is why the job often feels less like optimization in the abstract and more like investigation under pressure. And it sits somewhere between development, systems administration, and deep system analysis. It is important to note that many performance engineering tasks overlap with the responsibilities of a database administrator. Query optimization, lock analysis, and tuning parameters such as WAL settings are traditionally part of a DBA’s role. The difference is that a performance engineer usually operates at a broader level. They analyze the performance of the entire system, including the application, database, operating system, network communication, and underlying hardware. While a DBA focuses on a specific database platform, a performance engineer evaluates the system as a whole production pipeline. In cloud environments, this broader view may also touch tools when workload, container, or configuration findings overlap with production behavior. A Practical Example Performance problems rarely have a simple playbook. The same symptom can appear in different environments while the underlying cause is completely different. Engineers, therefore, rely on ongoing microlearning, hands-on experimentation, and careful analysis of system metrics to expand their troubleshooting knowledge as part of everyday engineering practice. One example illustrates how these investigations unfold. A client was migrating data from Oracle to PostgreSQL. The migration process relied on massive parallel data loading using COPY. At first, everything seemed to work normally, but eventually the process slowed dramatically. The investigation showed that the bottleneck was due to WAL (Write-Ahead Log) writes. In PostgreSQL, every change generates a WAL record that must be flushed to disk before the transaction commits. This mechanism guarantees durability and crash recovery, but under heavy write workloads, it can become a limiting factor. Initially, the team suspected that disk throughput was the problem. Developers even suggested a patch intended to speed up WAL writing. The patch did not improve performance. The client’s internal specialists were also unable to find a clear explanation. At that point, the performance team started analyzing the system in more detail. They noticed that many database sessions were waiting on the PostgreSQL wait event LWLock:WALInsert. That observation changed the direction of the investigation. It meant the system was not actually saturated by CPU or disk throughput. Instead, multiple processes were competing for internal synchronization while inserting WAL records. The migration workload involved hundreds of concurrent COPY operations. Each process attempted to reserve space in WAL buffers, which created contention around WAL insertion locks. The team experimented with several configuration parameters and eventually increased wal_buffers and wal_writer_flush_after. This allowed PostgreSQL to accumulate larger WAL batches in memory before flushing them to disk. The result was a significant reduction in contention around WAL insertion and about a 30 percent improvement in migration throughput. It is important to note that these changes are not universally safe defaults. Larger WAL buffers and less frequent flushing can increase the amount of data at risk during an unexpected crash. In this case, the workload was a migration. If the process stopped, it would have to be restarted anyway. Under those conditions, the temporary trade-off between reliability and performance was acceptable. The real lesson from this case is not a specific configuration value but the investigation process: identify where the system is actually waiting, test hypotheses, and verify improvements with measurements. Of course, real PE cases are often more complex than the simplified example shown here. In practice, investigations can take days and may involve analyzing internal database behavior, operating system limits, and network interactions to identify the true bottleneck. Common Performance Engineering Rules In performance engineering, there is an informal set of principles that experienced engineers tend to follow: 1. Proactivity Is the Best Prevention A performance engineer does not wait for a system to fail under load. The work starts earlier: analyzing the architecture of new services, anticipating how the system will behave as load grows, and identifying potential bottlenecks before they become production incidents. 2. Trust Metrics A common mistake, especially among less experienced engineers, is optimizing by eyeballing results. Someone changes a configuration or piece of code and says, “It seems to run three to five seconds faster.” That is not acceptable. Improvements must be confirmed with measurable data. Engineers compare metrics before and after a change: transactions per second (TPS), latency, CPU utilization, disk and memory usage, and queue lengths. Only these measurements can demonstrate whether performance has actually improved. Metrics matter more than subjective impressions, although experience still plays a role. Experienced engineers often use intuition to form an initial hypothesis about the cause of a problem. However, every hypothesis must be verified with measurements. Intuition helps guide the investigation, while metrics confirm the correctness of the solution. 3. Be Careful With Quick Fixes Sometimes incidents must be resolved immediately. A common example involves the max_connections parameter in PostgreSQL, which defines the maximum number of concurrent connections. When a system slows down under load, some developers try to increase max_connections. This can help temporarily, but it often creates new problems. A sudden increase in connections raises contention for internal database resources such as shared memory structures and locks. As contention grows, performance can degrade significantly due to locking and resource pressure. A quick fix can easily turn into a larger failure. A good performance engineer will highlight these risks and recommend a more systematic solution. Becoming a Performance Engineer Few people start their careers aiming for performance engineering. More often, they drift into it through a difficult problem that refuses to stay contained. That is how it happens in practice. A developer helps compare database options for an important project. Then the questions start multiplying. What should be measured? On physical servers or virtualized infrastructure? Which metrics matter? What changes under load? What changes only in production? One question leads to ten more. Before long, the person who thought they were helping with a tactical decision is working at the boundary between software, systems, and operational behavior. That path is common. People grow into it from development, systems administration, or operations. Nobody really graduates as a ready-made performance engineer, yet those who grow into the role often reach compensation levels that can support stronger long-term financial outcomes than many other career paths. Core Skills of Performance Engineers Because performance engineering sits at the intersection of software and infrastructure, practitioners usually combine skills from several technical disciplines. What does it take to move into this field? Programming A performance engineer needs to understand how software is written, how developers think, and what challenges they face. In many cases, the engineer works with tools built for other developers. Linux Strong Linux knowledge is very important: how the kernel works, how the user space operates, how processes are managed, and which operating system metrics can be measured. Algorithms Understanding algorithms and their complexity is essential for proposing efficient solutions. Math Another important but often missing skill is mathematical statistics. When an improvement is not dramatic but only one to two percent, engineers must prove that the change is meaningful and not just measurement noise. Concepts such as quantiles, percentiles, data distributions, and multimodal behavior help separate real improvements from measurement noise. Communication Performance engineers must clearly and carefully communicate findings to developers, testers, and business stakeholders. Explaining that a problem originates in someone’s code can be sensitive, so it must be done constructively. Attention to Detail Attention to detail is critical. An unusual spike in a graph or a repeating system pattern may point to the root cause of a problem. Persistence also matters. Test results can fluctuate due to environmental factors, so identifying the real issue often requires patience and careful investigation. The same applies to client-side performance work, where telemetry, crash analytics, and app data collection can help explain how the product behaves on real devices, networks, and usage patterns. The Future of Performance Engineering Systems are becoming more complex, and no single engineer can be an expert in every layer. As a result, the field is moving toward greater specialization. We are likely to see performance engineers focused on specific areas: application-level performance, operating system behavior, or hardware-level optimization, such as selecting the right CPU for a workload and tuning CPU frequency settings. What about AI? So far, there are no real tools capable of replacing performance engineers. AI can help engineers find information faster, although its output still needs verification. It does not yet solve complex analysis and optimization tasks. Automated tuning systems also do not currently appear capable of replacing human expertise. There is some expectation that AI will at least automate routine work. For now, most performance engineers see AI as an assistant rather than a threat. Final Thoughts: The Hunt Continues Performance engineering is a constant challenge. It is an intellectual puzzle with real operational impact. The work involves identifying hidden patterns, uncovering non-obvious relationships, and finding effective solutions where others see only complexity or system limits. Many engineers remember their first major optimization success. A query that once ran for minutes or even hours suddenly runs hundreds of times faster. Moments like this leave a lasting impression and often keep engineers in this field for many years. These experiences are what make performance engineering a difficult yet highly engaging profession, where the hunt for CPU cycles and response-time improvements never really ends.
I needed a job to run once a day, remember what it did yesterday, and cost nothing to operate. The obvious answer is a small VM with cron, or a Lambda plus DynamoDB. I did not want to pay for either, and I did not want a server to patch. So I pushed the whole thing onto GitHub Actions and used a JSON file committed back to the repo as the database. It has now run 139 times in production on the free tier, tracking just over 1,000 records, and the operating bill is still zero. Here is the part that took the most thought: keeping state across runs that are, by design, completely stateless. "The daily digest the pipeline sends, with new postings badged." The Constraint That Shapes Everything GitHub Actions gives you a cron trigger for free: Shell on: schedule: - cron: "0 16 * * *" # 09:00 EST daily workflow_dispatch: # manual button That solves scheduling. It does not solve memory. Every run starts on a fresh ubuntu-latest runner with a clean checkout. Anything you write to disk during the run is gone when the job ends. For my use case (a daily digest that must not re-send jobs it already sent), that is the entire problem. The script has to know what it saw yesterday. The standard fix is an external store. But for a workload that writes a few kilobytes once a day, standing up a database is more operational surface than the actual task. The repo is already there, the runner already has a checkout, and the workflow already has a token. So the store is the repo. Git as the Database The pattern is three lines at the end of the workflow: stage the state files, commit if they changed, push. Shell permissions: contents: write # the default token is read-only; you must opt in # ... run the script, which writes seen_links.json and job_history.json ... - name: Commit updated history files run: | git config user.name "GitHub Actions Bot" git config user.email "[email protected]" git add seen_links.json job_history.json 2>/dev/null || true git diff --staged --quiet || git commit -m "Update job history [skip ci]" git push One automated commit per day. The repo's own history is the database, and the audit log comes for free. Two details here are not optional, and I learned both the slow way. First, permissions: contents: write. The GITHUB_TOKEN handed to a workflow is read-only by default. Without this block, the git push fails with a 403, and the failure is at the very end of the run, after the real work succeeded, so it looks like everything worked until you check tomorrow and the state never persisted. Second, git diff --staged --quiet || git commit. This commits only when something actually changed. Committing an unchanged tree is an error, and a daily job that finds nothing new is a normal Tuesday. The || makes "nothing to commit" a no-op instead of a red X. The result is that the database lives in git history. Every state change is a commit. I can read yesterday's seen_links.json by checking out yesterday's commit. That is free audit logging I did not have to build. The Infinite-Loop Trap Here is the gotcha that will bite anyone who copies this pattern: a workflow that pushes a commit can trigger a workflow that runs on push, which pushes a commit, which triggers the workflow. The guard is the [skip ci] token in the commit message: git commit -m "Update job history [skip ci]" GitHub treats [skip ci] in a commit message as "do not start workflows for this commit." My scheduled workflow uses it. I also had a second, older workflow file in the repo whose commit message was a plain "Update seen links" with no skip token. Because that workflow only ran on schedule (not on push), it never actually looped, but it was one: push line away from a runaway. If your state-committing workflow has any push trigger, the skip token is the difference between a daily job and a billing incident. Put it in from the start. Decoupling "New" From "Still Worth Showing" The other decision I am glad I made early was separating two ideas that look like one: a record being new today, and a record being relevant today. A naive version sends only what is new since the last run. That breaks the moment a run finds nothing, or the moment the user skips a day. So state is two files with two jobs. seen_links.json is a flat set of every URL ever processed, used purely for deduplication. job_history.json is a rolling window: each entry carries a first_seen timestamp, and a record stays in the window for ten days regardless of how many runs happen in between. Shell def cleanup_old_jobs(history, max_days): today = datetime.now().date() cleaned = {} for category, jobs in history.items(): cleaned[category] = [] for job in jobs: first_seen = job.get("first_seen") seen_date = datetime.fromisoformat(first_seen).date() if (today - seen_date).days <= max_days: cleaned[category].append(job) return cleaned So "new" is computed per run (anything not in seen_links.json), and "relevant" is the trailing ten-day window. The daily output is never empty, nothing is ever sent twice, and a record ages out on a fixed schedule instead of vanishing the first quiet day. Two files, two responsibilities. Trying to make one structure do both is where this kind of project usually rots. The Dependency I Refused to Add The source data is two different table formats from upstream pages: one uses GitHub-flavored markdown tables, the other uses raw HTML tables inside the same document. The clean answer is a parsing library. I chose regex and the standard library instead, and I want to be honest about why and what it costs. The script tries markdown first, then falls back to HTML: Shell parsed_jobs = parse_markdown_table(text) if len(parsed_jobs) == 0: parsed_jobs = parse_html_table(text) # SimplifyJobs uses HTML The upside is a requirements.txt with exactly one line (requests), which means the install step on a cold runner is near-instant, and there is no transitive dependency that can break a 9 a.m. job. The downside is real, and I will not pretend otherwise: regex table parsing is brittle. When an upstream source changed its column layout, my parser silently returned zero rows for that source. It did not crash. It just quietly stopped finding jobs from one feed, which is the worst failure mode because nothing alerts you. For a personal tool with one user, that trade is fine: I notice within a day and patch a regex. For anything with real users, I would add a parser and, more importantly, a "parsed zero rows from a source that normally returns dozens" alarm. The lesson is not "regex bad." It is that a zero-result parse should be treated as a failure signal, not a valid empty result. Cheap Correctness Wins Two small filters do more work than their size suggests. Deduplication is a set membership check, which makes the whole pipeline idempotent. Running the workflow twice in one day produces the same output as running it once, because the second pass finds everything already in seen_links.json. For a cron job that you will inevitably trigger manually while debugging, idempotency is what lets you mash the button without consequences. Link quality is an allowlist of known applicant-tracking domains (Greenhouse, Lever, Workday, Ashby, and friends). Upstream rows mix real application links with company homepages and image badges. Filtering to known ATS hosts drops the noise without trying to validate every URL: Shell JOB_HOST_HINTS = ("greenhouse.io", "lever.co", "myworkdayjobs.com", "ashbyhq.com", "smartrecruiters.com", "icims.com", ...) def looks_like_job_link(url): return any(h in url.lower() for h in JOB_HOST_HINTS) An allowlist is the right default here because the failure mode is asymmetric. Letting through a dead homepage link wastes a click; an allowlist that occasionally drops a valid but unusual ATS is a one-line addition when I notice it. I would rather under-include than ship dead links. What it Actually Costs The numbers from production: 139 scheduled runs committed back to the repo, 1,062 unique links tracked in the dedupe set, three Python files, one runtime dependency, and one YAML workflow. Infrastructure cost is zero, because GitHub Actions' free tier covers a once-a-day job comfortably and Gmail's SMTP handles the delivery. There is no server, no database, no secret rotation beyond an app password, and nothing to wake up to at 3 a.m. When is This Pattern the Right Call? Reach for git-as-a-database when the write volume is low (you are committing on a human timescale, not a request timescale), the state is small and serializable, a single writer is doing the writing (the scheduled job), and you actively want the change history. A daily digest, a status snapshot, a slowly-changing config, a scoreboard: all good fits. Do not reach for it when you have concurrent writers (two runs racing to push will collide and one will fail the non-fast-forward push), when the state is large enough to bloat the repo, or when you need sub-minute reads or transactions. At that point you have outgrown the trick and a real datastore earns its keep. For everything in the first bucket, the calculus is hard to beat: the scheduler, the runtime, the storage, and the audit log are all things you already have for free. The only code you write is the part that does the work.
The Moment It Gets Real At some point in the last year, every data engineer had the same experience. You opened a copilot tool, typed a rough description of what you needed, and watched it generate a working ETL pipeline in about thirty seconds. Not a skeleton. Not pseudocode. Actual, runnable PySpark with joins, transformations, and a DAG scaffold. And for a moment, the question that the industry had been treating as hypothetical became very concrete: if AI can do this, what exactly am I here for? That question deserves a serious answer — not the dismissive "AI is just a tool" reassurance, and not the catastrophist "engineers are obsolete" take. The honest answer is more nuanced, more interesting, and more actionable than either of those. What AI Can Actually Do Today Let's be precise about what has changed, because the hype runs in both directions. AI copilots in 2026 are genuinely impressive at a specific class of data engineering tasks. Give a well-prompted model a schema and a business requirement, and it will produce SQL that would have taken a competent engineer thirty minutes to write. Ask it to scaffold a dbt model with tests and documentation, and it delivers something you can actually work from. Point it at a slow query and ask for optimization suggestions, and it identifies the right indexes and join strategies most of the time. The work that once defined the day-to-day of data engineering — writing transformations, building pipeline boilerplate, generating unit tests, documenting schemas — is now legitimately acceleratable by an order of magnitude. That compression is real. A pipeline that took a week to build from scratch now takes a day. A day's worth of dbt model work now takes a morning. The cycle time has collapsed, and pretending otherwise is not a useful position. But Would You Actually Deploy It? Here is where the honest conversation has to happen. AI generates code that looks production-ready. It compiles. The DAG runs. The transformations return the right rows on the test dataset. And then you look closer. There are no retry semantics. There is no idempotency guarantee — run it twice, and you get duplicates. There are no data quality checks, no row count assertions, no schema drift detection. Observability is absent. The error handling catches exceptions and logs them to nowhere. Governance controls do not exist because the model has no idea what your data classification policies are. The code is impressively correct at the logic layer and completely unprepared for production reality. And that gap — between "AI generated it" and "it is actually deployable" — is not a small gap. It represents most of what makes data engineering genuinely hard. This is not a criticism of AI tooling. It is a precise description of where the boundary currently sits. And that boundary is exactly where the value of a skilled data engineer now concentrates. The Three-Bucket Reality Not all data engineering work is equally automatable, and the honest framework is to split it into three categories based on where AI sits today. What AI handles well. SQL and transformation generation, dbt model scaffolding, unit test generation, schema documentation, query explanation, code refactoring, and first-draft pipeline boilerplate. These tasks are high-volume, pattern-heavy, and well-represented in training data. AI performs them at a level that meets or exceeds what most engineers produce under time pressure. What AI assists but cannot own. Pipeline architecture decisions, root cause analysis on production failures, performance tuning for complex distributed jobs, and data modeling judgment for novel domains. AI is genuinely useful here as a thought partner and accelerant, but the decisions require context, business knowledge, and judgment that models do not reliably carry. What remains fundamentally human. Trade-off evaluation with real organizational constraints, governance and compliance decisions, architecture choices with long-term consequences, and anything requiring accountability. These require not just the right answer but the right answer for this company, this data, this regulatory environment, this team. That is irreducibly human work. The critical observation is that the boundary between these buckets is not static. Tasks that sat in the second bucket eighteen months ago have migrated into the first. The direction of travel is clear. Engineers who have concentrated their value entirely in automatable work are already exposed. Engineers who have built depth in judgment, architecture, and systems thinking are in an increasingly strong position. The Workflow Has Already Changed The before and after is not theoretical. It is visible in how high-performing data engineering teams actually operate today. The traditional workflow moved linearly through extraction, transformation, loading, and serving — each stage measured in hours to days, the full cycle measured in weeks. It was plagued by boilerplate, manual testing, documentation that was always out of date, and context-switching that fragmented deep work. The AI-enhanced workflow runs the same stages but with a fundamentally different time signature. StageTraditionalAI-EnhancedExtractHours — manual SQL, custom connectorsMinutes — AI-generated queries, auto connectorsTransformDays — dbt models, Spark jobsHours — AI-assisted modeling, auto schema detectionLoadHours — DAG authoring, schedulingMinutes — auto DAG generation, smart schedulingServeDays — dashboard building, documentationHours — auto documentation, natural language query The total cycle time compresses from weeks to days. That compression does not come from removing the engineer. It comes from removing the repetitive execution work so the engineer can focus on the decisions that actually require human judgment. What the Collaboration Actually Looks Like The AI-native data engineer workflow is not "prompt and deploy." It is a structured collaboration with a clear division of responsibility. AI accelerates the build. The engineer ensures it is correct, reliable, observable, and production-ready. The accountability for what ships belongs to the engineer, not the model. That accountability is not a burden — it is the source of professional value. The engineers who treat AI output as a draft to be critically evaluated and hardened will consistently outperform those who either ignore the tools entirely or treat generated code as finished work. Both of those failure modes are common. Neither is sustainable. The Skill Set Reorganizes, Not Disappears The skills required to be an excellent data engineer are shifting, but they are not evaporating. They are reorganizing around three pillars. Technical depth now centers on evaluating AI-generated code rather than writing all code from scratch. This requires strong fundamentals — you cannot spot the subtle join fanout in AI-generated SQL if you do not understand join semantics. It also means investing in observability, reliability engineering, and prompt crafting as first-class technical skills. A well-constructed prompt that produces deployable output in one iteration is genuinely more valuable than the ability to write the same code manually from scratch. Systems thinking becomes the primary differentiator. Architecture decisions, data modeling judgment, trade-off evaluation, and problem framing are tasks that compound in value as AI handles more execution work. The engineer who can look at a generated pipeline and immediately identify the three ways it will fail at scale is providing something no current model reliably provides. Engineering leadership expands to include guiding AI usage within a team, establishing review standards for AI-generated code, owning governance controls, and setting the quality bar that separates production-ready from impressive-looking. This is not a soft skill add-on — it is a core engineering responsibility in an environment where the output volume of any individual engineer has increased dramatically. The role is shifting from execution to judgment. That is an upgrade, not a downgrade, for engineers willing to make the transition deliberately. How to Actually Evolve The path forward is concrete, not abstract. Start by integrating AI into your daily work right now — not as an experiment but as a workflow change. Use it for SQL drafting, pipeline scaffolding, and test generation. Build the muscle of critically evaluating what it produces. Develop prompting habits that consistently get you to a usable first draft rather than something you have to rewrite from scratch. Level up by investing deliberately in the areas AI does not cover well. System design. Distributed systems fundamentals. Reliability and observability patterns. Data modeling for complex domains. These skills appreciate in value as AI handles more of the execution layer — the relative scarcity of strong systems thinkers increases as the supply of generated boilerplate becomes effectively infinite. Lead by taking ownership of AI quality standards on your team. Be the person who defines what "production-ready" means for AI-generated pipelines, who establishes review checklists, who sets governance guardrails. This is influence that compounds over time and is not replicable by a model. The Honest Bottom Line AI will not replace data engineers. But data engineers who treat their value as residing primarily in writing code — rather than in the judgment, architecture, and reliability thinking that makes code worth deploying — are taking a position that becomes harder to defend with each model release. The opportunity is real, and it is now. The engineers who learn to work with AI as a genuine collaborator, who develop the critical evaluation skills to close the gap between generated and production-ready, and who invest in the systems thinking that AI cannot replicate — those engineers are not threatened by this transition. They are the ones who define what data engineering looks like on the other side of it. Evolve deliberately. The alternative is not standing still — it is falling behind at an accelerating rate.
A Temporal Workflow that appears stuck is rarely “stuck” in the conventional process sense. Temporal persists Workflow state through Event History and resumes execution through replay, so an open execution can remain healthy while waiting for a timer, Signal, Activity, or external condition. The operational problem is therefore not simply lack of completion; it is lack of expected progress. Effective diagnosis starts by establishing what event should have happened next, why it did not happen, and whether remediation can preserve the Workflow’s business invariants. Temporal’s history model makes that analysis unusually tractable because commands, task transitions, Activity attempts, failures, timers, and external interactions are durably represented as Events. Progress Is Visible in the Event History The first diagnostic artifact should be the execution description and raw history, not application logs. temporal workflow describe exposes current execution information and pending Activity state, while temporal workflow show --output json returns Event History in a form suitable for programmatic replay or analysis. A Workflow Query can additionally expose application-defined state without mutating the execution. Shell temporal workflow describe --workflow-id order-7814 temporal workflow show \ --workflow-id order-7814 \ --output json History should be read as a state-transition trace. A WorkflowTaskScheduled event with no corresponding start suggests that work is waiting for a Worker. A started Workflow Task that repeatedly times out can indicate blocked Workflow code, Worker instability, or excessive work inside a task. Repeated WorkflowTaskFailed events can indicate replay or deterministic-compatibility failures after code deployment. Workflow Task failures are retried by Temporal rather than governed by an Activity-style Retry Policy, so a Workflow can remain open while repeatedly failing to make application-level progress. Activity sequences reveal a different failure surface. ActivityTaskScheduled without ActivityTaskStarted points toward dispatch capacity, missing pollers, queue mismatch, or backlog. Temporal persists Workflow and Activity Tasks in Task Queues, and worker-health guidance identifies Schedule-to-Start latency and approximate backlog count as key signals when tasks wait for Workers. ActivityTaskStarted without completion requires inspection of Start-to-Close and Heartbeat behavior because Temporal relies on Start-to-Close timeout to detect a Worker crash after an Activity has started. Not every long pause is pathological. A timer that has not fired, a Workflow waiting for a Signal, or an Activity still inside a valid timeout window can represent correct durable waiting. Conversely, very large histories can become an operational risk. Temporal warns after 10,240 events or 10 MB and enforces a limit of 51,200 events or 50 MB; Continue-As-New creates a new run with a fresh history while carrying forward relevant state. Triage Works Best as Deterministic Evidence Before Model Judgment LangGraph is useful for automating this analysis, but the safest design keeps Temporal facts deterministic and uses an LLM only for classification, hypothesis ranking, and explanation. LangGraph explicitly supports graphs that mix deterministic nodes with model-driven nodes, while structured output can constrain routing decisions into a defined schema rather than free-form text. A compact analyzer can first reduce raw history into evidence that is difficult to hallucinate: the last completed Workflow Task, consecutive Workflow Task failures, pending Activity IDs, the latest Activity attempt, the timeout type, the last Signal, the last timer, the history size, the task queue, and deployment/version metadata. The model then receives that normalized evidence instead of thousands of raw events. Python def extract_facts(state): events = state["events"] return { "facts": temporal_fact_extractor(events), "tail": events[-60:], } def classify(state): result = triage_model.with_structured_output(TriageResult).invoke({ "facts": state["facts"], "tail": state["tail"], "allowed_causes": [ "worker_unavailable", "activity_retrying", "workflow_task_failure", "intentional_wait", "history_pressure", "unknown", ], }) return {"triage": result} That separation matters operationally. Event parsing can enforce hard rules such as “scheduled but never started,” while the model can correlate several weak signals and produce an explanation. Conditional edges can then route low-risk cases to observation, ambiguous cases to deeper diagnostics, and recovery candidates to an approval gate. LangGraph’s graph API supports conditional routing, and persistence stores checkpoints so triage state survives interruptions or process failures. Recovery Must Preserve Temporal and Business Semantics Diagnosis and remediation should remain separate graph stages. A model-generated recommendation must not directly issue cancellation, reset, or termination. LangGraph interrupts provide a natural control boundary because execution can pause with persisted state and resume only after external approval. Python def approval_gate(state): decision = interrupt({ "workflow_id": state["workflow_id"], "cause": state["triage"].cause, "action": state["triage"].recommended_action, "evidence": state["triage"].evidence, }) return {"approved": decision == "approve"} The remediation choice depends on the failure mode. A transient Worker outage usually requires restoring Worker capacity rather than mutating Workflow state because queued tasks persist until Workers can process them. An Activity repeatedly failing on a recoverable dependency can often be left to its Retry Policy, while permanent errors should be made non-retryable in application design to avoid pointless retries. Activity side effects should be idempotent because Activity attempts may execute more than once under retry and recovery behavior. Cancellation is the preferred stop mechanism when Workflow cleanup logic must run. Temporal records a cancellation request and schedules a Workflow Task so Workflow code can react. Termination is forceful: Workflow code does not receive a chance to clean up, and the terminated event closes the history. That makes termination an escalation path for executions that cannot process cancellation normally. Reset is more powerful and more dangerous. Temporal terminates the current execution and creates a new execution that copies history through a selected reset point, then replays forward using current Workflow code. Progress after the reset point is discarded. Reset is therefore appropriate only after the underlying cause has been corrected and after downstream side effects are reviewed for possible re-execution beyond the reset boundary. Shell temporal workflow reset \ --workflow-id order-7814 \ --event-id 42 \ --reason "Recovered after deterministic-compatibility fix" For history pressure rather than a fault, Continue-As-New is generally the safer lifecycle mechanism because it preserves logical continuity under the same Workflow ID while starting a fresh Event History with a new Run ID. It should be designed into long-lived or high-volume Workflow logic instead of used as an improvised emergency action. Safe Automation Requires an Explicit Remediation Envelope A production triage graph should treat remediation as a constrained transaction. The evidence snapshot, selected run ID, candidate reset event, intended action, reason, approval identity, and execution result should all be persisted before any mutation. The action node should re-read the Workflow immediately before execution and reject the operation if the run has changed or the observed condition no longer matches the diagnosis. This is an engineering safeguard rather than a Temporal requirement, but it reduces time-of-check/time-of-use errors when active Workflows continue progressing during investigation. LangGraph’s checkpoint model supports durable approval state, but resumed graph nodes can re-execute from checkpoint boundaries. Its documentation therefore recommends isolating side effects and designing them to be idempotent. A remediation executor should consequently use an operation ID, record completion externally, and refuse duplicate destructive actions. Recovery Without Guesswork Reliable recovery of a stuck Temporal Workflow is fundamentally an event-history problem, not a process-restart problem. The strongest diagnostic path reconstructs expected progress from Workflow Tasks, Activity attempts, timers, Signals, queue state, timeouts, and history growth before considering mutation. LangGraph can turn that evidence into a durable triage pipeline by combining deterministic extraction, constrained model reasoning, conditional routing, and interrupt-based approval. Safe remediation then follows Temporal semantics: restore Workers when dispatch is the issue, allow bounded retries for transient Activities, cancel when cleanup matters, terminate only as a last resort, reset only after the root cause is fixed, and use Continue-As-New to control long-running history growth. The result is automation that accelerates incident response without allowing probabilistic diagnosis to become an unchecked control plane.
Convolutional neural network workloads rarely fail because the forward pass is mathematically difficult. They fail because modern training and inference pipelines are distributed systems: datasets arrive late, GPU workers disappear, validation jobs stall, model registration breaks halfway through, and long-running executions need to resume without corrupting state. Temporal is designed for exactly that class of problem. A Temporal Workflow Execution is durable, reliable, and scalable, and Temporal defines durable execution as the ability of a workflow to maintain state and progress through crashes or outages. That makes it a strong fit for CNN pipelines whose control plane must survive for hours, days, or even longer while the actual tensor computation runs elsewhere. Why Temporal Fits CNN Pipelines A useful way to think about Temporal in ML systems is as a durable control plane rather than a replacement for PyTorch, CUDA, or a model server. Temporal Workflows hold orchestration state, issue commands, wait on results, and recover by replaying event history. The workflow code itself must stay deterministic, while failure-prone and non-deterministic work belongs in Activities. Temporal’s own documentation is explicit on that boundary: API calls, file I/O, database access, and other external interactions belong in Activities, while workflows should remain replay-safe. That separation aligns naturally with CNN systems, where data staging, training job submission, checkpoint writes, validation, and registry updates interact with external systems constantly. That design matters because a CNN training run is not one step. Even a modest image-classification job usually has a training phase, a validation phase, best-model selection, a checkpointing loop, and a final artifact publication step. The PyTorch transfer-learning tutorial illustrates that pattern directly by alternating train and validation phases, persisting the best state_dict, and loading the best weights at the end; the quickstart tutorial likewise uses model.eval() and torch.no_grad() before prediction. Temporal does not change the math of those stages. It makes the sequence durable, observable, and restartable. A concise workflow can therefore stay almost entirely orchestration-focused: Python @workflow.defn class CnnTrainingWorkflow: @workflow.run async def run(self, req: TrainRequest) -> ModelArtifact: dataset = await workflow.execute_activity( prepare_dataset, req.dataset_ref, start_to_close_timeout=timedelta(minutes=20), ) trained = await workflow.execute_activity( train_cnn, TrainJob(dataset_uri=dataset.uri, config=req.config), start_to_close_timeout=timedelta(hours=8), heartbeat_timeout=timedelta(minutes=1), retry_policy=RetryPolicy(maximum_attempts=3), ) metrics = await workflow.execute_activity( evaluate_cnn, EvalJob(model_uri=trained.best_model_uri, dataset_uri=dataset.val_uri), start_to_close_timeout=timedelta(minutes=30), ) return await workflow.execute_activity( register_model, RegisterRequest(trained.best_model_uri, metrics), start_to_close_timeout=timedelta(minutes=5), ) The important detail in this snippet is not syntax but placement. The workflow issues durable commands, while every side effect lives inside an activity. That matches Temporal’s execution model, where workflows await activity results, activities carry retry policies, and activity timeouts define how long a unit of external work is allowed to run. For long GPU jobs, heartbeat_timeout is especially important because an activity heartbeat tells Temporal the worker is still alive and making progress. Making Long Training Jobs Resumable The natural temptation with CNN training is to keep a Python process alive for hours and hope that the host, container runtime, and storage path all behave. Temporal offers a more robust approach. Activities can be retried automatically after transient failure, and Temporal recommends Start-To-Close timeouts for activity executions. If heartbeats stop arriving inside the heartbeat timeout, the activity can be considered failed and retried according to policy. For training jobs that run on flaky GPU nodes or preemptible infrastructure, that is a meaningful improvement over ad hoc retry shells. The training activity itself should then follow framework-native checkpoint discipline instead of trying to serialize the full training loop into workflow state. PyTorch’s guidance is centered on saving and loading model state with state_dict, and the transfer-learning tutorial shows a production-relevant pattern: save the best model parameters during validation, then reload them at the end. Temporal’s own ML Ops example follows the same philosophy by highlighting checkpoint-aware fine-tuning and resumable inference, with deterministic orchestration in workflows and non-deterministic ML work in activities. A minimal activity therefore looks more like this: Python @activity.defn def train_cnn(job: TrainJob) -> TrainingResult: state = restore_checkpoint(job.resume_uri) model = build_model(job.config, state) best_acc = state.best_acc if state else 0.0 for epoch in range(state.next_epoch if state else 0, job.epochs): train_one_epoch(model, job.train_loader) val_acc = validate(model, job.val_loader) save_checkpoint(job.resume_uri, model, epoch, val_acc, best_acc) if val_acc >= best_acc: save_best_weights(job.best_model_uri, model) best_acc = val_acc activity.heartbeat({"epoch": epoch, "best_acc": best_acc}) return TrainingResult(best_model_uri=job.best_model_uri, best_acc=best_acc) This pattern is deliberately boring, which is a strength. The epoch loop stays in the activity, checkpoints remain in object storage or a shared filesystem, and each heartbeat advertises progress to Temporal. If a retry occurs, the activity can resume from the latest checkpoint rather than restarting from epoch zero. That is the same operational idea highlighted in Temporal’s ML sample repository, and it matches PyTorch’s recommendation to persist model parameters through state_dict-based saves and reloads. When training runs are submitted to an external batch scheduler instead of executing directly inside the worker process, Temporal’s asynchronous activity completion becomes especially useful. An activity can submit the job, capture the task token, and return without marking the activity complete. The external system can later heartbeat and complete the activity through a Temporal client. Temporal documents this explicitly and notes that asynchronous completion is preferable when the external process needs heartbeats or cancellation. Python @activity.defn async def submit_training(job: TrainJob): token = activity.info().task_token launch_gpu_job(job, token) activity.raise_complete_async() Treating Inference as a Workflow Only When It Is One Temporal is not a substitute for a low-latency online model server. A single image classification request that must return in milliseconds usually belongs in the serving layer. Temporal becomes valuable when inference is part of a larger durable process, such as nightly batch scoring, asynchronous document-image analysis, model fallback, approval gates, or multi-stage post-processing. That recommendation follows directly from Temporal’s model of workflows as durable, stateful executions that communicate through activities, child workflows, signals, queries, and schedules. Inside the inference activity, the framework rules remain unchanged. PyTorch examples set the model to evaluation mode and disable gradient tracking during inference. That is essential for CNNs that use dropout- or batchnorm-sensitive behavior and for avoiding unnecessary autograd overhead. Python @activity.defn def score_batch(req: BatchScoreRequest) -> BatchScoreResult: model = load_model(req.model_uri) model.eval() with torch.no_grad(): return predict_batch(model, req.input_uri) Temporal’s message-passing model then makes the surrounding orchestration easier to operate. The Python SDK documents that a workflow can act like a stateful web service receiving queries, signals, and updates. For a batch scoring workflow, a query can expose current shard progress, while a signal can switch the canary model version for the remaining work without restarting the execution. Recurring inference or retraining can be started through Temporal Schedules, which the documentation describes as a more flexible and user-friendly approach than cron jobs. Python @workflow.query def status(self) -> dict: return {"phase": self.phase, "completed": self.completed, "model": self.model_uri} @workflow.signal def switch_model(self, model_uri: str) -> None: self.model_uri = model_uri Keeping ML Workflows Operable as They Grow CNN platforms do not stay small for long. A single training run becomes a hyperparameter sweep, then a retraining program, then a fleet of region-specific models. Temporal scales that expansion through composition. A parent workflow can start child workflows for each experiment, fold, or dataset shard, and the child workflow APIs guarantee that the child has actually started before the call resolves. That makes fan-out training and batch inference easier to reason about than out-of-band job launchers with partial status tracking. Long-running ML control planes also run into two operational realities: event history growth and code evolution. Temporal addresses the first with Continue-As-New, which closes the current execution successfully and starts a new run with the same workflow ID and a fresh event history. It addresses the second with versioning support and, in production, Worker Versioning. Temporal recommends Worker Versioning as the default way to deploy changes safely, and pinned workflows can stay on the worker deployment version where they started. That is especially relevant for training or evaluation flows that may stay active across multiple application releases. Conclusion Temporal brings discipline to CNN systems by separating durable orchestration from non-deterministic GPU work. The workflow owns state, retries, waiting, composition, and observability. Activities own data movement, training submission, checkpointing, evaluation, and inference execution. PyTorch continues to provide the familiar mechanics of state_dict checkpoints, evaluation mode, and gradient-free inference, while Temporal turns those stages into a resilient end-to-end process that can survive infrastructure faults, external scheduler delays, and repeated deployment cycles. For teams building CNN platforms that have outgrown shell scripts and brittle job glue, that combination is not merely convenient. It is often the boundary between a model pipeline that occasionally works and one that can be operated confidently in production.