DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Prompt Engineering Wasn't Enough; Context Engineering Is What Came Next
  • A Comprehensive Guide to Prompt Engineering
  • Classification Never Left. It Just Got a New Home in LLMs.
  • The Context Window Trap: Why More Context Doesn’t Mean Better AI

Trending

  • Building an AI-Ready Data Layer Without Rebuilding the Enterprise
  • Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
  • Are Passphrases Still Secure in the Age of AI?
  • Mistaking Code Production for Engineering Progress: AI Productivity Myths Part 1
  1. DZone
  2. Data Engineering
  3. AI/ML
  4. Agentic Workflows Without LLMs? How I Cut a 14-Hour Engineering Workflow to 2 Hours

Agentic Workflows Without LLMs? How I Cut a 14-Hour Engineering Workflow to 2 Hours

Cut a 14-hour engineering workflow to ~2 hours — LLM agents only for reasoning, deterministic code for the rest. Good agentic design means knowing when to skip the LLM.

By 
Lucas Yoon user avatar
Lucas Yoon
·
Oct. 07, 26 · Tutorial
Likes (1)
Comment
Save
Tweet
Share
150 Views

Join the DZone community and get the full member experience.

Join For Free

Agentic software is often described in a fairly simple way. Give an LLM a set of tools, describe a goal, and let the model work through the problem. That works surprisingly well for prototypes. It becomes much less attractive when the workflow has to be repeatable, reliable, and maintainable.

I ran into this while automating a recurring software maintenance process that could take roughly 14 hours of engineering time. The work involved finding outdated AI model servers and models, checking upstream sources, understanding compatibility changes, rebuilding container images, updating templates, deploying them to OpenShift, validating the result, promoting artifacts, and preparing pull requests.

At first, putting an LLM in the middle of the entire workflow seemed like the obvious solution. The system improved when I started doing the opposite. I moved as much work as I could out of the LLM. Version comparison became Python. Repeated operational work became a custom CLI. APIs were called directly. JavaScript handled deduplication. Workflow state was stored outside the model context. Independent work ran concurrently. LLM agents were reserved for the parts of the process that actually required interpretation or engineering judgment. 

That change reduced the amount of engineer involvement from roughly 14 hours to about two. The interesting part is not that AI somehow completed 14 hours of engineering in two hours. The more useful explanation is that the workflow removed around 12 hours of repetitive execution from the engineer's critical path.

The implementation behind this workflow is available in my ai-template-updater repository [https://github.com/JslYoon/ai-template-updater], where the deterministic CLI, Claude Code workflows, and specialized agents are separated into distinct layers.

An agentic workflow does not have to be an LLM workflow.

The Engineering Work Hidden Inside a Jira Story

In an Agile workflow, this kind of work can begin with a Jira story that looks routine.

Update the supported AI software templates to the latest compatible versions.

The ticket is short. The work behind it is not. A complete maintenance cycle can require version discovery, dependency analysis, repository updates, image builds, staging, deployment, testing, and pull requests across several systems.

A lot of the time is not spent writing difficult code. It is spent rebuilding context while moving between GitHub, PyPI, Hugging Face, Quay, local repositories, container tooling, OpenShift, and Jira. That makes the workflow a good candidate for automation. It does not mean every part should be given to an LLM.

The First Question I Ask

For each step, I ask one question. Does this task require reasoning, or can software determine the answer directly?

Checking whether version 0.12.0 is newer than 0.11.0 does not require an LLM. Fetching image tags from a registry does not require an LLM. Checking whether two SHAs differ does not require an LLM. Deduplicating ten references to the same server update does not require an LLM.

Those are normal software problems. Other parts are different. A vLLM upgrade can affect PyTorch, Triton, xFormers, CUDA, and other packages. The right change can depend on release notes, the current Containerfile, repository conventions, and compatibility across several dependencies. That is where an LLM becomes useful.

task best implementation

Query container registry tags

API or code

Compare semantic versions

Code

Read Hugging Face metadata

API or code

Detect version drift

Code

Deduplicate updates

Code

Store exact artifact identifiers

Persistent state

Determine whether a command worked

Exit code

Understand breaking release changes

LLM

Analyze dependency compatibility

LLM

Modify an unfamiliar Containerfile

LLM

Adapt changes across a repository

LLM


The rule I use now is simple. If the answer can be computed, I compute it. If it needs to be interpreted, I consider using the model.

Building a Deterministic Core

The repository exposes a Python package through a custom CLI named agentic-template-ops. The CLI gives the workflow stable capabilities instead of forcing the model to understand every storage format, API, and repository detail.

Python
 
agentic-template-ops investigate
agentic-template-ops record-builds
agentic-template-ops list-built
agentic-template-ops configure


For server updates, deterministic code retrieves tags, parses versions, ignores prereleases, checks upstream sources, and compares the current release with the newest one. Model checks query Hugging Face metadata directly. The investigation step also runs independent checks concurrently.

An LLM should not become an expensive replacement for a version library, an HTTP client, or a thread pool.

The Main Setup Workflow

The /setup workflow is where most of the automation happens. It moves from discovery to a testable deployment in a series of explicit phases rather than asking one giant agent to figure everything out.

  • Pre flight configures permissions, reads the environment, verifies the custom CLI, and checks Quay authentication.
  • Investigate runs the drift scan across model servers and models, then writes the newest audit run to the Version Status sheet.
  • Deduplication happens in JavaScript before the build phase so the same server or model is not rebuilt for every template that references it.
  • Build dispatches unique server and model updates in parallel to specialized workers and pushes staging images to a personal Quay namespace.
  • Record persists the exact pushed image tags and build state. Later phases read those values back instead of reconstructing them.
  • Stage updates the ai-lab-template environment files on one update-all branch, regenerates the templates, and pushes the branch to the fork.
  • Deploy points the rolling demo at the staged branch and installs it on ROSA so the engineer can test the result.
    The /setup pipeline
Figure 1. The /setup pipeline. The deterministic layer does discovery and reduction, LLM workers handle contextual edits, state is persisted after staging, and human verification begins after deployment.


Why the Custom CLI Matters

Without the CLI, an agent might need to know how to locate the newest audit run, parse rows, interpret build flags, preserve exact image tags, and understand the shape of the Google Sheet. That is implementation knowledge the model does not need.

I need the successfully built artifacts
↓
agentic-template-ops list-built
↓
structured result

Now the storage logic and validation live behind a stable interface. If the way state is stored changes, I update the CLI. I do not have to rewrite every prompt.

Reduce the Amount of Reasoning

A weak agent prompt gives the model responsibility for discovery, planning, implementation, execution, and validation all at once. A better design uses normal software to reduce the problem first.

YAML
 
Current version   0.11.0
Latest version 0.12.0
Update required true
Component vLLM
Repository        known


By the time the worker receives the task, the model does not have to discover whether an update exists or which component is affected. It can focus on the part where reasoning is valuable, such as dependency changes, repository edits, and build troubleshooting.

Some of the Best Optimizations Use Zero Tokens

After drift detection, the workflow deduplicates updates with JavaScript. If ten templates reference the same vLLM update, the system creates one unique build item instead of ten model calls that rediscover and rebuild the same artifact.

10 template references
↓
deduplicate in code
↓
1 unique vLLM upgrade
↓
1 build task

Prompt caching, smaller models, and context compression can all help. But there is an earlier question worth asking. Does this need to be an LLM call at all? Removing an unnecessary call is usually better than optimizing it.

Where I Actually Want the LLM

Once deterministic code identifies a real update, the model becomes much more useful. Some upgrades only change a few known pins. Others affect the dependency graph, container build, or repository structure.

A vLLM update can involve vLLM itself, PyTorch, Triton, xFormers, CUDA compatibility, package constraints, and assumptions inside the image build. The worker may need to inspect existing code, read release information, edit several files, run the build, and react to failures.

That is no longer scraping. It is a bounded engineering problem. This is the part I want an LLM to solve.

Redeploying Without Rebuilding

Not every validation cycle needs to repeat investigation and image builds. The /stage-demo workflow exists for that case. It reuses the branch that /setup already created and focuses only on deployment.

  • Pre-flight reads the environment and configures access.
  • Find branch locates the most recent update-all branch on the fork, or uses a branch explicitly provided by the caller.
  • Deploy generates the rolling demo environment, points values.yaml at the staged branch, commits to development, and runs make install.
  • The deployment is only considered successful when the ROSA pods are running, the ArgoCD application is healthy, and the RHDH endpoint returns HTTP 200.

The /stage-demo path

Figure 2. The /stage-demo path avoids investigation and rebuilds. It reuses an existing staging branch and performs only the work needed to redeploy and verify the environment.


Structured Results and Durable State

Agents can reason in natural language, but workflow boundaries should be structured. A build worker returns fields such as success, component, version, and image_tag rather than a paragraph that another model has to reinterpret.

JSON
 
{
  "success": true,
  "server_type": "vllm",
  "component": "server",
  "version": "0.12.0",
  "image_tag": "..."
}


The same principle applies to state. If an image is pushed with an exact tag, later phases should read that tag from persisted state. They should not rely on the context window or reconstruct it from memory.

Context is useful for reasoning. State is useful for persistence.

Keep the Human at the Consequential Boundary

The goal was never to remove the engineer completely. The goal was to stop requiring the engineer to manually execute every reversible step.

The workflow reaches a staged RHDH environment automatically. That is where human judgment becomes valuable. The engineer tests the templates, checks functionality, and decides whether the work satisfies the Jira acceptance criteria and the Definition of Done.

Only after that verification does /promote run.

  • Config reads the environment and the exact built rows from persistent state.
  • The workflow deduplicates server and model work again before promotion.
  • Promote retags staging images into the official Quay namespace in parallel.
  • DevImages commits server version directories and opens upstream pull requests. These operations run sequentially because they share a Git working tree.
  • Templates reuses the staging branch, swaps personal tags for official tags, regenerates the templates, and opens the upstream ai-lab-template pull request.

The /promote workflow

Figure 3. The /promote workflow runs only after human verification. It promotes exact staged artifacts and then updates the two upstream repositories.


What Actually Changed From 14 Hours to 2

It would be misleading to say that an AI completed 14 hours of engineering in two hours. Before the automation, the engineer was involved in investigation, version checks, dependency analysis, implementation, container builds, staging, deployment, testing, and review.

After the workflow was introduced, the system took over most of the repetitive investigation and execution. The engineer spends the remaining time reviewing the results, testing the staged environment, handling unusual failures, and making the promotion decision.

The engineering responsibility did not disappear. The distribution of engineering time changed.

From an Agile perspective, the same recurring Jira work now consumes far less engineering capacity. The saved time can move toward product work instead of routine maintenance.

Agentic Does Not Mean LLM Everywhere

The biggest lesson from this project is that the quality of an agentic system should not be measured by the number of model calls it makes. A workflow can be highly autonomous while relying heavily on normal software.

In this project, APIs retrieve structured information. Python detects version drift. Version libraries compare releases. Concurrency handles independent checks. JavaScript removes duplicate work. A custom CLI exposes stable capabilities. Google Sheets persists workflow state. Structured schemas connect model workers back to orchestration code. The LLM is used where the work stops being fully deterministic.

What surprised me was that the architecture became better as I removed the LLM from more parts of it. In hindsight, that should not be surprising. Software engineers have always tried to use the simplest reliable tool that solves the problem. Sometimes that tool is an LLM. Quite often it is a function.

Agentic does not mean LLM everywhere. Sometimes the right way to make an agent more reliable is to move more of the workflow into code.

The goal of an agentic workflow should not be to make the LLM do more work. It should be to make the LLM do only the work that actually benefits from an LLM.

Command-line interface Engineering large language model

Opinions expressed by DZone contributors are their own.

Related

  • Prompt Engineering Wasn't Enough; Context Engineering Is What Came Next
  • A Comprehensive Guide to Prompt Engineering
  • Classification Never Left. It Just Got a New Home in LLMs.
  • The Context Window Trap: Why More Context Doesn’t Mean Better AI

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook