DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Your AI Coding Assistant Stopped Suggesting and Started Shipping. Now What?
  • Engineering as a Service Is What Happens When You Let Vibe Coding Win
  • Slopsquatting: A New Supply Chain Threat From AI Coding Agents
  • AI Is Making PHP Cool Again

Trending

  • Building a Zero-Cost Daily Job Alert Pipeline on GitHub Actions
  • dbt Meets Apache Flink: One Workflow for Data Engineers
  • Agentic System Design in Practice: The Technical Debt in Enterprise Agentic Systems
  • Part 1: Building Governed MCP Tool Services With Quarkus LangChain4j and Goose
  1. DZone
  2. Data Engineering
  3. AI/ML
  4. AI Coding Is Moving From Trusting the Model to Constraining What It Can Do

AI Coding Is Moving From Trusting the Model to Constraining What It Can Do

AI coding is shifting from trusting models to constraining them with permissions, tools, and deterministic checks. BUBAS applies the same idea to business logic.

By 
Peter Verhas user avatar
Peter Verhas
DZone Core CORE ·
Sep. 28, 26 · Analysis
Likes (0)
Comment
Save
Tweet
Share
90 Views

Join the DZone community and get the full member experience.

Join For Free

For the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening.

The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified.

This is a significant architectural shift. The emerging pattern looks less like:

Plain Text
 
prompt → LLM → source code → trust it


(the infamous vibe ding pattern) and increasingly like:

Plain Text
 
intent  ↓ LLM  ↓ restricted set of operations  ↓ deterministic tools and validation  ↓ result


Several recent developments in mainstream developer tooling point in exactly this direction.

Permission Is Becoming Separate From Intelligence

On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations.

Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals.

That distinction is more profound than it may initially appear. The question is no longer merely:

Can the agent perform this operation?

It is:

Is this agent authorized to perform this operation in this environment?

Capability and authority are different things.

A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so.

A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities.

Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering.

The Model Is Becoming Replaceable

Another development makes the separation even clearer.

GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation.

This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go.

Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear.

Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is:

Plain Text
 
stable environment stable tools stable constraints stable validation         ↑ 
interchangeable models



  stable environment
  stable tools
  stable constraints
  stable validation
          ↑
interchangeable models


The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system.

Agents Are Separating From Editors

VS Code's Agent Host architecture points in another related direction.

The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts.

Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code:

Plain Text
 
     human
       ↓
     editor
       ↓
language server
       ↓
    compiler


Agentic development introduces another participant:

Plain Text
 
         human intent
              ↓
            agent
              ↓
         semantic tools
              ↓
 compiler / tests / environment


The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important.

A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things.

Verification Is Moving From Opinion to Execution

A fourth development may ultimately be the most important.

GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct.

This should sound obvious.

We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense.

A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system.

A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes:

Plain Text
 
     generate
        ↓
     compile
        ↓
       test
        ↓
inspect diagnostics
        ↓
     repair
        ↓
      repeat


That is substantially more robust than asking a model to inspect its own output and say whether it looks right.

The role of the LLM is reasoning. The role of deterministic software remains enforcement.

The Interesting Convergence

These developments come from different parts of the development stack, but they point toward the same decomposition.

An AI programming environment increasingly contains at least four distinct elements:

  1. A model that reasons and generates
  2. A vocabulary of operations available to it
  3. A capability policy defining which operations it may use
  4. Deterministic mechanisms that decide whether the result is valid

None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design.

What If We Put the Boundary Into the Language?

Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted.

There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes?

This is the idea I have been exploring with an open-source project called BUBAS.

BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment.

There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary.

An order-processing application could, for example, expose operations such as:

Plain Text
 
LOAD_ORDER 
ORDER_TOTAL 
CUSTOMER_RISK 
APPROVE REJECT 
REQUEST_APPROVAL


An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed.

If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that.

Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem.

This Is Not a Sandbox

Make the distinction carefully.

A restricted language does not magically make its host application safe.

If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong.

The useful guarantee is narrower:

Generated business logic can only name operations that the application deliberately made part of its vocabulary.

That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment.

These operate at different semantic levels.

Domain Capabilities Are More Interesting Than Operating-System Capabilities

Operating-system permissions are necessary, but business applications eventually need a richer vocabulary.

Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful.

It still does not answer questions such as:

  • May this program approve an order?
  • May it request approval but not approve directly?
  • Can it read customer risk information?
  • Can it initiate a payment?
  • Can it calculate a premium but not change the underlying policy?

Those are domain capabilities.

General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction.

A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations.

We May Be Seeing the New Shape of the AI Programming Stack

None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software.

But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change.

The industry is gradually replacing:

Trust the model.

with:

Give the model precisely defined capabilities and verify what it produces.

That is a much more promising engineering principle.

BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.

AI Coding (social sciences)

Opinions expressed by DZone contributors are their own.

Related

  • Your AI Coding Assistant Stopped Suggesting and Started Shipping. Now What?
  • Engineering as a Service Is What Happens When You Let Vibe Coding Win
  • Slopsquatting: A New Supply Chain Threat From AI Coding Agents
  • AI Is Making PHP Cool Again

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook