AI Coding Is Moving From Trusting the Model to Constraining What It Can Do
AI coding is shifting from trusting models to constraining them with permissions, tools, and deterministic checks. BUBAS applies the same idea to business logic.
Join the DZone community and get the full member experience.
Join For FreeFor the last few years, much of the discussion around AI-assisted programming has concentrated on models. Which model generates the best code? Which one understands the largest repository? Which one makes fewer mistakes? Which one has the largest context window? Those questions still matter, but something more interesting is happening.
The infrastructure surrounding coding agents is starting to assume that the model is not the component that should ultimately be trusted. Instead, increasingly sophisticated systems are being built around models to control what they can access, what operations they can perform, how those operations are approved, and how their results are verified.
This is a significant architectural shift. The emerging pattern looks less like:
prompt → LLM → source code → trust it
(the infamous vibe ding pattern) and increasingly like:
intent ↓ LLM ↓ restricted set of operations ↓ deterministic tools and validation ↓ result
Several recent developments in mainstream developer tooling point in exactly this direction.
Permission Is Becoming Separate From Intelligence
On September 9, 2026, GitHub announced centrally managed permissions for GitHub Copilot agent operations.
Enterprise administrators can classify operations such as shell commands, file reads and edits, and access to network domains as blocked, requiring approval, or allowed. Importantly, centrally imposed restrictions cannot simply be weakened by workspace configuration or previously saved user approvals.
That distinction is more profound than it may initially appear. The question is no longer merely:
Can the agent perform this operation?
It is:
Is this agent authorized to perform this operation in this environment?
Capability and authority are different things.
A sufficiently capable model may know perfectly well how to run curl, change a configuration file, query a database, or invoke a deployment tool. That does not imply that it should have the ability to do so.
A day earlier, GitHub announced enterprise-managed sandboxing for Copilot in JetBrains IDEs. Administrators can control filesystem access, network access, developer tools, proxies, macOS Keychain access, and related capabilities.
Again, the interesting part is not the specific list of switches. The architecture assumes that the agent operates inside an explicitly defined capability boundary. This is becoming infrastructure rather than prompt engineering.
The Model Is Becoming Replaceable
Another development makes the separation even clearer.
GitHub's experimental Project HydraFusion for Copilot CLI does not require the developer to choose one model and use it for the entire task. It can route parts of a workflow between local, cloud, and compound models and can use different models for drafting, criticism, revision, or escalation.
This is an important direction even if HydraFusion itself changes or disappears. It treats the model as a replaceable execution resource. That is probably where AI development tooling has to go.
Today, we debate whether one particular Claude, GPT, Gemini, or another model performs best on a particular benchmark. Six months later the answer may be different. Models improve, prices change, some are retired, local models become practical, and new providers appear.
Building the semantics of a software-development process around the behavioral peculiarities of one model therefore creates an uncomfortable dependency. A more durable architecture is:
stable environment stable tools stable constraints stable validation ↑
interchangeable models
stable environment
stable tools
stable constraints
stable validation
↑
interchangeable models
The model provides reasoning and generation. The surrounding system defines what constitutes a valid action. This also changes what a programming interface for an LLM should look like. Instead of hoping that a model remembers what it is allowed to do from a long textual prompt, we can give it a smaller, mechanically discoverable set of operations. The vocabulary becomes part of the system.
Agents Are Separating From Editors
VS Code's Agent Host architecture points in another related direction.
The agent is no longer conceptually an autocomplete feature living inside an editor window. Agent sessions can persist independently of that window, and the open Agent Host Protocol provides a common interface between clients and agent hosts.
Different agent harnesses can sit behind the same client-facing protocol. That separation is important. Traditional programming tools are centered on a human editing source code:
human
↓
editor
↓
language server
↓
compiler
Agentic development introduces another participant:
human intent
↓
agent
↓
semantic tools
↓
compiler / tests / environment
The editor becomes one possible interface onto that process rather than necessarily its center. This makes machine-facing programming interfaces much more important.
A language implementation can no longer assume that diagnostics, type information, available operations, and documentation exist only for presentation to a human inside an IDE. An agent also needs to interrogate those things.
Verification Is Moving From Opinion to Execution
A fourth development may ultimately be the most important.
GitHub recently expanded Copilot code review so that the reviewing agent can use shell tools to validate the code it examines. The review process can run builds, tests, scripts, and other deterministic checks rather than relying exclusively on the model reading source and deciding whether it appears correct.
This should sound obvious.
We have spent decades constructing deterministic machinery for checking software: compilers, static analyzers, unit tests, type systems, linters, model checkers, integration tests, and executable specifications. Throwing those away because an LLM can read code would make little sense.
A model is useful for deciding what to try. A compiler is much better at deciding whether a program satisfies its grammar and type system.
A unit test is much better at determining whether a known input produces a required result. The resulting loop becomes:
generate
↓
compile
↓
test
↓
inspect diagnostics
↓
repair
↓
repeat
That is substantially more robust than asking a model to inspect its own output and say whether it looks right.
The role of the LLM is reasoning. The role of deterministic software remains enforcement.
The Interesting Convergence
These developments come from different parts of the development stack, but they point toward the same decomposition.
An AI programming environment increasingly contains at least four distinct elements:
- A model that reasons and generates
- A vocabulary of operations available to it
- A capability policy defining which operations it may use
- Deterministic mechanisms that decide whether the result is valid
None of these requires us to believe that the model is reliable in the conventional software-engineering sense. In fact, the architecture is useful exactly because it assumes otherwise. The model can be probabilistic, and the boundary around it can remain deterministic. That observation has interesting consequences for programming-language design.
What If We Put the Boundary Into the Language?
Most current agent systems constrain an AI from outside a general-purpose programming language. The agent may generate Python, Java, JavaScript, shell commands, or some combination of them, while the surrounding sandbox tries to control which resulting actions are permitted.
There is another possible approach. What if the generated program itself could express only the operations the host application intentionally exposes?
This is the idea I have been exploring with an open-source project called BUBAS.
BUBAS is a deliberately small orchestration language embedded in Java. It has ordinary control-flow constructs, variables, types, decisions, and loops, but it deliberately does not expose the host programming environment.
There is no import mechanism, reflection, eval, filesystem API, network API, or way for a script to name an arbitrary Java class. Instead, the application defines a vocabulary.
An order-processing application could, for example, expose operations such as:
LOAD_ORDER
ORDER_TOTAL
CUSTOMER_RISK
APPROVE REJECT
REQUEST_APPROVAL
An insurance application would expose a different vocabulary. The significant property is not the syntax. Many DSLs have domain-specific words. The interesting property is what happens to everything that is not in the vocabulary. It cannot be expressed.
If DELETE_DATABASE has not been exposed, asking the model to delete the database does not require the model to refuse. There simply is no program in the language that means that.
Inventing such an operation results in a compile error. That turns part of the AI safety problem into a programming-language problem.
This Is Not a Sandbox
Make the distinction carefully.
A restricted language does not magically make its host application safe.
If the host deliberately registers an operation called RUN_SHELL_COMMAND, the language can run shell commands. If an exposed Java function contains a vulnerability, the language does not repair it. Resource limits, isolation, authentication, and authorization still belong where they normally belong.
The useful guarantee is narrower:
Generated business logic can only name operations that the application deliberately made part of its vocabulary.
That is very similar to the direction we now see in agent tooling, except the boundary moves from the agent harness into the language presented to the generator. The two approaches are complementary rather than competing. An agent sandbox can determine whether the agent may access a repository. A domain vocabulary can determine whether the program it produces can approve a claim, request additional documents, or initiate a payment.
These operate at different semantic levels.
Domain Capabilities Are More Interesting Than Operating-System Capabilities
Operating-system permissions are necessary, but business applications eventually need a richer vocabulary.
Consider an agent whose process is prohibited from opening arbitrary files and making arbitrary network requests. That is useful.
It still does not answer questions such as:
- May this program approve an order?
- May it request approval but not approve directly?
- Can it read customer risk information?
- Can it initiate a payment?
- Can it calculate a premium but not change the underlying policy?
Those are domain capabilities.
General-purpose programming languages do not naturally provide such a boundary because their strength is precisely that a programmer can combine low-level facilities to implement almost anything. For human-written general-purpose software, that is a feature. For generated business logic, it may sometimes be the wrong abstraction.
A small language with an application-defined vocabulary gives us a different unit of authority: not files, sockets, and processes, but business operations.
We May Be Seeing the New Shape of the AI Programming Stack
None of this means that general-purpose languages are going away, nor that every AI-generated program should use a DSL. Java, Rust, Go, Python, C++, and JavaScript will remain the implementation languages for enormous amounts of software.
But the rapid evolution of agent tooling suggests a useful architectural separation. Humans write the machinery. Models orchestrate the machinery. Deterministic systems constrain and verify the orchestration. And the interface between those layers becomes increasingly explicit. GitHub's managed permissions, IDE sandboxing, multi-model orchestration, persistent agent hosts, and execution-based code review are all different manifestations of this broader change.
The industry is gradually replacing:
Trust the model.
with:
Give the model precisely defined capabilities and verify what it produces.
That is a much more promising engineering principle.
BUBAS is one experiment in taking the same principle into the programming language itself. It is open source, and the implementation, examples, tests, and current design documentation are available in the BUBAS GitHub repository.
Opinions expressed by DZone contributors are their own.
Comments