Governance · · 7 min read

An agent approval is only as strong as its execution boundary

How to tell a displayed decision from an enforced one across local, cloud, and hybrid coding agents.

Editorial illustrationCreated for this story. GrantTap screens below are separate captures made with sample data.

“Approve on your phone” sounds precise until you ask what is being approved. A coding agent might request a shell command, a file edit, a network connection, a secret-bearing tool call, or a deployment. The phone can display the request, but it rarely performs the action. The meaningful security boundary is the runtime that holds the repository and credentials. Its decision must happen before the side effect, and the record should distinguish a proposal, a human response, policy enforcement, and the final result. These are separate events even if a product presents them on one card.

The official documents linked below describe different environments. Claude Code Remote Control connects a mobile or web client to a local Claude Code session with its own permission modes. OpenAI describes mobile Codex approvals for connected work. Cursor documents run modes that affect interruptions for tool calls, while its cloud agents may run commands automatically inside dedicated machines. AWS AgentCore Policy is a different class of managed gateway control: it evaluates tool access through a gateway, not the local shell of every coding assistant. We checked these descriptions on October 3, 2026. A fair comparison begins by naming the tool path that each policy actually covers.

Approval chainEach stage needs its own evidence.
01ProposedWhich exact action and target?
02DecidedWho allowed or denied it?
03EnforcedWhich runtime applied the decision?
04ObservedWhat actually happened afterward?

Follow the action to its runtime

Imagine an agent wants to read a production credential. A push alert might show the proposed command and an Allow button. That visible prompt is useful only if the local process or cloud worker pauses and requires the decision before access. If the UI updates while the host is offline, the user may have recorded an intention rather than authorization that reached the tool. If the agent executes through another route, the policy may not apply at all. The safest evaluation uses a harmless stand-in credential and watches both sides of the boundary in a dedicated test environment.

In GrantTap's model, the phone is a controller for supported local executions. The computer applies global capability denies and task policy, with a global deny winning. This is a product rule about that integrated runtime, not a claim that GrantTap controls all possible provider tools everywhere. A capability can be available without being enabled, enabled without being used, and requested without a confirmed run. Keeping those states separate prevents a settings screen from becoming false evidence. For every provider adapter, inspect which actions actually route through the enforcing hook and which remain provider-native.

Editorial illustrationAI-generated depiction of a decision crossing a computer boundary, not an implementation diagram.

Permission modes are not interchangeable

Claude Code's Remote Control documentation describes a permission mode for sessions the server starts and says interactive sessions can be controlled remotely. Its native permission handling remains part of the Claude process on the computer. Codex's mobile preview brings approval requests into a mobile surface, but exact sandbox and approval settings belong to the connected Codex execution. Cursor's run modes govern when its agent interrupts the user for approval. Its documentation also says cloud agents run in their own dedicated machines and do not ask for an action-by-action approval in the same way. These differences do not form a simple “more secure” ranking.

Instead, choose a desired operating rule. For a personal scratch repository, broad auto-run within a restricted sandbox may be appropriate. For a repository with deploy credentials, explicit review of network and filesystem effects may matter more. For a mixed setup, decide which agent can reach which files and tools before comparing phone screens. A mobile button cannot strengthen a runtime that has already been given unrestricted credentials. Conversely, a well-configured native permission system can be sufficient without adding another control layer. Verify the policy configuration where execution happens and record its scope.

Managed gateways cover a different path

AWS AgentCore Policy illustrates a useful principle: enforce a rule outside the agent's prompt at a boundary it must cross. Its documentation describes policy engines attached to gateways, with enforcement of agent requests that pass through those gateways. That can provide deterministic decisions and logs for gateway-mediated tool calls. It does not imply that a command run directly by a local coding agent on a laptop is intercepted by AWS. A product comparison that treats “governance” as one undifferentiated checkbox hides this scope distinction and may invite unsafe assumptions.

The same principle applies to any local controller. A rule has value when the actual action must cross its enforcement point. Draw the route from proposed action to tool, write down which component can deny it, and test a denied case. Include alternate paths such as provider-native tools, MCP servers, shell access, and cloud workers. If one path bypasses a controller, say so explicitly rather than presenting a global guarantee. This scope map makes it possible to use multiple controls together without claiming that one interface is the universal authority over every tool.

Editorial illustrationAI-generated view of policy layers and evidence; no real provider interface is represented.

Record four different outcomes

A request can be displayed but never delivered to a phone. It can be delivered and approved but fail to reach an offline computer. The computer can receive it and deny the action under a stronger rule. The tool can run but fail for an unrelated reason. These paths have different meanings. A useful event trail records a proposed action, a decision, a host-side enforcement result, and a tool outcome with timestamps and identity. The UI should avoid collapsing them into one success badge. “Unknown” is a legitimate answer when an integration did not report the outcome.

The GrantTap screenshot in this article is a deterministic demo capture. It shows how governance status is presented, not proof of a production enforcement result. In a real evaluation, use one dedicated test Task and two benign tool requests: one permitted, one denied. Watch the host log and provider transcript, then disconnect the phone and repeat. Check whether a queued decision expires, whether policy remains effective while offline, and whether the final state is reported accurately. The test should be safe enough to repeat whenever an adapter or provider version changes.

Put the human at the right point

Requiring approval for every trivial read can overload a person; allowing every action can make the word “approval” meaningless. Group actions by consequence: read-only inspection, local reversible edits, network access, credential use, and irreversible publication. Configure the runtime's default limits first. Then use a phone to decide exceptional requests with enough detail to understand their target and scope. If the phone cannot show that detail, defer the action until you can inspect it at the computer. This is a workflow judgment, not a claim that one provider has discovered a perfect universal policy.

Finally, check what happens after the decision. An approved deployment still needs a successful command, a reachable destination, and verification of the deployed state. A denied operation should leave the protected resource untouched. A governance system earns trust by representing these outcomes honestly and by exposing where its authority ends. The user's strongest question remains simple: which process stopped or allowed this exact action, under which rule, and what evidence shows the outcome?

GrantTap governance view with deterministic sample status. A sample card does not prove a live tool call was allowed or blocked.
GrantTap screen · sample dataGrantTap governance view with deterministic sample status. A sample card does not prove a live tool call was allowed or blocked.

Sources

Next storyCortex Loom: fewer context tokens, with the gaps visible →