AI Agent Security Controls
From Permission Boundaries
to Deployment Readiness
An employee asks an AI agent to query their team’s metrics and project progress, then assemble a report. The task needs boundaries around the data it can read, the operations it can perform, its working files and the report’s intended recipients. This illustrative assessment workload connects those business boundaries to execution evidence: which calls are authorized, how sessions and files are separated, and what access remains after stopping.
Bound execution consequences outside the model
An agent uses a model to interpret a request and choose actions through tools. Misunderstood instructions or malicious content can influence those choices. Controls outside the model determine whether the resulting actions can execute and what resources they can affect.
The assessment separates three questions because each requires different evidence.
| Assessment area | Question | Required evidence |
|---|---|---|
| Reachable capability | What can the agent reach and do? | An inventory of accessible data, tools, credentials and execution paths. |
| Per-call authorization | Who authorizes this particular action? | Allowed and denied calls tied to verified identity, target and parameters. |
| Detection and containment | Can abnormal execution be reconstructed and stopped? | Correlated execution records and a stopping drill that checks residual access. |
A limited tool inventory establishes a capability boundary. It does not establish that every use of an available tool is authorized. A rejection record establishes a decision at a particular control point, while containment evidence establishes what happens after stopping. None substitutes for the others.
For the team-analysis assistant, these questions begin with an approved business use. The permitted task defines the boundaries that the platform must enforce.
Define the team-analysis assistant's permitted use
The assessment use permits an employee to query their own team’s metrics and project progress, assemble a report and save working files within the assigned scope. It excludes cross-team data access, modification of business data and delivery to unapproved destinations. This is an evaluation scenario, not a statement that a customer has deployed it or that the reference application implements the complete workflow.
The same query tool can be permitted for one employee and denied for another. Team membership comes from verified identity attributes. A prompt asking for another team’s information cannot expand the caller’s authority, and a model-generated team parameter cannot replace the authorization decision.
The capability inventory includes every enabled execution path. Model Context Protocol, or MCP, standardizes how applications connect to data and tools; it does not establish business permissions. Shell commands, HTTP requests, SQL access and direct use of credentials also belong in the inventory if they are available. For example, if an agent retains a database-write credential, removing a database-write tool from its catalog does not remove the equivalent direct-access capability. This is a conditional assessment example, not a capability attributed to the team-query demo.
Read-only business access is also distinct from writing a report file. The former limits operations on the source data; the latter requires file permissions and recipient controls.
Required controls must pass for the approved use. If a necessary boundary remains unverified or fails, the use must be reduced and retested, or deployment paused.
Identity, authorization and credentials in a tool call
One employee identity across several services
A tool call involves identity, authorization and credentials. Identity establishes who is making the request. Authorization decides whether that identity may perform the requested action. Credentials are the token or key presented to a receiving service.
The reference platform’s enterprise SSO example connects employee identity to team APIs, the application interfaces serving team data. Its get_report and list_projects tools return mock team metrics and active projects. They illustrate team-scoped queries rather than a production reporting application.
Enterprise single sign-on, or SSO, authenticates the employee through an identity provider, or IdP. The IdP issues a signed JSON Web Token, or JWT. The token’s sub identifies the subject and team carries this example’s team membership. Its audience, carried in the JWT’s aud claim, identifies the intended receiving service.
Amazon Bedrock AgentCore Runtime validates the inbound user JWT and passes the same token to AgentCore Gateway, which connects the agent to its tools. Gateway also validates the inbound token. For team-a and team-b, it obtains exchanged tokens for the backend calls. In this configured example, those replacement tokens preserve sub and team; that preservation is not automatic for every IdP.
Backend authorization
Gateway interceptor authorization
Place authorization where the user identity is understood
The team-a and team-b services understand employee identity. Their OpenID Connect, or OIDC, middleware verifies the IdP token, and their backend code checks team membership before executing tools/call. A cross-team invocation is rejected at the backend.
The distinction between tool discovery and execution matters. The source deliberately permits authenticated protocol initialization and catalog listing across teams so Gateway can aggregate the catalog. Seeing a tool in that catalog does not authorize its invocation. The team restriction applies when the caller attempts to execute it.
Team-c uses a different arrangement. Its backend has no SSO support and accepts an API key, a shared credential that does not identify the employee. The Gateway deployment code therefore attaches a Lambda REQUEST interceptor: a function that examines the request before forwarding. It reads the team claim from the inbound JWT already validated by Gateway and rejects unauthorized team-c tool invocations before they reach the backend. The interceptor does not independently re-validate the token’s signature.
For a forwarded call, Gateway injects the target API key. Team-c checks that key, while the interceptor makes the employee-level authorization decision. Authentication success at Gateway and possession of a backend credential remain separate from approval of the requested action.
Policy enforcement and evidence of rejection
AgentCore Policy provides an optional policy mechanism distinct from the reference’s Lambda interceptor. Its behavior depends on Gateway’s policyEngineConfiguration. In LOG_ONLY, policies are evaluated and decisions logged without blocking calls. In ENFORCE, allow and deny decisions are applied.
Authorization evidence needs the caller, target, action and decision, correlated with execution records. For team-a and team-b, a rejected request reaches backend authorization but must not execute the tool. For team-c, an interceptor rejection occurs upstream, so there can be no downstream execution log for that request. The upstream rejection and absence of corresponding execution must be reconciled.
Negative tests also cover direct calls and equivalent bypass paths. A denial through Gateway alone does not establish that the same action is unavailable through another enabled route.
Separate execution, file and model-credential boundaries
Saving a report introduces boundaries beyond the team query. A session comprises the running agent interaction and its corresponding platform records. Its workspace contains saved working files. Amazon S3, an object-storage service, can hold those files outside the execution environment.
AgentCore Runtime provides session execution in dedicated microVMs, small virtual machines that isolate CPU, memory and filesystem resources. The reference platform’s Dev Workbench adds interactive sessions, workspace credentials and model-access mechanisms. The team-query demo and Dev Workbench are different implementations. Their mechanisms illustrate separate control requirements; they do not demonstrate a single integrated query-to-report application.
Session A
STS credentials for session A’s prefix onlySession B
STS credentials for session B’s prefix onlyworkspaces/*. Separately issued session credentials scope S3 access.Verify compute isolation, session ownership and file permissions separately
AWS’s Runtime security documentation describes dedicated microVM isolation and assigns session-to-user mapping to the application backend. These are separate responsibilities. Compute isolation cannot correct a backend that connects an employee to somebody else’s session. The employee would be entering the wrong isolated environment.
The Workbench keeps session records partitioned by user. Its session endpoints retrieve the caller’s record before connecting or performing management operations. That ownership check governs access to the session; it is not a test of the microVM’s kernel isolation.
File authorization is separate again. Isolated execution environments can still access shared object storage through workload credentials. An AWS Identity and Access Management, or IAM, role defines the workload’s AWS permissions, independently of the employee’s login identity. If a shared execution role can read every workspace, microVM isolation does not narrow that storage access.
AWS also documents that code inside the microVM can access execution-role credentials. The role’s permissions therefore form part of the agent’s reachable capability.
Limit workspace credentials to the current session
The Workbench’s Runtime IAM configuration gives the interactive execution role no access to workspaces/*. A separate workspace-access role supplies the storage permissions.
After establishing session ownership, the backend selects the session’s runtimeSessionId. The workspace credential service uses AWS Security Token Service, or STS, to assume the separate role and issue temporary AWS credentials. An inline session policy narrows object access to workspaces/{runtimeSessionId}/*. The ListBucket permission, which allows listing object names, is constrained to the same session prefix.
This lets the session save its report and working files without receiving the role’s full workspace scope. The temporary credentials remain readable inside the container, so their permitted actions and resources must match that exposure.
A session policy narrows the assumed role’s permissions. It cannot repair an incorrect user-to-session binding or credentials delivered to the wrong session. If the backend selects another employee’s session, a correctly scoped policy would still grant access to the wrong files. Ownership, credential delivery and storage scope each require verification.
Model client → loopback shim
Treat material accessible inside the container as readable by the user.
Short-lived session credentialllm-edge
Read the backend grant; check credentials, model and routing.
Inject the real gateway key hereLLM gateway
Handle authorized model requests.
Keep the platform model key outside the agent container
Report assembly may require further model calls after retrieving team data. The model-access credential needs its own boundary.
Dev Workbench exposes a real terminal. Its model proxy implementation explicitly treats process environment, files, memory and execution-role credentials as accessible inside the session. Removing a terminal button from the interface would not establish that tool-executing code cannot read those resources.
The platform implements llm-edge, a proxy outside the agent container. It is not a managed AWS feature. The request path runs from the model client to a container-local loopback shim, then to llm-edge in the private virtual network, or VPC, and onward to the large language model gateway. The edge service’s role can read the real platform gateway key. The container instead retains a short-lived session credential.
The backend records the session’s authorization in a grant: a server-side record specifying permitted model access. This platform record is distinct from downstream OAuth tokens used to access other services. On every request, llm-edge reloads the grant, checks the presented token’s digest and expiry, and reads the upstream destination, secret name and model allowlist from server-side state. The container cannot nominate a different real upstream or secret by reporting new routing values.
The route implementation permits only these HTTP method and path pairs to be forwarded:
POST /v1/messages
POST /v1/messages/count_tokens
GET /v1/models
Requests naming a model are checked against the stored allowance. Management routes cannot use the proxy to reuse the platform key. The boundary depends on external key placement, server-side authorization and restricted forwarding, not on a promise that session credentials are invisible in memory.
Check private connectivity, encryption and alternative model paths
Private connectivity limits reachability. Transport Layer Security, or TLS, encrypts traffic. They are different controls. This implementation defaults to internal HTTP and supports configurable TLS; placement inside a VPC does not automatically encrypt a connection.
The Runtime IAM configuration also permits another model-access path: direct Amazon Bedrock invocation. The llm-edge allowlist applies to calls through that proxy and cannot restrict bypass calls through a different path. Every enabled model route needs its own permission and stopping checks.
Check residual access after stopping a session
An abnormal team query must be traceable to the employee, session, tool, authorization decision and actual backend output. Platform logs describe the interaction, authorization logs describe decisions, and backend records describe execution and results. An agent’s account of its actions does not establish all of those facts.
Stopping also involves distinct objects.
| Object | Required stop action | Evidence |
|---|---|---|
| New requests | Block subsequent access through applicable entry points. | A new query is rejected, with correlated authorization and execution records. |
| Running session | Terminate the running session when required. | Runtime state and streaming behavior confirm termination. |
| Issued credentials | Revoke access or apply verified compensating controls that block access. | Previously issued credentials fail the relevant access checks. |
| Submitted business tasks, if any | Check and handle tasks in the downstream system. | Downstream task records establish cancellation, completion or remaining effects. |
AWS’s StopRuntimeSession terminates the specified running Runtime session and stops ongoing streaming responses. It is not an OAuth-token revocation operation or a downstream task-cancellation API.
The final category applies only when the workload submits business tasks. For example, if a coding agent submits a pull request, or PR, that triggers continuous integration, or CI, the resulting automated build-and-test job requires separate handling in the downstream system. Stopping the agent does not cancel that job. The team-query demo does not submit CI jobs, and stopping does not promise to reverse irreversible effects.
What the platform's stop endpoint actually does
The Workbench’s POST /{session_id}/stop route checks ownership, marks the session dormant and calls llm_credentials_service.revoke. It does not call StopRuntimeSession; it expects Runtime to reclaim the idle execution environment.
The credential service’s revocation implementation attempts to delete the model grant. Deletion exceptions are caught and logged as warnings. A successful stop-route response therefore is not independent evidence that grant deletion succeeded. The stored authorization and subsequent model-access behavior must be checked.
Revoking the model grant does not automatically invalidate temporary S3 credentials or downstream OAuth tokens. Those credentials have separate lifecycle and revocation semantics. Where immediate invalidation is unavailable, the acceptance requirement is verified compensating controls that block access, rather than simply waiting for expiry.
Response deadlines depend on the workload’s risk. The agreed deadline and the measured drill time are different records; neither creates a universal stopping-time guarantee.
A stopping drill: the next query and residual credentials
A stopping drill begins with an identifiable session and its issued credentials. After triggering the stop action, another team query tests whether new requests are blocked. Runtime checks establish whether execution and streaming have ended, independently of the platform’s displayed session status.
Previously issued model, storage and downstream credentials are then replayed or checked against their respective access paths. The drill reconciles request outcomes, authorization decisions, backend execution and measured timings. Any residual access and the blocking applied to it remain explicit in the record.
If the workload has submitted downstream tasks, those tasks are checked separately. This is an acceptance method for the deployed chain, not evidence that the reference platform has already passed it.
Use execution evidence to approve deployment
Deployment approval rests on a capability inventory, evidence from allowed and denied calls, bypass tests, stopping records and clear ownership of the controls. The inventory establishes what the agent can reach. Call evidence establishes whether identity and action boundaries hold. Stopping evidence establishes what access remains after intervention.
Existing enterprise SSO, permissions, approvals and logging can support this chain. Their presence alone does not establish that they constrain the agent’s actual execution paths. The complete path from employee identity through authorization, credentials and backend execution needs verification.
The reference tests illustrate the distinction between implementation and deployment evidence. e2e_session_isolation.py uses moto, an AWS-service mocking library, and a stub Runtime to test application-level session binding and grant ownership. It does not exercise real microVM kernel isolation. e2e_team_auth.py defines own-team calls, cross-team rejection and direct-call negative tests for the authorization chain. A test definition is not a record of successful execution against a deployed environment. No platform AWS end-to-end test success is claimed here.
Required control gaps cannot be offset by unrelated strengths. If the necessary boundaries do not pass, the approved use must be narrowed and retested, or deployment paused. Readiness is decided by the business-defined scope and actual execution evidence.