Security Boundaries for Agents with Organizational Authority

Security Boundaries for Agents with Organizational Authority

Enterprises build agents to increase throughput. Scattered across many departments, these agents interact with internal employee systems, execute business processes, and collaborate with one another to complete increasingly complex tasks.

Credal has helped organizations with thousands of employees make agent building accessible to everyone and grow adoption to 25% of the workforce, with nontechnical employees accounting for 75% of builders. This was made possible without weakening enterprise controls by enforcing a clear permissions model: agents act on behalf of the invoking user and can access only the systems, data, and other agents that user is already authorized to access.

As enterprises pursue more complex and cross-functional use cases, this model becomes restrictive. A new employee may be allowed to request onboarding without having permission to provision accounts, approve purchases, or edit HR records. Requiring every onboarding agent to operate within the requester’s access would prevent this workflow from completing.

Supporting this work requires agents to exercise organizational authority through their own scoped permissions, explicit permission escalation, or a combination of both.

This changes the security problem. Once the invoking user’s permissions no longer bound everything a workflow can do, the organization loses clear bounds on what can go wrong as well as an owner when it does. The infrastructure running these agents must determine and enforce who may invoke an agent’s additional authority, under what conditions, and how those constraints survive collaboration. Governing this workforce requires understanding the threat model created by those interactions.

 

What are people building today?

Through our work with enterprises such as Flatiron, ACF, Wise, and Checkr, we see four recurring categories of agent use cases:

  1. Personal assistants help an individual do more. They typically use the user’s delegated authority, with actions bounded by that user’s permissions and intended task. Example: after an AE’s sales call, update CRM records and draft their follow-up.
  2. Employee service agents deliver internal services using permissions the requesting employee does not have. The employee is authorized to request the outcome; the agent is authorized to perform the underlying operations. Example: provision software access after verifying eligibility and required approvals.
  3. Business process agents carry organizational processes forward. They may require permissions spanning systems and teams, independent of any individual employee’s access. Example: when a target account enters Salesforce, research its buying committee and coordinate an ABM campaign.
  4. Customer-facing agents deliver services to external users. They connect customer requests to internal capabilities, requiring controls around customer identity, tenant boundaries, and disclosure. Example: investigate a customer’s order and arrange an eligible replacement.

Any of these agents can respond to a direct request, an application event, a schedule, or another agent. Running autonomously or automatically does not necessarily mean operating independently of a user’s authority.

Across these categories, there are two permission models: agents acting within the invoking user’s delegated authority, and agents exercising authority independently granted by the organization. In the second model, the agent may perform operations the requester cannot execute directly.

Credal currently uses the former model. This post focuses on the additional security problems created when an agent can exercise authority its requester does not possess. We set aside issues such as work outliving its authorization, duplicate execution, cancellation, and cumulative limits: these affect both user-delegated and organization-authorized agents. Organizational authority can increase their consequences, but they are not specific to the permission model examined here.

 

Problems that may arise from agents with organizational authority

Agents can make mistakes, encounter malicious instructions in documents or tool results, or be compromised. We therefore treat model-generated instructions, agent-to-agent messages, and external content as potentially erroneous or adversarial, even when they arrive through an authenticated identity.

The assets at risk are the integrity of privileged operations, the confidentiality of restricted information, and the integrity of the authorization state itself. A failure occurs if the system performs an operation the organization did not authorize, discloses information to a recipient who is not entitled to receive it, or expands a task’s authority without a valid grant.

Potential adversaries include malicious requesters, compromised agents or tools, and untrusted content that manipulates an otherwise honest agent. For the purposes of this model, the identity provider, policy engine, and controlled execution boundary form the trusted computing base. We assume these components are independently secured and correctly configured; compromise of them is outside the scope of this post.

Agents may interpret requests and propose operations, but they cannot be the final enforcement point. Every privileged operation and every release from a privileged context must be checked by the trusted boundary at the moment it occurs.

Readers with a security background will recognize this as a form of the confused-deputy problem. Capability-based authorization, complete mediation, provenance, and information-flow controls provide useful foundations. What changes with agents is that the deputy interprets natural language, transforms evidence, delegates dynamically, and may be created by people without security expertise.

 

Problem 1: An authenticated message establishes who sent it, not whether the requested action is authorized.

Permission to contact an agent and permission to invoke a particular use of its authority are separate decisions. An IT agent may be able to provision software, but another agent contacting it should not automatically gain that capability. Similarly, all employees may have access to an agent with Slack access, but an analyst should not be authorized to ask about what is lurking in the CEO’s DMs.

However, request screening cannot establish authorization because natural-language asks often leave operational boundaries undefined. “Handle this customer’s issue” leaves open whether an agent may inspect records, modify an account, or issue a refund. The employee may be entitled to request help without authorizing every operation the agent considers useful. Additionally, language does not reliably reveal intent. “Export these customer records” could describe an approved migration or attempted data theft.

Screening request language alone therefore creates a difficult tradeoff. Escalating every ambiguity produces interruptions and alert fatigue. Accepting plausible explanations allows someone to present an unauthorized action as ordinary business.

Request language can help the agent understand what the user wants, but it should not be enough to authorize an action. We propose that before a privileged operation is performed, the system should check who requested it, what exactly will happen, what it affects, and whether current policy allows it.

 

Problem 2: Delegation can introduce authority that the original task never had.

A user asks agent A to investigate a billing issue. A asks B to analyze transactions, and B asks C to modify the account. Each step may appear useful, however agent C’s ability to modify accounts does not establish that this task was authorized to use that ability.

Handoffs can also change the meaning of evidence. “A refund might be appropriate” can become “process the refund,” while a summary saying “approved” may omit who approved it, what they approved, and under which conditions. The downstream worker sees a plausible instruction but lacks the context needed to assess its authority.

Malicious instructions can travel the same route. An agent may extract a directive from an untrusted ticket and repeat it to another agent as its own request. The message now comes from an authenticated internal identity, even though the instruction originated outside the trusted workflow. Delegation can therefore amplify both ordinary misunderstanding and deliberate manipulation. Each handoff should preserve the original requester, the task’s scope, and where the instruction came from. Delegation can pass along or narrow authority the task already has, but it should not expand it without a separate grant or approval.

 

Problem 3: Permission to access information does not establish permission to disclose it.

Organizational authority creates a gap between what an agent may inspect and what its requester may learn. An IT agent may retrieve an employee’s recovery key using permissions the employee does not possess. The employee can be entitled to their own key without being entitled to inspect the underlying API responses or other information available to IT. The response leaving that privileged context crosses a trust boundary. Returning everything to a conversational agent and asking it to redact the answer expands exposure. Even a summary can reveal information the requester was never authorized to receive.

The same distinction applies to derived information and persistent memory. Authorization to retrieve information for one task does not authorize storing it in shared memory or reusing it during a later execution. Disclosure, persistence, and later reuse are separate authorization decisions unless a trusted policy explicitly permits the information to be declassified.

 

Problem 4: Human approval can be insufficient or ineffective.

Approval can clarify intent, grant authority, check correctness, or accept exceptional risk. Approvals also introduce permission questions of their own: which operations or boundary crossings require approval, which agents may initiate those requests, who may approve them, and who may change or bypass the requirement? If requesting approval depends on the agent deciding that a step is sensitive, a misled agent may never raise the request. Approval can also cover the wrong decision: permission to retrieve information does not necessarily authorize releasing it. Having a human somewhere in the workflow does not establish that the right decision reaches the right reviewer with the right evidence.

Approval should therefore be represented as a scoped authorization artifact, not as a generic signal of comfort. It should bind to the proposed operation, target, parameters, recipient, limits, and expiration. Policy should determine when approval is mandatory, and the execution boundary should enforce its presence. An agent may request additional review, but it should not be able to waive required review.

 

Problem 5: Authority to act does not establish who owns the outcome.

Under delegated authority, the requester provides a clearer anchor for responsibility because the agent is limited to actions available to that user. Even then, accountability may be shared among the requester, builder, platform operator, and policy owner. Organizational authority makes that allocation less direct.

An employee may initiate a task, a business owner define the policy, an approver authorize a sensitive step, and several agents execute using independently assigned permissions. Each controls a different part of the outcome. If the result causes harm, knowing which agent acted does not establish who was responsible for the decision.

The requester may never have seen the restricted information or controlled the privileged operations. An approver may have authorized one specific release, while an agent owner may have configured an entire class of releases to proceed automatically. Event-triggered work may have no human requester at all. An execution history can reconstruct who requested, configured, approved, and performed the work while still leaving unresolved who owns the outcome and is responsible for putting things right.

This is the many-hands problem. If agent owners own outcomes, they need to scope standing permissions tightly and see every use of them. If approvers own outcomes, approvals must bind to concrete operations, as above. If the organization owns outcomes through policy, the policy must be the thing actually enforced. An infrastructure that has not chosen which model to use, it cannot tell commuicate to its builders which of these controls they are relying on.

 

Security properties the infrastructure must preserve

Taken together, these problems suggest four simple rules:

  1. Knowing who sent a request is not enough to authorize it
  2. An agent’s general capabilities should not determine what it may do in every task
  3. Delegation should not silently expand authority
  4. Human approval checks should not be entirely orchestrated by the agents
  5. Permission to access information should not automatically permit its disclosure or later reuse

These rules should be enforced by the infrastructure, not left to the model to interpret.

 

How do we build the right infrastructure?

The product challenge is to combine power, governance, and accountability while keeping agent building accessible to nontechnical employees. People should be able to delegate useful work without having to design its security model from scratch.

We suggest that agent execution infrastructure should

  1. Make the secure path the default path. It should be easy for nontechnical users to build things quickly and correctly with minimal risk of security or privacy issues.
  2. Reserve human judgment for ambiguity and exceptions. Routine, authorized work should complete without interruption; reviewers should see only the decisions that require a person.
  3. Enforce explicit rules deterministically outside the model. Agents may interpret requests and propose actions, but trusted components must decide whether each privileged operation, disclosure, persistence event, or reuse is allowed. This means moving things like human approval message bodies into the agent execution layer.
  4. Make the controlled execution path unavoidable, so that an agent cannot bypass rules through shell access, exposed credentials, or a tool that reaches the backend directly.
  5. Represent task authority explicitly at runtime. Separate an agent's identity from the three things that determine what it may do in a given execution: its standing permissions, the task-specific authorization it received, and current policy. Effective authority is the intersection of the three. Delegation can only narrow it unless an explicit grant is permitted for otherwise.
  6. Bind approvals to concrete decisions. An approval should identify the operation, target, parameters, data recipient, scope, and expiration. Approving access to information should not automatically authorize its release.
  7. Make outcomes verifiable and responsibility clear. Choose an accountability model, show builders which one they are operating under, and enforce it.
  8. Let autonomy grow with evidence. Evaluate complete workflows, including indirect instructions, delegation chains, and concurrent requests, and measure unauthorized actions and disclosures, legitimate work incorrectly blocked, task correctness, review burden, and maximum damage before intervention. Passing an evaluation is evidence for the conditions tested, not proof against an adaptive adversary. Use it to reduce supervision by workflow, resource scope, and consequence while the deterministic boundaries stay in place, and re-run it when models, tools, or policies change
  9. Detect and surface anomalies and risky beheavior from the execution record rather than the agent's self-report, so operators can pause work, revoke authority, and recover, including from effects like disclosure that cannot be undone

 

Closing thoughts…

The opportunity ahead is to give teams the capacity to pursue work they could never take on before. Infrastructure that makes authority clear and collaboration dependable will help enterprises delegate more ambitious outcomes with confidence. These foundations can make AI employees a trusted part of the organization, expanding what its people are able to accomplish.

 

We are building this infrastructure at Credal, contact us at founders@credal.ai if you’re interested.

Give every team access to governed MCPs

One platform for all agents. Full visibility for admins, full access for teams.

Ready to dive in?

Get a demo