The Direct Answer: Treat AI Agents as Nonhuman Users with Controlled Authority

Enterprises secure autonomous AI agents by treating them as nonhuman users and software actors, not as ordinary chatbot features. A production agent needs an attributable identity, least-privilege access, constrained tools, explicit approval boundaries, complete activity records, continuous monitoring, and a tested response process. The central question is not whether an agent can browse a file, send an email, execute code, or call an API; it is who authorized that action, under which policy, using which data, with what resulting risk. Conventional identity and access management remains necessary, but it is insufficient when an agent can choose sequences of actions, interpret natural-language instructions, or operate across several systems.

Also worth reading: What is the definitive agentic AI regulatory compliance checklist for enterprises deploying autonomous AI systems in 2026? · What is an autonomous agent governance framework and how should enterprises implement one in 2026? · Which Liability Clauses Should Enterprises Use for AI Agents in 2026?

A defensible production model combines identity, zero-trust access, policy enforcement, data controls, tool security, human approval, and observability. Identity should distinguish the human sponsor, the agent, the model, the runtime, and the service account so that actions are not attributed only to a shared API key. Access should be scoped to particular repositories, folders, customers, transactions, and tools rather than granted through inherited administrator permissions. Every agent should also have a written purpose, an owner, permitted actions, prohibited actions, spending or data-transfer limits, and a revocation mechanism.

Frameworks such as SOC 2, ISO 27001, and HIPAA help organize assurance, but none certifies an AI agent as safe. SOC 2 evaluates controls relevant to an organization’s commitments and system description, while ISO 27001 certifies an information security management system. HIPAA applies to covered entities, business associates, and protected health information; it does not regulate AI agents generally. An organization can hold these certifications and still deploy an unsafe agent because certification does not test every tool call, prompt injection path, emergent behavior, or authorization decision.

The practical baseline is therefore stronger than asking whether an agent passes a questionnaire. Ask whether its identity can be traced, its permissions expire, its instructions can be logged, its outputs can be inspected, and its actions can be stopped. If the answer to any of those questions is no, the agent is not ready for consequential autonomous production work. A low-risk research agent with no external tools and no business data may be deployed under a lighter control model; an agent that executes payments, modifies production systems, or communicates externally requires substantially tighter restrictions.

Why Existing Enterprise Security Controls Are Not Enough

AI agents differ from conventional applications because instructions arrive indirectly through users, retrieved documents, websites, emails, application outputs, and model-generated plans. An attacker can place text such as a fake system instruction in a document and cause a tool-using agent to expose data or change its behavior. This is often called indirect prompt injection, and ordinary malware defenses may not detect it because the malicious instruction can be natural language rather than executable code. Perplexity’s scraping dispute illustrates a related boundary problem: automated clients may ignore website restrictions or user-agent policies, making access governance and authorization part of agent security.

Agents also create a chain of delegated authority. A human may authorize a broad objective, such as resolving a support case, while the agent independently retrieves records, invokes a CRM update, and sends a response. Each step can be individually valid and collectively harmful. Traditional access control may approve access to the CRM without understanding that repeated agent actions could exfiltrate sensitive fields, create fraudulent refunds, or silently alter customer records. Security teams therefore need policy decisions based on context, including data classification, action type, destination, transaction size, confidence, user role, and whether a human approved the specific action.

Another problem is hidden dependency risk. Agents commonly rely on models, retrieval stores, vector databases, plugins, MCP servers, browsers, code interpreters, identity providers, payment services, and third-party APIs. A compromise or policy change in any dependency can change the agent’s behavior without changing its source code. The research context’s reference to hidden dependencies is important because inventory must cover runtime components and data flows rather than only the model vendor. A prompt can be harmless in isolation but dangerous when combined with an email tool, a permissive service account, and access to regulated data.

Industry messaging increasingly treats “what did my agent do?” as a first-class security question. NVIDIA has promoted an open agent safety platform, DigiCert has connected agent identity to that ecosystem, and vendors such as Zscaler and 1Password have announced controls for agent discovery, zero-trust access, and credential brokering. These developments indicate where the market is moving, but announcements are not proof of efficacy. Buyers should demand evidence from adversarial testing, permission-denial tests, log-quality reviews, incident exercises, and a demonstration that compromised instructions cannot bypass policy.

A Production Control Model for Enterprise AI Agents

The first control layer is identity. Create a distinct identity for every production agent and, where possible, for every delegated tool session. Do not let agents use employee passwords, shared administrator keys, or unrestricted cloud credentials. Bind the agent identity to its owner, environment, workload, and approved purpose, then issue short-lived credentials through a broker or workload identity mechanism. Human sponsors should be able to suspend the agent without disabling unrelated services, and security teams should be able to answer which agents, humans, repositories, and service accounts participated in any action.

The second layer is policy-based authorization. Default-deny is the appropriate target for tools that write data, execute code, make purchases, change permissions, or transmit records externally. Policies can require a read-only role for discovery, a sandbox for code execution, and human approval for consequential writes. For example, an agent may search approved support tickets automatically, but it should not export a full customer history or issue a refund above a defined threshold. A $500 transaction threshold may fit one business and be meaningless for another, so risk limits must be tied to data sensitivity, reversibility, and business impact.

The third layer is constrained execution. Containers, microVMs, dedicated sandboxes, restricted shells, allowlisted packages, isolated credentials, and temporary storage reduce the blast radius of a compromised agent. Network access should be limited to named hosts and methods, while file access should be limited to approved locations. Agent instructions and retrieved content should be marked as data rather than trusted control instructions. A practical review test is to place a malicious instruction in a document, email, or web page and verify that the agent neither reveals secrets nor invokes unauthorized tools.

The fourth layer is observability and intervention. Log prompts, retrieved context, tool names, arguments, policy decisions, model and agent versions, approvals, outputs, errors, and final business effects. Logs must preserve enough information to reconstruct an action without recording unnecessary secrets or regulated data. Real-time alerts should cover denied actions, repeated failures, unusual destinations, bulk data access, privilege escalation, new tool registration, and attempts to override human instructions. Teams should be able to pause the agent, revoke credentials, preserve evidence, and notify affected owners within minutes rather than waiting for a nightly report.

Practical Steps for Securing an Agent Before Production

Start with an action inventory rather than a model evaluation. For every tool the agent can call, record what data it can read, what it can change, who benefits, how the action can be reversed, and who is accountable. Classify actions into low, medium, and high consequence, then assign controls appropriate to each class. A read-only internal search agent may operate autonomously with strict logging, while an agent that deploys code, transfers money, changes access rights, or sends legal commitments should remain bounded by human approval.

Next, create a threat model for instruction manipulation, credential theft, data exfiltration, excessive agency, tool poisoning, memory contamination, supply-chain compromise, and confused-deputy behavior. Test both direct attacks and indirect attacks carried through documents, web pages, email, shared chats, and application records. Include tests for attempted policy rewriting, request smuggling, malicious files, poisoned retrieval results, Unicode tricks, replayed approvals, and concurrent actions. Record the model, prompt, tool schema, policy, and test date because a passing result can expire after a model, plugin, or data-source change.

Then stage deployment through four gates: development, sandboxing, limited production, and expanded production. In development, use synthetic data and no production credentials. In sandboxing, reproduce realistic permissions and test prompt-injection and tool-abuse cases. In limited production, restrict users, data, transaction sizes, and execution time while analysts review every anomaly. Expanded production should occur only after a defined observation period, named owners accept residual risk, and incident-response exercises demonstrate that agents can be stopped safely.

Finally, assign operational ownership. The business owner defines acceptable outcomes, security owns cross-system controls, privacy evaluates data use, legal reviews regulated or contractual restrictions, and the agent platform team owns telemetry and availability. A quarterly review is a reasonable minimum for a stable low-risk deployment, while agents using external tools, sensitive data, or write access may need monthly or continuous reassessment. Any new model, tool, data source, permission, or autonomous workflow should trigger a documented change review.

Comparing Control Options, Frameworks, and Deployment Models

Enterprises can combine several control approaches, but the choice depends on consequence, data sensitivity, and the degree of autonomy. The table below compares common options rather than presenting a single certification or vendor category as sufficient. It also makes clear why a staged, defense-in-depth model is usually more defensible than relying on a dashboard, an agent framework, or a general-purpose identity provider alone.

FeatureModel or rule-based guardrailsHuman-approved agent workflowIsolated autonomous agent
Best fitRepetitive, predictable decisionsLegal, financial, HR, or customer-impacting actionsLow-risk internal research or analysis
Approval modelAutomatic policy checksHuman approves consequential actionsAutomatic within narrow limits
Main strengthConsistent and inexpensiveClear accountability for high-risk actionsGreater task flexibility and throughput
Main weaknessCannot understand every contextSlower and vulnerable to approval fatigueLarger blast radius if controls fail
Required evidencePolicy tests, denial logs, exception processApproval record, identity, action traceSandbox, red-team tests, monitoring, rapid kill switch
Typical cost profileLower engineering cost, moderate operations costHighest workflow and review costPotentially lower unit cost, highest assurance cost
Good deployment gateRead-only or low-consequence tasksExternal writes and regulated decisionsProven, bounded, low-impact internal tasks
SOC 2, ISO 27001, and HIPAA should be treated as assurance context rather than mutually exclusive product choices. A SOC 2 report can support customer assurance about selected trust-service criteria, while ISO 27001 provides a systematic risk-management and control framework across the organization. HIPAA compliance can be necessary where an agent processes protected health information, but it still requires access controls, audit controls, integrity safeguards, workforce policies, and business-associate arrangements. A regulated deployment may need all three frameworks plus agent-specific controls, yet obtaining a report or certificate does not eliminate prompt injection, unsafe tool use, or excessive permissions.

Open-source agent control planes can reduce configuration work and may be attractive for technically mature teams. Commercial platforms may offer integrated identity, policy, observability, and incident tooling, which can be valuable when an enterprise lacks engineering capacity. A managed model can simplify operations, but buyers must identify where prompts, traces, retrieved data, and logs are stored, which subprocessors receive data, and whether customers can export evidence for their own compliance program. A build-versus-buy decision should compare total operating cost over at least 24 months, not just license fees.

Common Mistakes That Create False Confidence

A frequent mistake is treating prompt instructions as the primary security boundary. System prompts can reduce accidental behavior, but they are not a reliable authorization mechanism because models may misinterpret, ignore, or be influenced by conflicting instructions. Permissions must be enforced outside the model. Another mistake is giving an agent broad access because a prototype used a convenient service account; production scope should be derived from an explicit task and periodically checked against actual behavior.

Teams also confuse successful functional testing with security validation. An agent that completes support tickets correctly under normal conditions may still expose an entire database after encountering a crafted instruction in one ticket. Testing should include hostile inputs, malicious retrieved content, excessive tool sequences, boundary values, and failure recovery. It should also test the control plane itself, including token issuance, approval bypass, log tampering, administrator changes, and recovery from a compromised model provider.

Overreliance on human approval creates a different failure. Approvers may click through hundreds of prompts, especially when the interface does not explain the consequence or show the exact action. Approval should be reserved for meaningful decisions, show a concise action summary, use durable records, and support configurable thresholds. Automation can handle low-risk steps, while a human reviews unusual data access, external transmission, legal language, financial movement, or privilege changes.

Finally, many organizations monitor model responses but not real-world effects. A response that looks correct may still send the wrong recipient, alter the wrong record, or trigger an irreversible downstream action. Logs need to connect the agent’s intent and tool call to the resulting email, API request, database update, deployment, or payment. If that chain cannot be reconstructed, incident response becomes guesswork and customer notification decisions become unreliable.

When to Act, What It May Cost, and What Success Looks Like

A new control program should begin before an agent receives production credentials, because retrofitting identity, logging, and approval boundaries is more expensive than designing them into a pilot. Organizations should act immediately when an agent can execute code, access confidential or regulated information, send external messages, change financial records, modify permissions, or take actions that cannot be reversed. Even read-only agents deserve prompt-injection tests and data-access monitoring if they can traverse internal systems or retrieve untrusted web content.

There is no honest universal market price for enterprise AI agent security because pricing depends on deployment scale, cloud consumption, data volume, model usage, integration work, compliance scope, and whether the buyer uses an existing platform. A small pilot may cost little beyond staff time and can use open-source components, while a regulated production program can require dedicated engineering, identity integration, policy development, red-team exercises, logging infrastructure, and legal review. Budgets should separate platform fees, model and infrastructure usage, integration, assurance, and ongoing operations. A low license price can be offset by custom integrations and a large increase in human review time.

Useful success measures are behavioral rather than cosmetic. For example, an organization can target 100% of production agents having a named owner and unique identity, 100% of external write actions requiring policy approval, and detection of test prompt injections within a defined interval such as five minutes. Teams can also measure credential lifetime, number of standing administrative permissions, percentage of tool calls with complete traces, time to revoke an agent, and frequency of control retesting after material changes. These figures should be set against the organization’s own risk, rather than copied from a vendor benchmark.

A 90-day initial program is plausible for many organizations, but it will not certify every agent as safe. The first month can establish inventory, ownership, and data classification; the second can implement identity, least privilege, sandboxing, and logging; and the third can run adversarial tests, a limited deployment, and an incident exercise. The broader research context reports an estimate that 85% of enterprises are running AI agents while only 5% trust them enough to ship, but that figure should be treated as an industry claim rather than a universal audited statistic. The practical lesson is that experimentation is widespread and assurance is not, which is precisely why production governance should precede broad autonomy.

The Enterprise Decision: Constrain Autonomy, Measure Evidence, Improve Continuously

The definitive enterprise answer is to grant agents only the authority required for a defined task, enforce those limits outside the model, and make every consequential action attributable and reversible. Start with read-only or low-impact agents, then introduce human approval before external writes or regulated processing. Give each agent a unique identity, short-lived credentials, restricted tools, isolated execution, and a rapid shutdown path. Record the instruction, retrieved material, model version, tool arguments, policy decision, approval, output, and business effect so that an investigator can reconstruct what happened.

Security teams should not use SOC 2, ISO 27001, or HIPAA as substitutes for agent-specific engineering. Use them as governance inputs: SOC 2 can document relevant control operation, ISO 27001 can structure risk management, and HIPAA can establish obligations for protected health information. Then add tests for prompt injection, indirect instruction manipulation, credential compromise, data exfiltration, excessive tool use, and downstream impact. Require vendors to show how they enforce policy, store logs, isolate tenants, rotate secrets, respond to incidents, and change models or tools without silently weakening existing controls.

The strongest operating posture is staged autonomy. A research assistant with synthetic data and no tools can be deployed quickly; an agent that reads customer records can operate with strict field-level access; an agent that sends messages, changes systems, or spends money can require a human decision at the action boundary; and a fully autonomous agent should be reserved for tasks whose impact is demonstrably low, bounded, observable, and reversible. Security is not achieved by eliminating agents, nor by assuming that a safety badge makes them trustworthy. It comes from controlling identity, authority, data, execution, evidence, and recovery as one connected system.