Treat the harness as the enforceable security boundary. Put permission checks, tool allowlists, sandboxing, and approval gates outside the model so a compromised model cannot override them. Limit each session to the minimum data, tools, and actions needed for the task, and log the full action chain so you can reconstruct what happened after the fact.
Why the harness, not the model, must be the boundary
The core design choice is to assume the model can be manipulated, but the environment must still refuse unsafe actions. That means the harness owns permissioning, tool access, and execution policy. If those checks live inside the model prompt alone, a successful prompt attack turns into a direct path to action. AI Agent Authorisation Guide is useful here because it frames least privilege and per-action decisions as external controls, not model promises.
Good harness design also separates intent from execution. The model can propose, draft, or rank steps, but the harness decides whether a tool call is permitted, whether the request fits the current session, and whether human approval is required. That separation matters most when the agent can reach real systems, because the failure mode is not just wrong text, it is wrong action.
For teams securing agentic systems, the practical question is not whether the model is trustworthy. It is whether the surrounding control plane can constrain a compromised model well enough that the worst possible output remains bounded. Zero Trust for AI Agents reinforces that the request, principal, and action all need verification before anything executes.
What to put outside the model
The strongest controls are the ones the model cannot edit. Permission checks should sit in the harness or policy layer, not in system prompts. Tool allowlists should define exactly which functions exist for that task, and sandboxing should limit what those tools can reach if they are misused. Agentic AI Security Guide is relevant because it treats tools and orchestration as part of the attack surface, not as a passive implementation detail.
Approval gates should also be explicit and conditional. High-impact actions, irreversible actions, cross-environment actions, and actions involving secrets or production data should require a human or policy checkpoint outside the model. This is especially important when the agent has conversational freedom but should not have operational freedom.
Limit the session to the smallest workable scope. That means short-lived access, narrow data exposure, task-specific tools, and no standing privilege beyond the current job. AI Agent Authorisation Guide and Zero Trust for AI Agents both support that pattern by treating every action as individually authorised rather than assuming broad session trust.
How to prove the control worked after the fact
When the model is tricked, recovery depends on observability. Teams need a full action chain that shows the original request, the intermediate reasoning or plan if available, the tool calls, the policy decisions, and the resulting side effects. Without that chain, you cannot separate harmless hallucination from an actual executed change. AI Agent Observability, Audit and Incident Response Guide directly addresses logging, attribution, and kill-switch design.
Log the decision points that matter most: what was requested, what was allowed, what was denied, and why. Also record the identity or principal under which the action ran, because agents often act through delegated or borrowed authority. That makes reconstruction possible when a tool call succeeds even though the model was operating on bad instructions.
Teams should also test the logging path as part of control validation, not as an afterthought. If you cannot reconstruct an agent session end to end, then the control is incomplete even if the permissions look good on paper.
Risk and Threat Considerations
An agent harness becomes a security boundary only if it can resist model deception, prompt injection, and overbroad delegation. The main risk is not that the model answers incorrectly, it is that the environment executes a wrong, destructive, or unauthorized action because the control layer trusted model output too much.
Failure mechanism: The agent produces a plausible but unsafe action plan, and the harness forwards it into a tool or workflow that has broader privileges than the task requires. If approval and policy checks are embedded in the model path instead of enforced externally, a compromised model can bypass the intended guardrails.
Impact: The result can be data exposure, destructive changes, secret misuse, lateral movement through connected tools, or hard-to-reconstruct operational damage. The larger the blast radius of the allowed tools, the more a single prompt compromise turns into a real incident.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent harness controls must stop unsafe delegated actions. |
| ASI02 — Tool Misuse | The question centers on preventing unsafe tool execution by a tricked model. | |
| ASI09 — Human-Agent Trust Exploitation | Approval gates address cases where the model tries to exploit operator trust. | |
| Recommendation — Enforce per-action policy checks outside the model and deny privilege escalation. Allowlist tools and sandbox execution so model output cannot invoke unsafe actions. Require human approval for high-impact actions and verify the requested change before execution. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | The harness should restrict agent sessions to the minimum required access. |
| AU-2 — Event Logging | The answer depends on recording the full action chain for reconstruction. | |
| AU-12 — Audit Record Generation | Agent execution needs trustworthy records of what the harness allowed and blocked. | |
| Recommendation — Constrain each agent session to the minimum permissions needed for the task. Log agent requests, policy decisions, tool calls, and outcomes end to end. Generate auditable records for every privileged agent action and denial. | ||
| NIST CSF 2.0 | PR.AA-01 — Identities and Credentials Are Issued, Managed, Verified, Revoked, and Audited | The harness must govern agent credentials and delegated access lifecycle. |
| PR.AA-05 — Least Privilege | The answer explicitly calls for task-scoped access and minimal permissions. | |
| Recommendation — Manage agent credentials through issuance, revocation, and audit controls. Apply least privilege so each agent session can only do what the task requires. | ||
Practitioner Guidance
What to prioritise: Put the permission decision outside the model first, then reduce the tool set and data scope to the minimum needed for the session. If a task does not need write access, production reach, or secret access, do not grant it.
What to verify: Confirm that every high-impact tool call has an explicit policy decision, that denials are enforced in the harness, and that the logs show both the request and the outcome. If you cannot explain why a call was allowed, the control is too implicit.
Practitioner takeaway: Treat the model as an untrusted planner and the harness as the enforceable control point, because only the harness can reliably limit blast radius when the model itself has been manipulated.
Related resources from NHI Mgmt Group
- How should security teams handle AI agent visibility?
- How should security teams monitor AI agent activity without disrupting developers?
- How should security teams implement AI agent controls on GKE without creating blind spots?
- How should security teams implement authorization controls for AI agent tool calls in production environments?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org