Prompt-level instructions are guidance, not enforcement. Harness-level controls are the real security boundary because they govern what tools the agent can call, what actions are permitted, and what happens when it exceeds policy. In practice, security teams should treat prompts as advisory context and rely on runtime constraints, budgets, and identity controls to contain agent behavior.
Why prompt text cannot be the security boundary
Prompt-level instructions shape the agent’s behavior, but they do not reliably constrain it. An agent can be redirected by prompt injection, conflicting tool outputs, hidden context, or simply a more permissive downstream policy. The practical difference is that prompts express intent, while the harness enforces authority. For AI agents, that distinction determines whether the system is merely guided or actually controlled.
The harness is where the enforcement layer lives: tool routing, action approval, resource limits, logging, and policy evaluation. That is why a prompt can tell an agent not to exfiltrate data, but only the harness can stop it from calling an exfiltration-capable tool or using a forbidden connector.
That separation is also why prompt hygiene alone is not a defensible control. A good prompt can reduce mistakes, but it cannot define the trust boundary for actions that have external effects, especially when the agent can browse, call APIs, write files, or trigger workflows.
What changes at the harness layer
Harness-level controls operate at runtime and can enforce decisions per action, per tool, or per request. That includes authorization checks, approval gates, budget thresholds, context filtering, and step-up controls when the agent tries to cross a policy boundary. In practice, the harness is where you decide what the agent is allowed to do, not just what it is told to do. AI Agent Authorisation Guide is a useful reference for least-privilege access, delegated authority, and approval gates.
This matters because many agent failures are not model failures in the abstract, they are authorization failures in the concrete. If an agent has broad tool scope, long-lived credentials, or no per-action policy check, a harmless-seeming prompt can still lead to destructive behavior. A runtime harness can bound the blast radius even when the model is uncertain, manipulated, or overconfident.
For that reason, the harness should be treated like a policy enforcement point, not an implementation detail. Zero Trust for AI Agents frames the right operating model: verify the request, remove standing privilege, and enforce policy continuously instead of trusting the prompt.
How practitioners should separate guidance from control
The cleanest mental model is to treat prompt text as advisory context and harness logic as mandatory control. Prompt text can describe goals, constraints, and style, but the harness must decide whether the action is permitted, what identity is acting, what tool is exposed, and whether the request exceeds budget or policy. That is especially important when agents can operate with delegated access or act on behalf of a user.
Harness design also needs to account for failure modes that prompts cannot solve on their own. If an agent can chain tools, retain memory, or inherit external context, the security decision must happen where the action is executed, not where the instruction was written. Agentic AI Security Guide is relevant because it ties tool use, orchestration, memory, and identity back to the actual attack surface.
A useful rule is simple: if the behavior would be unsafe when the prompt is ignored, the control does not belong in the prompt. Put the control in the harness, and use the prompt only to improve task quality and operator intent. Threat Modelling AI Agents helps teams map those trust boundaries before they turn into production incidents.
Risk and Threat Considerations
Prompt-only governance fails when an attacker can alter the model’s context, induce tool misuse, or exploit a permissive connector. The real risk is not that the agent disobeys a sentence, but that it faithfully follows a malicious or misleading instruction into an unsafe action path. Once the agent has standing access to tools or data, the prompt becomes too weak to contain abuse.
Failure mechanism: A malicious prompt, injected instruction, or compromised upstream input changes the agent’s behavior while the harness lacks a runtime check on the resulting tool call or action.
Impact: The agent can overreach its intended scope, expose data, trigger unauthorized operations, or amplify compromise across connected systems.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent prompts can be bypassed when runtime authority is too broad. |
| ASI02 — Tool Misuse | Harness controls govern which tools an agent may invoke and how. | |
| Recommendation — Enforce per-action authorization and minimize agent privileges. Gate tool execution with policy checks and approval steps. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Agent harnesses must bound what actions and tools are reachable. |
| IA-9 — Service Identification and Authentication | Runtime control depends on authenticating non-human actors and their calls. | |
| Recommendation — Limit each agent to the minimum permissions needed for its task. Authenticate agent tool and service calls before allowing execution. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The question is fundamentally about trusted prompt intent versus enforced runtime control. |
| Recommendation — Apply continuous verification to each agent action instead of trusting prompts. | ||
Practitioner Guidance
What to verify: Confirm that every tool call, external side effect, and privileged action is authorized at runtime, not just described in the prompt. If you cannot point to a policy decision point, the agent is operating on trust, not control.
Decision rule: If the action can change state, spend money, move data, or reach outside the sandbox, enforce it in the harness and require a denial path that works even when the prompt is bypassed. If the prompt is the only thing preventing harm, treat the design as incomplete.
What good looks like: The agent can be helpful with natural-language guidance, but its actual authority is bounded by policy, identity, and observable runtime controls. Prompt text improves task performance; the harness determines whether the task may happen at all.
Practitioner takeaway: Prompts shape intent, but only the harness can define enforceable trust boundaries for agent behavior, so security teams should design for runtime denial, not instructional compliance.
Related resources from NHI Mgmt Group
- What is the difference between managed identities and hardcoded secrets for AI agents?
- What is the difference between workload identity and API keys for AI agents?
- What is the difference between logging actions and logging intent for AI agents?
- What is the difference between prompt-level controls and runtime governance for agents?