Join our Newsletter — 33% off our NHI Course
Home› FAQ› Agentic AI & Autonomous Identity› What should security teams do when models can…
Agentic AI & Autonomous Identity

What should security teams do when models can emit shell commands or tool calls?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated October 11, 2026 Domain: Agentic AI & Autonomous Identity

They should treat decoded output as an execution boundary and apply policy checks to the final arguments that reach the executor. That reduces the chance that a safe-looking token mapping turns a routine model response into a destructive or exfiltrating action.

Why this becomes an execution-boundary problem

Once a model can emit shell commands or tool calls, the security question shifts from “is the text safe?” to “what reaches the executor?” The risky handoff is the last parse, mapping, or templating step before execution. That is where apparently benign output can become a privileged filesystem change, network call, or data export if policy is checked too early or too loosely.

The practical distinction is important: the model may generate useful intent, but the executor acts on concrete arguments, flags, paths, and targets. Security teams should therefore validate the final structured command, not just the intermediate tokens or natural-language explanation. That is the point where quoting errors, argument injection, and hidden option switches become operationally meaningful.

What controls need to sit at the boundary

Boundary controls should be designed to understand the actual command form that will be executed, including resolved arguments and destination context. A policy engine can compare the final request against allowlists, environment rules, user intent, and privilege constraints before anything is handed to a shell, runtime, or API client. This is especially important when the same model can invoke multiple tools with different blast radii.

Strong implementations keep the executor narrow. Prefer explicit schemas, typed parameters, and dedicated tool interfaces over free-form command construction, because every extra translation layer is another chance for unsafe intent to be disguised as valid syntax. Where shell execution is unavoidable, reduce ambient privilege and make the allowed command surface as small and deterministic as possible.

Teams should also treat tool calls and shell commands as audit-worthy actions, not just application logs. The control objective is to preserve traceability from model output to final side effect, so that blocked, modified, and approved actions can be reviewed after an incident or policy dispute.

How to reduce the chance of destructive or exfiltrating actions

The safest pattern is to separate generation from execution. Let the model propose, then have a policy layer approve, rewrite, or reject based on final arguments, target resources, and the identity or environment the action would affect. That is more reliable than attempting to infer safety from the wording of the model response itself.

Escalate when a proposed action crosses trust boundaries, reaches sensitive paths, or would expose secrets, production data, or privileged systems. A command that looks routine in isolation can still be high-risk if it touches backup stores, package repositories, credentials, or remote admin channels. Good practice is to default to denial whenever the executor cannot explain exactly what will run and why.

  • Use allowlisted tools and explicit argument schemas.
  • Validate the final command after all substitutions and prompt-derived fields are resolved.
  • Run with least privilege and segregate high-impact tools from low-impact ones.
  • Log the model proposal, policy decision, and executed arguments together.

Risk and Threat Considerations

When model output can reach a shell or tool runtime, the main risk is command injection by composition, not by text alone. A malicious prompt, poisoned context, or malformed tool request can steer a benign-looking response into file deletion, data theft, or credential exposure if the final executor trusts the model’s surface form.

Failure mechanism: Unsafe characters, option injection, argument smuggling, or overbroad tool permissions survive the translation from model output to executable arguments, then the executor performs the attacker-influenced action with legitimate privileges.

Impact: The result can be destructive system changes, unauthorized outbound requests, data exfiltration, or lateral movement through any downstream system the tool can reach.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI02 — Tool MisuseModels issuing tool calls can misuse tools through unsafe arguments or overbroad actions.
Recommendation — Constrain tool execution with explicit schemas and reject requests that exceed approved tool intent.
NIST SP 800-53 Rev 5AC-6 — Least PrivilegeExecutors need minimized privilege so a bad model output cannot cause excessive damage.
SI-10 — Information Input ValidationFinal command arguments require validation before execution to prevent injection and malformed input abuse.
Recommendation — Limit executor privileges to the minimum needed for each tool action. Validate decoded command inputs and arguments before they reach any executor.
OWASP ASVSV8 — AuthorizationTool invocation and command execution need authorization checks on the final action surface.
Recommendation — Authorize the final action, not just the model output, before execution.
MITRE ATT&CKT1059 — Command and Scripting InterpreterShell-command emission maps directly to command execution abuse and interpreter-based attack paths.
Recommendation — Monitor and restrict command-interpreter use paths exposed to model-driven inputs.

Practitioner Guidance

What to verify: Verify that policy is evaluated on the exact command string or structured arguments after decoding, interpolation, and normalization, not on the model’s natural-language explanation.

Decision rule: If the action cannot be reduced to a bounded, inspectable, least-privilege request, do not execute it automatically, route it to a safer tool path or require human approval.

Common mistake: Teams often harden prompts while leaving the executor unconstrained; that protects the conversation, not the side effect.

Practitioner takeaway: Treat the model as a proposal generator and the executor as the security boundary, because only the final arguments determine whether the action is safe.

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org