By NHI Mgmt Group Editorial TeamBased on Twine Security: “The Next Step in Agentic AI: Accountability” (January 26, 2026)

TL;DR: AI agents cannot earn trust through output quality alone, because accountability depends on groundedness, memory, discretion, interface visibility, and persistence, according to Twine Security. The governance gap is that many IAM and access controls still assume actions are human-paced, reviewable, and easy to explain after the fact.


At a glance

What this is: This article argues that accountable AI agents require groundedness, memory, discretion, interface visibility and persistence, not just fluent model output.

Why it matters: IAM and governance teams need to account for AI agents as operational actors whose trustworthiness depends on verifiable behaviour over time, not only on session-level output quality.


Context

AI agents create a governance problem when organisations treat them like faster chat interfaces instead of runtime actors that can initiate actions, retain context and continue working after the original request. The security issue is not whether the model sounds correct, but whether the system around it can verify, constrain and explain what the agent did.

For identity programmes, the shift matters because trust in an agent is no longer just an authentication or authorisation question. It becomes a lifecycle and accountability problem across memory, task scope, observation and follow-up, which means existing IAM patterns need to be evaluated against behaviour that persists beyond a single prompt.


Key questions

Q: How should teams govern AI agents that can act after the original prompt is gone?

A: Governance has to follow the agent beyond the initial request. Teams should require observable task state, persistent memory controls and a clear revocation path so the agent can be checked after action, not only before it starts. If the actor can resume later, the control model must assume continued responsibility, not one-time completion.

Q: Why is groundedness important for accountable AI agents?

A: Because trust in an agent depends on whether its actions can be verified against facts, tests or source systems. Groundedness reduces the gap between what the agent says and what the organisation can prove. Without it, even accurate-sounding output is only a claim, not evidence, and accountability breaks when decisions need to be defended.

Q: What breaks when AI agents do not have persistent memory?

A: When AI agents do not have persistent memory, they cannot reliably retain corrections, risk cues, or task-specific constraints across sessions. That breaks follow-through and makes earlier guidance disappear unless it is reintroduced every time. The result is inconsistent behaviour that looks stateful to users but is actually reassembled from fragments.

Q: What is the difference between an AI agent and an agentic workflow?

A: An AI agent is the autonomous software entity that decides and acts. An agentic workflow is the broader process that coordinates one or more agents through planning, execution, reflection, and reporting to reach a goal. The agent is the actor, while the workflow is the system that governs how work gets done.


Technical breakdown

Groundedness in AI agents is a system property, not a model claim

Groundedness means the agent’s outputs stay tied to verifiable facts, source data or testable system state. In practice, that is rarely a property of the model alone. The control surface sits around the model: retrieval constraints, citations, test harnesses, policy checks and domain-specific validation. For an agent handling code, tests can verify whether output matches the desired state. For an agent producing text, citations and source checking reduce drift. Accountability starts where output can be checked against reality, not where the model sounds convincing.

Practical implication: define verification points outside the model and make grounded evidence mandatory before agent output is treated as authoritative.

Memory and discretion determine whether agent behaviour is governable

Stateless agents do not naturally accumulate experience the way humans do, so memory has to be supplied through prompts, retrieval and retained context. That makes memory a governance control, not just a convenience feature. Discretion is the related capability to interpret task importance, push back on bad instructions and adjust effort to risk. Without both, agents will either forget important constraints or execute blindly. The result is an identity subject that can act, but cannot reliably explain why it acted or how earlier feedback changes later behaviour.

Practical implication: treat memory persistence and task discretion as part of the agent’s control design, then test whether feedback actually changes later decisions.

Visibility and persistence are prerequisites for AI agent accountability

Interface visibility is the ability for humans to see what the agent is doing in a form that maps to the task, not just to raw model traces. Persistence is the ability to continue checking outcomes after the immediate action is complete. Together, they close the accountability loop. A one-shot workflow assistant can finish a prompt, but an accountable agent must remain inspectable and able to revisit its own prior actions. That is a very different governance expectation from standard IAM, which often assumes access is granted, used and then reviewed later.

Practical implication: require observable task state and post-action follow-up for agents that influence production systems or business decisions.


Threat narrative

Attacker objective: The objective is to get the organisation to accept agent output or action as trustworthy even when the system cannot prove what happened or why.

  1. Entry occurs when an AI agent is assigned a task and granted the context, tools and permissions needed to act on behalf of the organisation.
  2. Escalation occurs when the agent relies on weak groundedness, shallow memory or poor discretion and then takes an action that exceeds the intended scope of the task.
  3. Impact occurs when the system cannot explain, verify or revisit the agent’s behaviour after the fact, leaving the organisation unable to trust the outcome or correct recurrence.

Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Accountability, not output quality, is the real trust boundary for AI agents. A fluent answer can still be operationally untrustworthy if the actor cannot be observed, constrained or asked to justify the path it took. In identity terms, that means the programme has to govern the execution record, not only the model response. Practitioners should treat accountability as a control objective in its own right.

Groundedness is a system-level control problem, not an LLM personality trait. The article is right to separate model fluency from reliable action, because many current controls still test plausibility instead of verifiability. That breaks the assumption that a capable model is a safe operator. The implication is that task success has to be proved externally before the agent is allowed to stand in for human judgment.

Memory changes the identity problem because the actor now carries state across time. Once an agent can retain prior instructions, exceptions and corrections, it stops behaving like a disposable request processor and starts behaving like a governed identity subject. That makes lifecycle, context retention and reviewability part of the security model. Practitioners should stop thinking of memory as a feature and start treating it as durable authority.

Persistence is the broken assumption behind many current governance workflows. Access review processes assume access persists long enough to be reviewed; stateless or intermittently running agents do not fit that cadence. When an agent can act, stop and resume with retained context, the review window no longer aligns with the control window. The implication is that governance has to move closer to issuance, observation and revocation events, not later certification cycles.

AI agent governance now spans human IAM, NHI control and autonomous behaviour. The same operational issue looks different depending on whether the organisation is judging a person, a service account or an agent that can make runtime decisions. That is why identity teams need one governance model that can describe action, memory and persistence across actor types. Practitioners should build controls that survive the actor changing form.

From our research library:

What this signals

Accountable agent design shifts the control point from output review to behaviour proof. Once a system can act, remember and continue, the organisation needs evidence that its decisions were grounded, visible and reversible. That pushes identity programmes toward runtime observation rather than post-hoc approval.

Persistent memory is the named concept that changes the governance model. When an agent retains task context across sessions, it starts to accumulate authority in the same way a long-lived identity does. Practitioners should assume that every memory store is part of the control plane, not a convenience layer.


For practitioners

  • Define the agent’s verification gates Require external checks for groundedness before an agent’s output can trigger downstream action, especially in code, text and policy workflows.
  • Inventory where memory creates durable authority Identify prompts, RAG stores and tool-backed memory that let an agent retain constraints, exceptions or instructions across sessions.
  • Add visible task-state controls Expose what the agent is doing in a human-readable interface so operators can see task progress, scope and stopping points.
  • Rework review timing for persistent agents Shift oversight to issuance, runtime observation and revocation events when an agent can continue checking its own work after completion.
  • Test discretion under real task pressure Measure whether the agent asks clarifying questions, pushes back on risky requests and adjusts effort to the importance of the task.

Key takeaways

  • AI agent accountability depends on more than model quality, because trust breaks when behaviour cannot be verified, explained or revisited.
  • Groundedness, memory, discretion, interface visibility and persistence are the controls that make agent action governable rather than merely impressive.
  • Identity teams need to move oversight closer to runtime evidence and memory governance instead of relying on later review cycles.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseAgent accountability hinges on how runtime authority is granted and used.
ASI06 — Memory & Context PoisoningThe article centers on memory as a governance surface for agents.
ASI09 — Human-Agent Trust ExploitationThe article is about trust signals and how humans decide whether to rely on agents.
Recommendation — Map agent permissions to ASI03 and constrain runtime authority to the minimum task scope. Review agent memory paths for stale, incomplete or untrusted context that can alter later actions. Design controls that prevent misleading confidence cues from substituting for verifiable behaviour.
NIST AI RMFGOVERN — AI Governance and AccountabilityAccountability, oversight and responsibility are the core subject of the post.
MANAGE — AI Risk ManagementThe article focuses on managing the risks of agent action, memory and persistence.
Recommendation — Establish governance roles and accountability checks for agents before granting production tasks. Treat agent persistence and memory as managed risks with explicit review and escalation paths.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe post implies agent authority must be bounded and reviewable.
Recommendation — Apply PR.AA-05 to keep agent entitlements tied to task scope and revocation rules.

Key terms

  • Groundedness: Groundedness is the degree to which an AI response can be supported by verifiable source material. In practice, it measures whether the model answered from evidence rather than inference, memory, or fabrication, which is critical for RAG systems and any workflow that drives decisions from model output.
  • Agent Memory: Agent memory is the stored context an AI agent uses across sessions or tasks. In governance terms, it is controlled state, because the memories an agent retains can influence future actions, permissions use, and the safety of subsequent decisions.
  • Persistence: Persistence is the ability to retain memory, state, or goals across sessions and time. In NHI governance, persistence matters because retained context can influence later access decisions, create hidden privilege, and extend the impact of a prior task beyond its intended window.
  • Discretion: Discretion is an agent’s ability to adjust its actions to task importance, risk and context rather than following instructions blindly. In governance terms, discretion is what lets an agent push back, ask clarifying questions or reduce over-commitment when a task is ambiguous or high risk.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 6, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org