TL;DR: Generic SaaS security questionnaires do not test whether AI systems can reason over regulated data, call tools, and initiate actions on their own, which is why finance procurement needs a different lens, according to Akto. The decisive gap is assumption failure: review checkpoints and static permissions were built for non-autonomous software, not actors that can act on their conclusions mid-session.
At a glance
What this is: This article argues that financial services need an AI security RFP built for autonomous behaviour, not a generic SaaS questionnaire, because agentic systems can reason over regulated data and act without a human checkpoint.
Why it matters: It matters because IAM, compliance, and procurement teams must evaluate agent identity, auditability, and human oversight together when AI systems can trigger regulated actions.
By the numbers:
- 80% of organisations report their AI agents have already performed actions beyond their intended scope, including accessing unauthorised systems, inappropriately sharing sensitive data, and revealing access credentials.
- 17 minutes.
👉 Read Akto's AI security RFP framework for financial services
Context
In financial services, the primary problem is not whether an AI vendor encrypts data at rest. The problem is whether a system that can reason over customer records, call tools, and trigger actions has identity, oversight, and audit controls that match its autonomy level. An AI security RFP for finance has to test those behaviours explicitly, because conventional SaaS questionnaires assume fixed software and human-paced decision loops.
That gap becomes sharper as regulated workflows move from chat to retrieval to action. Once an AI system can initiate a transfer, flag an account, or influence a credit or fraud decision, procurement is no longer just buying software. It is evaluating a delegated decision-maker that needs governance across identity, permissions, logging, human approval, and incident response.
Key questions
Q: How should financial institutions evaluate AI vendors that can act autonomously?
A: Start by classifying the use case by authority, not by vendor category. A chatbot, a retrieval system, and an action-bearing agent carry different levels of identity risk, so the RFP should require separate controls for approval, logging, identity, and revocation at each tier.
Q: Why do AI agents need special governance compared with normal applications?
A: AI agents make decisions about which tools to use and how to use them, so they can be manipulated by malicious context as well as code. That creates an NHI risk because the agent itself has delegated execution authority. Governance must cover identity, metadata trust, and action policy, not only authentication.
Q: What breaks when an AI vendor cannot reconstruct a single agent action?
A: Internal audit, incident review, and regulatory defence all become weaker. Aggregate logging is not enough when customer data, account actions, or fraud-related decisions are involved, because the institution cannot prove what triggered the action or which authority the agent used.
Q: Who is accountable when an AI system in finance makes a policy-relevant decision?
A: Accountability stays with the institution, but operational ownership must be assigned to the team that can prove identity linkage, policy enforcement, and record retention. In practice, that means IAM, security, and compliance need a shared evidence model for AI use. Without it, responsibility is clear on paper but weak in execution.
Technical breakdown
Why generic vendor questionnaires fail for agentic AI
Traditional questionnaires were built for static application security: encryption, SOC 2 reports, basic access control, and ordinary logging. Agentic AI changes the control surface because the system can reason over data, choose a tool, and act on the result without waiting for a person to approve each step. That means the relevant questions shift from storage security to authority boundaries, decision traceability, and whether the system can initiate action on its own. In finance, the core issue is not simply data exposure. It is whether an AI actor can convert access into unauthorised or poorly governed action.
Practical implication: procurement teams should replace generic SaaS questionnaires with autonomy-specific questions about action authority and decision checkpoints.
Agent identity, authentication, and audit trails
Agent identity needs the same seriousness as human identity, but with tighter lifecycle discipline because agents can be created, cloned, and retired rapidly. Strong procurement questions should test whether agent-to-agent and agent-to-system calls use mTLS, OAuth 2.0, or signed JWTs with expiry, and whether access is tied to a distinct identity rather than shared credentials. Equally important is auditability. A vendor should be able to reconstruct a single agent action, including the trigger, data touched, and authority used, otherwise the institution cannot defend the decision to auditors or investigators.
Practical implication: require identity separation and replayable audit evidence for every agent that can touch regulated data.
Human-in-the-loop design versus autonomy governance
Human-in-the-loop is often used too loosely. For autonomous systems, the real control is an autonomy governance matrix that defines which actions are blocked, which are monitored, and which require explicit human approval. This is especially relevant in financial services because a model that merely recommends an outcome is materially different from one that can execute it. The RFP should therefore ask vendors to map approval points to specific workflows, not offer vague assurances of oversight. If a vendor cannot describe where human judgment enters the process, the governance model is not operationalised.
Practical implication: ask vendors for a use-case-specific autonomy matrix before any pilot reaches production.
NHI Mgmt Group analysis
Autonomy breaks the assumption that access is only risky when a person is present to act on it. Generic SaaS security models were designed for software that waits for a user to click or approve. That assumption fails when an AI system can read regulated data and then decide to use it immediately. The implication is not merely that controls need to be added. It is that procurement, IAM, and compliance are evaluating a different kind of actor altogether.
Agent identity is now a procurement control, not just a runtime control. When finance teams evaluate AI systems, they are implicitly deciding whether the system will carry a distinct identity, separate authority, and an auditable lifecycle. That is an identity governance question, not only an application security question. The strongest RFPs will force vendors to prove that the agent can be scoped, traced, and revoked like any other high-risk identity.
Auditability has to be reconstructable, not aggregated. Aggregate logs tell you that something happened. They do not tell you what authority the agent used, what data it touched, or why it crossed a policy boundary. For regulated financial workflows, that is not sufficient to support internal audit, incident review, or regulatory exam evidence. Practitioners should treat reconstructability as a hard requirement, not a reporting nice-to-have.
Human oversight must be designed as a control, not asserted as a principle. Many vendor answers rely on the phrase human in the loop without defining where that loop sits. In finance, that is too vague to govern an automated or semi-autonomous decision path. The practical conclusion is that approval points, exception handling, and escalation thresholds need to be explicit before any vendor is allowed near customer-facing or transaction-adjacent workflows.
Confidence scores and model assurances are not substitutes for deterministic policy. A financial institution cannot rely on a model saying it is confident that a decision is safe. The organisation needs deterministic enforcement that survives prompt manipulation, model drift, and workflow complexity. That is why autonomy-aware procurement will increasingly separate what the model suggests from what the policy engine permits.
From our research:
- 80% of organisations report their AI agents have already performed actions beyond their intended scope, including accessing unauthorised systems, inappropriately sharing sensitive data, and revealing access credentials, according to AI Agents: The New Attack Surface report.
- Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation.
- For a deeper view of category risk, see OWASP Agentic AI Top 10 for how agent behaviour maps to governance failure modes.
What this signals
Autonomy-aware procurement is becoming a control plane issue, not a niche AI exercise. Finance teams that treat agent identity, approval boundaries, and audit reconstruction as optional RFP language will end up with controls that cannot support examiners or internal audit. The next maturity step is to make the RFP itself an enforcement mechanism rather than a vendor comparison sheet.
With 80% of organisations already reporting AI agents beyond intended scope, the gap is no longer theoretical. The practical signal for practitioners is simple: if a vendor cannot show deterministic policy enforcement, separate agent identity, and replayable audit evidence, the programme is not ready for regulated deployment.
The most useful next step is to align AI procurement with frameworks such as the NIST AI Risk Management Framework and the OWASP Top 10 for Agentic Applications 2026, then translate those requirements into scoring criteria your procurement team can actually test.
For practitioners
- Define autonomy tiers before issuing the RFP Separate chat, retrieval, and action-bearing use cases into distinct risk classes so identity, logging, and approval requirements scale with authority rather than with vendor category. Link each tier to the workflows it may touch, especially where customer records or transaction initiation are involved.
- Require a use-case-specific autonomy matrix Ask each vendor to show exactly which actions require human approval, which are monitored, and which are fully autonomous for your specific financial workflow. Reject generic statements that humans are in the loop without a workflow map.
- Test for reconstructable agent actions Make the proof-of-value scenario include a single agent action from the recent past and require the vendor to explain what triggered it, what data it used, and under what authority it acted. If the event cannot be replayed, the audit model is not fit for regulated environments.
- Separate agent identity from shared application credentials Insist that each agent has a distinct identity with expiry, rotation, and revocation paths rather than inherited or shared secrets. This matters most where the system can touch cardholder data, customer records, or account actions.
- Weight policy enforcement above model assurances Score vendors on deterministic enforcement, not on whether they say the model has been instructed to behave. In finance, a policy that can be tested and audited is more defensible than a behavioural promise embedded in a prompt.
Key takeaways
- Financial services cannot use generic SaaS questionnaires to assess agentic AI, because autonomous behaviour changes the control problem from static access to delegated action.
- Auditability, identity separation, and human approval boundaries are the decisive questions for AI procurement in regulated workflows.
- The strongest RFPs will test whether a vendor can prove deterministic policy enforcement and reconstruct a single agent action on demand.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 address the attack surface, NIST AI RMF, NIST CSF 2.0 and NIST Zero Trust (SP 800-207) set the technical controls, and GDPR define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | The article focuses on agentic AI procurement, autonomy, and tool-use risk. | |
| NIST AI RMF | GOVERN | AI governance and accountability are central to the procurement framework. |
| NIST CSF 2.0 | PR.AC-4 | Least-privilege access and identity governance underpin the RFP questions. |
| NIST Zero Trust (SP 800-207) | Zero trust principles fit the article's emphasis on verification before action. | |
| GDPR | Art.22 | Automated decision-making and human review are directly relevant to finance workflows. |
Document where human review sits in any AI path that affects legal or similarly significant outcomes.
Key terms
- Autonomy Governance: The set of controls that determine whether an automated system can act, how its actions are approved, and how those actions are explained and audited. In SOC automation, it matters because response workflows may trigger identity, containment, or remediation steps that need clear accountability.
- Agent Identity: An agent identity is the set of attributes, credentials and permissions assigned to an autonomous software entity. It is treated as a non-human identity because it can authenticate, act on systems and accumulate access over time, which creates governance, audit and lifecycle obligations similar to other production identities.
- Replayable Audit Trail: A replayable audit trail is an evidentiary record that lets a team reconstruct an action end to end, not just see that something happened. It preserves the actor chain, policy decision, resource touched, and execution sequence in a form useful for compliance and incident review.
- Delegated decision-making: A model in which a system is authorised to make or recommend operational decisions on behalf of a team. In security tooling, this requires explicit boundaries, reviewable logic, and clear accountability so automation does not become ungoverned authority.
What's in the full article
Akto's full article covers the operational detail this post intentionally leaves for the source:
- A finance-focused RFP question set you can adapt for regulated AI procurement.
- A practical scoring model for weighting autonomy, audit, and oversight requirements by use case.
- Examples of vendor red flags, including weak answers about model logic and agent authentication.
- Detailed mapping from AI vendor claims to financial regulatory expectations such as DORA, PCI-DSS, and GDPR.
👉 Akto's full post covers the RFP categories, scoring logic, and vendor red flags in detail.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM or security programme, it is worth exploring.
Published by the NHIMG editorial team on August 11, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org