TL;DR: Claude Sonnet, ChatGPT and Gemini aligned reasonably well with survey responses on external business risk, but predicted materially lower ratings than mobile security leaders gave themselves on maturity, investment and readiness, according to NowSecure. The gap matters because 65% of organisations that rated their programs advanced or highly effective still reported incidents, while 95% said they deploy AI in mobile apps but 37% do not monitor that AI behavior.
At a glance
What this is: NowSecure compared three AI models’ survey predictions with mobile app security leaders’ responses and found strong alignment on business risk, but wide divergence on internal maturity and AI visibility.
Why it matters: This matters because IAM and app security teams cannot treat confidence, policy, or maturity scores as proof of control effectiveness when runtime behaviour, data flows, and embedded AI remain partially unseen.
By the numbers:
- Every respondent in the survey rated mobile apps as very important or critical to their business, and 58% said a single day of downtime causes severe business damage.
- 37% said they do not monitor AI behavior
👉 Read NowSecure's analysis of AI predictions and mobile app security maturity gaps
Context
Mobile app security maturity is often described with policy, budget, and self-assessment, but those signals can diverge from what apps actually do at runtime. In this case, the primary issue is not whether leaders feel confident, but whether their confidence is backed by direct evidence across application behaviour, data movement, third-party dependencies, and AI-driven features.
The article uses AI model predictions as a benchmark against survey responses, which is useful because it exposes how external risk is easier to anticipate than internal control reality. For IAM and NHI practitioners, the same pattern appears when access reviews, governance attestations, or platform approvals substitute for continuous verification of what identities, secrets, and embedded services are really doing.
That gap is especially relevant wherever mobile apps use federated identity, API tokens, service integrations, or embedded AI features. In those environments, program maturity claims can coexist with unobserved access paths, and that makes the starting position in the survey more typical than exceptional for modern application security programs.
Key questions
Q: How should security teams validate mobile app security maturity claims?
A: They should test the claims against runtime evidence. Self-reported maturity can be useful for programme management, but it does not show what the app is doing in production. Teams need repeatable checks for data flows, dependency behaviour, and AI-enabled features so the assurance model reflects observed control performance, not just documented process.
Q: Why can AI predictions differ from internal security assessments?
A: AI models work best when the answer is anchored in widely available public information. They are much less reliable on organisation-specific control maturity, because those details depend on internal evidence, operating discipline, and incident history. The difference is a reminder that prediction and verification are not the same activity.
Q: What signals show that mobile app governance is failing?
A: A common signal is high confidence paired with continued incidents or unexplained exposure. Another is AI deployment without monitoring of behaviour in production. If teams can describe the programme but cannot evidence runtime control, the governance model is probably measuring intent rather than actual protection.
Q: How should identity and appsec teams share responsibility for mobile app risk?
A: Identity teams should govern the access paths, secrets, and service identities that mobile apps depend on, while appsec teams validate runtime behaviour and data movement. The two functions only work together if governance covers both who can access what and what the application actually does with that access.
Technical breakdown
Why AI models can approximate business risk but miss control maturity
Large language models can predict survey answers that reflect broadly documented business conditions because those signals are abundant in public text. Downtime impact, mobile app importance, and the strategic role of applications are the kinds of facts models can infer from common discourse. Internal maturity, however, is harder to model because it depends on organisation-specific evidence, control design, incident handling, and operational consistency that are rarely visible outside the enterprise. That means AI can mirror the narrative of risk without being able to validate the quality of the control environment behind it.
Practical implication: treat model-based predictions as directional context, not as evidence of app security maturity or readiness.
Why self-reported maturity can diverge from runtime reality
Security maturity scores describe programme structure, not runtime behaviour. A team can have policies, tooling, and governance processes in place while an app still leaks data, invokes unvetted third-party services, or exposes sensitive flows during normal use. This is the same governance gap that appears in identity programmes when approval workflows exist but actual access paths remain unverified. In mobile app security, the critical question is whether controls observe what the app does after release, not just whether the programme has formally approved the app before release.
Practical implication: pair maturity assessments with evidence from runtime testing and data-flow inspection.
AI visibility in mobile apps is a governance problem, not just a telemetry problem
The survey finding that most organisations deploy AI in mobile apps while a substantial minority do not monitor AI behaviour shows a governance boundary problem. If AI features are shipped into apps without visibility into prompts, outputs, or downstream data handling, then the organisation cannot demonstrate control over the application’s real behaviour. For identity teams, this is analogous to deploying workload identities or tokens without lifecycle telemetry. The issue is not only whether a capability exists, but whether its use is bounded, attributed, and continuously monitored.
Practical implication: define AI visibility requirements as part of app governance, alongside access, secrets, and data controls.
Threat narrative
Attacker objective: The practical objective is not a single intrusion but sustained access to sensitive data and application behaviour that governance layers fail to observe.
- Entry occurs when AI-enabled mobile apps, third-party components, or identity-linked services enter the environment without sufficient runtime visibility.
- Escalation happens when security teams rely on self-reported maturity and miss the actual behaviour of the app, including data exposure and embedded AI activity.
- Impact is manifested as persistent blind spots in application security assurance, allowing exposure to continue despite formal programme confidence.
NHI Mgmt Group analysis
Confidence is not control, and mobile app security keeps proving it. The article’s central value is that it exposes the gap between perceived maturity and observable security reality. That gap is familiar across identity and application governance: policies, budgets, and attestations can all be present while runtime exposure remains poorly measured. In practice, this means organisations should stop treating maturity scores as evidence and start treating them as hypotheses that require validation.
Programmes fail when they measure structure instead of behaviour. A mobile app security programme can be highly organised and still miss what happens inside the app after release. That is the core governance problem the survey surfaces, and it maps directly to IAM and NHI oversight where entitlement design says one thing and actual use says another. The right question is not whether a programme exists, but whether it can prove the application behaves within the intended boundary.
AI visibility inside mobile apps is becoming a control requirement, not an optional enhancement. When 95% of respondents say AI is deployed but a meaningful share do not monitor that AI, the issue is no longer adoption. It is evidence of an emerging accountability gap that belongs in security governance, appsec review, and risk reporting. For practitioners, this means treating AI behaviour as part of the application’s identity and trust surface.
Verification trust gap: organizations increasingly confuse documented maturity with verified operational control. That confusion is especially risky in mobile environments where runtime behaviour, secrets, and third-party calls can change faster than governance cycles. Security leaders should interpret confidence scores as signals to test, not as proof to trust.
The mobile app security conversation is moving toward continuous evidence, not static programme claims. That shift is aligned with modern identity governance, where access, secrets, and workload behaviour require continuous observation rather than annual review. Practitioners who cannot evidence what their apps and embedded AI do at runtime will struggle to defend their risk posture. The implication is clear: assurance must become observable.
What this signals
Verification trust gap: mobile app programmes that rely on maturity scores without runtime proof will keep overestimating control effectiveness. The next governance step is to bind policy, testing, and telemetry into one assurance loop rather than treating them as separate workstreams.
For identity and application teams, the practical signal is that secrets, service identities, and embedded AI now need continuous visibility. That aligns with the control logic in the NIST Cybersecurity Framework 2.0 and with evidence-based assurance in NIST SP 800-53 Rev 5 Security and Privacy Controls.
The broader implication is that mobile app security is moving from declarative governance to observable governance. Organisations that cannot prove what their applications and AI features do at runtime will struggle to defend risk posture, especially where identity-linked services are part of the delivery chain.
For practitioners
- Tie maturity claims to runtime evidence Require every mobile app security maturity rating to be backed by runtime testing, data-flow analysis, and observed third-party interactions rather than self-assessment alone.
- Add AI behavior monitoring to app governance Define what AI features must be logged, reviewed, and escalated in mobile apps, including prompts, outputs, and downstream data handling.
- Validate identity-linked app dependencies Inventory federated identity flows, service tokens, API keys, and external dependencies so teams can see which identities and services each app actually relies on.
- Rework assurance reporting around observable controls Replace purely declarative programme metrics with measures that show whether controls are operating in production, including testing coverage and unresolved exposure findings.
Key takeaways
- Mobile app security maturity can look strong on paper while runtime behaviour, data exposure, and AI visibility remain under-verified.
- The size of the gap matters because reported confidence, active incidents, and unmonitored AI can all coexist in the same programme.
- Security teams should move assurance from self-report to evidence by combining runtime testing, dependency visibility, and identity-aware monitoring.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, NIST SP 800-53 Rev 5, NIST AI RMF and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-1 | Runtime monitoring and evidence-based assurance are central to this article's control gap. Use DE.CM-1 to validate mobile app behaviour with continuous runtime monitoring, not self-reported maturity. |
| NIST SP 800-53 Rev 5 | SI-4 | Security monitoring is needed where app behaviour and AI activity are not being observed. Apply SI-4 to detect unexpected application and AI behaviour across production mobile environments. |
| NIST AI RMF | MEASURE | The article hinges on measuring AI behaviour rather than assuming it is controlled. Apply MEASURE to evaluate whether embedded AI in mobile apps is observable, testable, and bounded. |
| CIS Controls v8 | CIS-8 , Audit Log Management | Monitoring gaps make audit evidence and log coverage directly relevant. Map mobile app telemetry to CIS-8 so AI and runtime activity are captured for review and investigation. |
Use DE.CM-1 to validate mobile app behaviour with continuous runtime monitoring, not self-reported maturity.
Key terms
- Activation Trust Gap: The activation trust gap is the difference between trusting data because it is protected and governing it because it is being reused. It appears when organisations move data from backup or archival systems into AI pipelines without reapplying access, sensitivity, and consumer controls.
- Runtime assurance: Runtime assurance is the practice of validating how an application actually behaves after deployment. It matters because configuration, identity flow, and integration state can change security outcomes in ways that source code analysis alone cannot prove.
- AI visibility: AI visibility is the ability to identify which AI tools are active, who is using them, and what data they can reach. In security practice, it is the prerequisite for policy enforcement because you cannot control usage, data flow, or risk when the environment is opaque.
- Maturity Score: A programme-level rating that summarises how an organisation believes its security processes are designed and managed. It can be useful for reporting, but it is not the same as evidence that controls are operating effectively in live environments or that risk exposure is actually reduced.
What's in the full report
NowSecure's full analysis covers the operational detail this post intentionally leaves for the source:
- The full survey breakdown behind the AI Prediction Score and the largest response gaps across maturity, investment, and readiness.
- Industry and vertical comparisons showing how mobile app risk perceptions differ across finance, healthcare, high tech, and retail.
- Additional findings on AI adoption, monitoring gaps, testing practices, and third-party risk in mobile applications.
- The survey's strategic recommendations for closing the gap between policy, AI visibility, and evidence-based testing.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, workload identity, and secrets management. It helps security practitioners turn identity policy into operational control across modern application environments.
Published by the NHIMG editorial team on September 4, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org