TL;DR: Open source models are closing the capability gap with closed systems and pushing more organisations toward in-house inference as monthly spend rises into five figures, according to WorkOS's interview with Baseten. The governance question is no longer whether AI workloads will scale, but which identity, access, and infrastructure controls will govern them when they do.
At a glance
What this is: WorkOS's interview with Baseten argues that open source model quality is now good enough to shift more AI inference from external APIs to in-house infrastructure.
Why it matters: That shift changes how IAM, NHI, and platform teams govern model access, runtime privileges, and cost-sensitive AI infrastructure as workloads scale.
Context
Open source models now sit close enough to closed systems that the decision is no longer only about model quality. The governance problem is what happens when AI inference becomes an internal workload with its own access paths, operating costs, and deployment choices.
The article frames an industry transition from token-by-token experimentation to controlled in-house inference as spend grows. For identity and security teams, that means AI infrastructure starts to look less like a vendor consumption problem and more like a privileged workload management problem.
Key questions
Q: How should security teams govern in-house AI inference workloads?
A: Security teams should govern in-house AI inference workloads as non-human identities with scoped permissions, named ownership, and lifecycle controls. That means inventorying service accounts, separating deployment rights from infrastructure administration, and reviewing credentials whenever models move between environments or vendors. If the workload can call other systems, it needs a defined trust boundary.
Q: Why does open source model adoption change identity governance for AI platforms?
A: Because the control boundary moves from an external API provider to the enterprise's own runtime. Once the organisation hosts inference itself, it must govern model deployment rights, workload identities, and operational access across more systems. That expands the number of places where privilege can be over-scoped or left standing.
Q: What breaks when AI workloads scale without lifecycle controls?
A: When AI workloads scale without lifecycle controls, old credentials and broad privileges tend to remain in place after the system changes. That creates orphaned access, unclear ownership, and excessive runtime authority across deployment, observability, and integration layers. The result is a machine identity estate that grows faster than the controls that govern it.
Q: What is the difference between external AI APIs and in-house inference governance?
A: External APIs shift much of the runtime control to the provider, while in-house inference brings deployment, execution, and access decisions into the enterprise. That means the organisation must manage who can operate the model, who can change it, and which non-human identities have privileged access. The governance burden moves inward with the workload.
Technical breakdown
Why open source model parity changes deployment patterns
When model quality differences shrink to months rather than years, architecture decisions move from capability selection to control selection. Organisations can trade external API dependence for local deployment, which changes where authentication, authorisation, logging, and platform trust boundaries sit. In practice, the model is no longer the only strategic object; the inference layer becomes part of the enterprise control plane. That matters because the privilege model for accessing GPU capacity, model endpoints, and orchestration services now sits inside the organisation's own governance perimeter.
Practical implication: Treat in-house inference as a governed workload with explicit access boundaries, not as a simple hosting decision.
Inference infrastructure as a privileged AI workload
Inference infrastructure is the runtime layer that loads a model, allocates compute, serves requests, and returns outputs. In Baseten's framing, the hard part is not just running the model quickly, but securing the GPU, the deployment path, and the surrounding platform at scale. That makes this an identity problem as much as an infrastructure problem, because service accounts, deployment credentials, and orchestration permissions determine who can change the model, access the runtime, or move workloads between environments.
Practical implication: Map every inference action to a workload identity and remove any standing privilege from deployment and runtime paths.
Why scale changes the AI governance model
The article's spending thresholds matter because they mark the point at which AI usage stops being opportunistic and starts requiring durable governance. At low cost, teams can tolerate loose controls and external dependency. At $10,000, $20,000, or $50,000 a month, the business logic changes: teams seek cost control, customisation, and independence, which also increases the number of systems that must be trusted to manage model access, update paths, and compute allocation. The security challenge is that scale amplifies both operational complexity and the blast radius of misconfiguration.
Practical implication: Reassess access reviews, change control, and secrets governance once inference becomes a recurring production expense.
Breaches seen in the wild
- LiteLLM PyPI package breach: LiteLLM PyPI supply chain attack, credentials stolen from users.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Inference governance is becoming a workload identity problem, not just a model-selection problem. Once organisations bring AI workloads in-house, the security question shifts to which identities can deploy, tune, and execute those models. The governance boundary moves from the SaaS provider to internal GPU, orchestration, and deployment layers. Practitioners should treat inference platforms as privileged production systems with explicit ownership and control.
Open source model parity shortens the decision cycle faster than most governance programmes can absorb. The article's own timeline shows the capability gap collapsing quickly enough that architecture decisions now change inside normal procurement and review cadences. That compresses the window in which central teams can standardise controls before teams decentralise model choices. The implication is a governance drift risk: adoption moves faster than policy.
In-house inference creates a new form of identity blast radius. When model access, deployment automation, and GPU scheduling converge, a single over-scoped service account can affect far more than one application endpoint. That is a classic non-human identity pattern with AI-specific consequences, and it sits squarely in OWASP-NHI territory. Practitioners should think in terms of control planes, not just applications.
Open source model adoption is exposing the limits of external-provider dependence. The article shows that cost, customisation, and performance now pull AI workloads inward once usage crosses a certain scale. That shifts accountability for access, observability, and resilience onto the enterprise, where identity governance must be able to track both human operators and non-human execution paths. Teams should expect in-house AI infrastructure to demand the same governance discipline as other privileged production estates.
Named concept: inference governance gap: This is the space between model adoption and the controls needed to manage the runtime, identities, and privileges around it. The article shows that the gap appears as soon as teams decide to optimise cost and control by moving inference inside the enterprise. Practitioners should close it by treating AI runtime access as a first-class governance domain, not an engineering afterthought.
From our research library:
- Software supply chain attacks were projected to cost organisations $60 billion in 2025.
What this signals
Inference governance gap: The moment AI workloads move in-house, the enterprise inherits the access paths, privilege model, and operational accountability that a provider previously absorbed. That means workload identity, deployment credentials, and change control become part of the AI programme's core security architecture, not an adjacent operations concern.
Open source model adoption also compresses the time available to standardise controls. When model choice can change faster than normal governance cycles, organisations need a repeatable way to decide which AI workloads may run internally, who can operate them, and which identities may alter production behaviour.
For practitioners
- Define ownership for inference runtime access Assign a named owner for model deployment, GPU access, and orchestration permissions so AI runtime changes are governed like any other privileged production path.
- Inventory the non-human identities behind AI workloads List the service accounts, tokens, and deployment credentials that can start, stop, or modify inference workloads, then remove any unnecessary standing access.
- Separate development and production model access Keep experimentation, tuning, and production inference on distinct access paths so a lower-trust testing identity cannot modify live workloads.
- Revisit governance once spend crosses a scale threshold Use recurring inference spend as a trigger to reassess access reviews, secrets handling, and change control before internal control drift becomes normal.
Key takeaways
- Open source model quality is now close enough to closed systems that infrastructure and governance decisions are driving more AI workloads back in-house.
- The main security consequence is that AI inference becomes a privileged production workload with service accounts, deployment permissions, and GPU access that must be controlled.
- Identity teams should treat the shift as a trigger to reset access boundaries, review standing privilege, and formalise ownership for model runtime access.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-05 — Overprivileged NHI | In-house inference concentrates privileged access in deployment and runtime identities. |
| NHI-08 — Environment Isolation | Production inference needs separation from testing and model experimentation paths. | |
| Recommendation — Reduce standing access for inference service accounts and keep model runtime permissions narrowly scoped. Isolate production inference identities and prevent lower-trust environments from touching live workloads. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | Inference platforms depend on tokens and credentials that must be governed across their lifecycle. |
| Recommendation — Apply authenticator lifecycle controls to deploy, rotate, and revoke inference credentials on a defined schedule. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The article is fundamentally about who may access and operate AI runtime infrastructure. |
| Recommendation — Review AI platform entitlements regularly and remove any unnecessary privileges from runtime operators. | ||
| MITRE ATT&CK | TA0006;TA0008 — Credential Access; Lateral Movement | Compromised runtime identities could be used to move through AI infrastructure and adjacent systems. |
| Recommendation — Monitor inference infrastructure for credential access attempts and lateral movement from over-scoped service accounts. | ||
Key terms
- Inference Infrastructure: The computing, orchestration, and access layer that runs AI models in production. It includes GPUs, deployment systems, routing, telemetry, and the credentials that let the workload operate. In identity terms, it is a governed execution environment, not just a hosting layer.
- Workload Identity: The identity assigned to a software workload, such as a containerised application, serverless function, or microservice, enabling it to authenticate to other services without storing static credentials.
- Model Runtime: The live execution environment where a model receives requests and produces outputs. It is distinct from model development because it carries production access, operational dependencies, and privileged execution paths that must be managed like other critical workloads.
- Identity Blast Radius: The amount of damage a compromised identity can cause across systems, data, and infrastructure. In NHI environments, it is shaped by permissions, network reach, and administrative capability rather than by the credential alone. Reducing blast radius is a containment strategy that limits lateral movement and data exposure.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org