When organisations do not measure AI-generated code by developer and assistant, they lose the ability to identify power users, adoption patterns, and the real footprint of each coding assistant. That makes it difficult to compare productivity, detect risky usage, or explain AI impact to leadership. The result is weak accountability and controls that are too generic to be useful.
Why Per-User Measurement Changes the Governance Picture
Measuring AI-generated code by both developer and assistant turns a vague adoption question into an operational control question. It shows which engineers rely on which assistants, where assistance concentrates, and whether productivity gains are broad or isolated. That matters because AI coding tools can affect code quality, review burden, intellectual property handling, and exception management in different ways depending on how they are used. Without that split view, leaders often misread overall usage as uniform adoption and miss the teams that need tighter guardrails, training, or approval paths. For a control baseline, NIST SP 800-53 Rev 5 Security and Privacy Controls helps teams think about accountability, logging, and monitoring as governance mechanisms rather than after-the-fact reporting.
In practice, many security and engineering teams only discover their real AI tool footprint after inconsistent usage has already spread across multiple repositories and delivery teams.
How Missing Developer and Assistant Attribution Weakens the Control Model
When AI-generated code is measured only in aggregate, the organisation loses attribution at the level where action is actually possible. A single dashboard total may show that AI is being used, but it does not show who is generating code, which assistant is involved, or whether a small number of users are driving most of the output. That makes it harder to separate routine augmentation from concentrated dependency on a tool, workflow, or vendor. It also weakens the ability to compare tools fairly, because different assistants can produce different levels of review work, defect patterns, or policy friction.
Operationally, that missing attribution creates three common breakdowns. First, security and engineering leaders cannot distinguish experimentation from embedded use, so policy decisions are based on anecdote rather than evidence. Second, reviewers lose the ability to spot outliers, such as one developer producing unusually high volumes of assisted code that may warrant closer scrutiny. Third, governance teams cannot explain whether the organisation is gaining broad productivity benefits or relying on a narrow set of heavy users. If the data only says “AI-assisted code increased,” it hides whether the increase came from safe standard use or from a few high-risk workflows that need more control.
- Per-developer data supports accountability, but per-assistant data shows which tool is shaping the code base.
- Comparing the two reveals whether adoption is balanced, concentrated, or driven by a specific workflow.
- That distinction also helps teams decide whether to tighten review thresholds, retrain users, or adjust approved-tool lists.
Where this guidance breaks down is when organisations treat the metrics as a proxy for code quality without pairing them with review outcomes, defect data, or policy exceptions.
When Aggregated Metrics Hide the Real Risk and the Real Value
Tighter measurement often increases reporting and privacy overhead, so organisations have to balance visibility against unnecessary surveillance. That tradeoff becomes more complex where AI use is informal, because people may move between tools, accounts, or environments faster than governance can keep up. In those cases, the organisation should be clear about whether it is measuring for productivity, compliance, risk management, or all three, because each purpose implies a different level of detail.
There is also a genuine consensus gap in the industry about how much attribution is enough. Some organisations only need team-level trend data, while others need user-level visibility to control sensitive development paths or manage high-impact assistants. The right threshold depends on how much code generation affects regulated systems, protected data, or release authority. If the tooling touches sensitive repositories or production-adjacent workflows, coarse reporting is usually too weak to support oversight. If it is limited to low-risk experimentation, overcollection can create unnecessary friction without improving control.
For broader AI governance, the key issue is not just volume but traceability. Organisations that cannot distinguish developer behaviour from assistant behaviour will struggle to prove which part of the workflow is creating value and which part is introducing exposure.
Risk and Threat Considerations
The main risk is governance blind spots around AI-assisted software production. When AI-generated code is not measured by developer and assistant, organisations can miss concentration risk, uncontrolled adoption, and uneven oversight of higher-risk tool usage. That weakens accountability and makes it harder to recognise when a small set of users or assistants is creating a disproportionate share of the organisation’s code change surface.
Failure mechanism: Aggregated reporting hides attribution, so policy owners cannot see who is using which assistant, where usage is concentrated, or whether a tool is producing code in sensitive contexts without review. That creates a control gap in logging, exception handling, and oversight of potentially unsafe or non-standard development behaviour.
Impact: Leaders lose defensible evidence for productivity claims, risk teams lose visibility into tool-specific exposure, and engineering teams inherit generic controls that do not match actual usage patterns. Over time, that can produce unmanaged dependency on one assistant, inconsistent review standards, and weaker incident investigation when code quality or policy issues emerge.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, CIS Controls v8 and NIST AI RMF set the technical controls, while ISO/IEC 42001:2023 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.2 — Roles, Responsibilities, and Authorities | Attribution gaps weaken accountability for AI-assisted code use. |
| DE.CM — Security Continuous Monitoring | Measuring assistant usage is part of monitoring operational exposure. | |
| Recommendation — Define ownership for AI code metrics and assign clear reporting accountability. Track AI-assisted code patterns continuously to detect abnormal concentration or drift. | ||
| CIS Controls v8 | 14 — Security Awareness and Skills Training | Usage measurement informs training needs and unsafe assistant adoption patterns. |
| Recommendation — Use usage data to target training where AI coding behaviour is most risky. | ||
| ISO/IEC 42001:2023 | 5.2 — AI policy | Assistant-level measurement supports AI policy enforcement and oversight. |
| Recommendation — Tie AI code measurement to policy rules for approved assistants and usage scope. | ||
| NIST AI RMF | GOVERN 3.1 — Establish AI governance structures | Granular measurement is needed to govern AI use with traceable accountability. |
| Recommendation — Instrument AI code activity so governance decisions are based on attributable evidence. | ||
Practitioner Guidance
What to prioritise: Separate the measurement model into two questions: who is generating the code and which assistant is involved. That gives governance teams enough structure to identify concentration without turning reporting into a generic productivity exercise.
What to verify: Check whether the data can support three decisions: identifying heavy users, comparing assistants, and linking usage to sensitive repositories or higher-risk workflows. If it cannot support those decisions, the metric is too coarse to guide policy.
Common mistake: Treating total AI-generated code volume as a sufficient success measure. That can hide risky dependency on one tool or a small number of developers, which is exactly where control weaknesses tend to accumulate.
Practitioner takeaway: The useful unit of governance is not “AI use” in the abstract, but the relationship between a specific developer, a specific assistant, and the code context they are influencing.
Related resources from NHI Mgmt Group
- What breaks when organisations treat AI-generated code as automatically trusted?
- What breaks when organisations rely on static checks alone for AI generated code
- What breaks when organisations rely on post-code scanning alone for AI-generated code?
- How can organisations detect AI-generated passwords in source code?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 7, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org