When the main concern is accountability, drift, or investigation readiness, tracing comes first. A system that cannot show what it did, when it did it, and what it touched is already difficult to govern, even if its outputs appear acceptable.
Why tracing should outrank tuning when governance is the real problem
Activity tracing matters first when the organisation needs a reliable account of system behaviour, not just better outputs. If you cannot reconstruct what an AI system did, which actions it took, or which records it accessed, tuning may improve quality while leaving governance blind spots untouched. That is a control problem before it is a model-quality problem.
Tracing becomes the priority when the question is whether the system can be explained after the fact. Broader tuning can reduce error rates, but it does not by itself create auditability, investigation evidence, or a usable timeline for incident review. In practice, the right order is often visibility first, optimisation second.
That distinction matters because a tuned system can still be operationally opaque. For accountability and change control, teams need durable records of prompts, tool calls, outputs, decision paths, and side effects. Without that record, a model may look acceptable in testing yet remain difficult to govern in production.
When tracing is the better first investment
Prioritise tracing when the environment is already producing decisions or actions that may need to be justified later. This is especially true where the system touches regulated workflows, internal approvals, external communications, or anything with downstream business impact. The immediate need is to preserve evidence and attribution, not to squeeze out incremental model performance.
Tracing also comes first when drift is the concern. If behaviour changes over time, tuning can mask the symptom while leaving the underlying change undocumented. A tracing layer helps you compare runs, spot unexpected tool use, and determine whether a failure is a model issue, a prompt issue, a retrieval issue, or a process issue.
For agentic or tool-using systems, tracing is often the only practical way to understand delegated actions. An investigation usually needs a sequence, not a score. That sequence should show intent, inputs, intermediate steps, accessed resources, and the resulting action so that teams can separate normal variance from unsafe behaviour.
What tuning still does, and why it usually follows tracing
Broader tuning is valuable when the organisation already knows what the system is doing and wants to improve accuracy, consistency, or safety. It is a refinement control. Tracing is a diagnostic and accountability control. If the system is not observable enough to explain failures, tuning risks becoming guesswork.
That is why tracing often has a stronger governance return early in the lifecycle. It shortens root-cause analysis, supports review and recertification, and gives operators a factual basis for deciding whether the answer is better prompts, better data, better guardrails, or a model change. Tuning should be informed by trace evidence, not used as a substitute for it.
Where a system is already stable and the remaining problem is precision, tuning can move ahead sooner. But if the organisation cannot answer basic questions about traceability, tuning first tends to produce opaque improvement rather than controlled improvement. The more autonomous the workflow, the more costly that opacity becomes.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack surface, NIST CSF 2.0 sets the technical controls, and ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-01 — Monitoring for Anomalies and Events | Tracing creates the event visibility needed to detect abnormal AI behaviour. |
| GV.OV-01 — Oversight of Cybersecurity Risk Management Strategy | Traceability supports oversight, accountability, and review of system behaviour. | |
| Recommendation — Instrument AI activity so anomalous actions are observable in detection workflows. Require auditable traces before treating model behaviour as governed and reviewable. | ||
| ISO/IEC 27001:2022 | A.8.15 — Logging | Activity tracing is the logging basis for investigation and accountability. |
| A.5.15 — Access control | Trace data shows whether access and actions stayed within authorised bounds. | |
| Recommendation — Log model actions and access events with enough detail to support reconstruction. Use trace evidence to verify that actions remained within approved access boundaries. | ||
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Tracing is essential when agents can act with delegated authority. |
| Recommendation — Capture agent action trails to detect privilege misuse and unsafe delegated actions. | ||
| OWASP Non-Human Identity Top 10 | NHI-10 — Human Use of NHI | Traces help distinguish human instructions from non-human actions and responsibility. |
| Recommendation — Retain traces that show when humans initiated or influenced non-human actions. | ||
Practitioner Guidance
What to prioritise: Start with tracing when you need accountability, drift analysis, or post-incident reconstruction. If you cannot produce a credible activity trail, model tuning is unlikely to resolve the governance gap.
What to verify: Confirm that traces are detailed enough to support a timeline, including key inputs, outputs, tool actions, and touched resources. If the record cannot support investigation or review, it is not yet good enough to rely on.
What good looks like: Teams can answer who did what, when, and through which step or tool without reconstructing the event from logs spread across unrelated systems.
Practitioner takeaway: Tune for quality once you can observe behaviour with confidence, but trace first when you need to govern, explain, or investigate what the system actually did.
Related resources from NHI Mgmt Group
- When should organisations prioritise AI security posture management over broader detection tuning?
- When should organisations prioritise agent identity controls over model tuning?
- When should organisations prioritise output validation over model tuning?
- Should organisations prioritise external exposure or internal credential governance first?