Join our Newsletter — 33% off our NHI Course

When should organisations use query history and popularity scores to make data governance decisions?

Organisations should use query history and popularity scores when they need a fast, evidence-based way to separate critical data from idle data. These signals are most useful for deciding which assets to certify, which to improve, and which to archive. They also help align data product roadmaps with real consumption patterns instead of assumptions.

Why This Matters for Security Teams

Query history and popularity scores are useful because they turn data governance from a debate about ownership into a decision anchored in actual usage. That matters when organisations need to decide which datasets deserve certification, stricter controls, better documentation, or retirement. The risk is not just wasted effort on low-value assets. It is also blind spots around heavily consumed data products that appear stable on paper but are actually business-critical.

This approach aligns well with the control mindset in NIST Cybersecurity Framework 2.0, because evidence-based prioritisation helps teams focus protection where impact is highest. NHIMG research also shows why this matters operationally: the Ultimate Guide to NHIs — Key Research and Survey Results highlights persistent visibility and confidence gaps in identity governance, which is a useful parallel for data governance teams that are still managing by assumption rather than measurement.

In practice, many security teams only discover that a “low priority” dataset is mission-critical after access reviews, incident response, or audit findings expose how often it is actually consumed.

How It Works in Practice

Start by treating query history as an operational signal, not a final verdict. The most useful pattern is to combine frequency, recency, and breadth of access with business context. A dataset queried daily by multiple teams may deserve stronger certification, lineage review, schema stability, and tighter change control. A dataset that has not been queried for months may be a candidate for archival, but only after confirming that the absence of queries is not due to broken discoverability, hidden downstream jobs, or seasonal usage.

Popularity scores work best when they are transparent and explainable. Teams should document how the score is calculated, what time window it covers, and whether it reflects direct human queries, automated jobs, or both. That distinction matters because machine consumption can dwarf human usage and distort the picture if it is not separated cleanly. The governance decision should then map the score to action: certify high-value data, improve data quality on frequently used assets, and reduce support for idle or duplicated datasets.

For practitioners, this is strongest when paired with lifecycle governance. The Ultimate Guide to NHIs — Lifecycle Processes for Managing NHIs is a useful reference point for the broader principle that assets should be managed according to actual lifecycle state, not static labels. The same logic applies to data products: high consumption should trigger active stewardship, while low consumption should trigger review, not automatic deletion. Current guidance suggests popularity metrics should support governance decisions, not replace stewardship or domain ownership. These controls tend to break down when query logging is incomplete, when self-service tools fragment visibility across platforms, or when automated pipelines mask real consumption patterns.

Common Variations and Edge Cases

Tighter popularity-based governance often increases administrative overhead, requiring organisations to balance better prioritisation against the risk of overreacting to noisy metrics. A dataset can look unpopular for legitimate reasons, especially in regulated environments where access is intentionally restricted, in seasonal businesses where usage is cyclical, or in analytics stacks where consumption happens through indirect dashboards rather than direct SQL queries.

There is no universal standard for this yet, so best practice is evolving. Some teams weight recent query volume more heavily, while others normalise by department, product line, or data sensitivity. The important point is consistency: if the method changes from month to month, the score becomes politically useful but operationally weak. Organisations should also watch for “popularity inflation,” where repeated automated queries make a dataset appear more important than the business actually considers it.

For audit and assurance purposes, the Ultimate Guide to NHIs — Regulatory and Audit Perspectives reinforces a broader governance lesson: evidence matters, but it must be defensible. Teams should be ready to explain why a high-score asset was certified, why a low-score asset was archived, and what exception process exists when usage patterns do not tell the full story. In parallel, NIST SP 800-53 Rev 5 Security and Privacy Controls supports the need for documented control selection and review. The approach becomes unreliable when query data is fragmented across warehouses, BI tools, and API layers because the apparent usage picture is no longer complete.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and CSA MAESTRO address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST CSF 2.0 ID.AM-1 Asset inventory decisions depend on knowing which data is actually used.
NIST SP 800-53 Rev 5 CM-8 Configuration baselines need evidence for which datasets are active or retired.
OWASP Non-Human Identity Top 10 Popularity signals help prioritize governance for highly used identity-backed data assets.
NIST AI RMF GOVERN Governance requires explainable, evidence-based decision-making and accountability.
CSA MAESTRO GOV-01 Operational governance should align control intensity to actual workload importance.

Define score logic, ownership, and review cadence before using usage metrics in governance.