TL;DR: Data classification is most effective as a repeatable process that combines content, ownership, and movement context, and Orion argues that large language models plus identity signals can reduce false positives while surfacing real exposures. The shift matters because static pattern rules miss how data is actually being handled, which is where governance breaks down.
NHIMG editorial — based on content published by Orion: Data classification process guide and context-aware classification approach
By the numbers:
- Only 5.7% of organisations have full visibility into their service accounts.
- 96% of organisations store secrets outside of secrets managers in vulnerable locations including code, config files, and CI/CD tools.
- 79% of organisations have experienced secrets leaks, and 77% of those incidents resulted in tangible damage.
Questions worth separating out
A: Security teams should combine automated classifiers with business context, so detection reflects what matters to the organisation rather than generic data categories.
Q: Why do static pattern rules fail for sensitive data classification?
A: Static rules can recognise patterns, but they cannot tell whether the transfer is routine, delegated, or out of policy.
Q: What are the signs that a data classification process is breaking down?
A: Common warning signs include too many manual exceptions, labels that do not change access or routing, stale classifications after business changes, and users treating labels as decoration.
Practitioner guidance
- Define classification levels as handling rules Tie each sensitivity level to explicit rules for storage, sharing, encryption, and approved destinations so the label changes behaviour, not just metadata.
- Add identity and destination signals to verdicts Require classification systems to consider who is moving the data, where it is going, and whether that destination matches the expected business context.
- Review labels on a fixed and event-driven cycle Reclassify when the business changes, when a new SaaS or AI tool is introduced, or when data starts moving in ways the original label did not anticipate.
What's in the full article
Orion's full guide covers the operational detail this post intentionally leaves for the source:
- Step-by-step data classification workflow for setting levels, ownership, and review cadence
- Examples of how Orion applies large language models to classify data by content and context
- Operational guidance for handling labels when data moves into email, SaaS, or AI assistants
- Customer examples showing how teams reduced false positives while finding real exposures
👉 Read Orion's guide to data classification as a repeatable security process →
Data classification and identity signals: where static labels fall short?
Explore further
Data classification is becoming a contextual control, not a label management exercise. Static tagging cannot keep pace with modern data movement across SaaS, email, collaboration tools, and AI assistants. The article’s strongest contribution is that it treats classification as a repeated governance process that must survive real-world use. For practitioners, that means the label only matters if it continues to shape handling decisions after the file moves.
A question worth separating out:
Q: Should organisations classify data before or after adding DLP and AI controls?
A: Classify first, but do not stop there. DLP and AI controls need a sensitivity model they can enforce, otherwise they only see isolated events. The right sequence is to define levels and ownership, then connect those levels to the controls that govern movement and use.
👉 Read our full editorial: Data classification is shifting from static labels to context-aware control