Organisations need custom SQL data discovery when they must locate information that is not covered by built-in categories, such as unusual internal identifiers, specialised regulated fields, or narrowly defined subject access request records. The practical test is whether the data must be found, highlighted, and returned accurately under a legal or operational requirement. If standard templates cannot express the pattern, custom profiling is needed.
When custom SQL discovery becomes necessary
Standard sensitive data templates work well when the target is already described by familiar patterns like email addresses, card numbers, or common personal identifiers. Custom SQL data discovery becomes necessary when the organisation needs to detect data whose structure, naming, or storage pattern is specific to its own business process, regulatory obligation, or system design. The test is whether the data can be reliably found and returned with enough precision to support the required use case.
That matters because discovery is not just about finding anything that looks sensitive. It is about identifying the exact records that must be surfaced for access review, legal response, retention decisions, or security classification. In practice, custom queries are often needed for fields that standard content detectors miss, including internal reference codes, case-specific subject access request records, specialist operational identifiers, and other data that only exists in the organisation’s own schema.
What standard templates can miss
Built-in templates usually detect broad, reusable patterns. They are fast to deploy and useful for common data classes, but they are limited by design. If the information is encoded in a non-standard way, split across multiple columns, nested in application logs, or described by business logic rather than by a simple format, a template may not flag it accurately.
Custom SQL discovery is the better option when the pattern depends on context. A column might contain a harmless-looking identifier that becomes sensitive only when joined to another table, or a text field may hold regulated material that can only be identified by matching a combination of keywords, identifiers, and record-state conditions. The stronger the dependence on schema knowledge, the less likely a generic template will be sufficient.
- Use templates for common patterns with low ambiguity.
- Use custom SQL when the sensitivity depends on business meaning, joins, or exceptions.
- Prefer custom logic when false negatives would create legal, operational, or privacy exposure.
How practitioners decide whether to build custom SQL
The practical decision point is whether the data must be found accurately under a specific obligation. If the answer is yes, and the standard template cannot express the rule, custom SQL is justified. That is especially true when the organisation needs a repeatable discovery process across multiple systems, rather than a one-time search.
Discovery should also be tested against output quality. If the query is too broad, teams will drown in false positives and stop trusting the results. If it is too narrow, they will miss records that should have been included. The right custom rule is the one that aligns with the legal or operational definition of the data, not just its surface format. NHIMG’s Lifecycle Processes for Managing NHIs is a useful reminder that discovery must connect to governance, ownership, and ongoing review, not just initial detection.
Risk and Threat Considerations
When standard templates fail to detect sensitive records, organisations can create blind spots in legal response, retention, and exposure management. The risk is not only missed data, but also misplaced confidence in a discovery programme that appears complete while actually leaving special-case records unclassified or unreviewed.
Failure mechanism: The discovery rule set cannot express the real pattern, so the affected records remain invisible to scanning, classification, or reporting workflows. That commonly happens when sensitivity is defined by schema relationships, local naming, or business logic rather than by a canonical pattern.
Impact: The organisation may fail to locate records needed for access requests, investigations, deletion, or regulatory response, and it may also understate exposure during audit or incident review.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | RA-5 — Vulnerability Monitoring and Scanning | Custom SQL discovery is a targeted scanning method for locating sensitive records. |
| Recommendation — Define tailored discovery rules to find non-standard sensitive records your templates miss. | ||
| ISO/IEC 27001:2022 | A.5.12 — Classification of information | The question is about identifying information that must be classified and surfaced accurately. |
| A.5.34 — Privacy and protection of PII | Custom discovery is often needed to locate records subject to privacy obligations. | |
| Recommendation — Create classification rules that reflect the organisation's real sensitive data definitions. Use discovery queries that reliably identify records covered by privacy obligations. | ||
| NIST CSF 2.0 | ID.AM-01 — Physical devices and systems are inventoried | Discovery depends on knowing where data resides and how it is organised. |
| PR.DS-01 — Data-at-rest is protected | Finding sensitive data accurately is a prerequisite to protecting it at rest. | |
| Recommendation — Inventory the data locations and systems that custom rules must cover. Use discovery results to apply the right protection controls to sensitive data. | ||
Practitioner Guidance
What to verify: Validate the rule against real records, not just sample values. A good custom SQL discovery rule should return the expected sensitive rows, exclude obvious non-sensitive noise, and remain stable when the schema changes in ways that do not alter the underlying business meaning.
Common mistake: Treating template coverage as complete because the platform reports a scan succeeded. Success means the query ran, not that it captured the organisation’s unique sensitive data definitions.
Decision rule: If the data would be hard to explain in a generic template, or if a missed record would create legal or operational harm, build the custom discovery rule and review it with the team that owns the data definition.
Practitioner takeaway: Custom SQL is warranted when precision matters more than convenience, especially where the real sensitivity lives in business context, not in a standard data pattern.
Related resources from NHI Mgmt Group
- Why do organisations need real-time remediation instead of discovery alone for sensitive data risks?
- When should organisations add custom reporting capabilities instead of relying on standard analytics views?
- Why do data risk programs need custom actions instead of relying only on standard remediation workflows?
- How should security teams detect custom sensitive data without relying on regex?