A common schema standardizes columns, but it does not explain what those columns mean in each data type. SaaS audit logs, EDR telemetry, HTTP access logs, proxy events, and firewall flows have different security surfaces and different signals. If teams force them into one shape, they lose critical fields or flatten away the distinctions that make detection accurate.
Why a common schema breaks down across different telemetry types
A common schema can normalize column names, but it cannot normalize meaning. A field like user, source, action, or object carries different security context in SaaS audit logs, endpoint telemetry, HTTP logs, proxy records, and firewall flows. If teams map everything to the same shape too early, they preserve uniformity at the cost of the evidence analysts need to tell benign activity from suspicious behavior.
That failure is not just cosmetic. Detection logic depends on what a field actually represents, how reliable it is, and which signals are missing when the source is compressed into a generic model. One source may expose explicit actor and object semantics, while another only shows network tuples, response codes, or process activity. The schema may look clean, but the investigative value drops because context has been flattened.
The practical problem is that detections are usually built on relationships, not raw columns. Analysts need to understand who acted, what changed, from where, through which protocol, and with what security consequence. When a schema hides source-specific meaning, it becomes harder to express conditions such as unusual SaaS admin actions, endpoint child-process chains, API misuse patterns, or east-west traffic that violates an expected trust boundary.
What context-rich detection actually requires
Context-rich detection depends on preserving source semantics alongside normalization. That usually means keeping both the common fields and the source-specific fields that explain them, so a query can still distinguish an audit event from a process event, or a proxy decision from a network flow. A useful schema supports correlation, but it does not pretend every telemetry type can be interpreted the same way.
For SaaS, the analyst often needs account, tenant, permission, object, and admin-action context. For endpoint data, the important context may be process lineage, command line, image hash, parent-child behavior, and file or registry changes. For API logs, method, resource, authentication state, rate, and response code matter. For network telemetry, session direction, duration, destination reputation, and protocol behavior may be more important than a generic actor-object model.
This is why one-size-fits-all normalization is useful for storage and routing, but not sufficient for detection design. The schema should support cross-domain joins without erasing the provenance of each source. When analysts can still ask source-appropriate questions, they can build detections that are precise enough to reduce noise and rich enough to reveal attacker tradecraft.
How to design detections without flattening away the evidence
Design the schema around investigation, not around the convenience of a dashboard. Keep canonical fields for shared concepts such as time, actor, target, action, and outcome, then retain source-native fields for the details that make those concepts trustworthy. That lets engineering standardize ingestion while preserving the higher-fidelity signals detection engineers rely on.
The best test is whether an analyst can still answer the source-native question after normalization. If the schema makes it impossible to distinguish a SaaS permission change from an ordinary user action, or an endpoint execution chain from a routine launch, the model is too coarse. A good common schema supports comparison across sources, but it still lets each source speak in its own security language.
For teams building detection content, the rule is simple: normalize the transport of data, not the meaning of the event. Use the common schema to correlate, enrich, and search consistently, but preserve the original fields and labels that explain why the event matters. That is the difference between having interoperable telemetry and having telemetry that can actually support context-rich detection.
Risk and Threat Considerations
The main risk is analytic blind spots created by over-normalization. When distinct telemetry types are forced into one generic structure, attackers benefit from the lost nuance, because the environment becomes harder to detect, harder to tune, and easier to misread during investigation.
Failure mechanism: The schema strips away source-specific semantics, so detections key off shallow fields that do not capture the real security meaning of the event. That can hide abuse patterns, inflate false positives, and weaken correlation across SaaS, endpoint, API, and network sources.
Impact: Teams miss suspicious behavior, respond later, and lose confidence in cross-source detections. In practice, that can leave privilege misuse, unauthorized access paths, or lateral movement less visible than they should be.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP ASVS, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP ASVS | V16 — Security Logging and Error Handling | Source-specific log context is essential for detection quality. |
| Recommendation — Preserve log semantics and error detail so detections can distinguish meaningful events from generic records. | ||
| NIST CSF 2.0 | DE.CM-01 — Monitoring for Unauthorized Personnel, Connections, Devices, and Software | Cross-source telemetry must retain context to support effective monitoring. |
| Recommendation — Maintain monitoring data that preserves enough context to spot unauthorized or unusual activity. | ||
| NIST SP 800-53 Rev 5 | AU-3 — Content of Audit Records | Audit records need source-specific details to remain useful for analysis and correlation. |
| Recommendation — Include the event details needed to reconstruct security-relevant activity. | ||
| ISO/IEC 27001:2022 | A.8.15 — Logging | Logging controls depend on retaining useful event context across diverse systems. |
| Recommendation — Specify logging requirements that preserve source-specific meaning for investigation. | ||
Practitioner Guidance
What to prioritize: Preserve the source-native fields that explain the event before you standardize for analytics. If a field loses its detection value when translated into the common model, keep the original alongside the normalized version.
What to verify: Test representative detections against each telemetry type and confirm that the schema still supports the intended security question. If the query cannot distinguish context that analysts would use manually, the model needs refinement.
Common mistake: Treating a schema project as a data-modeling exercise instead of a detection-engineering decision. The right design is the one that helps analysts preserve meaning across sources, not the one that merely makes all records look alike.
Practitioner takeaway: Standardization should improve correlation, not erase provenance; if the normalized model cannot preserve source-specific meaning, it is too abstract for reliable detection.
Related resources from NHI Mgmt Group
- How should organisations govern SaaS discovery across finance, identity, and endpoint data?
- How should security teams connect identities across cloud, SaaS, and endpoint data?
- What breaks when data protection is split across SaaS, endpoint, browser, and AI tools?
- How should security teams prevent data exfiltration across endpoint, SaaS, and AI tools?