Join our Newsletter — 33% off our NHI Course
Home› FAQ› Architecture & Implementation› What is the difference between running data quality…
Architecture & Implementation

What is the difference between running data quality jobs in a separate compute plane and pushing them to the source system?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 26, 2026 Domain: Architecture & Implementation

A separate compute plane copies or moves data for processing, which adds latency, cost, and operational complexity. Pushdown generates queries that run directly against the source, so the data remains in place. For teams, that means faster checks, less infrastructure burden, and a more efficient way to validate data at scale.

Separate Compute Plane vs Pushdown: What Actually Changes

A separate compute plane changes where the work happens. Data is copied, staged, or moved into another engine, which creates an extra processing layer and a second place to operate and troubleshoot. Pushdown keeps the data in the source system and asks that system to do the filtering or validation work directly, so execution stays closer to the original data.

The practical difference is less about syntax and more about control boundaries. With a separate plane, you are managing data movement plus compute behaviour. With pushdown, you are relying on the source system’s query engine, permissions, and performance characteristics to execute the job efficiently.

That matters because the trade-off is not neutral: separate compute can be more flexible for heavy transforms, but it introduces transfer overhead and duplicated operational surfaces. Pushdown is usually better when the check can be expressed in source-native terms and the source can handle the query load without harming production workloads.

Why Pushdown Often Wins for Data Quality Checks

Pushdown is attractive when the job is mostly about reading, filtering, joining, or validating records rather than reshaping them. It avoids unnecessary replication, reduces latency, and can lower cost because the source system does not have to hand data off to another platform first.

It also changes failure modes. In a separate compute model, you may need to account for failed transfers, stale snapshots, and mismatches between source state and copied state. In pushdown, the main question becomes whether the source can safely absorb the query volume and whether the generated SQL or predicate logic is faithful to the data quality rule.

For large datasets, this difference is often decisive. A well-designed pushdown can make repeated checks more scalable because it filters early, returns only the needed results, and avoids moving full tables just to test a small condition. When the source is highly governed or rate-limited, though, that efficiency can come with contention risk if the jobs are too frequent or too expensive.

When a Separate Compute Plane Is Still the Better Choice

A separate compute plane remains useful when the validation logic is too complex for source pushdown, when the source does not support the needed functions, or when the workload requires heavy enrichment, complex windowing, or cross-source joins. It can also make sense when teams need isolation from production databases and want to run expensive checks without competing with application traffic.

The operational cost is that you now own more infrastructure and more lifecycle management. You need to monitor freshness, lineage, retry behaviour, and the effect of data movement on the overall pipeline. If the copied data is large or rapidly changing, the compute plane can become the bottleneck rather than the accelerator.

For governance-heavy environments, the choice should follow the job shape. Use pushdown when the rule is simple, source-native, and read-heavy. Use separate compute when the rule depends on broader transformation logic, cross-system context, or performance isolation that the source cannot provide safely.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.DS-01 — Data-at-rest is protectedData quality jobs should avoid unnecessary copying and preserve data handling boundaries.
Recommendation — Minimise data movement and protect copied datasets only when separate compute is required.
NIST SP 800-53 Rev 5CM-6 — Configuration SettingsPushdown depends on predictable source-side execution and controlled query behaviour.
Recommendation — Standardise source query settings and limits before relying on pushdown.
ISO/IEC 27001:2022A.8.13 — Information backupSeparate compute often introduces copied or staged data that needs controlled handling.
Recommendation — Treat staged datasets as governed information assets with defined retention and protection.
CIS Controls v8CIS-11 — Data RecoverySeparate compute pipelines can create additional copies and recovery dependencies.
Recommendation — Inventory and protect any copied datasets created for validation workloads.

Practitioner Guidance

What to verify: Before choosing pushdown, confirm that the source can execute the query pattern at the required frequency without degrading production performance. If the check depends on joins or transforms that the source cannot express cleanly, the separate compute plane is usually the safer design.

Decision rule: If the data quality rule can be evaluated with source-native filters, aggregates, or predicates, prefer pushdown. If the rule depends on copied snapshots, complex enrichment, or workload isolation, keep the separate compute plane and treat data freshness as a first-class control.

Trade-off: Pushdown reduces movement and cost, but it shifts dependency onto the source engine’s availability and capacity. Separate compute gives you more freedom in processing, but it adds latency, duplicated operations, and another place where things can fail.

Practitioner takeaway: The right choice is usually the one that minimizes unnecessary data movement while preserving trustworthy execution, and that means matching the compute model to the complexity and runtime impact of the check.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 26, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org