Join our Newsletter — 33% off our NHI Course

Approval Test

An approval test checks that a piece of code produces the expected output before and after a change. It is especially useful during refactoring because it protects behaviour while the internal structure is improved, giving teams confidence that modernization has not altered functionality.

What Approval Test Means in Practice

An approval test is a regression check that locks in expected output, so teams can refactor internal code with confidence that visible behaviour has not changed. It is strongest when the output is stable, meaningful, and easy to compare.

Because the test compares current output to an approved baseline, it shifts attention from implementation details to contract preservation. That makes it especially useful for legacy code, high-change modules, and transformations where ordinary unit tests would be too brittle or too narrow.

How Approval Tests Protect Behaviour

Approval tests work by capturing the output of a function, service, or workflow and comparing it with a stored “approved” result. If the output changes unexpectedly, the test fails and prompts a review of whether the change was intended.

This pattern is valuable when behaviour is complex, text-heavy, or otherwise hard to assert with many small expectations. Instead of testing every internal step, the team validates the complete observed result, which can include rendered text, serialized data, reports, or command output.

Approval tests are not a replacement for all other tests. They complement unit, integration, and contract tests by covering the broader outcome, while other tests continue to verify edge cases, internal logic, and integration boundaries.

Where Approval Tests Fit Best

Approval tests are most useful during refactoring, modernization, and characterization of legacy systems. They help teams make safe structural changes when the current behaviour is known but not yet well documented in a precise specification.

They also work well for code that generates human-readable or machine-readable artifacts that should remain consistent over time. Common examples include templates, API payload snapshots, generated code, and transformation pipelines where output drift would be expensive or disruptive.

The trade-off is that approved output can become noisy or brittle if it is too large, too volatile, or too dependent on incidental formatting. Good approval tests focus on meaningful output boundaries so that a genuine behaviour change is visible without turning the test into a maintenance burden.

Common Failure Modes and Good Test Design

An approval test is only as useful as the quality of the approved output. If the baseline includes unstable timestamps, random values, environment-specific paths, or incidental formatting, the test may fail for reasons that have nothing to do with the intended behaviour.

Well-designed approval tests normalize or filter out irrelevant variation before comparison. They also keep the approved artifact small enough to review, so a failure produces an understandable diff rather than an unreadable wall of text.

Because the test captures observable behaviour, it can expose unintended changes that unit tests miss, especially when a refactor affects formatting, ordering, or composition rather than core logic. That makes the approval file both a safety net and a living record of what the code was expected to do at the time it was approved.

Risk and Threat Considerations

Approval tests reduce the risk of accidental behaviour drift during refactoring, but they can also create false confidence if the approved output is incomplete, overly permissive, or full of unstable noise. In those cases, a failing test may be ignored, or a passing test may miss a real regression in important output.

Failure mechanism: Weak baselines, poor normalization, or review fatigue can let unintended output changes slip through, while non-deterministic fields can create churn that hides real defects.

Impact: Teams may ship broken behaviour after modernization, miss output regressions in user-facing artifacts, or spend excessive time maintaining tests that no longer signal meaningful change.

Practitioner Guidance

What to watch for: Use approval tests where the important question is “did the output stay the same?” rather than “did each internal step happen?” They are strongest when the output is stable enough to review and the change risk is concentrated in behavioural drift.

Common misunderstanding: Approval tests are often treated as a catch-all regression strategy, but they work best as one layer in a broader test suite. The most effective teams pair them with more granular tests so a baseline diff is informative, not ambiguous.

Practitioner takeaway: Treat the approved artifact as a contract for visible behaviour, and keep it focused on the output that matters most to the reader or downstream consumer.