Join our Newsletter — 33% off our NHI Course
Home› FAQ› AI Security› What should teams do first when they start…
AI Security

What should teams do first when they start using LLMs for test creation?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated October 11, 2026 Domain: AI Security

Start by locking the basics: role, stack, test architecture, output type, and expected assertions. If those elements are not fixed first, the model may produce code that is technically valid but operationally unusable. A narrow prompt produces more reusable tests than a broad one because it removes ambiguity before generation begins.

Why the first prompt should define the test shape, not just the test idea

The first step is to remove ambiguity before generation starts. LLMs can draft syntactically correct test code that still fails in practice if the prompt leaves role, stack, architecture, output type, or assertion style open to interpretation. For teams, the key decision is not “can the model write a test,” but “what exact test artifact should it produce consistently?”

That means the prompt should anchor the model to a single target context: the language, framework, runtime, test layer, and the kind of assertion expected. When those constraints are explicit, the model can stay inside the existing test architecture instead of improvising new patterns, helper structures, or output conventions that look plausible but do not fit the repository.

Clear scoping also improves reuse. A narrow prompt gives the model enough structure to produce tests that match local conventions, naming, fixtures, setup patterns, and assertion libraries. A broad prompt often creates code that is functionally related to the request but awkward to maintain because it does not align with the team’s test architecture or execution environment.

What has to be fixed before the model writes anything

Start with the test’s operating context: what codebase it belongs to, what framework it uses, where it will run, and what layer it targets. A unit test, component test, integration test, and end-to-end test all need different assumptions, different dependencies, and different levels of setup. If that boundary is missing, the model may generate a technically valid test that is wrong for the intended layer.

Next, define the expected output type and assertions. Teams should say whether they want a success-path test, a negative-path test, edge-case coverage, snapshot-style output, mocking behavior, or state verification. The model needs that signal to decide whether to assert return values, thrown errors, emitted events, side effects, calls to dependencies, or rendered output.

Then lock the assertion contract. If the team cares about exact values, ordering, error messages, call counts, or schema shape, those expectations should be explicit. Without that, the model may overfit to generic checks that pass locally but do not validate the business rule the test was meant to protect.

How to keep LLM-generated tests operationally useful

The best prompt pattern is to constrain the model to the existing test architecture rather than asking it to invent one. In practice, that means naming the framework, the file location or naming convention, the fixtures available, and the style of assertions the codebase already uses. When the prompt mirrors repository reality, the generated test is more likely to be dropped into the suite with minimal editing.

A useful test prompt also limits scope to one behavior at a time. LLMs tend to overproduce when given broad instructions, which leads to tests that combine multiple paths, multiple mocks, and multiple expectations in one block. Teams usually get more value from one narrow, readable test than from one dense test that is hard to diagnose when it fails.

If the test depends on setup or mock data, specify that too. The model should know which dependencies may be mocked, what inputs are fixed, and what should remain real. That reduces the chance of the test encoding unrealistic assumptions or hiding the behavior under brittle scaffolding.

Risk and Threat Considerations

LLM-generated tests create quality and assurance risk when teams let the model infer too much. The main failure mode is not obvious syntax errors, it is tests that compile, run, and still validate the wrong behavior, which can give false confidence in coverage, regressions, or release readiness.

Failure mechanism: Ambiguous prompts let the model choose its own stack, architecture, assertion style, or scope, so the generated test may miss the intended behavior while still looking plausible to reviewers.

Impact: Teams can accumulate brittle or misleading tests that waste review time, hide gaps in coverage, and make later refactoring harder because the test suite no longer reflects the application’s real contract.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP ASVS, CIS Controls v8, NIST SP 800-53 Rev 5 and OWASP SAMM set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP ASVSV15 — Secure Coding and ArchitectureTest generation must fit the app's architecture and test boundaries.
Recommendation — Define the target test layer and architecture before generating assertions.
CIS Controls v8CIS-16 — Application Software SecurityLLM-generated tests are part of application security assurance and validation.
Recommendation — Validate generated tests against the application's expected security and behavior requirements.
NIST SP 800-53 Rev 5SA-11 — Developer Testing and EvaluationThe question is about creating tests that correctly verify system behavior.
Recommendation — Specify acceptance conditions so developer testing validates the intended control or behavior.
OWASP SAMMBSR — Requirements and Design ReviewTeams need to define test expectations before using AI to generate code.
Recommendation — Review the required behavior and constraints before asking the model to draft tests.

Practitioner Guidance

What to prioritise: Lock the minimum viable test contract first, role, stack, architecture, output type, and the exact assertion target. If any of those are still undecided, the prompt is not ready for generation.

What to verify: Check that the generated test matches the repository’s actual testing layer and assertion conventions before you trust it. A test that passes locally but uses the wrong abstraction level is usually a prompt-quality problem, not a code-quality win.

Common mistake: Teams often ask for “a test” when they really need a unit, integration, or behavior-specific test. That ambiguity pushes the model toward generic code that is valid in isolation but awkward in the suite.

Practitioner takeaway: The first prompt should reduce degrees of freedom, not increase them, because test usefulness comes from alignment with the existing architecture and assertions, not from model creativity.

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org