Join our Newsletter — 33% off our NHI Course
Home› FAQ› Cyber Security› How should Java teams safely extract untrusted archives…
Cyber Security

How should Java teams safely extract untrusted archives without creating path traversal risk?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 27, 2026 Domain: Cyber Security

Treat every archive entry name as untrusted input and resolve it against a fixed extraction root before writing any file. Use canonical path checks to confirm the resolved destination stays inside the intended directory, and reject entries that escape it. This matters most in applications that process uploads, because a crafted path can overwrite server files and, in some cases, enable remote code execution.

How should Java teams prevent archive extraction from escaping the target directory?

Safe extraction starts with a strict boundary: every entry must be resolved against a fixed root, then checked after canonicalisation before any file is written. That protects the application from path traversal when an uploaded ZIP, TAR, or similar archive contains hostile paths such as parent-directory references, absolute paths, or tricky encoded names.

The practical rule is simple. Resolve the destination path, normalise it, and confirm it still begins with the intended extraction directory before creating directories or writing bytes. If the entry would escape, reject it outright. This is especially important in upload handlers and build pipelines, where archive contents are treated as data but can behave like instructions for the filesystem.

Java teams also need to treat archive metadata as untrusted, not just file contents. A safe extractor should verify each entry independently, handle symlinks and hard links carefully, and avoid trusting library defaults that may preserve the archive’s original directory structure without boundary checks. The protection is not the archive format itself, it is the discipline of constraining every write to a known root.

Risk and Threat Considerations

Path traversal turns archive extraction into a write primitive against the host filesystem. If the extractor accepts attacker-controlled entry names, a malicious archive can overwrite configuration files, drop web-accessible payloads, or place code where the runtime later executes it.

Failure mechanism: The code concatenates or writes archive paths without proving that the final canonical destination remains inside the extraction root, so parent-directory segments, absolute paths, symlinks, or encoded variants escape the sandbox.

Impact: The application can corrupt files outside the intended directory, break service integrity, expose secrets, or in the worst case enable remote code execution when the overwritten path is later loaded or executed.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP ASVS, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.

FrameworkControl / ReferenceRelevance
OWASP ASVSV5 — File HandlingArchive extraction is file handling with attacker-controlled paths.
Recommendation — Validate archive entry destinations before writing any extracted file.
NIST SP 800-53 Rev 5SI-10 — Information Input ValidationEntry names are untrusted input that must be validated before filesystem writes.
AC-6 — Least PrivilegeLimiting write authority reduces the blast radius of a malicious archive.
Recommendation — Validate and constrain archive paths before processing extracted content. Restrict extraction processes to the minimum filesystem privileges needed.
CIS Controls v8CIS-3 — Data ProtectionArchive traversal can expose or overwrite sensitive files on disk.
Recommendation — Protect sensitive file locations from application write access.
ISO/IEC 27001:2022A.8.24 — Use of cryptographyNo material fit for the exact extraction question.
Recommendation — Use secure file-processing controls for untrusted archive extraction.

Practitioner Guidance

What to verify: Check the post-normalisation destination, not the raw entry string. If the resolved path does not stay under a predeclared root, treat the archive entry as malicious and stop processing that entry.

Common mistake: Do not rely on filename filtering alone, because blacklist-style checks are easy to bypass with encoding, nested paths, or archive features such as links. The control has to be path-resolution based, not pattern based.

What good looks like: A secure extractor applies the same boundary test to every entry, creates only expected directories, and leaves a clear audit trail for rejected paths so operators can tell whether the archive was malformed or actively hostile.

Practitioner takeaway: The safest default is to assume archive names are attacker input, constrain every write to an allowlisted root, and fail closed whenever path resolution cannot be proven safe.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 27, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org