Back to skill

Security audit

batch-resume-screener

Security checks for vulnerabilities and agentic risk

Overview

This resume-screening skill is purpose-aligned, but it stores candidate data locally and needs care with untrusted ZIP files.

Use this only in an environment where candidate resume data is allowed to be processed and stored. Prefer an isolated virtual environment with pinned dependencies, avoid untrusted ZIP archives or add extraction limits first, restrict access to generated .txt/JSON/report files, and delete derived resume data when it is no longer needed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
step1_extract_resumes.py:23
Finding

Unbounded ZIP Extraction Enables Resource-Exhaustion Attacks

Content
View full analysis

Vulnerability Details

File Location: step1_extract_resumes.py, lines 23-27
Vulnerability Type: Unrestricted archive extraction
Risk Level: Medium

Vulnerable Code

python
def extract_zip_if_needed(zip_path, extract_dir):
    """Extract ZIP file if needed"""
    if zipfile.is_zipfile(zip_path):
        print(f"Extracting ZIP file: {zip_path}")
        with zipfile.ZipFile(zip_path, 'r') as zip_ref:
            zip_ref.extractall(extract_dir)
        return extract_dir
    return os.path.dirname(zip_path)

Technical Analysis

The script extracts every entry from a user-provided ZIP archive through ZipFile.extractall() without first enforcing limits on:

  • The number of archive entries
  • Total uncompressed size
  • Maximum size of an individual entry
  • Compression ratio
  • Directory nesting depth
  • Encrypted or otherwise abnormal archive entries

Resume archives are an expected untrusted input to this skill. A maliciously constructed ZIP bomb can have a small compressed size while expanding into a very large amount of data. Extraction occurs before the script filters for PDF files, so unsupported files can also consume storage even though they are never processed.

The subsequent PDF traversal and parsing can further amplify resource use if the archive contains many files or oversized PDF documents.

Attack Path

  1. An attacker creates a ZIP archive containing highly compressible data, many nested entries, or an excessive number of files.
  2. The attacker submits the archive as a resume package.
  3. The skill invokes step1_extract_resumes.py with the attacker-controlled archive.
  4. zipfile.is_zipfile() validates only that the input has a recognizable ZIP structure.
  5. zip_ref.extractall(extract_dir) expands every entry without resource limits.
  6. The archive consumes available disk space, I/O capacity, memory, or processing time.
  7. Resume processing or other work ...[truncated 730 chars]
Remediation
View remediation

Remediation Suggestions

Replace unrestricted extractall() usage with validation followed by controlled, per-entry extraction.

  1. Inspect all entries with ZipFile.infolist() before extracting anything.
  2. Enforce conservative limits on:
    • Total entry count
    • Total declared uncompressed size
    • Individual entry size
    • Compression ratio
    • Path depth
  3. Reject encrypted entries, unsupported file types, symbolic links, and suspicious metadata.
  4. Permit only expected resume extensions such as .pdf, .doc, and .docx.
  5. Resolve each destination path and verify that it remains under the intended extraction directory.
  6. Stream each accepted entry with a byte limit instead of using extractall().
  7. Apply filesystem quotas, execution timeouts, and process resource limits.
  8. Ensure temporary data is removed through a try/finally block even when extraction or parsing fails.

Example validation pattern:

python
MAX_FILES = 200
MAX_FILE_SIZE = 20 * 1024 * 1024
MAX_TOTAL_SIZE = 500 * 1024 * 1024
ALLOWED_EXTENSIONS = {".pdf", ".doc", ".docx"}

entries = zip_ref.infolist()
if len(entries) > MAX_FILES:
    raise ValueError("Archive contains too many files")

total_size = 0
for entry in entries:
    total_size += entry.file_size
    if entry.file_size > MAX_FILE_SIZE:
        raise ValueError("Archive entry exceeds the size limit")
    if total_size > MAX_TOTAL_SIZE:
        raise ValueError("Archive exceeds the total size limit")

These checks should be combined with canonical path validation and bounded streaming extraction.

T08 · Insecure Dependencies

Note
Location
README.md:135
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: README.md, lines 135-138
Vulnerability Type: Unpinned dependency and missing integrity verification
Risk Level: Low

Vulnerable Documentation

markdown
Install dependencies:
```bash
pip install pdfplumber
text

### Technical Analysis

The documented installation command retrieves the latest version of `pdfplumber` and its transitive dependencies without version constraints, a lock file, or cryptographic hash verification.

Consequently, two installations performed at different times may resolve different code. The project has no reviewed dependency baseline against which users can verify the installed package set. The risk becomes security-relevant if a future package or transitive dependency release is compromised, malicious, or contains a newly introduced vulnerability.

The audit did not identify a typographically deceptive package name or an explicitly untrusted package repository. This finding concerns mutable, unverified dependency resolution rather than evidence that the current `pdfplumber` package is malicious.

### Attack Path

1. A user follows the installation instructions in `README.md`.
2. `pip` queries the configured package index and resolves the latest available `pdfplumber` release and transitive dependencies.
3. A compromised, malicious, or unexpectedly vulnerable release is selected because no reviewed versions or hashes are specified.
4. The package is installed into the user's Python environment.
5. Dependency code may run during package build or later when the resume extraction script imports and uses `pdfplumber`.
6. Any resulting behavior occurs with the privileges of the user running installation or resume processing.

### Impact Assessment

The potential impact depends on the behavior of a compromised dependency and the privileges used for installation or execution. It could include:

- Execution of dependency-controlled Pyth
...[truncated 428 chars]
Remediation
View remediation

Remediation Suggestions

  1. Define direct and transitive dependencies in a reviewed lock file.
  2. Pin exact versions rather than allowing unconstrained latest-version resolution.
  3. Record and verify cryptographic hashes for every distributed artifact.
  4. Install dependencies in an isolated virtual environment under a non-privileged account.
  5. Use a trusted package index or an internally controlled package mirror.
  6. Add automated dependency vulnerability and update monitoring.
  7. Review and test dependency upgrades before modifying the lock file.

A hardened installation workflow can use a hash-locked requirements file:

bash
python -m pip install --require-hashes -r requirements.txt

The corresponding requirements.txt should contain exact, reviewed versions and SHA-256 hashes for pdfplumber and all resolved transitive dependencies.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill is presented as a resume screening/evaluation workflow, but the behavior described by analysis includes archive extraction and document text conversion rather than the promised screening engine. This mismatch is dangerous because users may consent to hiring analysis while the skill actually performs additional data-handling actions on uploaded files, creating privacy and trust risks around bulk candidate documents.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file contains imperative Chinese-language instructions that constrain how the skill should operate, and the rest of the template heavily assumes Chinese-language outputs and labels. There is no indication that the user can opt into another language or locale, which creates a language-policy issue for a general-purpose skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
90% confidence
Finding

The report templates, column headers, ratings, and summaries are all specified in Chinese, effectively forcing a Chinese locale for generated outputs. Because the skill description does not document this as a region-specific tool or provide a language choice, it violates the requirement not to impose a language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill processes highly sensitive personal data from resumes and explicitly describes extracting text, saving it to .txt files, and producing persisted evaluation outputs, but it provides no privacy, retention, access-control, or consent guidance. In a resume-screening context this increases the risk of unauthorized disclosure, over-retention, and mishandling of PII, which can create legal, compliance, and security exposure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill instructs the agent to extract files, read resumes, and write derived text/JSON/report files, but it declares no explicit tool scope or permissions. In an agent environment, missing scope boundaries can let a skill invoke file operations without clear user-visible authorization constraints, increasing the chance of over-broad access to sensitive candidate data.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest description says to invoke the skill when the user asks to 'batch screen resumes' or 'evaluate multiple candidates against multiple job requirements' without clarifying boundaries or exclusions. These phrases are broad enough to match ordinary recruiting-assistance requests, which could cause the skill to trigger when the user did not specifically want this structured batch-screening workflow.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The usage guidance lists triggers like 'evaluate multiple candidates' and 'screen against multiple job positions' but does not specify minimum conditions, exclusions, or non-trigger cases. Because this file is markdown, that ambiguity matters: the skill may be invoked for many commonplace HR tasks that only partially resemble the intended batch workflow.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill directs extraction of resume contents and saving each resume as separate text files without warning the user that uploaded personal data will be transformed and persisted. Because resumes contain sensitive PII, silent generation of derivative files materially increases privacy exposure, retention risk, and the blast radius of any later file access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The workflow performs large-scale candidate evaluation and generates JSON, Markdown, tabular, and comparison reports, but it omits a privacy warning for bulk handling of personal and potentially sensitive employment data. Aggregated reports concentrate PII from many candidates into new artifacts, making accidental disclosure, over-retention, and misuse significantly more damaging.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code extracts an entire ZIP archive into a temporary directory, which is a file-write operation affecting the user's filesystem. While it prints that extraction is happening, there is no warning in comments, docstrings, or usage text about unpacking archive contents onto disk, so the side effect is not clearly disclosed to the user.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script saves extracted PDF text into .txt files, which persists potentially sensitive personal data from resumes onto disk. Although it logs the save path after writing, the code does not provide an upfront warning in its usage text or documentation that resume contents will be stored as plaintext files.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The natural-language instructions and examples are presented entirely in Chinese, and there is no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking context. This can violate language/locale policy when a skill implicitly forces one language without opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.