Back to skill

Security audit

JD + 简历 → 面试题预测助手

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its interview-prep purpose, but it sends résumé and job-description content to a third-party LLM endpoint without clear user-facing disclosure or consent.

Review before installing. Use only with résumés and job descriptions you are comfortable sending to the configured LLM provider, consider redacting contact details and other unnecessary personal data, verify OPENAI_API_BASE points to a trusted HTTPS endpoint, and install parsing dependencies in a dedicated virtual environment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

other

Warning
Location
generate_questions.py:53
Finding

Undisclosed Transmission of Résumé Data to a Third-Party API

Content
View full analysis

Vulnerability Details

File Location: generate_questions.py, lines 17-18 and 53-75
Vulnerability Type: Undisclosed sensitive data transmission
Risk Level: Medium

Vulnerable Code

python
API_BASE = os.environ.get("OPENAI_API_BASE", "https://api.deepseek.com")
MODEL = os.environ.get("LLM_MODEL", "deepseek-chat")
python
def call_llm(prompt: str) -> str:
    if not API_KEY:
        return "[Error] Set the OPENAI_API_KEY or DEEPSEEK_API_KEY environment variable"

    payload = json.dumps({
        "model": MODEL,
        "messages": [
            {"role": "system", "content": "You are an experienced HR and career consultant who specializes in analyzing job requirements and résumés and predicting interview questions. Answer in Chinese."},
            {"role": "user", "content": prompt}
        ],
        "temperature": 0.7,
    }).encode("utf-8")

    req = urllib.request.Request(
        f"{API_BASE}/chat/completions",
        data=payload,
        headers={
            "Content-Type": "application/json",
            "Authorization": f"Bearer {API_KEY}",
        }
    )
    with urllib.request.urlopen(req, timeout=60) as r:
        data = json.loads(r.read())
    return data["choices"][0]["message"]["content"]

The quoted English system-message text above is a direct translation for report-language compliance; the executable source contains the equivalent Chinese string.

Technical Analysis

The application reads job descriptions and résumés, embeds up to 3,000 characters from each document into a prompt, and sends that prompt to an OpenAI-compatible endpoint. The default endpoint is DeepSeek, while OPENAI_API_BASE permits the destination to be changed.

Résumés commonly contain names, contact information, employment history, education details, and other personal or confidential information. The project documentation does not clearly disclose that this in ...[truncated 1757 chars]

Remediation
View remediation

Remediation Suggestions

  1. Clearly disclose before processing that document content will be transmitted to an external LLM provider.
  2. Identify the default provider, applicable privacy terms, expected retention behavior, and the exact categories of data transmitted.
  3. Require explicit user consent before sending résumé or job-description content.
  4. Provide a local-processing or redaction mode for users who cannot transmit sensitive data externally.
  5. Minimize transmitted data and automatically redact email addresses, telephone numbers, physical addresses, government identifiers, and other unnecessary personal information.
  6. Validate OPENAI_API_BASE against an explicit allowlist of approved HTTPS origins.
  7. Reject non-HTTPS URLs, embedded credentials, unexpected ports, redirects to unapproved hosts, and malformed endpoint values.
  8. Use provider credentials with minimal scope, restricted quotas, and isolated billing limits.
  9. Avoid logging prompts, authorization headers, or API responses containing personal information.
  10. Add documentation covering data flow, consent, retention, deletion, and incident-response procedures.

T08 · Insecure Dependencies

Note
Location
SKILL.md:106
Finding

Unpinned Runtime Dependency Installation Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 106-107
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Low

Vulnerable Code

text
- PDF parsing requires `pdfplumber`: `pip install pdfplumber`
- DOCX parsing requires `python-docx`: `pip install python-docx`

Technical Analysis

The setup instructions install third-party packages without fixed versions, integrity hashes, or a dependency lockfile. As a result, installation resolves whichever package versions and transitive dependencies are available from the configured Python package index at installation time.

This prevents reproducible installation and means that the code executed in future environments may differ from the dependency versions originally reviewed. Python packages can execute build or installation logic, and imported dependencies execute with the privileges of the invoking Python process. A compromised upstream release, compromised transitive dependency, unsafe package-index configuration, or unexpected future version can therefore introduce unreviewed code.

No evidence was found that pdfplumber or python-docx is currently malicious. The finding concerns the unsafe dependency-management practice rather than a confirmed compromise of either package.

Attack Path

  1. A user follows the documented installation command.
  2. pip queries the configured package index and resolves the latest compatible package and transitive dependency versions.
  3. The downloaded artifacts are accepted without project-specified hashes.
  4. Package build or installation code runs, or the installed package is subsequently imported by parse_file.py.
  5. If a resolved package or transitive dependency has been compromised, its code executes under the account running the installation or parser.

Impact Assessment

Exploitation requires a compromised dependency, package source, or dependency-resolution environment. If th ...[truncated 495 chars]

Remediation
View remediation

Remediation Suggestions

  1. Create a reviewed dependency manifest containing exact versions for direct and transitive dependencies.
  2. Generate and verify cryptographic hashes for all accepted distribution artifacts.
  3. Install with a command such as pip install --require-hashes -r requirements.txt.
  4. Use a lockfile or constraints file to ensure reproducible dependency resolution.
  5. Prefer an approved internal package mirror or explicitly trusted package index.
  6. Audit dependencies regularly for known vulnerabilities and unexpected ownership or release changes.
  7. Test dependency upgrades in an isolated environment before updating pinned versions.
  8. Recommend installation in a dedicated virtual environment without administrator or root privileges.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (8)

Tainted flow: 'req' from os.environ.get (line 49, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · generate_questions.py (reported line 57)May include surrounding context.

python
"Authorization": f"Bearer {API_KEY}",
        }
    )
    with urllib.request.urlopen(req, timeout=60) as r:
        data = json.loads(r.read())
    return data["choices"][0]["message"]["content"]

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code chunk is a generic file-text extraction utility. It reads documents from disk and converts them to plain text using pdfplumber/pypdf, python-docx, or direct text reading. While file parsing could be a supporting component of the described skill, the declared purpose presents a much broader AI interview-prep workflow. In this code, none of the core advertised behaviors are implemented: there is no model inference, no question categorization, no STAR guidance generation, no semantic comparison between JD and resume, and no export step. Therefore the actual behavior is materially narrower and different from the declared description.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises executable commands that invoke local Python scripts on user-supplied file paths, but it does not declare any tool restrictions or permissions. That makes the skill's actual capabilities broader and less governed than the manifest suggests, increasing the risk of unintended shell, file, environment, or network access during use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly encourages users to upload or paste resumes and job descriptions, which commonly contain sensitive personal data such as names, contact details, employment history, and sometimes compensation or location information, but it provides no privacy, retention, redaction, or handling warning. In combination with file parsing and report generation, this increases the chance of over-collection, accidental disclosure, or unsafe downstream storage of personal information.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · generate_questions.py (reported line 29)May include surrounding context.

python
script_dir = os.path.dirname(os.path.abspath(__file__))
        parser_path = os.path.join(script_dir, "parse_file.py")
        import subprocess
        result = subprocess.run(
            [sys.executable, parser_path, text_or_path],
            capture_output=True, text=True
        )

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill sends the full JD and resume content to an external LLM service, which likely includes sensitive personal data and possibly confidential job information, without any explicit consent prompt, privacy notice, or redaction step. In this skill's context, exfiltration risk is elevated because resumes routinely contain PII such as names, phone numbers, email addresses, employment history, and other personal details.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The system prompt explicitly instructs the model to answer in Chinese, imposing a fixed language behavior. The file does not provide a user option to select language or indicate that this locale restriction is optional or region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module docstring and all user-facing output strings are written in Chinese, including usage and error messages. For a general-purpose file parsing script with no documented region- or locale-specific constraint, this imposes a specific language on users without opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.