Back to skill

Security audit

简历评估

Security checks for vulnerabilities and agentic risk

Overview

This resume-evaluation skill appears purpose-aligned, but it can send sensitive candidate and company data plus an API bearer token to a configurable external model endpoint without strong endpoint controls or a clear user-facing privacy gate.

Install only if you are comfortable sending candidate resumes, optional job descriptions, and company background to the configured model provider. Before use, lock the model endpoint to an approved HTTPS provider, bind it to the intended API key, avoid untrusted replacement config files, and ensure recruiters treat the generated scores as advisory human-review material only.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/llm_client.py:37
Finding

Unrestricted Model Endpoint Can Receive API Credentials and Sensitive Recruitment Data

Content
View full analysis
str: return f"{self.base_url.rstrip('/')}/{self.chat_completions_path.lstrip('/')}" ``` ```python user_payload = { "candidate_id": candidate_id, "candidate_id_instruction": "Return exactly this candidate_id in the top-level candidate_id field.", "output_language": evaluation_config.get("output_language"), "role_family_override": evaluation_config.get("role_family_override"), "seniority_override": evaluation_config.get("seniority_override"), "scoring_weights": evaluation_config.get("weights"), "enterprise_hiring_calibration": [ "Score for enterprise recruiting judgment, not academic ranking or resume polish alone.", "Education is important, and papers or competitions can support technical depth, but they should not outweigh concrete project delivery evidence when the project is technically plausible and has visible real-world effects.", "For projects, weigh scenario constraints, personal ownership, implementation details, production/pilot/customer delivery status, measurable impact, scale, cost, quality, efficiency, reliability, and post-launch iteration.", "When no JD is provided, still distinguish landed enterprise projects, pilot delivery, PoC, academic research, coursework, and vague claims; use holistic judgment without rigid caps.", "When JD is provided, prioritize JD-relevant landed projects and technology fit over impressive but unrelated academic or competition signals.", "Do not infer missing facts; use deductions and interview validation points when delivery status, personal contribution, or metrics a ...[truncated 2618 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/llm_client.py:57
Finding

Untrusted Resume Content Is Not Isolated from Model Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (46)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The code chunk is not implementing the end-user skill behavior described; instead, it tests lower-level infrastructure used to call an LLM service. Some message-building content aligns with resume/JD/company evaluation, so it appears related to the broader system. However, the declared purpose emphasizes batch resume evaluation and Markdown/CSV reporting, while this code specifically exercises model config loading, request formatting, retries, malformed JSON handling, and secret redaction. Those are materially different from the declared primary functionality, and the described report-generation capability is absent from this chunk. Therefore this chunk does not accurately represent the declared purpose on its own.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description promises an end-to-end recruiter-oriented resume review capability, including batch evaluation, quality checks, optional JD/company matching, and report generation. The actual code chunk only tests parsing-related helper functions: file discovery, text extraction from multiple document formats, candidate ID generation, whitespace/BOM normalization, and error handling. These are supporting ingestion details, but the chunk does not implement or demonstrate the core declared behaviors. Because the observed code’s primary purpose is materially narrower and different from the declared purpose, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

This skill explicitly forces Chinese output without user opt-in, which can override recruiter expectations and alter the usability of evaluation results. In hiring and screening contexts, this is more dangerous because reviewers may miss risk signals, scoring rationale, or disqualifying evidence if reports are generated in an unexpected language, creating process integrity, fairness, and compliance risks.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares executable behavior involving environment variables, file access, network access, and shell execution, but it provides no explicit tool-scope or permission constraints. In an agent setting, that weakens least-privilege controls and can allow broader-than-necessary access to local resumes, API secrets, and command execution if the skill is invoked in an unsafe runtime.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The display name, description, and default prompt are all specified only in Chinese, and the file does not indicate that language selection is optional or that the skill is intentionally limited to a Chinese-speaking context. This can violate a language/locale policy when users are not given an explicit opt-in or alternative locale.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The configuration hard-codes output_language: "zh" for all evaluations without any visible user-triggered condition, scope limitation, or fallback behavior. In an HR resume evaluation workflow, forcing output into a language the recruiter did not request can mislead reviewers, reduce auditability, and cause downstream screening or reporting errors, especially when evidence-backed reports and CSV outputs are expected to be consumed consistently.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The markdown explicitly summarizes and previews output_language: "zh", which indicates a fixed language setting. This is a natural-language policy concern because it enforces a specific locale without any visible user choice, opt-in, or justification in the file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script sends full resume text, and optionally job description and company context, to an external model via evaluate_with_model without any consent, warning, minimization, or privacy gate in this workflow. Because resumes routinely contain sensitive personal data, this creates a real data-exposure risk and is especially concerning in an HR skill where processing PII is central to normal operation.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This code transmits highly sensitive resume-evaluation inputs to an externally configured LLM endpoint, including full resume text, JD text, company context, and a bearer API token. Because the destination is controlled by configuration and there is no allowlist, scheme validation, or data-minimization control, the skill can exfiltrate candidate PII and employer-sensitive data to an untrusted service, which is especially risky in an HR workflow handling regulated personal information.

Content

Scanner excerpt · scripts/llm_client.py (reported line 197)May include surrounding context.

python
for attempt in range(max_attempts):
        attempts = attempt + 1
        try:
            response = requests.post(
                config.endpoint,
                headers=headers,
                json=payload,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This file is a markdown file, so SQP-2 applies under the markdown criteria. The summary states that the skill uses requests and includes functions for model evaluation plus an api_key lookup, but the description does not disclose that resume or system data may be sent to an external model endpoint or that credentials are required.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Line L72 writes output_language: zh-CN into the evaluation config, and the test later asserts Chinese output. This natural-language constraint forces a specific locale without any indication of user opt-in or a documented region-specific justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This test hard-codes assertions for Chinese strings in the generated report, which indicates the skill may force a specific output language/locale. Under the policy, locale constraints should either be user-selectable or clearly documented and justified; no such opt-in or justification appears in this file.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_evaluate_resumes.py (reported line 499)May include surrounding context.

python
def test_direct_script_help_runs():
    result = subprocess.run(
        [sys.executable, "scripts/evaluate_resumes.py", "--help"],
        cwd=PROJECT_ROOT,
        capture_output=True,

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The test data repeatedly sets evaluation_config={"output_language": "zh"}, which reflects a natural-language constraint forcing Chinese output. Under the policy, language-specific behavior should offer user choice or be clearly justified; this file shows a fixed locale with no opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This test again asserts behavior using output_language set to zh, reinforcing a skill design that appears to force a specific language. Because the file contains no indication that users can opt in or choose another language, it meets the natural-language policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The repeated use of output_language: zh across tests indicates the underlying skill likely enforces Chinese output by default. A forced locale without user choice or documented justification is a policy concern even when it appears in code literals and test fixtures.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · tests/test_llm_client.py (reported line 171)May include surrounding context.

python
return FakeResponse()

    monkeypatch.setattr("scripts.llm_client.requests.post", fake_post)
    config = _test_config(max_retries=0)

    result = evaluate_with_model(config, [{"role": "user", "content": "Return JSON"}])

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · tests/test_llm_client.py (reported line 373)May include surrounding context.

python
return FakeResponse()

    monkeypatch.setattr("scripts.llm_client.requests.post", fake_post)
    config = _test_config(max_retries=0)

    result = evaluate_with_model(config, [{"role": "user", "content": "Return JSON"}])

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · tests/test_llm_client.py (reported line 178)May include surrounding context.

python
assert result == {"ok": True}
    assert calls == [
        {
            "url": "https://api.example.com/chat/completions",
            "headers": {
                "Authorization": "Bearer secret-value",
                "Content-Type": "application/json",

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This test file encodes expected report content in Chinese strings such as section labels, evidence, deductions, and recommendation text. That implies the skill's rendered reports are forced into a specific language/locale with no visible opt-in or justification in this file, which matches the policy category for language or locale constraints.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The test explicitly requires Chinese output tokens like "总分" and "不得作为筛选依据" to be present. Enforcing one language in expected output suggests the underlying skill does not provide a user language choice, which is a natural-language policy concern unless clearly justified as locale-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The assertion for "批量简历评估汇总" shows the summary report is expected to use Chinese headings. A fixed language expectation without any visible opt-in or documented locale restriction can violate organizational language/locale policy requirements.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

This markdown file includes the linked folder index name "资料索引" in Chinese, which can indicate a fixed language choice in user-facing material. The file does not offer an alternative language or explain that the content is intentionally region- or locale-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The file includes specific Chinese decision-label terms alongside English instructions, which imposes multilingual policy content without stating whether the skill is locale-specific or user-selected. Under the language/locale policy rule, natural-language requirements should either offer opt-in choice or clearly justify the locale constraint.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: PyYAML has 8 known advisory(ies) (CVE-2019-20477 (Deserialization of Untrusted Data in PyYAML); CVE-2020-1747 (Improper Input Validation in PyYAML); CVE-2020-14343 (Improper Input Validation in PyYAML) +5 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.