Back to skill

Security audit

yindeng_ analyse

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Yindeng crawler, but it needs review because optional AI analysis can send financial document text and an API key to any configured LLM endpoint.

Install only if you are comfortable with the skill downloading public Yindeng PDFs, retaining local financial-report files, and sending analyzed PDF text to the configured LLM provider. Keep analyze disabled for confidential workloads unless the endpoint is trusted, restrict LLM_API_BASE to approved HTTPS provider domains, use a scoped API key, and pin dependencies before production use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
main.py:80
Finding

Arbitrary LLM Endpoint Can Receive API Credentials and Financial Document Content

Content
View full analysis

Vulnerability Details

File Location: main.py, lines 80–133
Vulnerability Type: Unvalidated external service endpoint and sensitive-data disclosure
Risk Level: High

Vulnerable Code

python
api_key = os.getenv("LLM_API_KEY")
api_base = os.getenv("LLM_API_BASE", "https://api.deepseek.com/v1")
model_name = os.getenv("LLM_MODEL", "deepseek-chat")
provider = os.getenv("LLM_PROVIDER", "deepseek").lower()

# Provider specific default configurations
if provider == "siliconflow" or provider == "硅基流动":
    # SiliconFlow default
    if not os.getenv("LLM_API_BASE"):
        api_base = "https://api.siliconflow.cn/v1"
    if not os.getenv("LLM_MODEL"):
        model_name = "deepseek-ai/DeepSeek-V3" # Common model on SiliconFlow
elif provider == "qwen" or provider == "dashscope" or provider == "通义千问":
    # DashScope (Aliyun) default - usually requires compatible OpenAI client or specific URL
    # Here we assume using OpenAI-compatible endpoint if available, or user sets BASE_URL
    if not os.getenv("LLM_API_BASE"):
        api_base = "https://dashscope.aliyuncs.com/compatible-mode/v1"
    if not os.getenv("LLM_MODEL"):
        model_name = "qwen-plus"

if not api_key:
    log_cb("Skipping LLM analysis: LLM_API_KEY environment variable not set.")
    return None

headers = {
    "Authorization": f"Bearer {api_key}",
    "Content-Type": "application/json"
}

data = {
    "model": model_name,
    "messages": [
        {"role": "system", "content": "You are a helpful assistant that extracts data into JSON."},
        {"role": "user", "content": prompt}
    ],
    "temperature": 0.1,
    "response_format": {"type": "json_object"}
}

try:
    response = requests.post(f"{api_base}/chat/completions", headers=headers, json=data, timeout=120)
    response.raise_for_status()
    result = response.json()
    content = result['choices'][0]['message']['content']

The prompt transmitted in this request includes the following raw document excerpt:

`` ...[truncated 2197 chars]

Remediation
View remediation

Remediation Suggestions

  1. Allowlist the documented provider endpoints and bind each endpoint to its corresponding provider:
    • api.deepseek.com
    • api.siliconflow.cn
    • dashscope.aliyuncs.com
  2. Parse the endpoint with a standard URL parser and reject:
    • Schemes other than HTTPS
    • Embedded usernames or passwords
    • Unexpected ports
    • IP-literal destinations
    • Loopback, link-local, private, and reserved network addresses
    • Hosts not associated with the selected provider
  3. Require a separate, explicit unsafe opt-in before permitting custom endpoints. Display the destination and explain that credentials and document contents will be transmitted.
  4. Associate credentials with specific providers instead of using one generic bearer token for arbitrary destinations.
  5. Redact or minimize document content before transmission. Extract only necessary fields locally where feasible.
  6. Add an explicit data-processing disclosure identifying what content is sent, the destination provider, and applicable retention implications.
  7. Avoid logging response bodies because provider responses may reproduce sensitive document content.
  8. Add automated tests confirming that hostile schemes, private addresses, redirects to untrusted hosts, and unapproved domains are rejected.

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Third-Party Dependencies Create Supply-Chain and Reproducibility Risk

Content
View full analysis

Vulnerability Details

File Location: requirements.txt, lines 1–13
Vulnerability Type: Unpinned dependency versions
Risk Level: Medium

Vulnerable Code

text
requests
beautifulsoup4
lxml
openpyxl
pandas
pytesseract
pdf2image
Pillow
opencv-python
numpy
PyPDF2
xlsxwriter
pdfminer.six

Technical Analysis

All third-party dependencies are specified without exact versions or integrity hashes. Consequently, each installation resolves whatever package versions are currently available through the configured Python package index.

This makes builds non-reproducible and prevents the project from guaranteeing that installed code matches versions reviewed during the audit. A future compromised, malicious, or incompatible package release could be installed automatically. The risk is amplified because these dependencies process complex, remotely obtained PDF, HTML, image, and spreadsheet content.

The audit did not identify a specific malicious or typosquatted dependency in the current list. The confirmed issue is the absence of version and integrity controls, not evidence that any listed package is currently compromised.

Attack Path

  1. A deployment runs pip install -r requirements.txt as documented.
  2. Pip resolves the latest versions from the active package index or an environment-configured mirror.
  3. A dependency release or package-index account is compromised, or an untrusted mirror serves altered artifacts.
  4. The installer retrieves the affected package because no reviewed version or artifact hash is enforced.
  5. Malicious package installation or import-time code executes with the privileges of the installation or Skill process.

Impact Assessment

The resulting privileges depend on how the Skill is installed and executed. A compromised dependency could access files, environment variables such as LLM_API_KEY, downloaded financial documents, generated reports, and network resources available to the Python proces ...[truncated 290 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to an exact, reviewed version.
  2. Generate and commit a lock file that includes transitive dependencies.
  3. Enforce artifact hashes, for example with pip's --require-hashes workflow.
  4. Install only from a trusted, explicitly configured package index and disable unintended fallback indexes.
  5. Run dependency vulnerability and license scanning in continuous integration.
  6. Review and update dependencies through controlled pull requests rather than resolving mutable latest versions during deployment.
  7. Build dependencies in an isolated environment and run the Skill under a non-privileged account with restricted filesystem and network access.
  8. Revalidate PDF, image, HTML, and spreadsheet parsing libraries regularly because they process attacker-influenced remote content.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (37)

Tainted flow: 'headers' from os.getenv (line 105, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · main.py (reported line 121)May include surrounding context.

python
}
    
    try:
        response = requests.post(f"{api_base}/chat/completions", headers=headers, json=data, timeout=120)
        response.raise_for_status()
        result = response.json()
        content = result['choices'][0]['message']['content']

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code creates a dated download directory and prepares Excel output paths, then later saves downloaded PDFs and crawl records there. Although the writes are part of the crawler's functionality, this file does not include a docstring, comment, or explicit user-facing disclosure near the file-writing behavior to warn that local files will be created and populated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

For result announcements, the code processes downloaded PDFs with OCR and stores recognition results in JSON and a combined Excel workbook. This can capture and retain document contents, but the file contains no explicit disclosure that downloaded materials will be analyzed and their extracted data saved locally.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill extracts text from local PDFs and sends it to a third-party LLM service for analysis, but there is no manifest justification, trust boundary declaration, or explicit security rationale for transmitting potentially sensitive financial-document contents off-host. In this context, the PDFs appear to contain loan, borrower, litigation, and pricing data, so unauthorized external disclosure could expose confidential business or personal information.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Natural-language content in the prompt and output schema is fixed to Chinese, which imposes a specific language/locale behavior on users. The file does not offer an opt-in language selection or explain that the tool is intentionally restricted to a Chinese-language regional workflow.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The prompt embeds up to 15,000 characters of raw PDF text directly into a remote LLM request, which can disclose sensitive financial and possibly personal data in plain form to a third-party service. Because the extracted fields include borrower counts, balances, litigation, guarantees, and pricing, the skill context makes the exposure more serious than generic document processing.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 65)May include surrounding context.

md
# Get API configuration from environment variables
    api_key = os.getenv("LLM_API_KEY")
    api_base = os.getenv("LLM_API_BASE", "https://api.deepseek.com/v1")
    model_name = os.getenv("LLM_MODEL", "deepseek-chat")
    provider = os.getenv("LLM_PROVIDER", "deepseek").lower()

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · main.py (reported line 82)May include surrounding context.

python
# Get API configuration from environment variables
    api_key = os.getenv("LLM_API_KEY")
    api_base = os.getenv("LLM_API_BASE", "https://api.deepseek.com/v1")
    model_name = os.getenv("LLM_MODEL", "deepseek-chat")
    provider = os.getenv("LLM_PROVIDER", "deepseek").lower()

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 73)May include surrounding context.

md
if provider == "siliconflow" or provider == "硅基流动":
        # SiliconFlow default
        if not os.getenv("LLM_API_BASE"):
            api_base = "https://api.siliconflow.cn/v1"
        if not os.getenv("LLM_MODEL"):
            model_name = "deepseek-ai/DeepSeek-V3" # Common model on SiliconFlow
    elif provider == "qwen" or provider == "dashscope" or provider == "通义千问":

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · main.py (reported line 90)May include surrounding context.

python
if provider == "siliconflow" or provider == "硅基流动":
        # SiliconFlow default
        if not os.getenv("LLM_API_BASE"):
            api_base = "https://api.siliconflow.cn/v1"
        if not os.getenv("LLM_MODEL"):
            model_name = "deepseek-ai/DeepSeek-V3" # Common model on SiliconFlow
    elif provider == "qwen" or provider == "dashscope" or provider == "通义千问":

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This code performs an external network transmission of document-derived content to an LLM API. External transmission is not inherently malicious, but in this skill it carries real confidentiality risk because locally extracted PDF contents are sent to a third party and the destination can be influenced by environment configuration.

Content

Scanner excerpt · main.py (reported line 121)May include surrounding context.

python
}
    
    try:
        response = requests.post(f"{api_base}/chat/completions", headers=headers, json=data, timeout=120)
        response.raise_for_status()
        result = response.json()
        content = result['choices'][0]['message']['content']

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code transmits extracted PDF text to an external API without any user-facing warning at the point of use, despite processing financial announcement documents that may include sensitive borrower and transaction information. Users or operators could reasonably assume analysis is local, leading to accidental privacy breaches and compliance violations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The comment at L088 states that recognition of sensitive identity-related fields has been removed, suggesting reduced extraction scope. However, the surrounding code still parses named counterparties, agreement dates, project identifiers, and debt amounts, so the documentation implies a narrower extraction behavior than the implementation actually performs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill persistently writes OCR-extracted document data to local Excel/CSV files, which may contain sensitive financial and organizational information. In environments where users do not expect local retention, this can create unintended data exposure through shared folders, backups, or later unauthorized access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation describes how to run the skill but does not clearly warn users that execution will download PDFs and create local directories and Excel files. This is a real transparency/safety issue because agents or users may trigger filesystem writes without informed consent, which can be problematic in restricted or sensitive environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The OCR call forces lang='chi_sim', which constrains processing to a specific language/locale. Under the policy, locale-specific behavior should either be user-selectable or clearly justified as region-specific; this file provides no such opt-in or justification in natural-language comments or interface.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
99% confidence
Finding

The dependency manifest leaves requests unpinned, so builds may pull different versions over time, including versions with known security defects. This weakens supply-chain integrity and makes it impossible to verify whether a safe release is consistently installed.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests
beautifulsoup4
lxml
openpyxl

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
98% confidence
Finding

requests has multiple published advisories, but the manifest does not pin a version, so there is no way to verify whether installations avoid affected releases. Because requests often handles outbound HTTP, a vulnerable version could leak credentials or weaken transport/security behavior depending on usage.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
99% confidence
Finding

beautifulsoup4 is unpinned, allowing non-reproducible installs and increasing exposure to future or existing vulnerable releases. Even if the current environment is safe, the manifest does not guarantee that outcome elsewhere.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests
beautifulsoup4
lxml
openpyxl
pandas

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
99% confidence
Finding

Unpinned lxml is risky because parsers frequently receive security fixes for memory corruption, XXE, and sanitization issues. Without a version constraint, deployments may resolve to affected builds and make the attack surface unpredictable.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
requests
beautifulsoup4
lxml
openpyxl
pandas
pytesseract

Unverifiable Dependency: lxml has 14 known advisory(ies) (CVE-2021-43818 (lxml's HTML Cleaner allows crafted and SVG embedded scripts to pass through); CVE-2014-3146 (lxml Cross-site Scripting Via Control Characters); CVE-2021-28957 (lxml vulnerable to Cross-Site Scripting ) +11 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
99% confidence
Finding

lxml has numerous historical security issues and is used for parsing structured content, making version uncertainty more dangerous than a generic utility package. If this skill processes untrusted HTML/XML, an affected release could expose XSS-like sanitization bypasses, XXE-style issues, or parser-level memory problems.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

openpyxl is unpinned, which permits drift across environments and may introduce vulnerable versions handling spreadsheet/XML content. For packages that process complex file formats, version control is important to avoid regressions and known parser flaws.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
requests
beautifulsoup4
lxml
openpyxl
pandas
pytesseract
pdf2image

Unverifiable Dependency: openpyxl has 2 known advisory(ies) (CVE-2017-5992 (Improper Restriction of XML External Entity Reference in Openpyxl); CVE-2017-5992 (Openpyxl 2.4.1 resolves external entities by default, which allows remote attack)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
97% confidence
Finding

openpyxl has known XML entity-related advisories, and this skill appears to include document-processing dependencies, which makes spreadsheet parsing a realistic attack path. Without a version pin, the environment could install an affected release and expose the application to crafted workbook attacks.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

pandas is listed without a version, so installations are not reproducible and may include insecure or behaviorally incompatible releases. This is a supply-chain hygiene issue that can complicate patch verification and incident response.

Content

Scanner excerpt · requirements.txt (reported line 5)May include surrounding context.

text
beautifulsoup4
lxml
openpyxl
pandas
pytesseract
pdf2image
Pillow

Unverifiable Dependency: pandas has 1 known advisory(ies) (CVE-2020-13091 (** DISPUTED ** pandas through 1.0.3 can unserialize and execute commands from an)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

pandas has at least one disputed advisory, and the lack of version pinning means the installed package cannot be verified against that risk. The direct danger is less certain than parser libraries here, but the unverifiable dependency state is still a security weakness.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.