Back to skill

Security audit

Education Filing Form Ocr

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it says, but it uploads sensitive education documents and an API token to a configurable OCR endpoint without validating the destination.

Install only if you are comfortable sending the selected education documents to Scnet's OCR service. Keep config/.env private, do not paste the API key into chat, and avoid changing SCNET_API_BASE unless you fully trust and control the destination. Prefer a pinned dependency install or isolated virtual environment before using it with sensitive records.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:78
Finding

Configurable OCR Endpoint Can Exfiltrate API Credentials and Sensitive Documents

Content
View full analysis

Vulnerability Details

File Location: scripts/main.py:78-112
Vulnerability Type: Unrestricted destination for sensitive network transmission
Risk Level: High

Vulnerable Code

python
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
return config
python
api_base = config['SCNET_API_BASE']
api_key = config['SCNET_API_KEY']
url = f"{api_base}/ocr/recognize"

if not os.path.isfile(file_path):
    sys.exit(f"File does not exist: {file_path}")

mime_type, _ = mimetypes.guess_type(file_path)
if mime_type is None:
    mime_type = 'application/octet-stream'

headers = {
    'Authorization': f'Bearer {api_key}'
}

try:
    with open(file_path, 'rb') as f:
        files = {
            'file': (os.path.basename(file_path), f, mime_type)
        }
        data = {
            'ocrType': ocr_type,
            'channelTag': "scnetSkills"
        }
        response = requests.post(
            url,
            headers=headers,
            data=data,
            files=files,
            timeout=60
        )
except Exception as e:
    sys.exit(f"Network request failed: {str(e)}")

Technical Analysis

The destination of the OCR request is derived directly from the configurable SCNET_API_BASE value. The implementation does not validate the URL scheme, destination hostname, port, or resulting request path before attaching the bearer credential and uploading the selected document.

Although uploading the document to SCNet is necessary for the declared cloud OCR functionality, permitting an arbitrary destination is not required. Anyone able to modify config/.env can redirect the request to an attacker-controlled HTTP or HTTPS server. The request then discloses both:

  • The SCNET_API_KEY bearer credential in the Authorization header.
  • The complete user-selected education document in the multipart request body.

Education filing ...[truncated 1754 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove SCNET_API_BASE configurability if only the official SCNet service is supported.
  2. If endpoint configuration is operationally required, parse the URL and enforce:
    • The https scheme.
    • An explicit allowlist of trusted hostnames, preferably only api.scnet.cn.
    • The expected port and API path.
    • No embedded user information or ambiguous URL components.
  3. Disable automatic redirects with allow_redirects=False, or independently validate every redirect destination before resending credentials or file content.
  4. Refuse loopback, link-local, private-network, and non-HTTPS destinations unless a separately documented enterprise mode explicitly requires them.
  5. Display a clear disclosure and obtain user authorization before uploading documents containing personal information.
  6. Protect config/.env with restrictive permissions and verify that it is owned by the expected user before loading sensitive configuration.
  7. Use narrowly scoped, revocable API tokens and rotate a token immediately if endpoint tampering is suspected.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:66
Finding

Unpinned Third-Party Dependency Installation Creates Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:66-71
Vulnerability Type: Unconstrained third-party package installation
Risk Level: Medium

Vulnerable Code

bash
pip install requests

The metadata also declares the dependency without a version constraint:

yaml
dependencies:
  - python3
  - requests

Technical Analysis

The installation instructions request the latest package version selected by pip from the user's configured package index. No reviewed version, integrity hash, lock file, or trusted package source is specified.

Consequently, installation behavior can change after the Skill has been audited. A compromised package index, compromised upstream release, malicious index configured in the user's environment, or future dependency-chain compromise could introduce attacker-controlled code. Python packages may execute code during installation and are subsequently imported by scripts/main.py, so compromise can occur with the privileges of the user installing or invoking the Skill.

No evidence was found that the current requests package is malicious. The vulnerability is the absence of reproducible and integrity-verified dependency resolution.

Attack Path

  1. A user follows the documented pip install requests instruction.
  2. pip resolves the dependency using the user's configured package indexes and selects an unconstrained available release.
  3. An attacker compromises the selected distribution, its dependency chain, or a configured package index.
  4. The malicious package executes during installation or when scripts/main.py imports requests.
  5. The package runs with the installing or executing user's privileges and can access files, environment data, network resources, and the Skill's API credential available to that process.

Impact Assessment

A successful supply-chain compromise could execute arbitrary Python code with the privileges of the user ...[truncated 354 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin requests and all transitive dependencies to reviewed versions in a requirements or lock file.
  2. Include cryptographic hashes and install with hash verification, for example through pip install --require-hashes -r requirements.txt.
  3. Specify and document a trusted package index rather than relying silently on arbitrary user configuration.
  4. Periodically update pinned versions after vulnerability and provenance review.
  5. Use an isolated virtual environment with only the dependencies required by the Skill.
  6. Add automated dependency vulnerability and integrity scanning to the release process.
  7. Keep skill.yaml dependency metadata consistent with the pinned dependency manifest.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (20)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"

# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 63)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 125)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 29)May include surrounding context.

python
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documentation indicates capabilities that read local files, invoke Python from the shell, and send data to an external OCR API, but it does not declare any explicit tool scope or allowed-tools restrictions. In an agent environment, missing scope declarations can let the skill run with broader-than-necessary privileges, increasing the chance of unintended file access or network exfiltration of sensitive document contents.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file contains natural-language instructions and metadata entirely in Chinese, including the user-facing description. The policy requires flagging language or locale constraints when a skill effectively forces a specific language without user opt-in or justification.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

The skill is designed to transmit user-provided local files containing education certificate data to an external third-party API endpoint. Because the content includes highly sensitive personal information such as name, ID number, school, and enrollment data, external transmission materially increases privacy and data protection risk if users are not clearly informed and controls are not enforced.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The example user phrasing and trigger guidance are only given in Chinese, so users of other locales are not offered an alternative or opt-in. This is a natural-language policy concern because the skill's activation guidance is language-specific without justification.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The documented configurable API base allows sending OCR payloads to an external endpoint, and because the skill processes identity-bearing education records, this creates a meaningful exfiltration surface. The risk is heightened by the presence of an overrideable base URL, which can enable redirection of sensitive documents to unintended servers if misconfigured or tampered with.

Content

Scanner excerpt · SKILL.md (reported line 107)May include surrounding context.

md
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |

### 输出

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The document instructs the skill to send uploaded files to an external OCR endpoint, which creates a real data exfiltration path for highly sensitive personal documents such as education certificates and ID-bearing forms. In this skill context, the risk is elevated because the sample response includes personal identifiers and the API supports image, PDF, and archive uploads, meaning users may transmit regulated personal data to a third-party service.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring, CLI usage text, error messages, and operational guidance are all written only in Chinese, which imposes a specific language on users. There is no indication that the skill is region-specific or that users can opt into this locale, so this is a natural-language policy issue under the language/locale rule.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 45)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 152)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This skill transmits user-supplied document contents to an external OCR API endpoint at api.scnet.cn. Because the documents include highly sensitive personal and educational data such as name, ID number, school, and enrollment dates, sending them off-host creates privacy, compliance, and data-handling risk if users are not clearly informed and protected.

Content

Scanner excerpt · scripts/main.py (reported line 78)May include surrounding context.

python
)
        sys.exit(error_msg)

    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
    return config

def recognize_with_retry(ocr_type, file_path, config, retry_count=0):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language description at L03 is entirely in Chinese, and the manifest does not indicate that the skill is region-specific or that users may choose another language. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The changelog content is entirely in Chinese, including the title and feature description, with no indication that language selection is optional or that the skill is specifically intended for a Chinese-speaking or region-specific context. This can violate language/locale policy where user opt-in or explicit justification is required.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

SKILL.md 在 L131 明确声明该技能会在 429 限流时自动等待并重试最多 3 次。这属于对运行行为的具体承诺;若实际技能代码未实现重试逻辑,则文档会误导调用方对错误处理和幂等性的预期,构成文档与实现意图不一致。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file's user-facing headings and descriptions are entirely in Chinese, which imposes a specific language/locale on users. The content does not indicate that Chinese is optional or that the skill is intentionally limited to a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown file presents all usage instructions, parameter descriptions, and warnings exclusively in Chinese. Under the policy rule, forcing a specific language without user opt-in or a documented locale justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.