Back to skill

Security audit

Train Ticket Ocr

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real OCR wrapper, but it can upload any readable local file to a configurable external endpoint and does not clearly warn users about sending sensitive ticket data off-device.

Install only if you are comfortable sending ticket images and their personal data to Scnet's OCR service. Before use, restrict inputs to intended ticket images or PDFs, do not pass arbitrary local paths, keep the API key private, and avoid changing SCNET_API_BASE unless you fully trust and verify the destination.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:91
Finding

Arbitrary Local File Upload to a Remote OCR Service

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:53
Finding

Unrestricted API Endpoint Override Can Expose Files and Bearer Credentials

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:62
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (20)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill processes local ticket images and extracts highly sensitive personal data such as ID numbers, then sends the source image to a third-party OCR API, but the markdown does not clearly warn users about this external transmission. Because train tickets commonly contain personally identifiable information, lack of notice meaningfully increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"
# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数
RETRY_BACKOFF_FACTOR = 2   # 退避因子,每次重试等待时间翻倍

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 61)May include surrounding context.

md
INITIAL_RETRY_DELAY = 1    # 初始等待时间(秒)
# --------------------
def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 123)May include surrounding context.

md
INITIAL_RETRY_DELAY = 1    # 初始等待时间(秒)
# --------------------
def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 27)May include surrounding context.

python
INITIAL_RETRY_DELAY = 1    # 初始等待时间(秒)
# --------------------
def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 41)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documentation declares capabilities that imply local file access, shell execution, and outbound network use, but it does not define any explicit tool scope or permission boundaries. In an agent environment, this can cause the skill to be invoked with broader privileges than necessary, increasing the chance of unintended file access or data exfiltration.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

清单描述明确限定为从火车票中提取出发站、到达站、车次、座位号等字段,但文档在功能特性中写为“覆盖火车发票”,并在示例中让用户“提取这张发票的信息”。这不是单纯信息缺失,而是把技能目标从车票识别表述成了发票识别,构成意图与文档的主动矛盾。

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

This finding reflects real external data transmission to a remote OCR endpoint. In context, outbound transmission is expected for a cloud OCR skill, but it is still security-relevant because the transmitted files may contain sensitive personal information from train tickets.

Content

Scanner excerpt · SKILL.md (reported line 53)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation guidance is broad enough to trigger on generic document-extraction requests, including invoices or other personal documents outside the declared train-ticket scope. In an agent setting, this can cause accidental processing and transmission of unrelated sensitive local files to the external OCR provider.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 105)May include surrounding context.

md
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |

### 输出

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The skill relies on sending uploaded images and documents to https://api.scnet.cn, which is an external service outside the local trust boundary. In the context of train ticket OCR, the transmitted content and results can contain PII and financial/invoice data, so this creates data exposure, compliance, and third-party handling risks if not clearly disclosed and controlled.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation instructs users to upload train tickets and receive OCR results from an external API, but it does not warn that the files and returned fields include sensitive personal data such as names, ID numbers, ticket details, and invoice data. This omission can lead operators to transmit regulated personal data to a third-party service without adequate user notice, consent, or data-handling review.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring and user-facing CLI usage/messages are written only in Chinese, indicating the skill is designed to communicate in a single language without opt-in or fallback. Under the policy, locale or language restrictions should either be user-selectable or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 43)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 150)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

The skill uploads a user-supplied file containing train-ticket data to an external OCR API endpoint. Because train tickets can contain sensitive personal data such as ID numbers, travel dates, stations, and ticket details, this creates a real data-exfiltration/privacy risk if users are unaware, if the endpoint is untrusted, or if the base URL is modified through configuration.

Content

Scanner excerpt · scripts/main.py (reported line 76)May include surrounding context.

python
)
        sys.exit(error_msg)

    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
    return config

def recognize_with_retry(ocr_type, file_path, config, retry_count=0):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description is written only in Chinese, which imposes a language constraint in the skill's natural-language metadata without offering any user choice or documenting a justified region-specific scope. This matches the policy concern for language or locale restriction in natural-language content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The natural-language content of the file is predominantly in Chinese, and there is no indication that the skill offers multilingual support or that Chinese is a required locale for a region-specific tool. Under the policy rule for language/locale constraints, this can be considered a forced language choice without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file contains user-facing natural language content exclusively in Chinese, and there is no indication that users may choose another language or that the skill is intentionally restricted to a Chinese-speaking or region-specific context. Per the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.