Back to skill

Security audit

Medical Invoice Ocr

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to perform medical-invoice OCR as described, but it uploads sensitive documents with an API key to a configurable network endpoint without strong scoping or a clear consent warning.

Review this skill before installing if you will process real medical invoices. Use it only when you are comfortable sending the selected file to Scnet, keep SCNET_API_BASE fixed to the official HTTPS endpoint, protect the API key, and avoid relying on automatic invocation for sensitive documents unless the agent asks for confirmation first.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:50
Finding

Unrestricted API Endpoint Can Disclose Credentials and Sensitive Medical Documents

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:68
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (21)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"

# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 61)May include surrounding context.

md
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 123)May include surrounding context.

md
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 29)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 32)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 45)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 41)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documents use of environment variables, local file access, network calls, and shell execution but does not declare an explicit tool scope such as permissions or allowed-tools. This creates a transparency and policy gap: an agent may invoke capabilities broader than users expect, especially when handling sensitive medical invoice data and local file paths.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill processes medical invoices, which commonly contain highly sensitive personal and health-related information, and sends image contents to a third-party OCR API. The documentation explains token setup and API usage but does not clearly and prominently warn users that local files and medical data will leave the device and be transmitted to an external service.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This finding reflects an external transmission endpoint used by the OCR service. External transmission is expected for a cloud OCR skill, but in this context it is security-relevant because the transmitted content may include medical invoices and other sensitive personal data.

Content

Scanner excerpt · SKILL.md (reported line 53)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This is a second instance documenting the same external API base URL. While not malicious by itself, it confirms that the skill relies on third-party network transmission for processing sensitive medical invoice content, making privacy and consent controls important.

Content

Scanner excerpt · SKILL.md (reported line 105)May include surrounding context.

md
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |

### 输出

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The file explicitly references a remote API endpoint for OCR processing, which means uploaded documents and their extracted contents are transmitted outside the agent's local boundary. Because this skill handles medical invoices, the external transmission is more sensitive than ordinary OCR and can expose personally identifiable, billing, and healthcare-adjacent data if users are not properly informed or if the vendor is not appropriately vetted.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation instructs users to upload medical invoice files to a third-party OCR endpoint but does not clearly warn that sensitive financial and potentially medical personal data will leave the local environment. In the context of medical invoices, this increases privacy, compliance, and data-handling risk because users may unknowingly transmit regulated or confidential information to an external service.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The skill is designed to send files and metadata to an external network endpoint by default. In the context of medical invoice OCR, this transmission materially increases risk because sensitive documents leave the local environment and are processed by a remote service, potentially implicating confidentiality and compliance obligations.

Content

Scanner excerpt · scripts/main.py (reported line 55)May include surrounding context.

python
config['SCNET_API_BASE'] = env_api_base

    # 3. 设置默认值
    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')

    # 4. 检查必要配置
    api_key = config.get('SCNET_API_KEY', '')

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 71)May include surrounding context.

python
"2. 配置文件:\n"
            f"   mkdir -p {SKILL_ROOT}/config\n"
            f"   echo 'SCNET_API_KEY=你的真实密钥' > {ENV_FILE}\n"
            f"   chmod 600 {ENV_FILE}\n"
        )
        sys.exit(error_msg)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script uploads the user-supplied file directly to a third-party OCR API, and this skill is explicitly for medical invoices, which commonly contain highly sensitive personal and financial data. There is no explicit consent flow, privacy notice, destination allowlist enforcement, or data minimization, so users may unknowingly exfiltrate regulated data off-host.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description is written only in Chinese ("支持识别医疗发票识别"), which indicates a language-specific presentation without any stated user opt-in or justification for restricting the skill to that locale. Under the policy, language or locale constraints should either offer user choice or be clearly documented as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

This markdown file uses Chinese throughout, including headings and change descriptions, but does not state that the skill is China-specific or that Chinese is an intentional locale choice. Under the policy rule for natural-language violations, forcing a specific language without user opt-in or documented justification can be a concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

This markdown file references testing, CI/CD, integrations, and deployment workflows, including deployment to Kubernetes, EC2, or ECS, but provides no warning that these actions can affect live systems, infrastructure, or project state. For markdown files, skill descriptions should disclose behaviours that could affect system integrity or user data when such operational actions are described.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The documentation states the AI will automatically trigger the skill based on keywords, which can lead to external transmission of sensitive medical invoice images without clear, explicit opt-in at the moment of use. In the medical context, automatic activation increases the chance of surprise data sharing and weakens informed consent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

Natural-language strings throughout the script, including the module description, usage text, and error messages, are only presented in Chinese. This imposes a specific language on all users without opt-in or locale selection, which matches the policy's language/locale violation criteria.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.