T09 · Insecure Skill Coding Practices
- Location
scripts/main.py:91- Finding
Arbitrary Local File Upload to a Remote OCR Service
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a real OCR wrapper, but it can upload any readable local file to a configurable external endpoint and does not clearly warn users about sending sensitive ticket data off-device.
Install only if you are comfortable sending ticket images and their personal data to Scnet's OCR service. Before use, restrict inputs to intended ticket images or PDFs, do not pass arbitrary local paths, keep the API key private, and avoid changing SCNET_API_BASE unless you fully trust and verify the destination.
scripts/main.py:91Arbitrary Local File Upload to a Remote OCR Service
scripts/main.py:53Unrestricted API Endpoint Override Can Expose Files and Bearer Credentials
SKILL.md:62Unpinned Third-Party Dependency Installation
The skill processes local ticket images and extracts highly sensitive personal data such as ID numbers, then sends the source image to a third-party OCR API, but the markdown does not clearly warn users about this external transmission. Because train tickets commonly contain personally identifiable information, lack of notice meaningfully increases privacy and compliance risk.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"
# --- 新增:重试配置 ---
MAX_RETRIES = 3 # 最大重试次数
RETRY_BACKOFF_FACTOR = 2 # 退避因子,每次重试等待时间翻倍
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
INITIAL_RETRY_DELAY = 1 # 初始等待时间(秒)
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
INITIAL_RETRY_DELAY = 1 # 初始等待时间(秒)
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
INITIAL_RETRY_DELAY = 1 # 初始等待时间(秒)
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)
The skill documentation declares capabilities that imply local file access, shell execution, and outbound network use, but it does not define any explicit tool scope or permission boundaries. In an agent environment, this can cause the skill to be invoked with broader privileges than necessary, increasing the chance of unintended file access or data exfiltration.
清单描述明确限定为从火车票中提取出发站、到达站、车次、座位号等字段,但文档在功能特性中写为“覆盖火车发票”,并在示例中让用户“提取这张发票的信息”。这不是单纯信息缺失,而是把技能目标从车票识别表述成了发票识别,构成意图与文档的主动矛盾。
This finding reflects real external data transmission to a remote OCR endpoint. In context, outbound transmission is expected for a cloud OCR skill, but it is still security-relevant because the transmitted files may contain sensitive personal information from train tickets.
SCNET_API_KEY=your_scnet_api_key_here
SCNET_API_BASE=https://api.scnet.cn/api/llm/v1
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)
The activation guidance is broad enough to trigger on generic document-extraction requests, including invoices or other personal documents outside the declared train-ticket scope. In an agent setting, this can cause accidental processing and transmission of unrelated sensitive local files to the external OCR provider.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |
### 输出
The skill relies on sending uploaded images and documents to https://api.scnet.cn, which is an external service outside the local trust boundary. In the context of train ticket OCR, the transmitted content and results can contain PII and financial/invoice data, so this creates data exposure, compliance, and third-party handling risks if not clearly disclosed and controlled.
# Sugon-Scnet OCR API 文档摘要
## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`
## 请求头
- `Content-Type: multipart/form-data`
The documentation instructs users to upload train tickets and receive OCR results from an external API, but it does not warn that the files and returned fields include sensitive personal data such as names, ID numbers, ticket details, and invoice data. This omission can lead operators to transmit regulated personal data to a third-party service without adequate user notice, consent, or data-handling review.
The module docstring and user-facing CLI usage/messages are written only in Chinese, indicating the skill is designed to communicate in a single language without opt-in or fallback. Under the policy, locale or language restrictions should either be user-selectable or clearly justified as region-specific.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
" b) 配置文件:\n"
f" mkdir -p {SKILL_ROOT}/config\n"
f" echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
f" chmod 600 {ENV_FILE}\n"
"\n配置完成后重新运行。"
)
sys.exit(error_msg)
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
" b) 配置文件:\n"
f" mkdir -p {SKILL_ROOT}/config\n"
f" echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
f" chmod 600 {ENV_FILE}\n"
"\n配置完成后重新运行。"
)
sys.exit(error_msg)
The skill uploads a user-supplied file containing train-ticket data to an external OCR API endpoint. Because train tickets can contain sensitive personal data such as ID numbers, travel dates, stations, and ticket details, this creates a real data-exfiltration/privacy risk if users are unaware, if the endpoint is untrusted, or if the base URL is modified through configuration.
)
sys.exit(error_msg)
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
return config
def recognize_with_retry(ocr_type, file_path, config, retry_count=0):
The manifest description is written only in Chinese, which imposes a language constraint in the skill's natural-language metadata without offering any user choice or documenting a justified region-specific scope. This matches the policy concern for language or locale restriction in natural-language content.
The natural-language content of the file is predominantly in Chinese, and there is no indication that the skill offers multilingual support or that Chinese is a required locale for a region-specific tool. Under the policy rule for language/locale constraints, this can be considered a forced language choice without user opt-in.
This markdown file contains user-facing natural language content exclusively in Chinese, and there is no indication that users may choose another language or that the skill is intentionally restricted to a Chinese-speaking or region-specific context. Per the policy, forcing a specific language without user opt-in is a natural-language policy concern.
No suspicious patterns detected.