T09 · Insecure Skill Coding Practices
- Location
scripts/recognize_table.py:77- Finding
Long-Term Plaintext Storage of OCR Credentials
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a real cloud OCR helper, but it should be reviewed because it stores OCR credentials in plaintext and can upload and persist sensitive document contents beyond a recognition-only request.
Install only if you are comfortable sending selected documents and OCR credentials to NetOCR. Prefer environment variables or a managed secret store over config.json, avoid shared repositories or sync folders for generated outputs, and review or delete .table_result.json and exported files after processing sensitive documents.
scripts/recognize_table.py:77Long-Term Plaintext Storage of OCR Credentials
scripts/recognize_table.py:379Recognition Requests Trigger Export Without Explicit User Consent
scripts/recognize_table.py:308Unconditional Persistent Storage of Complete OCR Responses
The setup flow explicitly instructs the agent to solicit API key/secret from the user and write them to ./config.json for permanent reuse. This creates a clear secret-handling vulnerability: the agent becomes a credential collection and storage mechanism, and the stored plaintext secrets can later be read by other tools, leaked through filesystem access, or exfiltrated if the skill is triggered unexpectedly.
The skill describes capabilities to read/write local files, access environment variables, and make outbound network requests, but it does not declare any explicit tool scope or permission boundaries. This increases the chance that an agent executes sensitive actions without clear user awareness or platform-level restriction, especially because the skill handles local documents and credentials.
The skill states that API credentials are loaded from a local config file and, on first use, will be requested from the user. This is dangerous because it normalizes collection of secrets through the conversational channel and local persistence on disk, increasing the risk of credential exposure, reuse by other workflows, or accidental inclusion in logs, backups, or shared directories.
The trigger phrases are broad everyday expressions such as recognizing a document or reading image text, which can cause the skill to activate in contexts the user did not intend. In this skill, accidental activation is more dangerous because activation can lead to uploading user files to a third-party OCR service and potentially soliciting or using stored credentials.
The skill is designed to send documents, images, PDFs, and extracted text/table content to a third-party OCR service, but the documentation lacks a clear privacy and consent warning. Because the supported inputs include invoices, contracts, IDs, statements, and other sensitive records, users may unknowingly transmit regulated or confidential data off-platform.
This code sends Base64-encoded document contents plus API credentials to an external domain (netocr.com) for OCR processing. External transmission is inherent to the skill's function, but it remains security-relevant because highly sensitive document contents may be exfiltrated to a third party if users are not clearly warned and consent is not obtained.
'format': 'json',
**options # 可选参数透传
}
resp = requests.post('https://netocr.com/api/recog_table_base64',
data=payload, timeout=60)
return resp.json()
This example uploads full files directly to the external OCR service, creating a clear third-party data transfer path for potentially confidential documents. In the context of OCR for contracts, invoices, IDs, and reports, the main risk is privacy and data governance rather than code execution, but the exposure is still meaningful.
'format': 'json', **options } resp = requests.post('https://netocr.com/api/recog_table_file', files=files, data=data, timeout=60) return resp.json()
The documentation explicitly states that /api/download_file returns a presigned OSS URL, but the sample code writes the HTTP response body directly to disk as if it were the exported file. This mismatch can cause downstream implementations to save JSON or URL text instead of the intended document, and may lead callers to mishandle untrusted download URLs without proper validation or a second fetch step.
The code posts consumeId, page range, and export type to an external API to obtain a presigned download URL. This transmission is less sensitive than the OCR upload itself, but it still crosses a trust boundary and, combined with the flawed sample logic, may encourage unsafe handling of returned download locations.
'num': num, # 如 "1-1" 或 "1-5"
'type': file_type # 如 "xls", "md", "flowWord"
}
resp = requests.post('https://netocr.com/api/download_file',
data=payload, timeout=60)
if resp.status_code == 200:
The file instructs users to store long-lived OCR credentials in a local config.json in the skill directory without warning about filesystem exposure, accidental check-in, or permission hardening. In an agent/skill environment, this increases the chance that secrets are copied, logged, bundled, or committed, enabling unauthorized API use if the file is exposed.
The module docstring presents the script as a general document/table recognition tool, implying both generic OCR and table OCR. However, the implementation exclusively calls table-specific APIs (recog_table_base64, recog_table_file) with fixed typeId 3050 and formats results as table output, so there is no general document OCR path in code.
The user-facing docstring, usage instructions, and CLI descriptions are presented exclusively in Chinese. This imposes a language constraint without any opt-in, alternative locale, or documentation that the tool is intentionally Chinese-only for a justified regional context.
The script uploads entire documents/images and API credentials to an external OCR provider without an explicit privacy or data-transmission warning at the point of use. In this skill context, users may process invoices, contracts, IDs, and financial records, so silent transmission of sensitive content to a third party materially increases confidentiality and compliance risk.
The code takes a URL returned by a remote service and performs a follow-up request to it after only string-replacing the host and scheme, while still preserving attacker-influenced path and query data and explicitly setting the Host header from remote input. If the upstream service or response is compromised, this creates a server-side request/credentialed fetch primitive to an attacker-chosen object path on OSS and can lead to unintended outbound requests, downloading untrusted content, or abuse of signed URLs.
oss_https = oss_url.replace("http://", "https://").replace(
original_host, "oss-cn-beijing.aliyuncs.com"
)
resp = requests.get(
oss_https,
headers={"Host": original_host},
timeout=60,
文档写明 nLanguage 默认值为 0=简体中文,并在“未指定→0”中规定了未显式说明语言时的处理方式。这会在用户未选择语言时强制采用特定语言/locale,而不是先征求用户确认或提供选择。
The script saves a JSON result file next to the source image and may also download and write exported output files locally. While there are post-action print statements indicating files were saved, there is no clear upfront warning in the usage/help text that running recognition will create additional files on disk by default.
The CLI advertises a table-recognition tool and offers immediate export choices including Word/PDF/TXT/OFD, while the rest of the script's documented examples and result handling are centered on recognition and table output. This is not a security issue by itself, but the documentation around scope and behavior is inconsistent enough that user intent may be misled about what the recognition stage versus download/export stage actually does.
No suspicious patterns detected.