T01 · Skill Instruction Hijacking
- Location
scripts/accounting_parser.py:341- Finding
Untrusted OCR Output Is Interpolated into a GUI-Agent Instruction
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This bookkeeping skill has a coherent purpose, but it handles sensitive financial screenshots and can automate account-record changes without enough user confirmation or containment.
Review this before installing if you care about financial-data privacy or bookkeeping accuracy. Enable confirmation before saving records, use only trusted versions of the image and GUI dependencies, keep any local history in a private directory, and verify that automation cannot leave the intended accounting app.
scripts/accounting_parser.py:341Untrusted OCR Output Is Interpolated into a GUI-Agent Instruction
scripts/user_preferences.py:32Accounting Entries Can Be Saved Without User Confirmation by Default
scripts/accounting_history.py:53Financial History and Failure Records Are Stored as Unprotected Plaintext Files
package.json:16Security-Sensitive Automation Dependencies Use Floating Version Constraints
scripts/runtime_validator.py:86Runtime and Target-Application Validation Can Be Trivially Spoofed or Bypassed
This finding indicates runtime fingerprinting, app/package inspection, keyword inspection for competitor detection, and execution blocking outside approved environments. In context, the extensive anti-modification and anti-competition language increases suspicion that the skill may perform unauthorized environment surveillance and restrictive behavior beyond legitimate bookkeeping, which is dangerous when paired with access to screenshots and GUI automation.
This finding indicates runtime fingerprinting, app/package inspection, keyword inspection for competitor detection, and execution blocking outside approved environments. In context, the extensive anti-modification and anti-competition language increases suspicion that the skill may perform unauthorized environment surveillance and restrictive behavior beyond legitimate bookkeeping, which is dangerous when paired with access to screenshots and GUI automation.
This finding indicates runtime fingerprinting, app/package inspection, keyword inspection for competitor detection, and execution blocking outside approved environments. In context, the extensive anti-modification and anti-competition language increases suspicion that the skill may perform unauthorized environment surveillance and restrictive behavior beyond legitimate bookkeeping, which is dangerous when paired with access to screenshots and GUI automation.
The validator hard-blocks execution outside a single vendor/runtime and a narrow app allowlist while the skill metadata claims support for multiple mainstream platforms. This creates a deceptive capability gap that can mislead users and platform operators, and it can be used as a hidden lock-in or anti-competitive control that only becomes apparent at runtime.
The example code performs an external POST request to an image-understanding API, which indicates data leaves the local environment. In the context of an auto-accounting skill that may handle receipts, payment screenshots, and transaction records, such transmission can expose sensitive financial and personal data if not explicitly controlled and disclosed.
signature = generate_signature(API_KEY, timestamp)
# API 会验证签名,非授权调用会被拒绝
response = requests.post(
"https://xiaoyi.api/understanding",
headers={
"X-Timestamp": str(timestamp),
The document proposes sending user images to a server-side API for image understanding, but it does not mention any user consent, privacy notice, data minimization, retention policy, or transport/security controls. Because this skill processes accounting screenshots that may contain financial and personal information, undocumented external transmission materially increases privacy and compliance risk.
The README explicitly promotes automatic extraction of transaction details from screenshots and automatic entry into an accounting app, but it provides no warning about handling sensitive financial and personal data. This creates a real privacy and integrity risk because users may upload receipts, order pages, and payment screenshots containing names, account details, addresses, and purchase history without informed consent or guidance on safe handling.
The skill declares no explicit tool scope or permissions, yet the documented file structure and static analysis indicate environment access plus local file read/write behavior. In an automation skill that processes financial screenshots and may persist history, missing scope declarations reduce transparency and prevent users or platforms from understanding what data the skill can access or store.
Automatic activation on any received image is insufficiently constrained for a skill that handles privacy-sensitive financial screenshots and can drive a GUI agent. This can lead to unintended processing of unrelated images, accidental extraction of sensitive data, or unintended writes into the bookkeeping app without clear user intent.
The trigger keywords are generic terms like '记账', '账单', '消费', '支出', and '收入', which can appear in ordinary conversation. Broad triggers increase the chance of unintended activation, causing the skill to process screenshots or initiate bookkeeping actions when the user did not explicitly intend automation.
The skill processes screenshots containing financial and personal transaction details, but the description does not provide a clear privacy warning about extraction, storage, transmission, or retention of that data. Because this is finance-related and may involve local persistence plus GUI automation, the absence of a privacy notice materially increases the risk of uninformed consent and mishandling of sensitive data.
The manual trigger examples remain broad and do not define clear boundaries for what images or conversations should be processed. In a skill with image parsing and GUI automation, ambiguous commands can cause overcollection of financial information or accidental automated actions.
The manifest explicitly states that the skill will automatically recognize accounting information from images and record it into a bookkeeping app, but it does not disclose how sensitive financial data is handled, stored, or confirmed before submission. Because this skill operates on payment and purchase screenshots, missing user warnings and consent boundaries increase the risk of privacy exposure, mis-entry, and unintended side effects in a finance context.
The trigger keywords are very broad everyday finance terms such as '记账', '账单', '消费', and '支出', which can cause the skill to activate in contexts where the user did not intend to launch an automation that processes images and records financial data. In a finance automation skill, unintended activation is more dangerous because it can lead to accidental extraction of sensitive payment information and unintended bookkeeping actions in another app.
The configuration disables confirmation before saving accounting entries by default, which can cause the agent to persist incorrect or hallucinated OCR/LLM-extracted financial data without user review. In this skill's context, the model is parsing screenshots from payment and shopping platforms, so extraction mistakes can directly create false bookkeeping records and privacy-sensitive financial actions.
The manifest explicitly describes automated extraction of financial information from user images and automatic entry into a bookkeeping app, but it provides no notice about handling sensitive financial data, no consent model, and no warning about mistaken OCR-driven actions. In this context, the skill processes receipts/payment screenshots from platforms like WeChat and Alipay, so silent automation materially increases privacy and integrity risk.
The clear_history method erases all in-memory accounting history and immediately writes the empty state to disk via save_history, making the action effectively destructive. There is no confirmation prompt, visible warning, or explicit disclosure in this file that all stored records will be permanently cleared.
The module docstring and surrounding user-facing text are entirely in Chinese and tied to a specific app/environment, with no indication that language selection is optional. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is clearly justified.
The code explicitly checks the target app package and raises a PermissionError when it detects competing accounting apps, enforcing undisclosed vendor exclusivity at runtime. This is a genuine security/integrity issue because the skill behavior diverges from its general-purpose description and can be used to restrict user autonomy or sabotage interoperability in multi-app environments.
The header text states that the skill may only be used in an 'official' environment and a specific app, without giving users choice or technical justification. While not directly exploitable like code execution, this is a concerning policy signal because it indicates intentional restriction beyond the declared accounting function and increases suspicion when combined with runtime enforcement elsewhere in the file.
The file enforces use of an 'authorized' runtime and app, and exits when those conditions are not met, but the user-facing description does not disclose these restrictions. Hidden platform enforcement is risky because it changes the trust model of the skill, can coerce users into a specific ecosystem, and may conceal nonfunctional policy controls inside operational code.
The validator checks for platform-identifying environment variables and refuses to run outside a vendor-specific environment, even though the stated skill purpose is image-based accounting automation. This creates an unnecessary execution gate that can mislead users, reduce portability, and impose undisclosed platform lock-in unrelated to core functionality.
This Python file contains natural-language docstrings and user-facing messages entirely in Chinese, including the skill title and runtime restrictions, with no indication that users can opt into another language. Under the policy, forcing a specific language without user choice is a natural-language policy violation unless the locale constraint is clearly documented and justified as region-specific.
The header claims MIT-0 licensing but then adds usage restrictions that MIT-0 does not permit, creating a contradictory and potentially deceptive licensing posture. This can expose users and integrators to legal ambiguity and can be used to pressure or block legitimate analysis, redistribution, or deployment.
The code enforces a vendor-specific runtime requirement through environment-based checks that are not disclosed in the manifest description. Hidden execution constraints are risky because they undermine informed consent, portability, and trust, especially for an automation skill expected to run across multiple platforms.
No suspicious patterns detected.