T05 · Unauthorized Access and Privilege Escalation
- Location
halucatch/scanner.py:84- Finding
Scan-Root Escape Through Symbolic-Link Files
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is an offline local audit tool, but it needs review because its filesystem boundaries are broader than its safety text implies.
Review before installing. Use it only on trusted or pre-checked directories, avoid scanning folders that may contain symlinks to sensitive files, and run it in a sandbox when auditing untrusted Skills. Be careful with --output-dir because it can write reports outside the target folder.
halucatch/scanner.py:84Scan-Root Escape Through Symbolic-Link Files
The declared description promises a substantive reliability and safety assessment of AI skills along four audit dimensions. The supplied code does not perform any such evaluation; it merely inspects simple indicators in the provided info object and returns a binary skill type classification. This is a materially different primary purpose, not just an implementation detail. No evidence shows analysis of execution reliability, hallucination risk, reproducibility, or business scrutiny guardrails.
The code’s docstring explicitly states it evaluates complexity and maintainability using structural indicators without semantic understanding. The implemented functions only inspect markdown structure, links, script references, code/document ratios, repetition, tables, and instruction density. There is no logic for executing a skill, validating output correctness, testing reproducibility, detecting hallucinations, checking business logic ambiguity, or evaluating interpretation guardrails. While there is a loose overlap with 'code risk' in the sense of static structural heuristics, the primary purpose is materially different from the declared description. Therefore this is a clear description-behavior mismatch.
The description promises a broad reliability audit for AI Skills, including data pipeline integrity, code risk, business logic ambiguity, and interpretation guardrails, aimed at determining whether outputs are trustworthy and reproducible. The supplied code does something much narrower: it scans provided text fields for a few specific indicators such as existence of a .py file, hardcoded paths, a '--validate' flag, column-checking keywords, file discovery functions, and dependency mentions. This is related to code quality and some reproducibility hygiene, but it does not implement the broader multidimensional reliability evaluation described. The primary purpose is therefore materially narrower and different from the declared scope.
The description presents a comprehensive evaluator for AI Skill reliability spanning four dimensions. The supplied code chunk is materially narrower: it is a guardrails checker centered on interpretation safeguards and output-constraint signals extracted mainly from SKILL.md via regex, plus limited file/code heuristics for output determinism. This is not just an implementation detail; the primary purpose in the code is one sub-dimension of the declared system. Additionally, the code scans template files and Python source for rendering patterns, which is more specific behavior than the description states. Therefore, the declared description does not accurately represent what this code chunk actually does.
The declared description promises a broad reliability audit of AI Skills, including trustworthiness, reproducibility, business scrutiny, and four explicit dimensions: data pipeline integrity, code risk, business logic ambiguity, and interpretation guardrails. The actual code chunk only performs a lightweight methodology/documentation review of SKILL.md using regex and simple heuristics. It checks for structured steps, branch-density signals, output format mentions, examples, and whether files mentioned in the markdown exist. This is materially narrower than the declared purpose and misses several core promised dimensions, especially data pipeline integrity and code risk. While the code does contribute to one aspect of documentation/guardrail quality, its primary purpose is not an overall execution reliability evaluator as described.
There is a material mismatch between the broad declared description and the supplied code. The description promises a comprehensive execution reliability evaluation across four dimensions, including data pipeline integrity, code risk, and interpretation guardrails, with emphasis on trustworthy and reproducible AI Skill execution. In contrast, this code is limited to static text-pattern checks on SKILL.md business-rule wording and ambiguity. It does not inspect execution behavior, outputs, reproducibility, pipeline integrity, hallucinations, or deployment safety. The helper function for tool-skill classification is also not reflected in the declared purpose. This is not just an implementation subset; the primary purpose in code is much narrower and materially different from the declared capability.
The declared description is about assessing reliability, reproducibility, hallucination risk, and business scrutiny of an AI Skill's execution. The supplied code does not perform any such evaluation or scoring. Instead, it scans a folder, collects file metadata and contents, identifies likely SKILL.md content, separates test and non-test Python files, detects data files, and extracts version information. These are preparatory repository-scanning capabilities, not the described auditing logic itself. While such scanning could support a later reliability audit pipeline, this chunk's actual primary purpose is file discovery and content extraction, which is materially different from the declared purpose.
Referenced artifact was not completely inspected
| `SKILL.md` | ✅ | Skill 的主指令文件 |
Referenced artifact was not completely inspected
| `SKILL.md` | ✅ | Skill 的主指令文件 |
Referenced artifact was not completely inspected
| `SKILL.md` | ✅ | Skill 的主指令文件 |
Referenced artifact was not completely inspected
| `SKILL.md` | ✅ | Skill 的主指令文件 |
Referenced artifact was not completely inspected
| `SKILL.md` | ✅ | Skill 的主指令文件 |
The skill uses 'network' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.
The documentation states the tool will not access paths outside the target directory, but later allows writing to a custom output directory outside that boundary. This contradiction can mislead operators and downstream agents into granting broader filesystem trust than warranted, undermining sandbox assumptions.
Forcing automatic language selection and asserting the AI already knows the user's language can override explicit user preference or platform language policy. While not a classic security exploit, it can cause incorrect behavior, consent issues, or policy noncompliance in agent execution.
The skill claims writes are confined to the target skill's reports/ directory, but later permits an arbitrary --output-dir. That weakens containment guarantees and can cause an agent or user to write generated files outside the reviewed project boundary, including into sensitive or unexpected locations.
Broad everyday trigger phrases increase the chance that the skill is invoked unintentionally in unrelated conversations. In an agent setting with Write and Bash permissions, accidental activation can lead to unnecessary directory scanning, report generation, or local script execution that the user did not intend.
This markdown file presents all user-facing documentation in a single language and does not indicate that users may choose another language or that the skill is intentionally region-specific. Under the policy, forcing a specific language without user opt-in is a natural-language locale violation.
This code returns English as the default language whenever locale detection fails, which imposes a specific language preference rather than offering a user-selectable fallback. Under the policy, forcing a language without user opt-in is a natural-language locale violation.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
('无引号变量', r'(?:rm|mv|cp)\s+(?:\$\w+[^"\'\s]|[^"\'\s]*\$[^"\'\s]*)', '危险操作中使用未加引号的变量 — 空格/特殊字符导致意外行为'),
('静默吞错', r'\|\|\s*true\b', '|| true 静默丢弃所有错误 — 出问题时无信号'),
('参数缺失', r'\$\{[1-9]\d*\}(?!\s*[:-])|\$[1-9]\d*(?!\s*[:-])', '引用位置参数但无默认值 — 参数缺失时静默失败'),
('提权操作', r'\bsudo\b', 'sudo 提权 — 可能执行意料之外的高权限操作'),
]
# Go
This Python file contains a module docstring entirely in Chinese, and the rest of the user-facing comments/messages are also written in Chinese, indicating a fixed language/locale choice. Under the policy, forcing a specific language without user opt-in or clear region-specific justification is a natural-language policy concern.
The module docstring and function docstring specify behavior entirely in Chinese, and the user-facing issue messages throughout the function are also Chinese-only. This can amount to a language policy violation because the skill appears to force a specific language without user opt-in or a documented region-specific justification.
This Python file contains user-/developer-facing natural-language descriptions entirely in Chinese, including the module and function docstrings that describe the skill's behavior. Under the policy rule, forcing a specific language without opt-in is a locale/language policy concern, and there is no indication anywhere in this file that users may choose another language.
The module-level description is written only in Chinese, and the main report-generation entrypoint defaults to lang='zh-CN'. This imposes a specific language/locale unless the caller explicitly overrides it, which is a natural-language locale policy concern when no user opt-in is obtained.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
'缺少输出': 'No specification of what the output should look like. Different AIs may give completely different formats.',
'缺少结构化步骤': 'Instructions are like a running account. AI may skip key steps or mix up order.',
'缺少示例': 'No examples. AI can only guess and is prone to misunderstanding.',
'缺少验证': 'No validation steps. AI may confidently output incorrect content without verification.',
'未声明前提假设': 'Does not state under what conditions this Skill can work properly.',
'未定义错误回退': 'No fallback plan when execution fails. AI will get stuck.',
'未声明数据来源限制': 'Does not state where data comes from or what it covers. Different AIs may use different data sources, results not comparable.',
No suspicious patterns detected.