Back to skill

Security audit

Android Static Analyzer

Security checks for vulnerabilities and agentic risk

Overview

The skill is broadly coherent for Android test analysis, but it handles source code and possible passwords too loosely and should be reviewed before use.

Install only for Android projects whose source may be processed by an LLM and reviewed in generated profiles. Before use, confirm the exact project path, avoid running it on repositories with real credentials or proprietary secrets unless redaction is added, review static-profile.json before committing or uploading it, and treat generated profile text as untrusted input to downstream agents.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
scripts/analyze.py:143
Finding

Untrusted Android Source Code Is Embedded Directly into the LLM Prompt

Content
View full analysis

Vulnerability Details

File Location: scripts/analyze.py, lines 143–149 and 202–204; prompt output occurs at line 249
Vulnerability Type: Prompt injection through untrusted source-code content
Risk Level: High

Vulnerable Code

python
source_text = '\n\n'.join(
    f'// === {name} ===\n{code}'
    for name, code in source_code.items()
)

prompt = f"""你是一个资深 Android 测试专家。请仔细阅读以下 Android 应用源码(包名:{app_package}),输出一份供 AI 自动化测试 Agent 直接使用的**测试先验知识**。

# ... prompt instructions omitted ...

## 源码

{source_text}"""

return prompt

The resulting prompt is subsequently emitted for processing by the agent:

python
print("--- LLM_PROMPT ---")
print(prompt)

Technical Analysis

The analyzer treats Kotlin and Java source files as trusted prompt content. source_text is interpolated verbatim into the same instruction message that defines the LLM's role, analysis goals, and required output format.

There is no security boundary distinguishing trusted analyzer instructions from untrusted project content. The prompt also does not instruct the model to treat source-file text exclusively as evidence or to ignore instructions appearing inside it.

Although read_source_files() removes certain empty and comment-only lines, this does not prevent prompt injection. An attacker can place natural-language instructions in string literals, annotations, multiline strings, identifiers, or ordinary code lines that survive filtering.

This permits indirect prompt injection: a malicious Android project can attempt to override the analysis instructions, suppress or fabricate findings, violate the required JSON format, or insert attacker-selected content into the generated static profile.

Attack Path

  1. An attacker creates or modifies a Kotlin or Java file in the Android project.
  2. The attacker embeds an instruction such as a request to ignore the analyzer's prior requirements and emit attacker-controlled profile content.
  3. find_kotlin_files() di ...[truncated 1506 chars]
Remediation
View remediation

Remediation Suggestions

  1. Separate trusted instructions from source data

    • Submit analyzer policy in a trusted system or developer message.
    • Place source code in a distinct structured input field or user-data attachment rather than interpolating it into the instruction body.
  2. Explicitly classify project content as untrusted

    • Tell the model that all source text is evidence only.
    • State that instructions, requests, role declarations, or output directives found in source files must never be followed.
  3. Use strong content boundaries

    • Wrap each source file in clearly identified data blocks.
    • Use dynamically generated delimiters and ensure project content cannot terminate or imitate the boundary.
    • Include the canonical relative path and a content hash for provenance.
  4. Validate generated output

    • Parse the response strictly as JSON.
    • Validate it against an authoritative JSON Schema.
    • Reject extra fields, non-JSON prefixes or suffixes, invalid priority values, and instruction-like content in fields intended to hold factual data.
  5. Constrain downstream use

    • Treat generated summaries and profiles as untrusted data when consumed by another agent.
    • Do not place generated free text into privileged instruction messages.
    • Require user review before generated artifacts influence actions with side effects.
  6. Use deterministic analysis where possible

    • Extract manifests, navigation edges, identifiers, permissions, and hardcoded constants using parsers.
    • Reserve LLM processing for semantic interpretation and require claims to reference source locations.
  7. Add adversarial tests

    • Test payloads in Kotlin strings, Java annotations, multiline literals, identifiers, and code statements.
    • Confirm that embedded instructions cannot alter the output format, role, or analysis policy.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

代码的核心行为是一个“提示词构造器/源码采样器”,而不是完整的“Android 源码分析并生成测试先验知识文档并上报平台”的实现。它会查找 AndroidManifest.xml、提取包名、收集部分源码、按固定模板拼接出供外部 Agent/LLM 使用的 Prompt,然后打印 metadata 和 prompt。脚本注释也明确说明“脚本不调用 LLM,由 OpenClaw Agent 读取 LLM_PROMPT 后自行调用 LLM 分析”。因此,声明中的前半部分“分析 Android 项目源码”只部分成立;但更关键的生成知识文档、打包、上报测试平台等能力并未在该代码中实现,属于实质性能力差异。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

声明描述的是一个较完整的 Android 源码分析与测试先验知识生成/上报能力;而实际代码是一个非常专项的导航图提取脚本。它的功能局限于读取 navigation XML、构建路由图并导出 JSON,可作为更大分析流程中的一个辅助步骤,但单独来看,其主要目的与声明的核心能力存在明显差异。代码没有体现 LLM 调用、全项目源码分析、测试知识文档生成,也没有任何打包或上传行为,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/analyze.py (reported line 206)May include surrounding context.

python
{source_text}"""

    return prompt


def main():

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill instructs reading project files and writing a generated static-profile.json, but it does not declare any tool scope such as allowed-tools or permissions. This weakens containment and user transparency because file-system capabilities exist implicitly rather than being explicitly constrained and disclosed.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to activate on common Android analysis requests, which can cause the skill to run in contexts where the user did not specifically consent to its file-reading and file-writing workflow. In a skill that touches local source code and generates output files, overbroad activation increases the chance of unintended data processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill description and all prescribed prompts/output formats are written as mandatory Chinese instructions, with no indication that the user may choose another language. This can violate language or locale policy where skills should not force a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs saving generated analysis to the project root without clearly warning the user beforehand that a local file write will occur. Silent or weakly disclosed writes can alter repositories, leak derived sensitive information into tracked files, or create unexpected artifacts in developer environments.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The skill includes a concrete curl command to POST the generated static-profile.json to a local platform endpoint. Although it says not to auto-POST, the file may contain sensitive source-derived metadata, business logic, test secrets, hardcoded credentials, or internal app structure, so embedding an exfiltration path materially raises data-transfer risk.

Content

Scanner excerpt · SKILL.md (reported line 491)May include surrounding context.

md
然后告诉用户:
- 文件路径:`<project_path>/static-profile.json`
- 在测试平台「自动测试」页面,点「导入先验知识」按钮,选择此文件上传即可
- 或者平台 server 启动后执行:`curl -X POST http://localhost:2601/api/v1/sdk/static-profile -H "Content-Type: application/json" -d @<project_path>/static-profile.json`

**不要自动 POST 到平台**,由用户手动导入。

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The schema description and multiple field descriptions specify LLM-generated content in Chinese, such as 'AI 自动化测试所需' and Chinese-only examples/descriptions. This creates a natural-language locale constraint without any indication that users can choose another language or that the tool is intentionally region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The businessSummary field description defines the content in Chinese and implies a fixed-language output for an LLM-generated summary. Under the policy, forcing a specific language without user opt-in is a locale policy issue unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The LLM-generated fields businessRole and testHints are described entirely in Chinese and implicitly require Chinese-language content. This is a natural-language policy concern because the schema does not provide user choice or explain a region-specific restriction.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The prompt explicitly instructs the downstream LLM to extract hardcoded accounts, passwords, promo codes, and similar test data from source code and include them in the output JSON. This materially increases the chance that secrets, credentials, or sensitive business data embedded in the codebase will be surfaced, copied, retained, or transmitted to external systems during analysis.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script prints collected project metadata and the full constructed LLM prompt to stdout, and that prompt embeds substantial portions of the application's source code. In this skill's context, stdout is an exfiltration channel to a downstream agent/LLM and potentially logs, so proprietary code and embedded secrets can be disclosed without any explicit consent gate, redaction, or warning.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This Python file contains its module description, usage text, and later user-facing console output entirely in Chinese, with no indication that the tool is region-specific or that another language can be selected. The stated policy flags language/locale constraints when a skill forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
99% confidence
Finding

L054 明确写着“LLM 生成 8 份文档”,但后续实际列出了文档1到文档9,且 L017、L341、L432 也都围绕 9 份文档展开。这会让技能的实际产出范围与其执行说明不一致,属于文档描述与实际行为范围不匹配。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
73% confidence
Finding

The criticalFlows section includes Chinese-only descriptions and an example flow name in Chinese, suggesting a fixed locale for generated content. While less explicit than a hard requirement, it still nudges the skill toward one language without documenting user opt-in or justified locale scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The embedded prompt requires the model to produce its analysis and JSON field contents in Chinese, and the surrounding script documentation is also Chinese-only. This is a locale/language constraint presented without opt-in or any documented justification, which matches the policy category for forced language selection.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

文件头和函数说明将该工具描述为“提取”与“解析”Navigation Graph,但 main() 在完成分析后还会在输入目录旁生成 nav-graphs.json。对 Android 源码做分析本身是符合技能目的的,但额外的本地持久化写文件属于超出纯读取/提取语义的行为,应在描述中明确。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.