Back to skill

Security audit

SLS + ARMS 全链路问题排查

Security checks for vulnerabilities and agentic risk

Overview

The skill matches Alibaba Cloud troubleshooting, but it needs review because it can expose raw production logs and automatically expand into local source-code inspection with weak scoping.

Install only in an environment where you are authorized to query the configured Alibaba Cloud SLS/ARMS projects and inspect the relevant source repositories. Use least-privilege cloud keys, avoid broad production logstores, expect raw logs and stack traces to be shown, and review/redact sensitive request, response, token, user, and business fields before sharing output.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:47
Finding

Skill instructions override Agent autonomy and force automatic repository access

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/query_trace.py:204
Finding

Unescaped user input permits SLS query-language injection and unintended log retrieval

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

该代码与声明有部分一致之处:确实会查询阿里云 SLS 和 ARMS,构建调用链,输出异常、堆栈、慢调用以及建议,符合“查日志”“画调用链”“给出修复建议”的一部分。但声明中的关键能力“结合源码定位”“排查数据库”在代码中完全没有实现;代码没有搜索本地源码、没有读取仓库文件、没有调用数据库客户端或执行数据库查询。由于声明把这些步骤写成完整流程中的核心环节,而非可选描述,因此属于实质性能力不符。此外,代码实际还支持以 wusid/path 为入口检索日志并交互选择 trace_id,这比声明中的 trace_id 查询更宽,但这点是扩展而非主要问题。综合来看,声明显著高估了代码覆盖范围,因此应判定为不匹配。

Content

No source excerpt is available for this finding.

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · SKILL.md (reported line 3)May include surrounding context.

md
---
name: sls-trace-analysis
description: >
  查询阿里云SLS日志和ARMS调用链,结合源码和数据库进行全链路问题排查。
  完整流程:查日志 → 画调用链 → 定位源码 → 排查数据库 → 给出修复方案。
  Use when: 用户说「分析sls」「分析问题」或想排查业务服务/线上接口/用户请求的报错或异常。
  触发示例:「分析sls」「帮我查一下这个trace_id」「分析一下这个trace_id」
  「查一下这个用户的请求」「wusid 是 xxx」「uid xxx」
  「查一下 /path/to/api 这个接口的报错」「帮我排查一下这个�

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill mandates complete disclosure of all logs for each logstore and forbids deduplication or filtering. In a log-analysis context, this is dangerous because raw logs commonly contain secrets, tokens, personal data, internal identifiers, request payloads, and error details that should be minimized or redacted before being shown to a user.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The report template explicitly instructs inclusion of raw attach/request/response fields from logs. Those fields frequently contain user identifiers, authentication material, request bodies, business data, or internal responses, so rendering them directly in natural-language output creates a strong data exfiltration path.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 559)May include surrounding context.

md
**定位文件:** `{file_path}:{line_number}`

**关键代码段:**
​```{language}
// 标注问题所在行
{code snippet with comments}
​```

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 562)May include surrounding context.

md
**定位文件:** `{file_path}:{line_number}`

**关键代码段:**
​```{language}
// 标注问题所在行
{code snippet with comments}
​```

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 620)May include surrounding context.

md
**定位文件:** `{file_path}:{line_number}`

**关键代码段:**
​```{language}
// 标注问题所在行
{code snippet with comments}
​```

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 622)May include surrounding context.

md
**定位文件:** `{file_path}:{line_number}`

**关键代码段:**
​```{language}
// 标注问题所在行
{code snippet with comments}
​```

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README advertises access to SLS logs, ARMS traces, code repositories, databases, and cloud credentials without any warning about privacy, secrets exposure, or least-privilege handling. In this context, the skill operates across multiple high-sensitivity systems, so missing disclosure and safety guidance materially raises the risk of exposing personal data, tokens, SQL contents, stack traces, or proprietary code during normal use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger phrases are broad enough to match routine troubleshooting requests such as '分析问题' or 'SQL超时', which can cause the skill to activate in contexts the user did not explicitly intend. Because this skill can query cloud logs, traces, source code, and database-related data, overbroad activation increases the chance of unnecessary access to sensitive operational information.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill requires environment credentials, reads local files, and accesses remote cloud services, but it declares no explicit tool scope or permission boundaries. In an agent environment, that increases the chance of over-broad invocation and unauthorized access to sensitive logs, source code, or cloud data because operators and policy layers cannot clearly constrain what the skill may do.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill description does not clearly warn users that it may access cloud logs, tracing systems, local source code, and potentially sensitive request or user data. Without transparent notice and consent, users may unknowingly authorize retrieval and display of privileged operational information, increasing privacy and confidentiality risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to match ordinary troubleshooting requests, which can cause the agent to invoke a powerful skill in contexts where the user did not clearly consent to cloud-log, trace, or source-code inspection. Because this skill can access sensitive operational and user data, over-triggering materially increases the risk of unintended data exposure.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill directs the agent to use request/response business data from logs to reconstruct user context and triggering conditions. Even if intended for debugging, that practice can expose private user information and sensitive business context beyond what is necessary for root-cause analysis.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file-level documentation and manifest describe a full workflow including locating source code, investigating the database, and giving a repair plan. The actual code queries SLS and ARMS, formats traces/logs, and generates heuristic suggestions, but it contains no code access, repository lookup, database connection, or SQL inspection capability.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script reads ~/.openclaw/openclaw.json and uses its env section as a fallback source for credentials and configuration. That reaches into local OpenClaw internals despite the skill contract saying it is not for OpenClaw internal state, and it expands the trust boundary to a sensitive local file that may contain secrets not intended for this skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill queries production logs and tracing systems, then emits raw log messages, span metadata, exception IDs, and stack traces into its output without any minimization, masking, or user-facing warning. In this context, those artifacts can contain PII, tokens, internal hostnames, SQL fragments, or other secrets, so unrestricted retrieval and display increases data-exposure risk.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · LICENSE.md (reported line 12)May include surrounding context.

md
permit persons to whom the Software is furnished to do so.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED,
INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A
PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT
HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION
OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The natural-language interface, trigger phrases, and examples are presented entirely in Chinese, which suggests the skill may expect or operate in a single language by default. There is no statement that Chinese is optional, user-selectable, or required for a justified region-specific use case.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script reads access keys and secrets from environment variables and from ~/.openclaw/openclaw.json as a fallback source. While this is functionally expected for cloud API access, there is no visible warning in code comments, CLI help, or output that local credentials will be read from these locations.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/query_trace.py (reported line 716)May include surrounding context.

python
def get(obj, key):
        if use_sdk:
            return getattr(obj, _snake(key), None)
        return obj.get(key)

    def _snake(s):

Static analysis

No suspicious patterns detected.