Back to skill

Security audit

社区垃圾过滤(免费版)

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to be a community spam-filter helper, but it asks agents to run an unbundled local script with broad read/write/execute authority and nearby API credentials.

Review before installing. Only use this in a sandboxed project where you control filter.js, can inspect the script before running it, and can provide a minimally scoped API key. Avoid using callback_url unless you know exactly what endpoint receives the data and what payload is sent.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:14
Finding

Overprivileged Tool Permissions Without a Verifiable Implementation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 14-17; credential and execution workflow at lines 91-95
Vulnerability Type: T05: Unauthorized Access and Privilege Escalation
Risk Level: Medium

Vulnerable Code

yaml
tools:
  - read
  - exec
  - write

The workflow subsequently instructs the agent to verify a credential file, execute filter.js, and edit its isSpam() function. However, the audited package contains only SKILL.md; the referenced script is not included.

Technical Analysis

The Skill declares unrestricted file-reading, file-writing, and command-execution capabilities. These privileges exceed what can be justified from the package's auditable contents because no executable filtering implementation is supplied.

The documented workflow also directs the agent toward ./.config/platform/credentials.json, which is expected to contain a valid API key. Because filter.js is absent, reviewers cannot verify how that script would use, transmit, store, or log the credential. If an unrelated or attacker-controlled filter.js is present in the working directory, the documented command could execute it with the agent's ambient permissions and access to sensitive credentials.

This is a least-privilege failure and an unsafe trust-boundary design. The audit found no explicit instruction to disclose the API key and no bundled malicious script, so the evidence supports an overprivileged and suspicious configuration rather than confirmed credential exfiltration.

Attack Path

  1. An agent loads the Skill and grants its declared read, write, and exec capabilities.
  2. Following the documented workflow, the agent checks the local credential file containing the platform API key.
  3. The agent invokes node filter.js scan ... or node filter.js feed ....
  4. Because the audited package does not supply filter.js, command resolution depends on an externally supplied file in the cu ...[truncated 916 chars]
Remediation
View remediation

Remediation Suggestions

  1. Include the referenced filter.js implementation in the package so its command execution, network destinations, input handling, and credential handling can be audited.
  2. Remove write permission unless the Skill has a concrete, documented need to modify files. Rule customization should preferably occur through a constrained configuration file rather than arbitrary source-code editing.
  3. Restrict read permission to explicit feed inputs and configuration paths. Do not instruct the agent to open or display raw credential files.
  4. Obtain credentials through a platform secret provider or a narrowly scoped environment variable and ensure they are never printed, logged, embedded in command arguments, or written to generated output.
  5. Restrict execution to a packaged, integrity-verified script using an absolute or package-relative path rather than resolving filter.js from an uncontrolled working directory.
  6. Document and allowlist the exact remote API host, permitted HTTP methods, and required API-key scope.
  7. Validate subcommunity names and other user-controlled arguments before passing them to commands or network requests.
  8. Run the implementation in a sandbox with a read-only project filesystem, limited network access, no access to unrelated user files, and a minimally privileged API key.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (4)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 90)May include surrounding context.

md
不适用于:服务端垃圾防御、账户封禁、内容举报、ML垃圾检测.
## 使用流程

1. 确认凭证文件 `./.config/platform/credentials.json` 存在且API key有效
2. 确认 Node.js 运行时已安装
3. 用 `node filter.js scan [submolt]` 扫描目标子板块,查看垃圾率
4. 用 `node filter.js feed [submolt]` 获取过滤后JSON,管道到下游工具

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description repeatedly says the skill is suitable for general 'related development scenarios' and '需要moltbook filter相关能力的开发场景' without defining precise invocation phrases, scope boundaries, or exclusion conditions. In a manifest/markdown context, this is ambiguous enough to risk unintended activation because it does not clearly distinguish when the skill should or should not be invoked.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill repeatedly describes itself as a client-side community spam filter focused on local scanning and JSON feed filtering, and later says it is '客户端only' with '无服务端共享状态'. However, the documented callback_url parameter implies outbound notification or remote callback behavior, which contradicts the stated local-only/no-server-side model.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The documented callback_url parameter creates an outbound data flow to an external endpoint without any warning about what results may be transmitted. If users provide sensitive input, filtered content, or metadata, the skill could cause unintentional exfiltration to third-party infrastructure.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.