Back to skill

Security audit

Acceptance Test

Security checks for vulnerabilities and agentic risk

Overview

This skill looks like an acceptance-testing tool, but its scripts generate random demo results instead of testing the user's project.

Do not use this skill as evidence that software was tested or accepted. It appears safe from a system-access perspective, but its outputs are unreliable for release approval, customer sign-off, compliance, or quality gates unless the publisher replaces the demo/random logic with real deterministic validation.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/acceptance_runner.py:14
Finding

Acceptance and release decisions are fabricated using random values

Content
View full analysis
Dict: """执行验收测试""" import random stories_data = [ {'id': 'US-001', 'name': '用户注册', 'passed': random.random() > 0.1}, {'id': 'US-002', 'name': '用户登录', 'passed': random.random() > 0.1}, {'id': 'US-003', 'name': '商品购买', 'passed': random.random() > 0.1}, ] passed = sum(1 for s in stories_data if s['passed']) return { 'environment': env, 'stories': stories_data, 'total': len(stories_data), 'passed': passed, 'failed': len(stories_data) - passed, 'pass_rate': round(passed / len(stories_data) * 100, 1), 'approved': passed == len(stories_data) } ``` `scripts/bdd_runner.py:13-28` ```python def run_bdd(features: str, steps: str) -> Dict: """执行BDD测试""" import random scenarios = [ {'name': '成功登录', 'given': '用户已注册', 'when': '输入正确密码', 'then': '登录成功', 'passed': random.random() > 0.1}, {'name': '失败登录', 'given': '用户已注册', 'when': '输入错误密码', 'then': '显示错误', 'passed': random.random() > 0.1}, ] passed = sum(1 for s in scenarios if s['passed']) return { 'features_path': features, 'steps_path': steps, 'scenarios': scenarios, 'passed': passed, 'failed': len(scenarios) - passed } ``` `scripts/acceptance_report.py:12-23` ```python def generate_acceptance_report(results: str, signoff: str) -> Dict: """生成验收报告""" import random return { 'results_file': results, 'signoff_required': signoff == 'required', 'overall_status': 'PASSED' if ...[truncated 2576 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/acceptance_runner.py:49
Finding

Documented command-line inputs are silently ignored and fixed demonstration data is used

Content
View full analysis
0.2 else 'CONDITIONAL', 'risk_level': random.choice(['LOW', 'MEDIUM', 'HIGH']), 'blockers': random.randint(0, 3), 'recommendations': ['完善文档', '优化性能'], 'report_path': 'reports/acceptance-report.pdf' } ``` ### Technical Analysis The documented interface advertises six command-line options: `--stories`, `--env`, `--features`, `--steps`, `--results`, and `--signoff`. None of the three scripts implements command-line argument parsing. When invoked as docu ...[truncated 1922 chars]
Remediation
View remediation
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

描述将该技能定位为完整的验收测试助手,覆盖验证、测试、检查和客户验收支持等能力;但提供的代码只是在本地构造一份报告对象并打印,且关键字段(overall_status、risk_level、blockers)是随机生成的。虽然“验收报告生成”这一子能力与声明部分一致,但代码的实际主要行为远比声明范围窄,且不具备声明中的核心验证/测试功能,因此属于明显的描述与行为不一致。

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill metadata and function naming present this code as a business-requirement verification and acceptance reporting tool, but the implementation performs no real validation and returns randomized conclusions. Because this skill is explicitly meant to support customer acceptance and delivery decisions, the mismatch is especially dangerous: users may trust fabricated outputs as evidence of quality or compliance.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The function claims to generate an acceptance report but fabricates core assurance fields such as overall status, risk level, and blockers using random values instead of evaluating actual test evidence. In an acceptance-testing context, this can mislead stakeholders into approving or rejecting delivery based on false data, undermining governance and potentially hiding serious defects.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The implementation does not fulfill the skill's declared purpose of validating business requirements and user scenarios; it only emits simulated acceptance outcomes. Because the skill is presented as an acceptance-test assistant, this mismatch increases the risk of false assurance, audit failure, and unsafe release decisions based on non-evidence.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The function claims to run acceptance tests but instead generates pass/fail outcomes with random values, which can falsely indicate that business requirements were validated when they were not. In an acceptance-testing skill, this is especially dangerous because stakeholders may approve releases based on fabricated evidence, leading to deployment of untested or noncompliant functionality.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This skill is positioned as an acceptance-test assistant for business requirement validation, yet its core logic returns random pass/fail outcomes unrelated to the system under test. In this context, false assurance is especially dangerous because acceptance results are often used for release approval, customer signoff, and compliance evidence.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description and primary instructional content are written in Chinese, and the file does not indicate that users may choose another language or that the skill is intentionally restricted to a Chinese-language context. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file’s natural-language strings and docstrings are entirely in Chinese, including the function description and all user-facing output, with no indication that language is configurable or that the skill is region-specific. This can violate a language/locale policy when users are not given an opt-in or alternative language choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This Python file contains natural-language strings and documentation exclusively in Chinese, including the module docstring and all user-facing console output. For a general-purpose acceptance test runner, this imposes a specific language/locale without user opt-in or any documented regional justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module and function claim to execute BDD tests, but the implementation fabricates results using randomness instead of running feature files and step definitions. In an acceptance-testing skill, this can mislead users into believing requirements were validated when no verification occurred, undermining quality gates and enabling unsafe or noncompliant releases.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The docstring and all user-facing print strings are written in Chinese, and the file provides no option to select another language or indicate that the skill is intentionally region-specific. This creates a natural-language locale policy concern because the skill implicitly enforces a single language without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.