Back to skill

Security audit

Quick Test

Security checks for vulnerabilities and agentic risk

Overview

This skill is a simple diagnostic runner, but its default test can expose environment variables and its documentation promotes broad command execution without enough safety boundaries.

Review before installing. Run it only in a low-sensitivity environment, avoid using the environment-variable examples, and do not invoke it from shells or CI jobs that contain tokens or credentials. A safer version should remove the default env check, use shell=False, redact or avoid sensitive output, and replace broad command support with an explicit diagnostic allowlist.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/run_tests.py:10
Finding

Environment Variable Disclosure Through Unsafe Shell Invocation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The stated purpose is a simple environment verification tool, but the described behavior extends to environment inspection, directory listing, and file writes. That mismatch is dangerous because users may authorize a benign-seeming diagnostic skill that actually exposes sensitive system context or modifies the filesystem.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

Using shell=True in a tool-execution helper is a classic tool parameter abuse issue because it delegates argument interpretation to the shell. In a reusable helper that runs multiple commands, this creates a dangerous primitive that can be repurposed for command injection or shell metacharacter abuse, especially in future edits or if command inputs ever become dynamic.

Content

Scanner excerpt · scripts/run_tests.py (reported line 14)May include surrounding context.

python
def run_command(cmd, description=""):
    """Run a command and return result."""
    try:
        result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=30)
        return {
            'command': ' '.join(cmd),
            'description': description,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises shell-command execution but does not declare any explicit tool scope or allowed-tools restrictions. In an agent ecosystem, undocumented shell capability increases the chance the skill is invoked with broader execution power than users or orchestrators expect, weakening policy enforcement and review.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill description is broad enough to match many generic debugging or troubleshooting prompts, increasing the chance it is selected in contexts where shell execution is unnecessary. Overbroad invocation language is risky because it can cause accidental activation of a powerful skill for low-risk tasks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The markdown promotes custom command execution and system inspection without warning about privacy, credential exposure, or system-modification risks. Users may reasonably assume diagnostic commands are harmless, leading them to run commands that disclose secrets or affect system state.

Content

No source excerpt is available for this finding.

Unrestricted Tool Access

Medium
Category
Excessive Agency
Confidence
96% confidence
Finding

Advertising that the skill can 'Run any command' indicates effectively unrestricted shell access. In practice, that grants the skill the ability to execute destructive, persistence-establishing, or data-exfiltrating commands far beyond a simple verification workflow.

Content

Scanner excerpt · SKILL.md (reported line 105)May include surrounding context.

md
- ✅ **Python Availability Check** / **Verificação de Disponibilidade do Python** - Confirms Python 3.x installed
- ✅ **System Command Execution** / **Execução de Comando do Sistema** - Runs and validates system commands
- ✅ **File System Access** / **Acesso ao Sistema de Arquivos** - Verifies directory access and permissions
- ✅ **Custom Command Support** / **Suporte a Comandos Customizados** - Run any command with validation
- ✅ **Working Directory Check** / **Verificação de Diretório de Trabalho** - Confirms current location
- 📝 **Detailed Logging** / **Log Detalhado** - Comprehensive output for debugging

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly supports arbitrary custom shell commands, which turns a health-check utility into a general command-execution wrapper. This is dangerous because it can be used to read secrets, alter files, install software, or pivot into broader host compromise under the guise of troubleshooting.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation encourages environment-reading commands like env | head -10, which normalizes exposure of environment variables during routine testing. Environment variables frequently contain tokens, API keys, hostnames, and internal configuration, so even partial output can leak sensitive information.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The code invokes subprocess.run with shell=True while passing commands that are represented as argument lists, which is an unsafe and error-prone pattern. In an agent skill context, shell execution increases the risk of command injection, unexpected shell parsing, and unintended execution if any command components ever become user-influenced or are modified later.

Content

Scanner excerpt · scripts/run_tests.py (reported line 14)May include surrounding context.

python
def run_command(cmd, description=""):
    """Run a command and return result."""
    try:
        result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=30)
        return {
            'command': ' '.join(cmd),
            'description': description,

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill presents itself as a simple system verification tool, but it also enumerates environment variables and workspace contents, which expands its data access beyond what users would reasonably expect. This mismatch increases the chance of unauthorized disclosure of sensitive operational data during a supposedly harmless test run.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Reading environment variables during a quick test can expose secrets such as tokens, credentials, service endpoints, or internal configuration. Even though the command appears intended to limit output, the implementation is flawed and still represents unnecessary sensitive-data access for the stated purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The environment inspection lacks any warning that command output may contain sensitive information. In an agent environment, even a partial dump of environment variables can leak secrets into logs, transcripts, or downstream systems, making this more dangerous than a normal local diagnostic script.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script performs a filesystem write to /tmp as part of a quick verification flow without clearly disclosing that side effect. Unexpected writes can violate least surprise, interfere with shared environments, and create persistence artifacts that users did not authorize.

Content

No source excerpt is available for this finding.

Excessive Permissions

Low
Category
Privilege Escalation
Confidence
80% confidence
Finding

Skill requests more permissions than appear necessary for its stated functionality. Review if elevated access is justified.

Content

Scanner excerpt · SKILL.md (reported line 266)May include surrounding context.

md
## Limitations / Limitações

**User Permissions:** Requires read and execute access to directories
- **Permissões do Usuário:** Requer acesso de leitura e execução a diretórios
- System commands must be in PATH / Comandos do sistema devem estar no PATH

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file write test has a side effect that is not explicitly disclosed to the user, which is a policy and safety concern even if the write target is only /tmp. In an automation context, undisclosed writes can surprise users and complicate forensic or compliance expectations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.