Back to skill

Security audit

Security

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed remote safety-check wrapper that sends operation details to a backend but does not itself execute, persist, or mutate the checked operations.

Use this skill only when you are comfortable sending the instruction text plus any context and target identifiers, such as file paths, service names, database names, or URLs, to Modeio's safety backend or the SAFETY_API_URL you configure. Do not include secrets unless your organization allows that backend to receive them, and treat the backend result as guidance that the caller still must enforce.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description promises a substantive security/safety guardrail capability that evaluates instructions for risky operations. The supplied code chunk does not implement that functionality; it merely defines an empty package initializer with a generic docstring. This is a material mismatch in primary purpose and implemented capability.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

Core commands

bash
python3 scripts/safety.py -i "Delete /tmp/cache/build-123.log" \
  -c '{"environment":"local-dev","operation_intent":"cleanup","scope":"single-resource","data_sensitivity":"internal","rollback":"easy","change_control":"none"}' \
  -t "/tmp/cache/build-123.log" --json

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 39)May include surrounding context.

python
class TestSafetyContract(unittest.TestCase):
    def _run_cli(self, args, env=None):
        merged_env = os.environ.copy()
        if env:
            merged_env.update(env)
        return subprocess.run(

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 263)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 263)May include surrounding context.

python
"risk_level": "high",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /", "--json"])

        self.assertEqual(code, 1)
        payload = json.loads(stdout)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documentation describes execution of a Python script that depends on network access and likely shell invocation, but the manifest does not declare any explicit tool scope such as permissions or allowed-tools. That mismatch weakens reviewability and policy enforcement because callers may not realize the skill can make external requests or invoke code with side effects.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The code performs an outbound HTTP POST to a remote API with the supplied instruction payload, creating an external data exfiltration path by design. While this is intended functionality for a backend-backed safety service, it is still a real security/privacy concern because the transmitted content may include sensitive operational details and the endpoint is environment-variable overridable.

Content

Scanner excerpt · modeio_guardrail/cli/safety.py (reported line 88)May include surrounding context.

python
last_exc = None
    for attempt in range(1 + MAX_RETRIES):
        try:
            resp = requests.post(url, json=json_payload, timeout=timeout)
            if resp.status_code in (502, 503, 504) and attempt < MAX_RETRIES:
                time.sleep(RETRY_BACKOFF * (2 ** attempt))
                continue

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The tool sends user-provided instruction text, and optionally context and target metadata, to a remote backend for analysis, but the CLI does not provide a clear user-facing disclosure or consent step before transmitting that potentially sensitive content. In a safety-check skill, this is especially relevant because users may submit operational plans, file paths, service names, or compliance-sensitive instructions that can contain confidential data.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 42)May include surrounding context.

python
merged_env = os.environ.copy()
        if env:
            merged_env.update(env)
        return subprocess.run(
            [sys.executable, str(SCRIPT_PATH)] + args,
            capture_output=True,
            text=True,

Static analysis

No suspicious patterns detected.