Back to skill

Security audit

Security

Security checks for vulnerabilities and agentic risk

Overview

This is a real safety-checking CLI, but it sends operation details to a configurable external backend that can be redirected and trusted for decisions, so it needs review before use.

Review before installing in workflows that gate destructive or production actions. Keep SAFETY_API_URL unset or locked to a trusted HTTPS endpoint, control the process environment, avoid sending secrets in instruction/context/target fields, and treat the returned decision as advisory unless your caller enforces fail-closed policy.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
modeio_guardrail/cli/safety.py:52
Finding

Unrestricted Backend Override Permits Sensitive Data Exfiltration and Forged Safety Decisions

Content
View full analysis

Vulnerability Details

File Location: modeio_guardrail/cli/safety.py:52-53, 111-118
Vulnerability Type: Unrestricted external endpoint configuration and cleartext transport
Risk Level: Medium

Vulnerable Code:

python
# Backend API URL, overridable via SAFETY_API_URL environment variable
URL = os.environ.get("SAFETY_API_URL", "https://safety-cf.modeio.ai/api/cf/safety")
python
def detect_safety(instruction: str, context: str = None, target: str = None) -> dict:
    """
    Call the Modeio safety backend and return the full response JSON.
    Response includes: approved, risk_level, risk_types, concerns, recommendation, etc.
    """
    payload = {"instruction": instruction}
    if context:
        payload["context"] = context
    if target:
        payload["target"] = target
    resp = _post_with_retry(URL, json_payload=payload)
    return resp.json()

The HTTP request is issued without endpoint validation at modeio_guardrail/cli/safety.py:88:

python
resp = requests.post(url, json=json_payload, timeout=timeout)

Technical Analysis

The SAFETY_API_URL environment variable can replace the trusted safety backend with an arbitrary URL. The implementation does not enforce HTTPS, validate the hostname, restrict ports, or maintain an allowlist of trusted endpoints. The test at tests/test_safety_contract.py:92-96 confirms that a plain HTTP URL is accepted:

python
result = self._run_cli(
    ["--input", "Delete all log files in production", "--json"],
    env={"SAFETY_API_URL": "http://127.0.0.1:9"},
)

Requests may contain instruction text, operational context, file paths, database identifiers, service names, URLs, data-sensitivity classifications, and change-control information. If an attacker can control the process environment, these values can be redirected to an attacker-controlled server or exposed through unencrypted HTTP.

The clie ...[truncated 2187 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require the configured URL to use https and reject cleartext HTTP or other schemes.
  2. Allowlist the production safety backend hostname and expected port. If custom endpoints are needed for development, require an explicit development-only option.
  3. Reject URLs containing embedded credentials, fragments, unexpected ports, or untrusted redirect destinations.
  4. Disable redirects or validate every redirect target against the same scheme and hostname policy.
  5. Consider certificate or public-key pinning where the operational environment supports it.
  6. Validate backend responses against a strict schema. Require approved to be a Boolean and risk_level to be one of the documented values; reject unknown, missing, or malformed decision fields.
  7. Minimize transmitted data and redact secrets, credentials, tokens, sensitive query parameters, and unnecessary resource details before submission.
  8. Document that endpoint overrides are security-sensitive and ensure service managers, CI systems, and wrappers prevent untrusted users from modifying the Skill's environment.
  9. Add tests proving that non-HTTPS URLs, unapproved hosts, malformed responses, and redirects to unapproved hosts are rejected.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Yes, this is a mismatch. The declared description claims a concrete safety-checking capability with specific operational scope, but the provided code does not implement any such behavior. It is effectively an empty package initializer with only a docstring. While absence of implementation could be due to an incomplete code sample, based on the supplied chunk alone the actual behavior does not substantiate the declared purpose.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

Core commands

bash
python3 scripts/safety.py -i "Delete /tmp/cache/build-123.log" \
  -c '{"environment":"local-dev","operation_intent":"cleanup","scope":"single-resource","data_sensitivity":"internal","rollback":"easy","change_control":"none"}' \
  -t "/tmp/cache/build-123.log" --json

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 39)May include surrounding context.

python
class TestSafetyContract(unittest.TestCase):
    def _run_cli(self, args, env=None):
        merged_env = os.environ.copy()
        if env:
            merged_env.update(env)
        return subprocess.run(

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 263)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
100% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 246)May include surrounding context.

python
"recommendation": "review",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /tmp/cache", "--json"])

        self.assertEqual(code, 0)
        payload = json.loads(stdout)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 263)May include surrounding context.

python
"risk_level": "high",
            },
        ):
            code, stdout, _ = self._run_main(["--input", "rm -rf /", "--json"])

        self.assertEqual(code, 1)
        payload = json.loads(stdout)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill advertises and instructs use of a Python script that depends on network access and can process potentially dangerous operational instructions, but the manifest does not declare an explicit tool scope such as permissions or allowed-tools. That ambiguity can cause callers or hosting platforms to misjudge what capabilities the skill may exercise, weakening least-privilege controls and auditability.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · modeio_guardrail/cli/safety.py (reported line 88)May include surrounding context.

python
last_exc = None
    for attempt in range(1 + MAX_RETRIES):
        try:
            resp = requests.post(url, json=json_payload, timeout=timeout)
            if resp.status_code in (502, 503, 504) and attempt < MAX_RETRIES:
                time.sleep(RETRY_BACKOFF * (2 ** attempt))
                continue

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_safety_contract.py (reported line 42)May include surrounding context.

python
merged_env = os.environ.copy()
        if env:
            merged_env.update(env)
        return subprocess.run(
            [sys.executable, str(SCRIPT_PATH)] + args,
            capture_output=True,
            text=True,

Static analysis

No suspicious patterns detected.