Back to skill

Security audit

Supervision Layer

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local supervision helper, but it overstates its enforcement guarantees and can persist sensitive audit and crash data in plaintext.

Review this as a prototype or library helper, not a guaranteed non-bypassable security boundary. Install only if you are comfortable explicitly integrating every tool call through it and accepting local plaintext logs; avoid logging secrets or private prompts until redaction, retention, and file-permission controls are improved.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/supervision/audit.py:82
Finding

Insufficient redaction and insecure storage of potentially sensitive audit data

Content
View full analysis
Optional[str]: """Truncate args for audit log, stripping PII-prone values.""" if args is None: return None s = str(args) if len(s) > self.max_args_length: return s[:self.max_args_length] + "..." return s ``` ```python entry = AuditEntry( timestamp=time.time(), request_id=str(uuid.uuid4()), session_id=session_id, agent_id=agent_id, tool=tool, phase=phase, outcome=outcome, duration_ms=round(duration_ms, 2), tokens_used=tokens_used, estimated_cost_usd=round(estimated_cost_usd, 6), error=error, args_summary=self._truncate_args(args), metadata=metadata or {}, ) try: if self._should_rotate(): self._rotate() with open(self.audit_file, "a") as f: f.write(entry.to_json() + "\n") except IOError as e: logger.error(f"Failed to write audit log: {e}") ``` ```python crash_record = { "timestamp": now, "session_id": session_id, "task_name": task_name, "error": error, } state.crashes.append(crash_record) ``` ```python try: with open(self.state_file, "w") as f: json.dump(data, f, indent=2) except IOError as e: logger.error(f"Failed to save crash state: {e}") ``` ### Technical Analysis The audit logger states that argument processing strips PII-prone values, but `_truncate_args` only converts the supplied value to a string and limits its length. Truncation is not sanitization or redaction. Sensitive values appearing near the beginning of an argument—such as API keys, authorization headers, passwords, cookies, personal information, prompt content, or private file paths—remain i ...[truncated 2575 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the artifact mainly provides tests or examples rather than an actual supervision layer, then its stated purpose is materially misleading. While not necessarily malicious, this can still cause unsafe deployment decisions because consumers may assume the presence of non-bypassable protective controls that do not exist in the shipped skill.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the artifact mainly provides tests or examples rather than an actual supervision layer, then its stated purpose is materially misleading. While not necessarily malicious, this can still cause unsafe deployment decisions because consumers may assume the presence of non-bypassable protective controls that do not exist in the shipped skill.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the artifact mainly provides tests or examples rather than an actual supervision layer, then its stated purpose is materially misleading. While not necessarily malicious, this can still cause unsafe deployment decisions because consumers may assume the presence of non-bypassable protective controls that do not exist in the shipped skill.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the artifact mainly provides tests or examples rather than an actual supervision layer, then its stated purpose is materially misleading. While not necessarily malicious, this can still cause unsafe deployment decisions because consumers may assume the presence of non-bypassable protective controls that do not exist in the shipped skill.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the artifact mainly provides tests or examples rather than an actual supervision layer, then its stated purpose is materially misleading. While not necessarily malicious, this can still cause unsafe deployment decisions because consumers may assume the presence of non-bypassable protective controls that do not exist in the shipped skill.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The logic sets permanently_failed to true once the threshold is reached, but pruning old crashes never clears that flag when the crash window expires. As a result, an agent can become permanently disabled after transient failures, creating a persistent denial-of-service condition that contradicts the documented time-window semantics and can block recovery indefinitely.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The manifest declares no explicit tool scope even though the skill content implies capabilities involving environment access and file operations. In a security-sensitive agent ecosystem, missing scope declarations can cause the platform or reviewers to underestimate what the skill may access, enabling broader-than-expected file or environment interaction.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrase is broad and can activate the skill in loosely related contexts, increasing the chance of unintended invocation. For a skill that claims authority over supervision and enforcement, over-broad activation could interfere with unrelated workflows or cause users to trust protections that were not intentionally selected for that task.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code claims to produce an audit-safe summary by 'stripping PII-prone values', but the implementation simply converts the full arguments object to a string and truncates it. This can still log secrets, tokens, personal data, file paths, prompts, or other sensitive inputs in cleartext, especially when the sensitive fields appear near the beginning of the argument string. In a supervision/audit layer that wraps every tool call, this creates a broad and centralized data-exposure surface.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This is a real logic flaw in a security enforcement component. In HALF_OPEN, can_call() returns True for every caller and never tracks or enforces half_open_max_calls, so multiple concurrent or repeated calls can pass during the supposed probe phase, defeating the containment property of the circuit breaker and allowing a failing tool to continue consuming resources or causing repeated side effects.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The component persists crash records, including session_id, task_name, and error text, to a JSON file on disk even though crash-loop enforcement only requires bounded counters and timestamps. Those fields can contain sensitive operational or user-derived data, and storing them under a default workspace path increases the risk of unnecessary retention, disclosure, or later misuse if the file is read by other local processes or included in logs/backups.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill metadata promises an enforcement layer that wraps every tool call and cannot be bypassed, with audit logging and crash-loop protection, but this file only provides an optional timeout helper plus in-memory statistics. If operators rely on this module as a security boundary, agents or calling code can simply avoid using the wrapper, and there is no audit trail or crash-loop enforcement to detect or stop abuse.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · scripts/test_supervision.py (reported line 80)May include surrounding context.

python
assert result.succeeded is True
        
        # Should timeout with 0.05s timeout
        result2 = await wrapper.call("medium_tool", medium_fn, timeout=0.05)
        assert result2.succeeded is False
        assert result2.timed_out is True

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code performs pre- and post-execution audit logging of tool name, agent ID, session ID, optional argument summaries, task name, duration, and error details. While the module docstring mentions audit logging, there is no explicit user-facing warning, confirmation, or privacy disclosure about collecting and persisting this execution metadata, which is relevant because logging can affect user data visibility.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.