Back to skill

Security audit

Council v2

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly coherent, but its JSON mode can execute crafted Python code and it under-discloses sensitive data sent to external model providers.

Review this carefully before installing. Do not use `--format json` with untrusted input until the heredoc serialization is fixed. Avoid submitting secrets, proprietary code, customer data, or security-sensitive plans unless every configured model provider is approved for that data. Run reviewer sessions with minimal read-only context and require stricter validation of reviewer JSON before trusting synthesis results.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/council.sh:118
Finding

Arbitrary Python Code Execution Through Unsafe Heredoc Interpolation

Content
View full analysis
' target.txt --format json ``` 4 ...[truncated 1158 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
scripts/council.sh:98
Finding

Prompt Injection Through Untrusted Review Content

Content
View full analysis
5. On Full Council, present raw reviewer outputs along with synthesis. 6. Synthesizer narrates result but does not override vote count. Decision options: ${OPTIONS:-N/A} Content under review: --- $CONTENT --- EOF ) ``` ### Technical Analysis The Skill inserts arbitrary reviewed content directly into an orchestration prompt and directs the orchestrator to give that content to every spawned reviewer. The prompt does not explicitly state that the reviewed content is untrusted data, that instructions found inside it must not be followed, or that it must not alter the reviewer role and output contract. Markdown delimiters such as `---` provide visual separation but do not establish a security boundary for a language model. A malicious source file, plan, or architecture document can therefore contain instructions such as requests to disregard the reviewer role, return a predetermined verdict, omit findings, reveal contextual information, or invoke tools. The role prompts require JSON output, but that requirement alone does not prevent an injected document from directing a reviewer to produce attacker-selected JSON that still conforms to the expected shape. ### Attack Path 1. An attacker adds hidden or visible natural-language instructions to a file that will be submitted for council review. 2. ...[truncated 1617 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/synthesize.py:39
Finding

Incomplete Runtime Validation of Reviewer-Controlled JSON

Content
View full analysis
dict: required = ["reviewer", "model", "verdict", "confidence", "findings", "summary"] for field in required: if field not in review: raise ValueError(f"missing field: {field}") if review["verdict"] not in VERDICT_POINTS: raise ValueError(f"invalid verdict: {review['verdict']}") if not isinstance(review["findings"], list): raise ValueError("findings must be a list") return review ``` ### Technical Analysis The project documents a restrictive JSON Schema in `references/schema.md`, but the runtime validator does not enforce that schema. The implementation validates only: - Presence of six top-level fields - Membership of `verdict` in `VERDICT_POINTS` - Whether `findings` is a list It does not validate: - That the top-level input is an object - Types of `reviewer`, `model`, `summary`, or `confidence` - The documented confidence range of `0.0` through `1.0` - That each finding is an object - Required finding fields - Finding field types - The allowed `critical`, `warning`, and `note` severity values - Additional properties prohibited by the documented schema - Maximum field lengths or maximum numbers of reviews and findings - Whether the review set is empty, complete, duplicated, or from expected reviewer identities Downstream functions assume stronger types. For example, `minority_report` converts confidence with `float(...)`, `derive_conditions` calls `.strip()` on recommendations, and `strongest_finding` calls `.get()` on every finding. Malformed but initially accepted JSON can therefore produce synthesis errors. More importantly, the synthesis mechanism trusts reviewer-supplied severity values. Any accepted review containing a `critical` finding causes an automatic block, witho ...[truncated 1515 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (6)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · README.md (reported line 85)May include surrounding context.

md
|------|---------|
| `SKILL.md` | Operational skill — workflow, when to use, interpreting results |
| `references/review-types.md` | Review type definitions and tier recommendations |
| `references/role-prompts.md` | Reviewer role prompts and shared output instructions |
| `references/schema.md` | JSON schemas for reviewer and synthesis output |
| `references/synthesis-rules.md` | Mechanical synthesis protocol and edge cases |
| `scripts/council.sh` | Orchestration script |

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 193)May include surrounding context.

md
|------|---------|
| `SKILL.md` | Operational skill — workflow, when to use, interpreting results |
| `references/review-types.md` | Review type definitions and tier recommendations |
| `references/role-prompts.md` | Reviewer role prompts and shared output instructions |
| `references/schema.md` | JSON schemas for reviewer and synthesis output |
| `references/synthesis-rules.md` | Mechanical synthesis protocol and edge cases |
| `scripts/council.sh` | Orchestration script |

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims to provide multi-model review, orchestration, and synthesis, but the static finding indicates the actual behavior diverges substantially and instead performs local filesystem enumeration and retrospective scaffolding. This kind of description-behavior mismatch is dangerous because users may authorize or trust the skill for one purpose while it performs materially different actions, enabling unexpected data exposure and undermining informed consent.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · references/role-prompts.md (reported line 16)May include surrounding context.

md
- Override via command line: `--opus claude-sonnet-4 --gpt gpt-4.1` etc.
- If a provider is unavailable, drop that reviewer rather than doubling up on another provider — a 4-of-5 diverse council beats 5-of-5 with duplicate bias

## Shared output instruction block

Use this block in every reviewer prompt:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README explicitly promotes sending review content to 3-5 external AI providers, but it does not clearly warn users that proprietary code, plans, or security-sensitive material may be transmitted to third parties. In a code-review and architecture-review skill, that omission materially increases the risk of unintended data disclosure, especially when users may assume a local or single-provider workflow.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.