Back to skill

Security audit

Squid

Security checks for vulnerabilities and agentic risk

Overview

The skill is mainly a Squid workflow authoring guide, but its examples and templates include high-impact command, repository, GitHub, Docker, and Kubernetes actions with inconsistent approval enforcement.

Review this skill carefully before installing. Do not run the bundled examples unchanged in a real repository or deployment environment. Pin and verify the Squid installation, prefer sandbox or dry-run first, add explicit approval conditions after every gate, validate or safely pass all shell arguments, and avoid using it where authenticated GitHub, Docker, Kubernetes, or agent credentials can affect production resources.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
examples/lobster-migration.yaml:39
Finding

Shell Command Injection Through Unvalidated Pipeline Arguments

Content
View full analysis
&1 retry: maxAttempts: 2 ``` `examples/multi-agent-dev.yaml:188-205`: ```yaml - id: create-pr type: run description: Create pull request run: | cd ${args.repo} && \ git checkout -b feat/$(echo "${args.feature}" | tr ' ' '-' | tr '[:upper:]' '[:lower:]') && \ git add -A && \ git commit -m "feat: ${args.feature}" && \ gh pr create --title "feat: ${args.feature}" --body "$(cat <<'EOF' ## Summary ${args.feature} ## Architecture ${architect.json.summary} ## Review ${reviewer.json.summary} EOF )" when: $deploy-approval.approved ``` `examples/simple-deploy.yaml:19-45`: ```yaml steps: - id: build type: run run: docker build -t ${args.image} . retry: 2 - id: test type: run run: docker run --rm ${args.image} npm test retry: maxAttempts: 3 backoff: fixed delayMs: 5000 - id: approve type: gate gate: "Deploy ${args.image} to ${args.env}?" - id: deploy type: run run: kubectl set image deployment/myapp app=${args.image} -n ${args.env} when: $approv ...[truncated 3347 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
examples/multi-agent-dev.yaml:44
Finding

Rejected Approval Gates Do Not Guard Downstream Mutating Steps

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
references/patterns.md:103
Finding

Mandatory Workflow Templates Demonstrate Ungated Deployment and Rollback Operations

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:26
Finding

Unpinned Remote Repository Installation and Dependency Script Execution

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (29)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 11)May include surrounding context.

md
- `references/step-types.md` — Full config for every step type (run, spawn, gate, parallel, loop, branch, transform, pipeline). Contains exact field names, type

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 137)May include surrounding context.

md
- `references/step-types.md` — Full config for every step type (run, spawn, gate, parallel, loop, branch, transform, pipeline). Contains exact field names, type

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 453)May include surrounding context.

md
- `references/step-types.md` — Full config for every step type (run, spawn, gate, parallel, loop, branch, transform, pipeline). Contains exact field names, type

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 357)May include surrounding context.

md
| `examples/observability.yaml` | Event hooks, OTel spans, audit trails, chat notifications |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 453)May include surrounding context.

md
| `examples/observability.yaml` | Event hooks, OTel spans, audit trails, chat notifications |

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The pipeline goes beyond workflow authoring and performs deployment-adjacent repository actions including branch creation, committing, and pull-request creation. In an agent skill whose stated purpose is to create and modify Squid YAML workflows, this materially expands the trust boundary and can cause external side effects in a user repository and GitHub account.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The example invokes gh pr create, which performs an external network action against the user's GitHub account and can publish code or sensitive project metadata. This capability is not justified by the stated skill purpose and increases risk of unintended disclosure, unauthorized change submission, or abuse if the repo or generated content is malicious.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description includes broad trigger terms such as 'workflows', 'agents', and 'pipelines', which can match many ordinary user requests outside the narrow Squid-specific context. Over-broad activation increases the chance that this skill is invoked unexpectedly, causing it to steer the agent toward repository operations, pipeline generation, or command suggestions in situations where those actions are irrelevant or unsafe.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This manifest file describes the skill as 'Write code, review, refine until quality threshold is met' without specifying the activation context, allowed trigger phrases, or boundaries on when it should be invoked. In a manifest, such broad language can overlap with many general coding requests and may cause unintended invocation because no negative examples or narrowing constraints are provided.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The comment 'Step 5: Commit' and description 'Apply the code' actively suggest that the step performs a commit or code-application side effect. In reality, the implementation is just an echo command with no commit, file modification, or application of generated output. This is a direct contradiction between inline documentation and actual behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file description says the pipeline will 'Write code, review, refine until quality threshold is met,' and the final step is labeled 'Apply the code.' However, the actual run command only prints 'Code accepted with score ...' and does not write files, commit changes, or otherwise apply the generated code. This is a semantic mismatch between the documented pipeline behavior and the implemented action.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The pipeline directs multiple agents to modify the repository, run tests, and ultimately create a branch, commit, and PR, but it does not prominently warn the user about these side effects. Lack of explicit disclosure undermines informed consent and increases the chance that users trigger broad code changes or external actions unexpectedly.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The pipeline executes npm test inside an arbitrary user-supplied repository, which runs project-defined scripts with the user's local privileges. Because npm test hooks are fully attacker-controlled by repository contents, this can execute unintended commands and exceeds the narrowly described YAML-pipeline authoring role of the skill.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This manifest file applies to vague-trigger checks, and the description "End-to-end video content creation with agentic sub-steps" is a broad capability statement without any explicit trigger phrases, scope limits, or exclusion conditions. That ambiguity can cause unintended invocation because it does not specify what user requests should or should not activate this skill.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/patterns.md (reported line 205)May include surrounding context.

md
preview: $build.json           # data shown to approver
    items: $test.json.results      # list of items shown
    timeout: 3600                  # auto-reject after N seconds
    autoApprove: false             # only for dev/CI

    # Structured input (form fields, not just yes/no)
    input:

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/step-types.md (reported line 67)May include surrounding context.

md
preview: $build.json           # data shown to approver
    items: $test.json.results      # list of items shown
    timeout: 3600                  # auto-reject after N seconds
    autoApprove: false             # only for dev/CI

    # Structured input (form fields, not just yes/no)
    input:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file includes observability examples that send gate prompts to Slack and errors to PagerDuty, and describes emitted event data including fields like data, runId, and stepId. The surrounding documentation does not warn users that enabling these integrations may transmit workflow or user-derived data to external services, which fits missing user warnings for markdown files.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/testing.md (reported line 53)May include surrounding context.

md
|------|-----|-------|------|-----|
| `run` | Execute | Real agent | Halt | `squid run pipeline.yaml` |
| `dry-run` | Skip | Skip | Skip | `squid run --dry-run` |
| `test` | Execute | Mocked | Auto-approve | `squid run --test` |
| `sandbox` | Mocked | Mocked | Mocked | `squid test` (YAML tests) |
| `integration` | Execute | Mocked | Mocked | `squid test` (YAML tests) |

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 420)May include surrounding context.

md
**Spawn**: always mocked in both modes. Unmocked spawns return `{ mocked: true }`.

**Gates**: unmocked gates auto-approve. Set `gates: { stepId: false }` to reject.

## CLI

Static analysis

No suspicious patterns detected.