Back to skill

Security audit

wip

Security checks for vulnerabilities and agentic risk

Overview

This task-tracking skill is partly coherent, but it also includes hook-style enforcement, agent transcript/config access, persistent local state, automatic deletion of task state, and automatic external checks that warrant review before installation.

Install only if you want a WIP system that can actively enforce task registration, inspect agent transcripts or task artifacts, mutate task/checklist state, and run external status checks. Review or disable the hook resources and remote-check guidance if you only need a passive checklist tracker, and be especially cautious in environments where ~/.claude, ~/.gemini, GitHub CLI, ssh, or deployment endpoints are available.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (44)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Embedding or depending on hook-script tests and subprocess validation for blocking Edit operations is inconsistent with a benign progress-tracking skill and indicates hidden operational complexity. Users may unknowingly enable logic that enforces or rejects actions based on transcript inspection, which is a meaningful security and trust boundary issue.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Embedding or depending on hook-script tests and subprocess validation for blocking Edit operations is inconsistent with a benign progress-tracking skill and indicates hidden operational complexity. Users may unknowingly enable logic that enforces or rejects actions based on transcript inspection, which is a meaningful security and trust boundary issue.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Embedding or depending on hook-script tests and subprocess validation for blocking Edit operations is inconsistent with a benign progress-tracking skill and indicates hidden operational complexity. Users may unknowingly enable logic that enforces or rejects actions based on transcript inspection, which is a meaningful security and trust boundary issue.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Embedding or depending on hook-script tests and subprocess validation for blocking Edit operations is inconsistent with a benign progress-tracking skill and indicates hidden operational complexity. Users may unknowingly enable logic that enforces or rejects actions based on transcript inspection, which is a meaningful security and trust boundary issue.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 27)May include surrounding context.

md
| resume | Environment-agnostic task cleanup + remaining-work workflow | [resume.md](./resume.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 69)May include surrounding context.

md
| resume | Environment-agnostic task cleanup + remaining-work workflow | [resume.md](./resume.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 83)May include surrounding context.

md
| resume | Environment-agnostic task cleanup + remaining-work workflow | [resume.md](./resume.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 122)May include surrounding context.

md
| resume | Environment-agnostic task cleanup + remaining-work workflow | [resume.md](./resume.md) |

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
88% confidence
Finding

The skill directs unconditional file deletion via rm -f on a path in the user's home directory, which is a destructive command outside the narrow scope of work-progress tracking. Even though the target path is specific, embedding deletion behavior in this skill normalizes shell-level state mutation and could be misapplied or generalized by an agent following the pattern.

Content

Scanner excerpt · claude.md (reported line 327)May include surrounding context.

exit 2

text

4. **Cleanup responsibility**: the hook (step 3) is the primary cleanup path. As a backup, any session that detects `reset_at < now` while reading the file (e.g., during a `/wip` lookup or consolidate Step 2.4) should also `rm -f ~/.claude/copilot-rate-limit.json` so the next Copilot invocation is not gated by stale data.

## Compact recovery

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · resources/block-wip-register-before-execute.py (reported line 34)May include surrounding context.

python
"was a task registered since /wip" check is reinterpreted as
                  "is THIS write targeting task.md itself" (the write to
                  task.md IS the registration act).
  Same script is registered in ~/.claude/settings.json AND
  ~/.gemini/config/hooks.json (precedent:
  consolidate/resources/block-noncompliant-review-comment.sh).

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · resources/block-wip-register-before-execute.py (reported line 35)May include surrounding context.

python
"is THIS write targeting task.md itself" (the write to
                  task.md IS the registration act).
  Same script is registered in ~/.claude/settings.json AND
  ~/.gemini/config/hooks.json (precedent:
  consolidate/resources/block-noncompliant-review-comment.sh).

  NOTE (unverified): Antigravity's transcript-equivalent (conversation log

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · resources/block-wip-register-before-execute.py (reported line 193)May include surrounding context.

python
"is THIS write targeting task.md itself" (the write to
                  task.md IS the registration act).
  Same script is registered in ~/.claude/settings.json AND
  ~/.gemini/config/hooks.json (precedent:
  consolidate/resources/block-noncompliant-review-comment.sh).

  NOTE (unverified): Antigravity's transcript-equivalent (conversation log

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This hook injects a system reminder telling the agent to verify and then delete tasks automatically when completion-like text is detected. That directly bypasses the skill’s own required per-item user-direction workflow, so a vague or misdetected completion phrase could cause destructive state changes without explicit user confirmation.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · tests/test_block_wip_register_import_safety.py (reported line 38)May include surrounding context.

python
def test_import_does_not_block_or_execute():
    """Importing the module must return promptly without reading stdin or exiting.

    stdin is an OPEN pipe with no data: a top-level json.load(sys.stdin) blocks
    forever (this is the task-528 hang mode). A main()-wrapped module imports

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares no explicit tool scope even though the behavior described and referenced companion guides imply file access, environment access, and possibly shell usage. In a skill system, missing scope boundaries increases the chance that the skill can read, write, or execute more than users expect, especially when paired with broad triggers and workflow enforcement logic.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Broad triggers like "resume" and "track progress" can cause accidental activation during ordinary conversation, leading the skill to inspect task state, ask workflow questions, or alter task artifacts unexpectedly. In a skill with write and enforcement semantics, unintended invocation meaningfully increases operational risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The "When to Use" section is broad enough that the skill may self-apply in many normal task scenarios, not just explicit task-tracking requests. That raises the likelihood of unexpected task registration, cleanup, or deletion behavior, especially in environments where the skill can write files or influence workflow routing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill directs immediate deletion of stale completed or in-progress entries without user confirmation or a visible warning. Destructive state changes can erase task history, reduce auditability, and accidentally discard important work-tracking data if the cleanup logic misclassifies entries.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The 'Called from resume.md' section defines Step 2 as a per-item direction ask offering proceed/split/merge/hold/delete-style decisions for each remaining item. Later, the 'Interactive Task Selection (Optional)' section shows a different ask.md workflow with options like 'Implement now', 'User edits manually', and 'Skip', which does not match the mandated per-item direction model and reframes the ask as optional. That is an intent-level contradiction within the documentation for /wip behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file says resume Step 2 labels are mapped from the full set 'proceed / split / merge / hold / defer-to-checklist / delete', but the actual example and instruction immediately below only provide Proceed, Split, Defer to checklist, and Delete. Because this section defines the concrete Antigravity implementation, omitting Merge and Hold actively contradicts the documented mapping and can cause the skill to fail the manifest-required per-item direction behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This section expands a WIP-tracking skill into operational GitHub, CI, deploy, and merge-state interrogation actions, including network and shell commands unrelated to merely recording progress. That scope creep is dangerous because it authorizes external side effects and data retrieval under a benign-seeming trigger, increasing the chance an agent performs unintended actions beyond the user's expected intent.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
90% confidence
Finding

The instruction to 'run the action immediately without asking' explicitly permits autonomous execution based on inferred task meaning rather than direct confirmation. In a skill whose purpose is task tracking, that autonomy is misplaced and increases the risk of unintended tool use, especially when the downstream actions query external systems.

Content

Scanner excerpt · claude.md (reported line 212)May include surrounding context.

md
### Auto-proceed — verification/lookup tasks need no ask

If a task subject contains any of the following keywords, run the action immediately without asking and reflect the result:

| Keyword | Auto-run command |
|---------|-----------------|

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The auto-run rule matches on very broad keywords such as 'check', which commonly appear in ordinary task titles and can therefore trigger automatic command execution unexpectedly. Because the matched behavior includes gh, curl, and ssh operations, a trivial or ambiguous task label can lead to unintended external actions without clear user authorization.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Embedding curl/ssh deploy checks in a WIP skill authorizes network and remote-system interaction that is not necessary for progress tracking. In context, this makes the skill more dangerous because innocuous task text can trigger access to production-like systems or leak operational data under the guise of status verification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The Copilot rate-limit cache and hook-management logic introduces persistent cross-session state and enforcement behavior that exceeds the stated purpose of in-session progress tracking. A tracking skill should not write policy files under ~/.claude or influence later command execution, because that creates hidden environment mutation and durable behavior changes outside the immediate session.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_block_wip_register_import_safety.py:49