Back to skill

Security audit

SkillForge — Self-Evolving AI Skills

Security checks for vulnerabilities and agentic risk

Overview

This skill is coherent and local-only, but it asks to automatically analyze broad work logs, long-term memory, current conversation history, and installed skills with limited scoping controls.

Review this before installing if your WorkBuddy logs, MEMORY.md, conversation summaries, or installed skills may contain secrets, client data, internal procedures, or personal information. Prefer disabling auto-scan and realtime detection until you can choose exactly which sources and dates it may read, and periodically delete generated pattern archives, reports, and health records you do not want retained.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:213
Finding

Overbroad Automatic Access to Work Memory and Conversation History

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:144-147, 158-163, 213-220; references/ALGORITHM.md:14-19, 27-38; references/CONFIG_FULL.md:102-105
Vulnerability Type: Overbroad access to sensitive Agent memory and conversation data
Risk Level: Medium

Vulnerable Code Snippets

SKILL.md:144-147:

markdown
1. **Never auto-install** — All generated Skills require explicit user confirmation
2. **Never delete existing Skills** — Only suggests archiving; you decide
3. **Read-only on work memory** — SkillForge reads your logs but never modifies them
4. **Privacy-first** — Reports contain pattern summaries, never raw quotes from your work logs

SKILL.md:158-163:

markdown
| `scout.lookback_days` | 14 | How far back to scan daily logs |
| `scout.similarity_threshold` | 0.65 | How similar two actions must be to cluster together |
| `scout.min_cluster_size` | 3 | Minimum occurrences to count as a "pattern" |
| `smith.merge_threshold` | 0.70 | When to merge into existing Skill vs. create new |
| `sensei.zombie_threshold_days` | 30 | Days of zero usage before flagging as zombie |
| `general.realtime_detection` | true | Detect patterns during live conversations |

SKILL.md:213-220:

markdown
SkillForge works best with WorkBuddy's daily log system:

- **Required**: Daily log files at `{workspace}/.workbuddy/memory/YYYY-MM-DD.md`
- **Optional**: Long-term memory at `{workspace}/.workbuddy/memory/MEMORY.md`

**Don't have daily logs yet?** No problem — SkillForge can also analyze your current conversation history. The more data it has, the better the pattern detection. Daily logs just give it a longer memory.

**First-time setup**: Just start using SkillForge. It will create its working directory (`{workspace}/.workbuddy/skillforge/`) automatically on first run.

references/ALGORITHM.md:14-19:

markdown
| Data Source | Path | Purpose |
|--------|------|------|

...[truncated 3360 chars]
Remediation
View remediation

Remediation Suggestions

  1. Set general.auto_scan_enabled and general.realtime_detection to false by default.
  2. Require explicit user authorization before every scan and clearly enumerate the sources, paths, and date range to be accessed.
  3. Obtain separate opt-in consent for daily logs, long-term memory, conversation history, and installed Skill definitions.
  4. Default to the narrowest source necessary and provide file-, directory-, topic-, and time-range allowlists.
  5. Exclude conversation history and long-term memory unless the user explicitly enables them for the current operation.
  6. Detect and redact credentials, tokens, private keys, personal data, and other secrets before fingerprint extraction or report generation.
  7. Define retention periods and provide controls to inspect and delete fingerprints, reports, drafts, and health records.
  8. Store only minimized derived metadata and avoid retaining source excerpts or reversible summaries.
  9. Display a pre-scan preview and a post-scan record identifying which sources were read and which artifacts were created.
  10. Enforce workspace boundaries and reject symlinks or paths that resolve outside user-approved directories.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The quick-start and surrounding description emphasize convenience but do not clearly foreground that routine use involves ongoing accumulation and analysis of work logs and possibly conversation-derived data. Because the skill is explicitly designed to watch work patterns over time, insufficient upfront disclosure can undermine informed consent and lead users to expose sensitive operational history they did not realize was being retained and mined.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad terms such as "pattern," "evolve," "review," and the skill name itself, which could be mentioned in ordinary conversation and unintentionally invoke the full pipeline. In this skill's context, accidental activation is meaningful because the pipeline scans local work logs and conversation history and may generate drafts or reports based on private data without a clearly intentional invocation boundary.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file presents all user-facing instructional content in Chinese and states it is for developers/advanced users, but it does not offer any alternative language or indicate that Chinese is optional. This can violate language/locale policy when a skill or its documentation implicitly requires a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file’s explanatory text and inline comments are written in Chinese throughout, with no indication that language is configurable or that this is a region-specific document. This creates a natural-language locale policy concern because it imposes a specific language on users without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The quick-start instructs users to say a Chinese phrase or one English alternative, and the broader trigger section is heavily centered on fixed Chinese terms. This can create a language/locale policy concern when activation depends on prescribed language-specific phrases without an explicit user choice or opt-in model.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.