Back to skill

Security audit

Private Knowledge Base

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent local document knowledge base, but its ingestion script has a real code-execution risk and it retains private document text and source paths without strong controls.

Review this skill carefully before installing. Only ingest PDFs and folders you trust, avoid sensitive documents unless you can protect and delete the KB_ROOT directory, set KB_ROOT to a private location with restrictive permissions, and do not run the ingestion script on files or paths supplied by untrusted parties until the Python interpolation and JSON/search handling are fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/ingest.sh:31
Finding

Arbitrary Python Code Execution Through Unsafely Interpolated Paths

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ingest.sh:46
Finding

Metadata JSON Injection Through Unescaped Filename and Path Values

Content
View full analysis
"$KB_ROOT/index/${DOC_ID}.json" << EOF { "id": "$DOC_ID", "name": "$DOC_NAME", "source": "$PDF_PATH", "ingested": "$(date -Iseconds)", "text_file": "$KB_ROOT/docs/${DOC_ID}.txt" } EOF ``` ### Technical Analysis The script generates JSON through an unquoted here-document and directly inserts `DOC_NAME`, `PDF_PATH`, and `KB_ROOT` without JSON encoding. Filesystem names can contain quotation marks, backslashes, newlines, and most other control characters. Such characters can terminate JSON strings, corrupt the generated document, or introduce attacker-selected properties. Later scripts compound the issue by extracting the `name` field with `grep` and `cut` instead of using a conforming JSON parser. Although shell command substitution contained in variable values is not recursively executed during parameter expansion, the unescaped content still controls the structure and displayed contents of the metadata file. ### Attack Path 1. An attacker supplies a PDF with a filename containing JSON structural characters, embedded newlines, or terminal-control content. 2. `scripts/ingest.sh` derives `DOC_NAME` from that filename. 3. The value is inserted verbatim into the JSON here-document. 4. The resulting metadata is malformed or contains attacker-selected fields and values. 5. Search or summary scripts subsequently parse the metadata using regular-expression-based extraction and display the resulting value. ### Impact Assessment Exploitation can corrupt the integrity of an ingested document’s metadata, misrepresent document names or source information, disrupt later searches, and cause misleading output. Crafted control characters in displayed metadata may also manipulate terminal presentation. This issue does not directly provide code execution, ...[truncated 90 chars]
Remediation
View remediation
"$KB_ROOT/index/${DOC_ID}.json" ``` Alternatively, use Python’s `json.dump`. All consumers should parse metadata with `jq` or a JSON library rather than `grep` and `cut`. Terminal-bound values should also be treated as untrusted and sanitized or safely escaped before display. ]]>

T09 · Insecure Skill Coding Practices

Note
Location
scripts/search.sh:28
Finding

Search Query Interpreted as Grep Options and Regular Expressions

Content
View full analysis
/dev/null; then MATCHES=$(grep -in "$QUERY" "$txt" | head -3) ``` ### Technical Analysis The user-controlled query is supplied to `grep` without the `--` end-of-options marker. A query beginning with `-` can therefore be interpreted as an additional command-line option rather than a search pattern. The query is also treated as a basic regular expression even though the interface presents it as an ordinary search concept. Invalid or expensive patterns can trigger errors, produce unexpectedly broad results, or consume disproportionate processing time on large extracted documents. Shell quoting prevents shell word splitting and expansion, but it does not disable `grep` option parsing or regular-expression semantics. ### Attack Path 1. A user or upstream Agent supplies a query beginning with `-`, an invalid expression, or a computationally expensive regular expression. 2. `scripts/search.sh` passes the value directly to `grep`. 3. `grep` interprets the value as an option or evaluates it as a regular expression. 4. Search behavior changes, fails, returns unintended passages, or consumes excessive local resources. ### Impact Assessment The issue can cause incorrect search results, disclosure of passages beyond the user’s intended literal match, command failure, or local denial of service through resource-intensive matching. No direct shell command execution was identified. ]]>
Remediation
View remediation
/dev/null; then MATCHES=$(grep -Fin -- "$QUERY" "$txt" | head -3) fi ``` Also consider: - Limiting query length. - Rejecting control characters. - Handling `grep` errors separately from a legitimate “no match” result. - Documenting regular-expression behavior if regex searches are intentionally supported. - Applying execution time or resource limits when searching large collections. ]]>

T09 · Insecure Skill Coding Practices

Note
Location
scripts/summarize.sh:29
Finding

Summary Concept Interpreted as Grep Options and Regular Expressions

Content
View full analysis
/dev/null; then echo "📄 $NAME" grep -i -A2 -B2 "$CONCEPT" "$txt" | head -10 ``` ### Technical Analysis The user-controlled concept is passed to `grep` without an end-of-options marker and is evaluated as a regular expression. Values beginning with `-` may alter `grep` behavior, while invalid, broad, or expensive expressions can fail the operation, expose unintended passages, or consume excessive resources. Quoting the variable is necessary for shell safety but does not make it a literal `grep` pattern and does not prevent option interpretation. ### Attack Path 1. A user or Agent invokes `scripts/summarize.sh` with a concept beginning with `-` or containing a malicious or pathological regular expression. 2. The script passes the concept directly to both `grep` invocations. 3. `grep` interprets it as an option or evaluates its regular-expression syntax. 4. The script fails, prints unintended passages, or spends excessive resources processing document text. ### Impact Assessment The issue can affect the availability and integrity of summary input by causing failures or selecting passages outside the intended literal concept. It may also create a local denial-of-service condition against the invoking process. It does not directly allow shell command execution. ]]>
Remediation
View remediation
/dev/null; then echo "📄 $NAME" grep -Fi -A2 -B2 -- "$CONCEPT" "$txt" | head -10 fi ``` Further hardening should include: - Enforcing a reasonable maximum concept length. - Rejecting terminal-control characters. - Distinguishing search errors from an absence of matches. - Adding resource limits if large or attacker-controlled document collections are supported. ]]>

T09 · Insecure Skill Coding Practices

Note
Location
scripts/ingest.sh:6
Finding

Private Knowledge-Base Files May Be Created with Permissive Filesystem Modes

Content
View full analysis
"$KB_ROOT/index/${DOC_ID}.json" << EOF ``` ### Technical Analysis The Skill describes the resulting store as private but does not establish a restrictive `umask` or explicitly set permissions on the knowledge-base directories and files. Under a common `022` umask, newly created directories can be mode `755` and files can be mode `644`. On a multi-user system, this may permit other local users to traverse the knowledge base and read extracted document text or metadata containing absolute source paths. The precise resulting permissions depend on the invoking process’s existing umask and any pre-existing directory permissions. Therefore, exposure is environment-dependent but is not prevented by the Skill. ### Attack Path 1. A user runs the ingestion script on a multi-user system with a permissive umask such as `022`. 2. The script creates knowledge-base directories and output files without explicit restrictive modes. 3. Extracted PDF text and metadata become readable to other local accounts, depending on parent-directory permissions. 4. Another local user traverses the knowledge-base directory and reads the stored content or source-path metadata. ### Impact Assessment The issue can disclose personal document contents, document names, ingestion timestamps, and absolute filesystem paths to other local users. It does not provide remote access or privilege escalation and is limited to accounts that can access the relevant filesystem hierarchy. ]]>
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs users to ingest PDFs and other documents into a persistent local knowledge base, but it does not warn that personal files, extracted text, metadata, and derived embeddings will be stored for later retrieval. This creates a privacy and consent issue because users may provide sensitive documents without understanding that the content will be retained and indexed beyond the immediate session.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
91% confidence
Finding

The skill explicitly describes persistent storage in kb/index.json and kb/docs/, along with extracted metadata and embeddings, which means user-provided document contents survive beyond the current interaction. In a personal knowledge-base context this persistence is expected functionality, but it still introduces security and privacy risk if users are not clearly informed or if sensitive documents are ingested without retention controls.

Content

Scanner excerpt · SKILL.md (reported line 39)May include surrounding context.

md
When user provides new PDFs or papers:

1. Create document entry in `kb/index.json`
2. Extract text and metadata
3. Generate embeddings for semantic search
4. Store in `kb/docs/` with normalized name

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The file comment at L02 presents the script as a straightforward PDF ingestion utility, but the implementation also records the caller-supplied source path into a JSON metadata file at L50. This is not merely omitted detail about extraction; it changes the retained data footprint of the operation beyond just ingesting document contents.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.