Back to skill

Security audit

prompt-archaeology

Security checks for vulnerabilities and agentic risk

Overview

The skill is coherent, but it deserves review because it searches private session history, persists copied transcript contents, and loads index files with unsafe pickle deserialization.

Install only if you are comfortable letting the agent search selected session/export directories. Do not point it at broad home directories, shared profiles, credential folders, or other users' transcripts. Treat generated .idx files like sensitive transcript copies, do not share or commit them, and do not load index files from anyone else unless the implementation is changed away from pickle.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/excavate.py:271
Finding

Arbitrary Code Execution Through Unsafe Pickle Index Deserialization

Content
View full analysis
"ArchaeologyIndex": with open(path, "rb") as fh: data = pickle.load(fh) idx = cls() idx.docs = data.get("docs", []) return idx ``` The user-facing `query` command passes the supplied index file into this method: ```python idx = ArchaeologyIndex.load(args.index_file) ``` ### Technical Analysis Python's `pickle` format is capable of encoding instructions that invoke arbitrary Python callables during deserialization. Consequently, `pickle.load()` is not a data-only parser and must never process an index whose integrity and provenance are not guaranteed. The `query` subcommand accepts an arbitrary filesystem path, checks only that it refers to a file, and then deserializes it without authentication, integrity verification, type validation, or a trust warning. The documented `.idx` extension does not provide any protection. A malicious pickle can use a crafted reduction operation, such as an object implementing `__reduce__`, to invoke a command-execution callable while `pickle.load()` reconstructs the object. Execution occurs before `data.get("docs", [])` or any subsequent validation can run. This is not remote code execution by itself because the program does not download indexes. However, it becomes arbitrary local code execution whenever an attacker can convince a user or agent to query an attacker-provided index, replace an existing shared index, or modify an index in a writable location. ### Attack Path 1. An attacker constructs a malicious pickle whose deserialization routine invokes an arbitrary Python callable or operating-system command. 2. The attacker distributes it as a plausible archaeology index, such as `sessions.idx`, or replaces an index in a shared or attacker-wr ...[truncated 1121 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/excavate.py:265
Finding

Sensitive Session Contents Persisted in an Unencrypted Index with Default File Permissions

Content
View full analysis
None: with open(path, "wb") as fh: pickle.dump( {"version": self.FORMAT_VERSION, "docs": self.docs}, fh, protocol=pickle.HIGHEST_PROTOCOL ) ``` ### Technical Analysis The project is designed to process conversation histories, which may contain credentials, tokens, private source code, internal URLs, personal information, operational commands, and confidential business decisions. The saved index duplicates the original corpus by serializing: - The original unredacted `raw` session content. - A normalized copy of the session text. - Extracted code blocks. - Original source paths and timestamps. Although pickle uses a binary representation, it does not provide confidentiality. Its contents can be recovered with standard Python tooling or binary inspection. The output file is created using the process's normal `open(..., "wb")` behavior, so effective permissions depend on the user's umask. In permissive environments, the resulting index may be readable by other local users or services. The index also creates an additional copy of sensitive information outside the source corpus. Removing or securing the original logs does not remove the indexed copy. The documen ...[truncated 1640 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill exposes deserialization as a first-class feature by providing a load path for reusable index files backed by pickle, despite the data model being simple enough for safer formats. This unnecessarily creates a remote/local code execution sink whenever an attacker can convince a user or agent to load a crafted index file.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill explicitly instructs the agent to read local session logs and markdown/JSON exports via excavate.py and idx.scan("./sessions"), but it declares no tool scope or permissions boundary. In an agent system, that mismatch is dangerous because it can enable unscoped file access to historical data, including sensitive transcripts, without an upfront limitation on what paths may be read.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill is designed to search and extract from past conversation sessions, exported logs, JSONL transcripts, and even mentions 'another agent's transcripts,' but it does not present a strong upfront privacy warning or mandatory consent gate. This creates a real risk of surfacing secrets, credentials, personal data, or cross-user confidential information from historical records that the current user may not be authorized to access.

Content

No source excerpt is available for this finding.

File System Enumeration

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

The example idx.scan("./sessions") # walk the directory once demonstrates recursive filesystem enumeration over a session corpus. While this is consistent with the skill's purpose, enumeration over local directories containing transcripts can expose unexpected files, expand access beyond the intended dataset, and increase the chance of collecting sensitive material if directory scope is not tightly constrained.

Content

Scanner excerpt · SKILL.md (reported line 140)May include surrounding context.

from excavate import ArchaeologyIndex

idx = ArchaeologyIndex() idx.scan("./sessions") # walk the directory once for hit in idx.search("kafka rebalance", top=5, explain=True): print(hit.score, hit.path, hit.extraction)

text

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation explicitly states that the reusable index is stored in pickle format and later loaded, but it does not warn that Python pickle deserialization is unsafe for untrusted files. If a user is induced to load a malicious index file, arbitrary code execution can occur at load time, which is especially risky for a tool intended to consume artifacts from past sessions or shared corpora.

Content

No source excerpt is available for this finding.

File System Enumeration

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code scans file system directories looking for sensitive files. This could be reconnaissance for credential theft.

Content

Scanner excerpt · references/cli-reference.md (reported line 122)May include surrounding context.

md
from excavate import ArchaeologyIndex, SearchHit

idx = ArchaeologyIndex()
idx.scan("./sessions")                    # walk directory, parse files
idx.save("sessions.idx")                  # serialize
idx2 = ArchaeologyIndex.load("sessions.idx")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The index command stores full parsed session contents, including raw conversation text and extracted code, into a reusable file without warning users that sensitive logs may be copied into a persistent artifact. Given this skill’s purpose—mining past conversation/session history—the indexed data is especially likely to contain secrets, credentials, proprietary code, or personal information, increasing confidentiality risk if the file is shared, backed up, or read by other local users.

Content

No source excerpt is available for this finding.

Insecure deserialization: pickle.load()

Medium
Category
Dangerous Code Execution
Confidence
99% confidence
Finding

The code deserializes an attacker-controlled file with pickle.load(), which can execute arbitrary code during loading. In this skill, the query subcommand accepts any existing index file path, so a user opening an untrusted .pkl index could trigger local code execution.

Content

Scanner excerpt · scripts/excavate.py (reported line 274)May include surrounding context.

python
@classmethod
    def load(cls, path: str) -> "ArchaeologyIndex":
        with open(path, "rb") as fh:
            data = pickle.load(fh)
        idx = cls()
        idx.docs = data.get("docs", [])
        return idx

Static analysis

No suspicious patterns detected.