Back to skill

Security audit

Coding Pronoun Prompt Resolver

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed local Claude hook that scans prompts for ambiguous references and stores limited local metadata, with no evidence of external transfer or destructive behavior.

Install this only if you are comfortable with a Claude Code hook running on every submitted prompt and keeping local analytics plus a resolution ledger under ~/.claude/skills/pronoun-resolver/.claude. Review the stored fields, use the documented disable sentinel for projects where you do not want it active, and delete the ledger or analytics files when you no longer want retained metadata.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (23)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · README.md (reported line 151)May include surrounding context.

md
- **pronoun-ledger.json** — resolution metadata (pronoun, resolved referent, confidence tier, correction status). Prompts are hashed, never stored as text.
- **pronoun-resolver-analytics.jsonl** — per-message stats (timestamp, flag count, word count). No message content is stored.

No data is sent externally. Both files are local-only and can be deleted at any time without affecting functionality. To clear all stored data: `rm ~/.claude/skills/pronoun-resolver/.claude/pronoun-ledger.json ~/.claude/skills/pronoun-resolver/.claude/pronoun-resolver-analytics.jsonl`

## File Structure

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · README.md (reported line 230)May include surrounding context.

md
- **pronoun-ledger.json** — resolution metadata (pronoun, resolved referent, confidence tier, correction status). Prompts are hashed, never stored as text.
- **pronoun-resolver-analytics.jsonl** — per-message stats (timestamp, flag count, word count). No message content is stored.

No data is sent externally. Both files are local-only and can be deleted at any time without affecting functionality. To clear all stored data: `rm ~/.claude/skills/pronoun-resolver/.claude/pronoun-ledger.json ~/.claude/skills/pronoun-resolver/.claude/pronoun-resolver-analytics.jsonl`

## File Structure

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · README.md (reported line 192)May include surrounding context.

ln -s "$(pwd)/coding-pronoun-prompt-resolver" ~/.claude/skills/pronoun-resolver

text

### 2. Add the hook to `~/.claude/settings.json`

```json
{

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · README.md (reported line 224)May include surrounding context.

Re-enable

bash
rm .claude/pronoun-resolver-disabled

Reset the ledger

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description substantially overstates the skill's scope. The code only tokenizes a short input message and applies fixed word-list heuristics to identify certain imperative forms such as a lone verb ('Fix') or verb plus adjective/filler ('Make good', 'Clean up'). It explicitly does not perform pronoun/reference ambiguity detection, does not resolve anything using prior conversation context, and contains no mechanism for learning from corrections or assigning adaptive confidence tiers. While the declared description includes bare imperative detection, that is only one subset of the claimed functionality, and even that subset is implemented narrowly. Therefore the declared purpose does not accurately represent the actual code behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose emphasizes real-time linguistic detection and contextual flagging of ambiguous language in user messages, with self-learning confidence behavior. This code chunk does not do detection at all. Its primary function is persistence: recording already-resolved pronoun resolutions to a JSON ledger, sanitizing free text, validating prompt hashes, and marking corrections. While the ledger relates to the declared 'self-learning via correction ledger' concept, that is only a supporting subsystem and not the advertised primary behavior. The code also writes to local files and uses locking/atomic replacement, which is a meaningful resource interaction absent from the declared permissions. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The declared purpose describes the core runtime behavior of a pronoun/vagueness detection hook with contextual resolution and adaptive learning. However, the provided code chunk only implements a stats utility for previously logged analytics. While analytics may be a supporting component of such a skill, this specific code does not match the declared primary purpose: it neither scans user messages nor flags ambiguity, nor does it resolve references using conversation context. It also accesses a local analytics file, which is an undeclared resource relative to the empty permissions list. Therefore this code chunk is materially different from the declared behavior and should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description is for an operational runtime skill that detects and flags ambiguity in user messages with zero-latency hooking and adaptive learning. The supplied code chunk does not implement that runtime behavior. Instead, it is clearly an evaluation script used to test a pronoun-resolver via predefined cases and a prompt template. It calls an external Claude CLI through subprocess, measures latency, scores correctness/confidence/tier predictions, and writes results to a file. Those are materially different primary behaviors from the declared purpose. While evaluation infrastructure can support the skill, this code chunk itself is not the described detector/resolver and introduces undeclared capabilities such as external process execution and file output. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code chunk does not implement or test pronoun ambiguity detection, vague referent analysis, bare imperative detection, conversation-context resolution, or hook-based zero-latency message analysis. Instead, it focuses on a separate logging/ledger subsystem: sanitizing sensitive values, validating confidence/hash fields, persisting resolution records to disk, handling corrupt JSON files, and ensuring concurrency safety. While a 'resolution ledger' and 'confidence' concepts loosely overlap with the description, the primary behavior shown here is storage and sanitization infrastructure, not linguistic ambiguity detection. That is a material description-to-behavior mismatch.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
95% confidence
Finding

The skill instructs installation into the agent's configuration directory and modification of ~/.claude/settings.json to register an automatic hook. Access to agent config is highly sensitive because it enables persistent behavior on every future prompt and can materially change agent execution without per-use visibility.

Content

Scanner excerpt · SKILL.md (reported line 135)May include surrounding context.

Install

  1. Symlink or copy this directory to ~/.claude/skills/pronoun-resolver
  2. Add the hook to ~/.claude/settings.json:
json
"UserPromptSubmit": [
  {

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · evals/run_evals.py (reported line 11)May include surrounding context.

python
python3 evals/run_evals.py                    # Run all cases
    python3 evals/run_evals.py --case high-context-single-file  # Run one case
    python3 evals/run_evals.py --tier1-only       # Skip council tests (faster)
    python3 evals/run_evals.py --dry-run          # Show prompts without calling LLM
"""

import json

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · evals/run_evals.py (reported line 315)May include surrounding context.

python
python3 evals/run_evals.py                    # Run all cases
    python3 evals/run_evals.py --case high-context-single-file  # Run one case
    python3 evals/run_evals.py --tier1-only       # Skip council tests (faster)
    python3 evals/run_evals.py --dry-run          # Show prompts without calling LLM
"""

import json

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · evals/run_evals.py (reported line 48)May include surrounding context.

python
prompt = prompt.replace(
        "{{CONTEXT_RELIABILITY}}", json.dumps(case.get("context_reliability", {}))
    )
    return prompt


def call_haiku(prompt):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The README states that flagged messages trigger automatic local logging of resolution metadata, but this behavior is not prominently disclosed up front before installation/use. Even though the data is local-only and partially redacted, it still creates persistent records derived from user prompts and model behavior, which can surprise users and retain sensitive context longer than expected.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
86% confidence
Finding

The README describes persistent session data being written to a ledger across runs, including resolution metadata and analytics. Persistence itself is intentional, but it increases security and privacy risk because prompt-derived data can accumulate over time, potentially exposing sensitive workflow context or user behavior if the local environment is later accessed by another process or user.

Content

Scanner excerpt · README.md (reported line 123)May include surrounding context.

md
`bin/log-resolution.py` owns all writes so logging is safe by construction:

- **Locked, atomic writes.** The whole read-modify-write runs under an exclusive
  `flock`, so concurrent hook fires (e.g. parallel agents) never drop each
  other's entries; the file is replaced atomically and never left half-written.
- **Secret/PII redaction.** Every free-text field (`resolved_to`, the pronoun

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares shell execution and file-write capabilities but does not define an explicit tool scope such as permissions or allowed-tools. That makes the operational boundary unclear and can cause an agent to invoke shell and filesystem actions more broadly than a user would reasonably expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill installs a user-prompt-submit hook that runs on every message and writes per-message analytics locally, but the user-facing description does not prominently warn that all submitted messages are automatically analyzed. Even with hashing and local-only storage claims, silent continuous processing of every prompt creates a privacy and consent risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest frames this skill as a detector that flags ambiguous language for later resolution in conversation context. In addition to detection, the script creates a .claude analytics directory and appends per-message telemetry to a JSONL file, which is a separate persistent analytics behavior not implied by simple zero-latency detection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The script derives analytics from stdin, which is explicitly the user's message, and appends metadata including word count and whether ambiguity was detected to a persistent JSONL file. Although the write is commented as analytics, there is no user-facing prompt, print, or warning in this file disclosing that message-derived data is being logged locally.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The hook injects an instruction telling the agent to persist resolved pronouns, referents, correction state, and confidence into a local ledger on every resolution. That creates a hidden secondary workflow that stores user-derived semantic data outside the conversation, increasing privacy risk and violating the stated expectation that resolution stays in-conversation.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The logging directive tells the agent to record resolved references, confidence tiers, and corrections into a persistent ledger. Even without raw prompt text, resolved referents can contain sensitive personal, project, or operational information inferred from the conversation, creating durable data retention and possible cross-session privacy leakage.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · evals/run_evals.py (reported line 52)May include surrounding context.

python
def call_haiku(prompt):
    result = subprocess.run(
        ["claude", "-p", "--model", "haiku", "--output-format", "json"],
        input=prompt,
        capture_output=True,

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This JSON file contains a natural-language example, "delete it," where the referent is ambiguous between an API endpoint and a Redis cache. While this is test data rather than executable behavior, it still embeds organizationally risky language around destructive action without clarifying scope.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_log_resolution.py:19