Back to skill

Security audit

hekouwang-claude-md-doctor-skill

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real AGENTS.md/CLAUDE.md checker, but its documentation includes unsafe CI execution guidance and optional suite tooling that expands into local skill and environment inspection.

Install only if you are comfortable with a branded Chinese-language checker that can read and, with confirmation, edit AGENTS.md/CLAUDE.md-related project files. Do not copy the provided CI curl-to-python example as written; vendor the reviewed script or pin it to a full commit and verify its SHA-256. Treat run-all-doctors.sh as a broader local-environment diagnostic, not just a document linter, and run it only after reviewing what the separate skill-doctor and env-doctor tools inspect.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:57
Finding

Mandatory Branding and Commercial Steering Manipulate Agent-Generated Reports

Content
View full analysis
Remediation
View remediation

T03 · Remote Payload Retrieval and Execution

Error
Location
README.en.md:83
Finding

CI Instructions Download and Execute Mutable Remote Python Code Without Integrity Verification

Content
View full analysis
Remediation
View remediation
" run: | curl --fail --show-error --location \ --output check.py \ "https://raw.githubusercontent.com/huiyonghkw/hekouwang-claude-md-doctor-skill/${CHECKER_COMMIT}/check.py" ``` 3. Publish a trusted SHA-256 digest and verify it before execution: ```yaml run: | echo " check.py" | sha256sum --check - python3 check.py . ``` 4. Prefer checking the reviewed script directly into the consuming repository so updates are visible in code review. 5. If distributed as a GitHub Action, pin the action by full commit SHA rather than by branch or floating tag. 6. If a container is used, pin the image by digest rather than a mutable tag: ```yaml image: ghcr.io/example/checker@sha256: ``` 7. Set minimum workflow permissions: ```yaml permissions: contents: read ``` 8. Do not expose deployment, package-publishing, or repository-write credentials to the checker job. 9. Separate auditing from build and release jobs so a compromised checker cannot tamper with trusted release artifacts. 10. Apply the same corrections to both `README.en.md` and `README.md`. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (50)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill metadata and opening description present this as a focused AGENTS.md/CLAUDE.md auditor, but the content also instructs broader behavior such as evaluating skill directories and invoking other doctor-style environment checks. That scope expansion can cause the agent to access or reason about files and system state the user did not intend to include when invoking what appears to be a narrow document linter.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 202)May include surrounding context.

md
注意: 本脚本只做"机器能确定的部分"。是否真的是"图书馆内容""规则是否可执行"
这类判断需要人/模型读正文定夺——交给 SKILL.md 的定性复核环节。脚本绝不读取
任何 .env / *.key / *.pem 等密钥文件。
"""

import os

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · check.py (reported line 18)May include surrounding context.

python
注意: 本脚本只做"机器能确定的部分"。是否真的是"图书馆内容""规则是否可执行"
这类判断需要人/模型读正文定夺——交给 SKILL.md 的定性复核环节。脚本绝不读取
任何 .env / *.key / *.pem 等密钥文件。
"""

import os

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · check.py (reported line 347)May include surrounding context.

python
注意: 本脚本只做"机器能确定的部分"。是否真的是"图书馆内容""规则是否可执行"
这类判断需要人/模型读正文定夺——交给 SKILL.md 的定性复核环节。脚本绝不读取
任何 .env / *.key / *.pem 等密钥文件。
"""

import os

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · check.py (reported line 494)May include surrounding context.

python
注意: 本脚本只做"机器能确定的部分"。是否真的是"图书馆内容""规则是否可执行"
这类判断需要人/模型读正文定夺——交给 SKILL.md 的定性复核环节。脚本绝不读取
任何 .env / *.key / *.pem 等密钥文件。
"""

import os

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

An overly broad trigger phrase can cause the skill to activate during ordinary conversation or unrelated review requests, leading to unintended analysis or suggested edits. In an agent context, ambiguous invocation boundaries are dangerous because they can cause tools or workflows to run without the user's specific intent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README claims the audit is read-only while elsewhere stating the tool can apply fixes, creating a safety-boundary contradiction. Users may grant trust or approvals based on the read-only claim and then be surprised by write behavior, which increases the risk of unintended file modification in sensitive repositories.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger text says essentially any request about reviewing or optimizing AGENTS.md or CLAUDE.md quality should invoke the skill. Overbroad activation increases the chance the skill runs in contexts where the user wanted a simple answer, causing unnecessary file inspection, shell execution, or modification proposals under a broad interpretation.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 19)May include surrounding context.

md
allowed-tools:
  - Bash
  - Read
  - Write
  - Edit
  - Glob
  - Grep

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

L126 instructs the skill to produce output in Chinese unconditionally. This is a language policy concern because it forces a specific language regardless of the user's preference and does not offer a choice or opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code treats phrases like "汇报用中文" and "回复用中文" as desirable persona markers, which encodes a preference for Chinese-language output. This is a natural-language locale policy issue because it promotes a fixed language choice without any visible user opt-in or alternative.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This markdown file contains user-facing instructions and descriptions exclusively in Chinese, and it does not provide an opt-in, alternative language, or justification that the skill is region- or locale-specific. Under the stated policy, forcing a specific language without user choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script materially expands scope beyond the advertised purpose of checking AGENTS.md/CLAUDE.md by also auditing local skills and scanning the host environment. That creates a capability mismatch: a user invoking a config-file doctor may unintentionally trigger broader inspection of project contents and machine state, which can expose sensitive information or violate least-privilege expectations.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Invoking a host environment scanner is not justified by the stated function of auditing runtime configuration markdown files. Even if the scanner is legitimate, running it under this skill context can collect system details the user did not intend to expose, making the behavior unexpectedly invasive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script executes the environment scan directly with no confirmation, warning, or preview of what host data will be inspected. In an agent/tooling setting, silent host inspection is dangerous because users may assume they are only linting repository documentation while the script is actually probing local development environment state.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.