Back to skill

Security audit

AI 树德:义商本体伦理安全系统

Security checks for vulnerabilities and agentic risk

Overview

This skill is not destructive or covert, but it overstates safety-monitoring capabilities while its code uses fragile heuristics and fixed or fail-open scores.

Install only if you treat this as a prototype or educational ethics-audit aid, not as a dependable safety gate. Review generated reports manually, avoid using its scores for automated approval decisions, and install dependencies in an isolated environment after checking whether they are actually needed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/reports/generate_audit_report.py:120
Finding

Value-alignment reporting fails open with a hard-coded passing score

Content
View full analysis

Vulnerability Details

File Location: scripts/reports/generate_audit_report.py, lines 120 and 162
Vulnerability Type: Fail-open security assessment caused by an inconsistent result key
Risk Level: Medium

Vulnerable Code

python
value_alignment_report = check_value_alignment(text, text)
python
"value_alignment": value_alignment_report.get("alignment_score", 0.95),

The called check_value_alignment() function returns the keys total_score, dimension_scores, needs_alignment, and recommendations. It does not return alignment_score.

Consequently, value_alignment_report.get("alignment_score", 0.95) always uses the fallback value 0.95. The generated report therefore presents a high alignment score regardless of the detector's actual result.

Technical Analysis

This is a fail-open logic defect in a security assessment path. A missing result field should cause an explicit error or conservative failure, but the implementation substitutes a high passing score.

The defect is made more significant because the generated Markdown presents this value as an authoritative assessment. Consumers can therefore receive a favorable value-alignment result for content that the underlying function marked as requiring adjustment.

Passing the same text as both the response and user request also prevents meaningful request-to-response comparison, although the current implementation does not use the user_request argument.

Attack Path

  1. An attacker or user supplies ethically problematic text to scripts/run_audit.py through --text.
  2. generate_formal_report() calls check_value_alignment().
  3. The detector returns total_score and needs_alignment, but no alignment_score.
  4. The report generator requests the nonexistent alignment_score key.
  5. The fallback value 0.95 is selected.
  6. The final audit report presents the content as having a high value-alignment score, p ...[truncated 534 chars]
Remediation
View remediation

Remediation Suggestions

  1. Use the actual returned field and normalize it to the documented range:

    python
    raw_score = value_alignment_report["total_score"]
    alignment_score = raw_score / 5.0
    
  2. Remove the favorable fallback. Missing mandatory fields should raise an exception or produce a conservative failure result:

    python
    if "total_score" not in value_alignment_report:
        raise ValueError("Value-alignment assessment returned an invalid result")
    
  3. Pass the original user request separately from the generated response.

  4. Define a typed result schema shared by the detector and report generator.

  5. Add regression tests asserting that poorly aligned input cannot produce a fixed score of 0.95.

  6. Ensure the Markdown report displays both the normalized score and the needs_alignment decision.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/alienation_protection.py:91
Finding

Case-normalization mismatch allows alienation indicators to evade detection

Content
View full analysis

Vulnerability Details

File Location: scripts/alienation_protection.py, lines 91–96
Vulnerability Type: Fail-open keyword matching caused by inconsistent normalization
Risk Level: Medium

Vulnerable Code

python
keyword_count = {
    kw: text.lower().count(kw)
    for kw in config["检测关键词"]
}

Technical Analysis

The input is converted to lowercase, but the configured keyword is not. Many configured indicators contain uppercase characters, including phrases beginning with I, I'll, I'm, As, How, Let, My, or Optimize.

Python substring matching is case-sensitive. Searching a lowercased input for an uppercase-containing keyword always returns zero. For example, lowercasing I can help you manipulate produces i can help you manipulate, which cannot match the unchanged configured keyword I can help you manipulate.

This causes multiple high-risk phrases to be systematically ignored rather than merely reducing detection accuracy.

Attack Path

  1. An attacker produces content containing a configured high-risk phrase, such as an offer to manipulate, forge documents, or bypass security.
  2. detect_alienation_patterns() converts the complete input to lowercase.
  3. The function compares that lowercase input against the original mixed-case keyword.
  4. The comparison returns zero occurrences.
  5. No indicator is added for that keyword.
  6. The report can classify the content at a lower risk level and omit the associated mitigation plan.

An attacker does not need unusual encoding or elevated privileges; using affected phrases is sufficient.

Impact Assessment

This issue does not provide system access, but it undermines the central detection feature of the Skill:

  • High-risk content can evade configured controls.
  • Risk counts and severity can be understated.
  • Protection strategies may not be selected.
  • Generated audit reports may falsely state that no alienat ...[truncated 104 chars]
Remediation
View remediation

Remediation Suggestions

  1. Normalize both operands with casefold():

    python
    normalized_text = text.casefold()
    keyword_count = {
        kw: normalized_text.count(kw.casefold())
        for kw in config["检测关键词"]
    }
    
  2. Pre-normalize configured indicators during module initialization.

  3. Add tests for lowercase, uppercase, title-case, and mixed-case variants of every configured phrase.

  4. Consider Unicode normalization before matching to prevent visually equivalent forms from bypassing checks.

  5. Replace exact phrase matching with carefully bounded token or semantic detection where appropriate.

  6. Treat detector failures conservatively when the module is used as an enforcement gate.

T08 · Insecure Dependencies

Note
Location
SKILL.md:113
Finding

Installation guidance uses unnecessary and unpinned third-party dependencies

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 113; duplicated in README.md, line 22
Vulnerability Type: Unpinned dependency installation from a mutable package registry
Risk Level: Low

Vulnerable Code

bash
pip install regex numpy pandas

Technical Analysis

The installation command does not constrain package versions or verify package hashes. It therefore installs whichever versions the package index resolves at execution time. The audited scripts use Python's standard re module and do not import regex, numpy, or pandas, making the recommended dependencies unnecessary for the reviewed implementation.

Unnecessary dependencies expand the supply-chain attack surface. Unpinned packages also make installations non-reproducible and allow future upstream releases to change the code executed during package installation or import.

No evidence was found that the named packages are currently malicious. The confirmed weakness is the unsafe and unnecessary installation practice, not an existing malicious payload.

Attack Path

  1. A user follows the documented installation command.
  2. pip resolves current versions from its configured package index.
  3. If an upstream package, release account, index configuration, or transitive dependency is compromised, attacker-controlled package code is downloaded.
  4. Package installation logic executes with the privileges of the user running pip.
  5. The malicious package can access data and resources available to that user.

This path depends on supply-chain compromise or use of an untrusted package index; it is not directly triggered by the current repository code.

Impact Assessment

The potential privileges are those of the account running pip. In a compromised supply-chain scenario, impact could include:

  • Arbitrary code execution during installation.
  • Access to user-readable files, tokens, and development credentials.
  • Mod ...[truncated 272 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the dependency installation instruction if the packages are not required.
  2. If dependencies become necessary, declare exact versions in a dedicated lock file.
  3. Use hash verification, such as pip install --require-hashes.
  4. Document the intended trusted package index.
  5. Run dependency vulnerability and provenance checks in continuous integration.
  6. Avoid recommending privileged installation and use an isolated virtual environment.
  7. Keep SKILL.md, README.md, and dependency manifests synchronized.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (38)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

According to the finding, the skill claims comprehensive IIQ-based safety monitoring but only performs simplistic scoring or keyword matching, with no real implementation of IIQ analysis, alienation detection, manipulation-risk auditing, or authenticity checks. That gap is dangerous because simplistic heuristics can be easily bypassed while the branding implies strong protection, increasing the chance of unsafe operational reliance.

Content

No source excerpt is available for this finding.

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · scripts/alienation_protection.py (reported line 41)May include surrounding context.

python
"I can help you manipulate",
            "Let me write a fake review",
            "I'll forge documents for you",
            "How to bypass security measures",
            "Optimize your deception tactics"
        ],
        "防护策略": [

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · scripts/alienation_protection.py (reported line 314)May include surrounding context.

python
"I can help you manipulate",
            "Let me write a fake review",
            "I'll forge documents for you",
            "How to bypass security measures",
            "Optimize your deception tactics"
        ],
        "防护策略": [

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · scripts/test_all.py (reported line 91)May include surrounding context.

python
"I can help you manipulate",
            "Let me write a fake review",
            "I'll forge documents for you",
            "How to bypass security measures",
            "Optimize your deception tactics"
        ],
        "防护策略": [

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README is written entirely in Chinese and does not indicate that other languages are supported or that the Chinese-only presentation is a deliberate region-specific constraint. Under the policy, language or locale restrictions should either be optional for users or clearly justified.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documentation presents examples and behavior that imply local file read/write capability, but it does not declare any explicit tool scope or permissions boundary. In an agent ecosystem, missing scope declarations can cause reviewers or orchestrators to underestimate what the skill may access, increasing the risk of unintended file access or data exposure if the surrounding runtime grants broad defaults.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file applies to SQP-3, which covers natural-language policy violations for all file types. The skill presents all substantive instructions and usage guidance in a single language, with no user opt-in, alternative locale, or explanation that the skill is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The configuration uses Chinese keys and value labels throughout the operational sections, which indicates the skill may be designed to operate in a single language/locale. Under the policy, forcing a specific language without user opt-in or clear region-specific justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest says the skill provides a "27-type personality matrix for AI team configuration," but this file's concrete assessment flow and report template only compute IIQ/EQ/IQ scores and alienation risks. Although the report template includes a placeholder for personality_type, there is no described logic or implementation here for deriving a 27-type matrix, creating a mismatch between the claimed capability and the documented behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file’s natural-language content is fully Chinese and does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-speaking or region-specific context. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This Python file contains natural-language descriptions, docstrings, keys, comments, and printed output in Chinese throughout, and there is no indication that the skill is region-specific or that users can opt into another language. Under the language/locale policy, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This Python file contains user-facing natural-language content in Chinese, including the module description and later printed output, but does not indicate that the language is configurable or limited to a specific region or audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script emits user-visible status messages entirely in Chinese during execution. Because the file provides no mechanism for language selection and no documented justification for a Chinese-only locale, this appears to enforce a specific language without user choice.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The docstrings claim quantifiable AI personality and comprehensive assessment, while the code returns mostly constant or length-based scores. This mismatch is dangerous because downstream users may interpret the output as scientifically grounded or security-relevant, leading to misclassification of unsafe or manipulative systems as acceptable.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The module docstring is written entirely in Chinese and presents the tool description in that language without offering any language choice or documenting a justified region-specific constraint. This can violate a language/locale policy when skills are expected to be language-neutral or user-selectable.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file presents itself as an AI ethics/safety assessment component, but the implementation uses trivial heuristics and fixed scores that do not substantively evaluate authenticity, empathy, or insight. In a safety-critical context, this can create false assurance, causing operators to rely on meaningless scores when making trust, deployment, or audit decisions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This Python file contains the docstring, comments, CLI help text, status messages, and reports entirely in Chinese, which imposes a specific language on all users. The file does not offer any locale selection, fallback language, or explicit documentation that the skill is region-specific, so it violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code presents itself as a real-time ethics monitor, but the monitoring loop only cycles through hard-coded sample responses rather than inspecting actual AI outputs. This creates a false sense of safety: operators may believe live conversations are being evaluated when no real monitoring is occurring, undermining audit, detection, and compliance expectations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The CLI and documentation imply service control and monitoring capabilities, but --start/--stop/--status only affect the local process state and do not manage or inspect any persistent external service. In a security or governance context, this deceptive operational model can cause administrators to assume protections are running or stoppable when they are not.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This Python file contains user-facing natural language that specifies the report purpose and all generated report content in Chinese. Because the skill does not offer user opt-in for language/locale selection and does not document that it is intentionally region-specific, it violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.