Back to skill

Security audit

scl-90

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it needs review because it stores sensitive mental-health answers locally and presents incomplete demo scoring as a professional assessment.

Review carefully before installing. Treat it as a Chinese-language demo or reference tool, not a professional or clinical assessment. Be aware that running an assessment saves full answers in plaintext under $HOME/.scl90_history, so do not use it for private mental-health responses unless that retention is acceptable and protected on your device.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scl90.sh:351
Finding

Unprotected Plaintext Storage of Sensitive Mental-Health Responses

Content
View full analysis

Vulnerability Details

File Location: scripts/scl90.sh, lines 9–11 and 351–360
Vulnerability Type: Plaintext storage of sensitive health information
Risk Level: Medium

Vulnerable Code

bash
HISTORY_DIR="$HOME/.scl90_history"

mkdir -p "$HISTORY_DIR"
bash
local timestamp=$(date +%Y%m%d_%H%M%S)
local result_file="$HISTORY_DIR/result_$timestamp.txt"

echo "Saving results to: $result_file"
{
echo "SCL-90 assessment result - $(date)"
echo "Total score: $total_score"
echo "Positive item count: $positive_items"
echo "-------------------"
echo "Answers: ${answers[*]}"
} > "$result_file"

The textual labels above are English renderings of the original localized output strings; the shell operations and variables are unchanged.

Technical Analysis

The application automatically creates a persistent history directory and writes the complete response set, total score, and positive-item count into an unencrypted text file. These records constitute sensitive mental-health information.

Neither the directory nor the result file is created with an explicit restrictive mode. Their effective permissions therefore depend on the invoking user's umask and the accessibility of the home directory. For example, with a permissive configuration, the directory or files may be readable by other local accounts, backup agents, support tools, or unrelated processes executing under the same user.

The application also saves the information automatically without requesting explicit consent, providing a no-retention mode, or defining deletion and retention controls.

Attack Path

  1. A user runs the assessment and submits 90 sensitive responses.
  2. The application automatically creates $HOME/.scl90_history.
  3. The complete response set and derived scores are written to result_YYYYMMDD_HHMMSS.txt.
  4. A local actor or process with access to the user's home directory enumerates that directory ...[truncated 774 chars]
Remediation
View remediation

Remediation Suggestions

  1. Default to not retaining assessment answers.
  2. Obtain explicit informed consent before saving any result.
  3. Set a restrictive process mask before creating storage:
    bash
    umask 077
    
  4. Create the directory with an explicit owner-only mode:
    bash
    mkdir -p -m 700 -- "$HISTORY_DIR"
    chmod 700 -- "$HISTORY_DIR"
    
  5. Create result files atomically with mode 600, rather than relying only on shell redirection and the ambient umask.
  6. Consider storing only aggregate results instead of the complete answer set.
  7. Encrypt retained records using authenticated encryption and protect the key separately from the data.
  8. Add commands to list and securely delete saved records.
  9. Document the storage location, retention behavior, and privacy implications before assessment collection begins.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scl90.sh:247
Finding

Invalid Questionnaire Prompts Can Produce Misleading Mental-Health Conclusions

Content
View full analysis

Vulnerability Details

File Location: scripts/scl90.sh, lines 247–261 and 314–340
Vulnerability Type: Assessment-integrity failure caused by incomplete implementation
Risk Level: Medium

Vulnerable Code

bash
for i in {1..90}; do
local question="Question $i"
echo "────────────────────────────────────────"
echo "Question $i / 90"
echo ""
echo " $question"
echo ""
echo " 0=None 1=Very light 2=Moderate 3=Heavy 4=Severe"
echo ""

while true; do
read -p " Your selection (0-4): " answer
if [[ "$answer" =~ ^[0-4]$ ]]; then
answers+=("$answer")
total=$((total + answer))
break
bash
local crisis_flag=0
if [ "$total_score" -ge 160 ]; then
crisis_flag=1
echo " Total score exceeds the cutoff"
fi
if [ "$positive_items" -ge 43 ]; then
crisis_flag=1
echo " Positive item count exceeds the cutoff"
fi

echo ""
echo "Factor analysis"
echo "────────────────────────────────────────"

echo " [This demonstration version does not implement detailed factor analysis]"

The textual labels above are English renderings of the original localized output strings; the control flow, variables, thresholds, and generic-question behavior are unchanged.

Technical Analysis

The project bundles actual SCL-90 question text in data/questions.json, but the assessment loop does not load that file. Instead, each prompt is generated as only a generic question number. Users therefore assign symptom-severity values without being shown the symptom to which each answer should apply.

Despite collecting semantically meaningless answers, the script applies SCL-90 aggregate cutoffs and reports either that the result may require attention or that the overall indicators are within the normal range. The factor-analysis section explicitly admits that the detailed implementation is incomplete.

This is an integrity and safety defect in health-related software. The declared behavior presents the pro ...[truncated 1587 chars]

Remediation
View remediation

Remediation Suggestions

  1. Load every prompt from data/questions.json and display the correct text for question IDs 1 through 90.
  2. Validate the JSON before beginning:
    • Require exactly 90 unique question IDs.
    • Require IDs to be within the expected range.
    • Validate option values and labels.
    • Abort safely if data is absent, duplicated, or malformed.
  3. Load data/factors.json and implement the documented factor mappings and factor averages.
  4. Add automated tests using known response sets and expected aggregate and factor results.
  5. Do not display a normal-range conclusion while scoring functionality is incomplete.
  6. Detect high-risk individual responses, including the suicide and violence-related items, and display immediate crisis guidance independently of aggregate thresholds.
  7. Clearly label the application as an unvalidated demonstration until its questionnaire and scoring behavior have undergone professional review.
  8. Verify cited norms and cutoff interpretation with qualified mental-health professionals before describing the tool as professional or clinically meaningful.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims to be a professional psychological assessment tool, but the finding indicates it stores sensitive mental-health responses and history locally without clearly declaring that behavior, while also exposing a history feature not disclosed in the description. In a mental-health context, undeclared collection and retention of raw answers is especially sensitive, and incomplete/demo-like assessment logic can mislead users into trusting inaccurate results, creating both privacy and safety risks.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script creates a persistent history directory under the user's home directory for a mental-health assessment workflow without clear necessity or informed consent. Because the tool handles highly sensitive psychological responses, local retention materially increases privacy and disclosure risk if the account, device, backups, or shared environment are accessed by others.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is advertised as a professional SCL-90 assessment, but the implementation only asks placeholder '题目 i' prompts and explicitly states factor analysis is a demo-only stub. In a mental-health context, presenting incomplete or inaccurate scoring as an assessment can mislead users about psychological risk, including false reassurance or unnecessary alarm.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The tool writes raw answers and scores for a mental-health assessment to a history file without warning users before they begin answering. This undermines informed consent and can expose intimate psychological information to anyone with local access to the account, terminal history context, backups, or synced home directories.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JSON file contains all user-facing questionnaire text and response labels exclusively in Chinese, but provides no indication that the skill offers a language choice or that the Chinese-only locale is intentional and justified. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

All user-facing prompts, help text, disclaimers, and crisis guidance are presented in Chinese, and the script does not offer users any language or locale selection. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script claims Chinese norms are part of the assessment, but only prints static reference text and does not apply those norms to the user's results. This is misleading in a clinical-style tool because users may believe their scores were normalized or interpreted against population data when they were not.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script stores full assessment answers in plain text in a local history log, creating a sensitive data disclosure risk. Mental-health responses are especially sensitive, so plain-language retention without protection or minimization is dangerous even if the data never leaves the machine.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description explicitly states the tool includes a China-specific norm ("含中国常模"), which indicates a locale-specific framing without offering user choice or clarifying that this is intended only for users who want that locale context. The policy requires either opt-in or clearly justified region-specific constraints.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JSON contains multiple natural-language description fields entirely in Chinese, while no field documents that the skill is China/Chinese-specific or that users can opt into this locale. Under the policy rule for language/locale constraints, forcing a single language without user choice can be a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This file contains user-facing natural-language strings such as "中年常模", "全国常模", and "中国心理卫生杂志" entirely in Chinese, with no accompanying indication that the skill is China-specific or that users can opt into this locale. Per the policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

SQP-3 applies to all file types and covers language or locale policy violations. This document presents all guidance in a single language without user opt-in or an explicit statement that the skill is intended only for Chinese-speaking users, which can conflict with organizational language-choice expectations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.