Back to skill

Security audit

SPSS Data Cleaning Assistant

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward SPSS data-cleaning skill, with ordinary data-processing risks but no hidden execution, persistence, or unrelated access.

Install dependencies in an isolated environment, prefer pinned package versions, and run cleaning on a copy of the original dataset. Do not upload personal, regulated, confidential, or proprietary research data unless you are authorized and have removed identifying fields where appropriate. Review the proposed cleaning plan carefully before approving deletions, recoding, imputations, or outlier changes.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:147
Finding

Unpinned Third-Party Dependencies

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 147–149
Vulnerability Type: Unpinned and unhashed third-party dependencies
Risk Level: Medium

bash
pip install pandas scipy openpyxl pyreadstat statsmodels

Technical Analysis

The documented installation command installs five third-party packages without version constraints, integrity hashes, a lockfile, or an explicitly approved package index. Consequently, package versions and transitive dependencies are resolved dynamically at installation time.

This does not demonstrate that any named package is currently malicious. However, it creates a supply-chain risk because a compromised upstream release, dependency, package repository, or index configuration could introduce malicious installation or runtime code. It also permits future incompatible releases to alter behavior without review.

Attack Path

  1. A user or automated environment follows the installation command in SKILL.md.
  2. pip contacts its configured package index and resolves the latest eligible versions and their transitive dependencies.
  3. An attacker compromises an upstream package or dependency, publishes a malicious release through a compromised maintainer account, or manipulates the environment's package-index configuration.
  4. The uncontrolled resolver selects the malicious component.
  5. Malicious package installation or import-time code executes with the privileges of the account running pip.

Impact Assessment

Successful exploitation could execute arbitrary code with the privileges of the installing user. Depending on those privileges and the environment, this may expose accessible datasets, credentials, environment variables, generated reports, and other local files. It could also modify the Python environment or affect subsequent data-processing operations. If installation is performed as an administrator or inside a privileged build pipeline, the potential scop ...[truncated 24 chars]

Remediation
View remediation

Remediation Suggestions

  • Replace the inline installation command with a reviewed and version-pinned dependency manifest.
  • Pin direct and transitive dependencies to known-good versions.
  • Generate and verify cryptographic hashes, such as by using pip-compile --generate-hashes and installing with pip install --require-hashes.
  • Use a controlled package index or internal artifact repository and explicitly configure the trusted index URL.
  • Regularly scan dependencies for known vulnerabilities and review updates before changing pinned versions.
  • Install packages in an isolated virtual environment using a non-privileged account.
  • Add a reproducible lockfile or constraints file to source control.

Example hardened installation approach:

bash
python -m venv .venv
. .venv/bin/activate
python -m pip install --require-hashes -r requirements.txt
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill advertises destructive cleaning operations such as deletion, recoding, type conversion, and outlier handling, but does not explicitly warn users that these actions can permanently alter data and reduce sample size. In a data-cleaning context this is a real safety issue because users may upload valuable research data and approve operations without understanding the risk of data loss or analytic bias.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly requests users to upload research datasets in formats like SAV, CSV, and Excel, but provides no privacy notice or guidance about sensitive data handling. Research datasets often contain personal, health, educational, or proprietary information, so omission of a privacy warning increases the risk that users submit regulated or confidential data without informed consent or sanitization.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language description is presented entirely in Chinese, which can effectively force a specific language experience without indicating that users may choose another language. Under the policy, language constraints should be opt-in or explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.