Back to skill

Security audit

Automl Skill

Security checks for vulnerabilities and agentic risk

Overview

This is a documentation-only PyCaret AutoML skill with optional save and deployment examples, and I found no hidden code or automatic unsafe behavior.

Before running examples, review any local output names to avoid overwriting model files, and treat cloud deployment or app/API/Docker generation as actions that may create files, use cloud credentials, and expose model artifacts outside the local environment.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (17)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill includes examples for deploy_model, create_app, create_api, and create_docker without any warning that these actions can publish services, write files, or interact with cloud resources. In an agent setting, users may treat these snippets as safe defaults and trigger deployment or artifact generation without understanding the operational and security consequences, increasing the chance of unintended exposure or infrastructure changes.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This JSON eval file contains natural-language prompts and expected outputs exclusively in Chinese across the dataset. Because the file does not indicate that the skill is region-specific or provide an opt-in language choice, it appears to impose a specific language/locale in a way that may violate language policy requirements.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest describes an AutoML skill centered on PyCaret workflows, with statistical testing mentioned as a use case keyword, but these evals explicitly expect direct use of statsmodels/scipy APIs for OLS significance testing, t-tests/ANOVA, ARIMA/SARIMAX, regression diagnostics, and GLM. That behavior is not AutoML via PyCaret and materially broadens the implemented intent from low-code AutoML to manual statistical analysis and classical econometric modeling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document title is bilingual, but the substantive documentation and instructions are presented in Chinese only. This imposes a specific language on users without indicating any option to choose another locale, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The file presents the module reference and explanatory text primarily in Chinese, with no indication that users can select another language or that the locale restriction is intentional for a region-specific audience. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file is primarily written in Chinese, including section headings, parameter explanations, and workflow guidance, while only the title is bilingual. Under the stated policy, forcing a specific language without user opt-in or a documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file is primarily written in Chinese and does not indicate that language selection is optional or that the document is intentionally limited to a Chinese-speaking audience. Under SQP-3, forcing a specific language without user opt-in can be a natural-language policy violation unless the locale constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes this skill primarily as a low-code AutoML modeling tool for tasks like classification, regression, clustering, forecasting, NLP, and experiment workflows. This reference file additionally documents cloud deployment to AWS/GCP/Azure and generation of Streamlit/FastAPI/Docker artifacts, which are software deployment capabilities rather than core automated model building or analysis behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The deployment examples encourage pushing models to AWS, GCP, and Azure without warning that model artifacts, metadata, or configuration may leave the local environment and be stored in external services. In an AutoML skill context, users may deploy sensitive models or derived data artifacts without understanding the data exposure, compliance, or credential-handling implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file includes a Quick Start example that writes a model to local storage with save_model(final, 'best_model'), but the surrounding documentation does not warn users that running the workflow will create files on disk. For markdown files, user-facing descriptions should disclose behaviors that may affect user data or system state.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The manifest emphasizes fast AutoML experimentation, model selection, tuning, ensembling, and feature engineering, but this eval expects cloud deployment via deploy_model and application creation via create_app. Deployment and app scaffolding are adjacent capabilities, but they are not clearly stated in the manifest description of the skill's purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The document presents the skill instructions and descriptions primarily in Chinese, with only a partial English heading, and does not indicate any user language choice or opt-in. This can violate a language/locale policy when skills are expected to respect user language preferences rather than unilaterally enforcing one language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

Most of the instructional content is presented in Chinese, and the file does not state that users may choose another language or locale. Under the policy rule, forcing a specific language without opt-in can be a natural-language policy concern unless the locale restriction is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file includes a workflow step that calls save_model(ensemble, 'best_model'), which creates artifacts on the user's filesystem. The document does not provide any warning or note that the example writes files locally, which is a user-impacting behavior covered by the markdown-specific missing-warning rule.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This is a markdown file, so SQP-2 applies to missing warnings in documentation. The workflow example includes save_model(ensemble, 'best_regression_model'), which writes files to disk, but the surrounding documentation provides no warning that running the example will create persisted model artifacts and may overwrite existing files with the same name.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

Most headings, explanations, and inline comments throughout the file are written in Chinese, with no note that the user can choose another language or that the content is intentionally region-specific. Under the policy, a skill artifact that effectively enforces a specific language without opt-in can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

Natural-language guidance throughout the file is predominantly Chinese, and the file does not indicate that users may choose another language or that the locale is intentionally restricted. Under the stated policy, forcing a specific language without user opt-in is a language/locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.