Back to skill

Security audit

Auto Arena

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent model-benchmarking workflow, but users should treat its external model calls, saved evaluation artifacts, and unpinned Python installs carefully.

Install in an isolated virtual environment, pin reviewed package versions where possible, and only benchmark prompts or responses that you are authorized to share with the configured model providers. Review the output directory because queries, responses, rubrics, comparison details, reports, and checkpoints may be saved locally.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:24
Finding

Unpinned Third-Party Dependencies Are Installed and Executed

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 24–29 and 51–63
Vulnerability Type: Supply-chain exposure through unpinned dependencies
Risk Level: Medium

Vulnerable Code

bash
# Install OpenJudge
pip install py-openjudge

# Extra dependency for auto_arena (chart generation)
pip install matplotlib

The installed package is subsequently executed:

bash
# Run evaluation
python -m cookbooks.auto_arena --config config.yaml --save

# Use pre-generated queries
python -m cookbooks.auto_arena --config config.yaml \
  --queries_file queries.json --save

# Start fresh, ignore checkpoint
python -m cookbooks.auto_arena --config config.yaml --fresh --save

# Re-run only pairwise evaluation with new judge model
# (keeps queries, responses, and rubrics)
python -m cookbooks.auto_arena --config config.yaml --rerun-judge --save

Technical Analysis

The installation commands do not pin reviewed package versions, verify package hashes, or require a lockfile. Consequently, the code installed and executed can change over time without any modification to this Skill.

The supplied project contains only SKILL.md; it does not include the referenced cookbooks.auto_arena implementation. The security-sensitive behavior of that module and its transitive dependencies therefore cannot be verified from the audited artifact. Package installation may execute build hooks, while later module invocation executes code obtained through the Python package supply chain.

This is a supply-chain weakness rather than evidence that the currently published packages are malicious. Exploitation requires compromise or substitution of a package, one of its dependencies, or the configured package index.

Attack Path

  1. An attacker compromises a future release of py-openjudge, one of its transitive dependencies, or a package source used by the victim.
  2. The victim follows the documented pip install commands without a version or hash constraint.
  3. Pip res ...[truncated 999 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to an explicitly reviewed version.
  2. Maintain a lockfile or hash-locked requirements file that includes transitive dependencies.
  3. Require cryptographic hash verification during installation, for example:
bash
python -m pip install --require-hashes -r requirements.txt
  1. Install packages only from a trusted, explicitly configured package index and disallow unexpected fallback indexes.
  2. Review dependency provenance and monitor pinned versions for security advisories.
  3. Execute the evaluation in an isolated virtual environment or container under a non-privileged account.
  4. Expose only the API credentials required for the current run and rotate them if dependency compromise is suspected.
  5. Where practical, include or vendor the security-relevant pipeline source so its behavior can be audited together with the Skill.
  6. Add reproducible installation instructions and integrity-verification procedures to SKILL.md.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (8)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs users to benchmark multiple external models and a judge model, but it does not clearly warn that user prompts, generated test queries, model outputs, and evaluation artifacts may be transmitted to third-party endpoints and saved locally by default. This can expose sensitive data from the task description or responses without informed user consent, especially because the workflow fans data out to multiple services and persists it in output files.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

md
task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 109)May include surrounding context.

md
task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 140)May include surrounding context.

md
task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 144)May include surrounding context.

md
task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 149)May include surrounding context.

md
task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

Defaulting report language to Chinese without clear user opt-in can cause unintended transmission of user content to a model for transformation into a different language, and may surprise users who expect outputs in their own language. While not a direct exploit primitive, it increases the chance of mishandling or misunderstanding evaluation artifacts and reduces transparency in a workflow already involving external model processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The configuration reference sets the report language default to Chinese without explicit user consent, which can lead to surprising output transformations and lower transparency. In a system that already sends and stores generated content, silent language defaults can contribute to data-handling confusion and operational mistakes.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.