Back to skill

Security audit

agiza_agents

Security checks for vulnerabilities and agentic risk

Overview

This engineering bundle is mostly legitimate tooling, but it includes under-disclosed automation that can run shell commands, spawn/background agents, copy env files, and rewrite Git history.

Review before installing. Use this only in repositories where you can tolerate automated branches, worktrees, local env-file copies, evaluator command execution, and possible Git history rollback. Start from a clean worktree, avoid untrusted .autoresearch or .agenthub configs, run experiments in disposable branches/worktrees, pin npx tools where possible, and do not enable recurring loops unless you want background automation.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
autoresearch-agent/scripts/run_experiment.py:195
Finding

Repository-wide destructive rollback can erase unrelated work

Content
View full analysis
Remediation
View remediation
--staged --worktree -- ``` 6. Do not invoke repository-wide `checkout` or `reset --hard` automatically. 7. Require explicit user confirmation before any operation that rewrites branch history. 8. Preserve recovery metadata, including the pre-reset commit ID, and create a backup reference before rollback. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
autoresearch-agent/scripts/setup_experiment.py:46
Finding

Persisted evaluation commands are executed through a system shell

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:34
Finding

Unpinned npx commands download and execute mutable third-party packages

Content
View full analysis
Remediation
View remediation
add alirezarezvani/claude-skills/engineering npx --yes prisma-erd-generator@ ``` 2. Add the required tools to a reviewed `package.json` and commit the corresponding lockfile. 3. Use `npm ci` so dependency resolution follows the lockfile exactly. 4. Verify package provenance, publisher identity, release history, and registry source. 5. Use registry integrity metadata and retain hashes for reviewed artifacts. 6. Disable or tightly control dependency lifecycle scripts where compatible with the selected tools. 7. Prefer locally installed, reviewed binaries over runtime download and execution. 8. Run third-party generators in an isolated environment without credentials and with restricted filesystem and network access. 9. Document that these commands may download executable code so users can make an informed trust decision. ]]>
Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (395)

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description presents a general-purpose bundle of advanced engineering agent skills and plugins across many domains. The supplied code chunk instead implements a specific local message-board utility for AgentHub. Its primary behavior is CRUD-style board management over files in .agenthub/board, not agent design, RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, or platform ops. This is a materially different purpose and introduces concrete filesystem-based messaging capabilities that are not reflected in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The supplied code is not implementing engineering agent skills such as RAG, CI/CD, database design, observability, security auditing, or platform ops. Instead, it is a maintenance/QA utility that validates the structure and contents of the plugin repository itself. Its primary purpose is dry-run verification of packaging/docs consistency, including file existence, schema restrictions, markdown/frontmatter checks, and running local scripts with --help. That is a materially different function from the declared description of a collection of advanced engineering agent skills and plugins. While such a script could support the plugin package, this specific code chunk's actual behavior is a repository validator, so the description does not accurately represent what this code does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description describes a broad package of 25 advanced engineering agent skills and plugins spanning many domains. The supplied code chunk does not implement that broad functionality; it only initializes a local AgentHub session scaffold inside a git repository. Its behavior is limited to filesystem writes, simple git-repo detection, branch-name inference from .git/HEAD, argument parsing, and emitting text/JSON output. There is no evidence in this chunk of the claimed capabilities such as RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, or platform ops. Therefore the code chunk's actual behavior is materially narrower and different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a large, multi-domain collection of 25 engineering agent skills and plugins spanning many capabilities and platforms. The supplied code chunk instead implements one narrowly scoped session manager for AgentHub. It reads local session metadata, enforces a state machine, reports status, and performs git worktree cleanup. That is materially different from the declared primary purpose and omits any evidence of the advertised areas such as RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, or platform ops. While session management could be adjacent to agent tooling, this specific code is not accurately represented by the broad declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a large multi-skill engineering toolkit spanning many domains, but the supplied code chunk is a narrow benchmarking utility for measuring artifact sizes. Its actual behavior is limited to local filesystem inspection and optional subprocess execution for builds or Docker image inspection. This is materially different from the declared purpose and includes specific operational capabilities (shell command execution, Docker inspection) that are not represented in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broad collection of engineering agent skills and plugins across many domains, but this specific code chunk implements a narrow benchmark-speed evaluator. Its primary behavior is to run a configured command with shell=True, capture output, enforce a timeout, and compute latency statistics. That is materially different from the stated purpose areas such as RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, or platform ops. The command-execution and benchmarking capability is undeclared in the provided description, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a suite of engineering-focused agent skills and plugins covering software/infrastructure topics. The supplied code instead implements an LLM-based judge for marketing copy such as Twitter, LinkedIn, Instagram, email, and ads. Its primary purpose is content quality scoring, not engineering assistance, security auditing, CI/CD, RAG, MCP, database design, or platform ops. While using a CLI tool is just an implementation detail, the actual capability and domain are materially different from the declared purpose, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description is a high-level catalog of engineering agent skills and plugins, but this code chunk specifically implements a memory-usage evaluator for a local command. Its primary behavior is operational benchmarking: launching a configured command, relying on OS-specific tooling, parsing peak resident set size, and printing memory metrics. That is much narrower and materially different from the declared purpose, and it includes executable command-running capability not reflected in the description. While observability/platform ops are loosely adjacent domains, the actual code is a concrete measurement utility rather than an agent skill/plugin matching the stated description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents a broad set of engineering agent skills/plugins across many domains, but this specific code chunk is narrowly a test evaluation utility. Its primary purpose is to invoke pytest, inspect the results, and calculate pass-rate metrics. That behavior is materially different from the declared skill set and introduces an undeclared capability: executing a shell command to run tests. While test measurement could loosely relate to CI/CD, the code itself is not implementing a general CI/CD/plugin skill; it is a standalone evaluator script with a much narrower purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad collection of engineering agent skills and plugins across many domains, but this specific code chunk implements a narrow autoresearch results viewer/reporting tool. Its primary function is to read local experiment metadata and results, compute summary statistics, and render/export dashboards. That behavior is not represented in the declared purpose, and none of the listed areas like RAG, MCP servers, CI/CD, database design, observability, security auditing, or platform ops are implemented here. This is a material description-behavior mismatch rather than a mere supporting detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad collection of engineering agent skills and plugins, but the supplied code chunk is narrowly focused on running one autoresearch experiment iteration. Its primary function is experiment orchestration: reading config, running evaluations, parsing metrics, logging outcomes, and resetting git history when results are worse or crash. That is a materially different purpose from the declared broad skill/plugin suite. The code also performs concrete capabilities not conveyed by the description, especially executing configured shell commands and mutating repository state with git reset/checkout. This is more than an implementation detail and indicates the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a large, multi-skill engineering toolkit spanning many domains and platforms. The supplied code chunk instead implements one narrow operational utility: setting up and validating an experiment workflow for an autoresearch agent. Its primary behavior is filesystem manipulation, git branch management, evaluator copying, and execution of a user-specified eval command. Those are materially more specific and different than the broad declared purpose, and important capabilities like shell command execution and git operations are not reflected in the description. This is a clear description/behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is a high-level catalog of engineering agent skills/plugins across many domains, but this code chunk implements a specific commit-message linter. Its concrete behavior is focused on release/CI hygiene via Conventional Commits validation, which is not accurately or explicitly represented by the broad declared purpose. While commit linting could loosely fit under CI/CD or release management, the description does not meaningfully identify this actual capability, and the code does not implement the broader agent/plugin functionality suggested by the declaration. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk does not implement a broad engineering-agent skillset or the listed domains such as RAG, MCP servers, database design, observability, security auditing, or platform ops. Instead, it is a narrow release-engineering utility for generating changelog content from commit history. While changelog generation could loosely relate to release management, the declared description is materially broader and does not accurately represent this specific code’s primary purpose and concrete capabilities. The code also accesses git history and writes changelog files, which are specific operational behaviors not reflected in the generic declaration.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description describes a broad multi-skill engineering toolkit with numerous advanced capabilities across several domains. The supplied code chunk instead implements a narrow, standalone repository analyzer for onboarding summaries. It walks a directory tree, ignores common build/vendor folders, detects languages from file extensions, checks for common config files, summarizes file sizes and structure, and outputs text or JSON. This is materially different from the declared purpose and does not implement the advanced agent/plugin capabilities claimed.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description presents a general package of many engineering agent skills and plugins across multiple domains, but this code implements one concrete operational script for Git worktree cleanup. Its primary function is local repository inspection and optional destructive maintenance via git worktree remove. That behavior is much narrower and materially different from the declared high-level purpose. While worktree management could loosely fit platform/dev workflow tooling, the declaration does not accurately represent this script’s actual behavior, especially its ability to delete worktrees and its direct interaction with local Git and filesystem resources.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is a high-level catalog of many engineering agent skills and plugins, but this code chunk implements one concrete operational tool: a git worktree manager. Its actual behavior is materially narrower and different from the declared purpose, focusing on local repository manipulation, environment file syncing, port assignment, and optional package installation. Those capabilities are not accurately represented by the supplied description, which does not mention git worktree management or these specific local automation actions. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill set is for engineering agent skills and plugins related to developer/platform domains such as RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, and platform ops. The supplied code instead implements an HR-focused hiring calibration analyzer. Its primary purpose is materially different: it processes interview result JSON, assesses interviewer consistency and demographic disparities, generates coaching/process recommendations, and outputs hiring calibration reports. This is not a supporting implementation detail of the declared engineering-agent/plugin functionality; it is a separate, unrelated capability.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code does not implement agent skills for Claude Code/Codex/Gemini/Cursor/OpenClaw, nor functionality related to RAG, MCP servers, CI/CD, database design, observability, security auditing, release management, or platform ops. Its clear primary purpose is designing and outputting structured interview loops for various job roles and seniority levels. While reading/writing files is a normal implementation detail, here it supports an entirely different domain—recruiting/interview planning—than the declared description. Therefore the description materially misrepresents the code chunk's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description describes engineering agent skills/plugins focused on software engineering operations and assistant tooling domains such as RAG, MCP servers, CI/CD, observability, security auditing, and platform ops. The supplied code instead implements an interview-system designer utility for generating structured interview questions and evaluation materials. Its primary purpose is hiring/interview assessment content generation, which is materially different from the declared agent-skill/plugin suite. The code does not implement the described engineering-agent capabilities; instead it reads role/input data, generates questions, rubrics, and follow-up probes, and writes them to files.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description describes a collection of advanced engineering agent skills and plugins focused on technical infrastructure and software engineering workflows such as RAG, MCP, CI/CD, observability, security auditing, and platform operations. The provided code does none of that: it defines static interview round templates for junior through staff candidates and outputs an interview plan with suggested questions. This is a materially different primary purpose, with capabilities unrelated to the declared domains. There are no notable resource-access or permission issues, but the behavior is clearly outside the stated description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description describes a broad suite of advanced engineering and agent-related skills, but the actual code chunk only implements a basic text processor. Its capabilities are limited to reading local text files, computing simple statistics, transforming text (upper/lower/title/reverse), recursively finding text files by extension, and writing formatted output. There is no evidence of agent/plugin orchestration, RAG, MCP servers, CI/CD workflows, database design features, observability, security auditing, release management, or platform ops. This is a material mismatch in primary purpose and represented capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a large, multi-skill engineering toolkit spanning many domains, but the supplied code chunk implements a specific script-testing utility. Its primary function is to inspect and execute Python files in a skill directory, including running them with no args, --help, and sample files, then generating JSON or human-readable test reports. That behavior is materially narrower and different from the declared purpose. The subprocess execution and filesystem inspection are notable concrete capabilities absent from the description. This is not just a supporting detail of an engineering skill pack; the code shown is a QA/testing harness for Python scripts, so the description does not accurately represent this chunk.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · agent-designer/tool_schema_generator.py (reported line 303)May include surrounding context.

python
if format_name in self.format_validators:
                rules.update(self.format_validators[format_name])
        
        return rules
    
    def _generate_string_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]:
        """Generate string-specific validation rules"""

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · agent-designer/tool_schema_generator.py (reported line 341)May include surrounding context.

python
if format_name in self.format_validators:
                rules.update(self.format_validators[format_name])
        
        return rules
    
    def _generate_string_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]:
        """Generate string-specific validation rules"""

Static analysis

Detected: suspicious.dangerous_exec, suspicious.dynamic_code_execution, suspicious.exposed_secret_literal (+1 more)

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
skill-security-auditor/scripts/skill_security_auditor.py:161

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
skill-security-auditor/scripts/skill_security_auditor.py:154

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
api-design-reviewer/references/api_antipatterns.md:441

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
api-design-reviewer/references/rest_design_rules.md:350

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
helm-chart-builder/references/values-design.md:218

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
tech-debt-tracker/assets/sample_codebase/src/user_service.py:14

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
terraform-patterns/scripts/tf_security_scanner.py:93

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
skill-security-auditor/references/threat-model.md:75

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
skill-security-auditor/SKILL.md:60