Back to skill

Security audit

eval-integrity

Security checks for vulnerabilities and agentic risk

Overview

This skill is a read-only benchmark audit workflow whose repository searches, subagent use, and GitHub PR checks fit its stated purpose.

Install only for repositories you are comfortable letting the agent inspect. Expect it to read benchmark files, CI config, results, and open PR metadata, but the artifact instructs the agent not to modify the repo or rerun evaluations unless the user later approves fixes.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep

Static analysis

No suspicious patterns detected.