Security audit
Aa Benchmarking Framework
Security checks for vulnerabilities and agentic risk
Overview
This is a draft benchmarking skill with only descriptive Markdown and no hidden execution, network, credential, or persistence behavior in the inspected artifact.
Reasonable to install as a draft reference skill. It is not yet feature-complete, so users should re-review any future version that adds LangFuse integration, CSV/JSON export implementation, dependencies, or network access.
Vulnerability Patterns
- Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
- Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
- Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
- Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
- Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
- Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
- Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
- Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
- Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
- Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Static analysis
No suspicious patterns detected.
