Back to skill

Security audit

Agent Evaluation Benchmark Engine

Security checks for vulnerabilities and agentic risk

Overview

This is a documentation-only benchmarking skill whose evaluation, regression testing, and reporting guidance matches its stated purpose.

Installers should expect an Indonesian-language benchmarking and evaluation guide. It may recommend security tests, regression gates, and low-risk reversible rollback decisions, so users should keep human review around important deployments.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Confidence
94% confidence
Finding
The phrase 'Use when the conditions in the description are met' does not identify concrete trigger phrases or boundaries for when the skill should be invoked. Because it refers back to the description rather than specifying clear activation criteria, it can lead to inconsistent or unintended invocation decisions.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
This markdown file contains the operational guidance predominantly in Indonesian, beginning at L0034 and continuing through the rest of the document, while the skill does not state that it is intentionally limited to Indonesian users or provide an opt-in language choice. That can violate language/locale policy because it implicitly forces one language for users who may expect English or another locale.

Static analysis

No suspicious patterns detected.