Back to skill

Security audit

AutoSpec

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent spec-writing skill; the serious scan alerts mainly point to benchmark output examples, not hidden install or runtime behavior.

Installers should treat this as a lightweight coding/spec assistant. Review any generated specs or code before use, especially for privacy-sensitive history summarization or network-facing URL validation, because the included examples show places where stronger security constraints would be needed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
evals/iteration-1/forward-spec-then-code/with_skill/outputs/implementation.go:46
Finding

Server-Side Request Forgery in Skill-Generated Image URL Validator

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
evals/iteration-1/forward-spec-then-code/without_skill/outputs/implementation.go:100
Finding

Server-Side Request Forgery Through Direct and Redirected URL Validation

Content
View full analysis
= 10 { return http.ErrUseLastResponse } return nil }, ``` ### Technical Analysis The function verifies that the URL has an HTTP or HTTPS scheme and a non-empty host, but it does not determine whether the destination is public and authorized. Direct URLs targeting loopback, private, link-local, or other sensitive network ranges are therefore accepted. The HTTP client also follows up to ten redirects without validating each redirect destination. This permits an attacker to provide an apparently public URL whose server responds with a redirect ...[truncated 2255 chars]
Remediation
View remediation
= 10 { return http.ErrUseLastResponse } return validatePublicDestination(req.URL) }, ``` 4. Do not rely only on pre-request DNS validation. Add a custom `DialContext` that validates the resolved address immediately before connecting. 5. Reject loopback, private, link-local, multicast, unspecified, reserved, and metadata address ranges for both IPv4 and IPv6. 6. Restrict allowed destination ports and prefer HTTPS. 7. Consider disabling redirects entirely if redirect support is not required for the declared validation behavior. 8. Apply egress network controls so the process cannot contact internal networks or metadata services. 9. Maintain strict request timeouts, response-header timeouts, connection limits, and overall validation concurrency limits. 10. Add security tests for direct internal addresses, public-to-private redirects, IPv6 destinations, alternative IP encodings, internal DNS names, and DNS rebinding. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a developer-assistance skill focused on specification authoring and reverse-spec analysis. The supplied code does not do any spec-related analysis or assistance. Its primary purpose is operational URL processing: deduplication, syntax validation, and live reachability checking over the network. The outbound HTTP requests are a significant undeclared capability and materially different from the declared intent. Therefore, the description does not accurately represent the code's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a development-assistance skill focused on writing or deriving behavioral specs from code. The supplied code does not perform any spec generation, code understanding, intent alignment, or user-facing analysis. Instead, it is an operational utility function for URL deduplication and remote accessibility validation, including outbound network requests, timeout handling, redirects, concurrency control, and logging. That is a materially different primary purpose and includes undeclared network-access capability, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 66)May include surrounding context.

md
| A feature/capability | Feature-level | Complete behavior of "quota management" |
| A system/service | System-level | Overall behavior of "order system" |

Don't ask the user to specify granularity — infer from context. When uncertain, go one level higher than what the user described, then drill down into key parts within the spec.

### Spec Template

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The design automatically forwards prior conversation content to a summarization LLM and persists a transformed version of that history without any user-facing notice, consent model, or documented data-handling boundary. That creates privacy and compliance risk because sensitive user content may be reprocessed by another model endpoint and retained in a modified form outside the user's expectations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The spec is internally inconsistent about storage semantics: one section says the summarized sequence is persisted back to storage, while another says raw message storage is unchanged. This ambiguity can cause implementers to overwrite canonical conversation records with lossy summaries or deploy a design that mishandles audit, recovery, privacy, or legal-retention requirements.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This function performs live outbound HTTP HEAD requests as part of 'validation', which goes beyond passive spec/understanding behavior and can be abused to probe attacker-controlled or internal URLs. In agent or backend contexts, this creates SSRF-like risk, unexpected network side effects, and data leakage through request metadata and logs.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The implementation actively probes remote resources instead of only deduplicating or analyzing inputs, introducing network interaction not implied by the skill's spec-oriented purpose. Because the URLs are attacker-influenced, this can trigger requests to internal services, third-party endpoints, or tracking infrastructure, causing SSRF-style exposure and operational side effects.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The spec explicitly directs the implementation to issue HTTP HEAD requests to user-supplied URLs, which creates server-side request forgery and privacy exposure risk if untrusted input is accepted without restrictions. It also specifies logging invalid or failed URLs, which can disclose sensitive user-provided URLs or internal endpoints into logs; the skill context makes this more dangerous because the spec is likely to be implemented directly by an agent or developer without additional security guardrails.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file documents behavior that uses request context containing user and session information, refreshes profile info, persists message history, and calls an external LLM service. Under the markdown-specific warning rule, the description should disclose privacy-relevant behavior because it affects user data and external processing.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

L21-L30 describe concrete decision logic: branching tool sets by client version and disabling or rewriting tool descriptions via ConfigOverrides. That contradicts L52's claim that the module is "purely responsible for assembling and starting the agent" and contains no business logic.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

L07 says the file is the "main entry point" of the AI assistant agent, while L14 says Generate is "the main external entry point." These statements are not merely incomplete; they present conflicting descriptions of what the true entry point is.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
evals/iteration-1/reverse-spec-module-level/with_skill/outputs/spec_output.md:94