T09 · Insecure Skill Coding Practices
- Location
references/examples.md:94- Finding
Untrusted LLM Output Controls Generated Agent Prompts and Tool Declarations
- Content
View full analysis
Vulnerability Details
File Location:
references/examples.md:94-122, 198-217, 328, 629-643
Vulnerability Type: Improper validation of untrusted model output in an agent-generation pipeline
Risk Level: MediumVulnerable Code
typescript const prompt = ` Analyze the following agent requirements and extract: 1. Primary domain and expertise area 2. Required capabilities and skills 3. Tool requirements 4. Constraints and limitations 5. Expected input/output patterns 6. Performance requirements Requirements: ${JSON.stringify(requirements, null, 2)} Provide analysis in JSON format. `; const response = await this.openai.chat.completions.create({ model: 'gpt-4', messages: [ { role: 'system', content: 'You are an expert in AI agent design and analysis.' }, { role: 'user', content: prompt }, ], response_format: { type: 'json_object' }, }); const analysis = JSON.parse(response.choices[0].message.content!); return { domain: analysis.domain, expertise: analysis.expertise, capabilities: analysis.capabilities, tools: analysis.tools, constraints: analysis.constraints, patterns: analysis.patterns, performance: analysis.performance, };The unvalidated output is subsequently propagated into the generated prompt and agent metadata:
typescript const context = { domain: analysis.domain, expertise: analysis.expertise, capabilities: capabilities.primary.map(c => ({ name: c.name, description: c.description, examples: c.examples, })), supportingCapabilities: capabilities.supporting.map(c => ({ name: c.name, description: c.description, })), tools: analysis.tools, constraints: analysis.constraints, bestPractices: this.generateBestPractices(analysis, capabilities), approach: this.generateApproach(analysis, pattern), }; return template(context);typescript const metadata: AgentMetadata = { name: this.generateAgentName(components.domain, components.expertise), description: ...[truncated 3505 chars]- Remediation
View remediation
Remediation Suggestions
-
Define and enforce a strict runtime schema for the complete analysis response. Reject missing fields, unexpected fields, invalid enum values, oversized strings, and malformed arrays.
-
Select tools exclusively from a server-controlled allowlist. Intersect model-suggested tools with tools authorized for the selected pattern and deployment context:
typescript const AnalysisSchema = z.object({ domain: z.string().min(1).max(100), expertise: z.array(z.string().max(100)).max(20), capabilities: z.array(z.string().max(100)).max(20), tools: z.array(z.enum(['Write', 'Read', 'Grep', 'Glob'])).max(10), constraints: z.array(z.string().max(500)).max(20), patterns: z.array(z.string().max(100)).max(20), performance: z.object({ responseTime: z.number().positive().optional(), accuracy: z.number().min(0).max(1).optional(), reliability: z.number().min(0).max(1).optional(), }), }).strict(); const parsed = AnalysisSchema.parse( JSON.parse(response.choices[0].message.content!) ); const authorizedTools = parsed.tools.filter(tool => pattern.supportedTools.includes(tool) );-
Do not treat model-generated tool declarations as authorization decisions. Apply least privilege in downstream infrastructure and require an explicit policy decision before granting each tool.
-
Place user-controlled data in clearly delimited sections and instruct the analysis model to treat it as data rather than instructions. This reduces risk but must not replace output validation.
-
Add semantic policy checks for generated prompts. Reject instructions that attempt to override governing policies, disclose secrets, modify persistent state, execute unrelated commands, or expand privileges.
-
Require human review before activating generated agents that request sensitive capabilities such as shell execution, filesystem writes, network access, credential access, or administrative APIs.
-
Add adversarial tests covering prompt injection, un ...[truncated 168 chars]
-
