Back to skill

Security audit

Schema Markup Generator

Security checks for vulnerabilities and agentic risk

Overview

This schema generator is mostly coherent, but it needs review because it fetches arbitrary URLs and can produce unsafe HTML for websites.

Install only if you are comfortable reviewing URL inputs and generated HTML before use. Avoid running the URL or sitemap features on untrusted links, and do not paste HTML output into a live site unless script-tag breakout characters are escaped or the JSON-LD is otherwise sanitized.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_schema.py:284
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis
` value extracted from the sitemap can cause outbound requests. A malicious sitemap therefore provides a convenient mechanism for issuing multiple requests to destinations accessible from the Agent environment. ### Attack Path 1. An attacker persua ...[truncated 1511 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_schema.py:329
Finding

Unescaped JSON-LD HTML Output Enables Script-Tag Breakout and Cross-Site Scripting

Content
View full analysis
{json_str} """ ``` ### Technical Analysis The function embeds serialized JSON directly inside an HTML `` terminates a script element even when it appears inside a JSON string. An attacker-controlled value such as: ```text ``` can therefore close the JSON-LD element and create a new executable script element. Attacker-controlled values can enter the schema through: - Interactive prompts. - JSON input supplied with `--file`. - Metadata extracted from a URL, including page titles and description or author meta tags. The generated HTML is explicitly intended to be pasted into a webpage's ``, making browser execution a realistic downstream consequence. ### Attack Path 1. An attacker places a script-breakout sequence in a field consumed by the generator, such as a title, description, author, organization name, FAQ answer, or JSON input property. 2. The operator generates output using `--output html`. 3. `json.dumps` retains the literal `<` characters and the `` sequence. 4. `output_html` places the serialized value directly inside the JSON-LD script element. 5. The operator follows the documented workflow and deploys the generated markup to a webpage. 6. A visitor's browser terminates the JSON-LD element at the injected `` sequence and executes the attacker-created script. ### Impact Assessment ...[truncated 789 chars]
Remediation
View remediation
` with `\u003e`. - Replace `&` with `\u0026`. - Escape Unicode line and paragraph separators where browser compatibility requires it. 3. Ensure that no literal case-insensitive ``, mixed-case closing tags, nested script elements, and HTML comments. 6. Apply a restrictive Content Security Policy as defense in depth, while not relying on CSP as the primary correction. 7. Clearly distinguish raw JSON output from HTML-safe output in documentation and require validation before deployment. A hardened implementation can serialize first and then replace the relevant characters: ```python def output_html(schema): json_str = json.dumps(schema, indent=2, ensure_ascii=False) json_str = ( json_str .replace("&", "\\u0026") .replace("<", "\\u003c") .replace(">", "\\u003e") .replace("\u2028", "\\u2028") .replace("\u2029", "\\u2029") ) return f'' ``` ]]>
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

If the implementation is actually validator-only or otherwise substantially narrower than the stated generation behavior, the skill can mislead users and upstream agents into invoking the wrong workflow. While this is more of an integrity and trust issue than a direct exploit primitive, it can still cause unsafe reliance on nonexistent functionality and accidental mishandling of user content.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the implementation is actually validator-only or otherwise substantially narrower than the stated generation behavior, the skill can mislead users and upstream agents into invoking the wrong workflow. While this is more of an integrity and trust issue than a direct exploit primitive, it can still cause unsafe reliance on nonexistent functionality and accidental mishandling of user content.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation is actually validator-only or otherwise substantially narrower than the stated generation behavior, the skill can mislead users and upstream agents into invoking the wrong workflow. While this is more of an integrity and trust issue than a direct exploit primitive, it can still cause unsafe reliance on nonexistent functionality and accidental mishandling of user content.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest promises 'complete, validated Schema.org JSON-LD markup for any content type,' but the code supports only a fixed list of schema types and mostly fills templates from user input without validating required fields, URL structure, or Schema.org conformance. Returning unknown input types as-is from generate_from_file further shows the tool does not ensure completeness or validation.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill advertises and documents operations that imply reading local files, writing output, and fetching remote URLs, but it declares no explicit tool scope or permission boundaries. This is dangerous because an agent may invoke broader file or network capabilities than users expect, increasing the chance of unintended data exposure, uncontrolled outbound requests, or writing artifacts to sensitive locations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation text uses very broad trigger phrases such as any request about schema, structured data, rich snippets, or AI understanding, which can cause the skill to be selected in contexts beyond its actual competence or safety envelope. In an agent system, over-broad routing increases the likelihood of unintended execution, unnecessary file/network operations, and user confusion about what the skill will do.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest describes support for Organization, FAQPage, Article, BlogPosting, Product, HowTo, BreadcrumbList, WebSite, VideoObject, and ImageObject. However, the code detects and emits "LocalBusiness" and defaults to "WebPage", which are outside the stated set of supported outputs. This means the implemented behavior exceeds and diverges from the manifest's declared schema coverage.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest description says the skill creates Schema.org markup for ImageObject, but neither SCHEMA_TEMPLATES nor the CLI choices include ImageObject as a top-level schema type. The implementation only embeds ImageObject as a nested publisher logo object, so the advertised capability is not actually provided.

Content

No source excerpt is available for this finding.

Tainted flow: 'url' from input (line 232, user input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
95% confidence
Finding

The script takes a user-supplied URL and fetches it directly with requests.get() without any allowlisting, scheme restriction, or private-address filtering. In an agent/automation context, this creates a server-side request forgery surface that can be abused to access internal services, cloud metadata endpoints, or other network resources reachable from the runtime environment.

Content

Scanner excerpt · scripts/generate_schema.py (reported line 290)May include surrounding context.

python
import requests
            from bs4 import BeautifulSoup
            
            resp = requests.get(url, timeout=10)
            soup = BeautifulSoup(resp.text, 'html.parser')
            
            template = SCHEMA_TEMPLATES.get(schema_type, {}).copy()

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The example availableLanguage value explicitly limits support to English and Spanish, which can be read as prescribing a fixed language set. In a reference document, this is not framed as an optional example or justified as region-specific, so it may conflict with language/locale choice expectations.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest frames the skill as generating validated Schema.org JSON-LD for content, but this script performs sitemap fetching, multi-page crawling, page-type inference, and bulk file generation across many URLs. While related, this is a broader operational capability than the manifest describes, especially as it introduces autonomous site traversal from a sitemap rather than generating markup for provided content.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The implementation includes a full LocalBusiness template and interactive generation flow, but the manifest's enumerated supported types does not mention LocalBusiness. This creates a mismatch between documented capabilities and actual behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.