Back to skill

Security audit

Report Creator

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent report creator, but generated reports load mutable third-party browser code by default, which can expose report contents if that code is compromised.

Before installing, decide whether your reports may contain confidential data. For sensitive reports, use or require bundled/offline assets, pin and verify any CDN scripts, and avoid opening generated reports where third-party network loads are not acceptable. Only combine this skill with Telegram or other messaging tools when you explicitly intend to send the output externally.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T03 · Remote Payload Retrieval and Execution

Warning
Location
templates/en/dark-tech.html:7
Finding

Generated reports execute mutable third-party JavaScript without integrity verification

Content
View full analysis
``` ```javascript /* Preload html2canvas eagerly — fires while user reads, so first export is instant */ let libPromise = null; function loadLib() { if (libPromise) return libPromise; libPromise = new Promise(resolve => { if (window.html2canvas) { resolve(); return; } const s = document.createElement('script'); s.src = 'https://cdn.jsdelivr.net/npm/html2canvas@1/dist/html2canvas.min.js'; s.onload = resolve; document.head.appendChild(s); }); return libPromise; } loadLib(); /* start loading immediately */ ``` ### Technical Analysis The template loads and executes JavaScript directly from public CDNs. The Chart.js and html2canvas URLs select mutable major-version ranges (`@4` and `@1`) rather than exact, audited releases. None of the external scripts use Subresource Integrity. Consequently, the effective executable code can change after the Skill package has been reviewed. The html2canvas dependency is injected eagerly as soon as the report is opened, even if the user never requests image export. This network and code-execution behavior is not the minimum privilege necessary for viewing a generated report. Third-party JavaScript executes with the same browser-page privileges as the report’s own scripts. It can read the report DOM, including confidential report text, KPIs, tables, and user edits; observe browser interactions; modify rendered content or exported images; and initiat ...[truncated 1704 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (266)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The README describes automatic sending of generated output to a Telegram channel without any privacy, authorization, or data-transmission warning. In an agent environment, this is a significant exfiltration risk because reports may contain sensitive business data, and users may not realize content is leaving the local environment for a third-party platform.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The description claims a user-facing content generation skill for creating reports and related artifacts, including multiple generation/review/export-style flags. The actual code instead implements a QA/validation utility for ensuring documentation consistency around review-related contracts. Its behavior is limited to reading specific local markdown files and checking for required/forbidden snippets. While some checked strings reference review/report concepts, the script itself does not perform those capabilities. This is a material mismatch in primary purpose and functional behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Yes, this is a mismatch. The declared description describes a user-facing content generation skill with substantial functionality for creating reports and related artifacts. The provided code chunk, however, is only an init.py file containing a docstring: 'Evaluation helpers for kai-report-creator.' That indicates at most package/module labeling for evaluation utilities, not the declared generation behavior. There is no implemented functionality in the supplied chunk matching the described purpose, inputs, flags, or outputs. While this may be a partial repository file, based on the supplied code alone the actual behavior is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code chunk is focused on contract checking: regex-based validation of placeholders, dates, KPI values, timelines, chart/diagram block structure, frontmatter parsing, and counting/collecting document components. This is related to report content quality control, but it is not the declared primary function of creating or generating reports, dashboards, business summaries, or research documents. None of the prominent declared features such as --plan, --generate to HTML, --review automatic refinement, --themes preview styles, bundling, or image export are implemented here. This is therefore a material description-behavior mismatch rather than a mere supporting implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

This code does not implement report creation, planning, HTML generation, review, theming, bundling, or export behavior described in the declared purpose. Its primary purpose is operational/testing support: changing into the script directory, selecting test mode based on --fast, and running pytest over specific test files or the full tests directory. That is materially different from the user-facing generation workflow in the description, so this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose describes a content/report generation skill with bilingual support and multiple generation/review/export workflows. The supplied code instead implements a maintenance script for cleaning or checking Python cache files in a repository. Its primary purpose, CLI surface, and effects on the filesystem are entirely unrelated to report or dashboard generation. This is not a minor implementation detail; it is a different tool altogether.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents an end-user content generation skill: create/generate reports, dashboards, and research documents, including planning, HTML rendering, review, styling previews, bundling, and export-image workflows. The supplied code instead implements internal analysis utilities for context isolation and IR snapshot comparison. Its primary behavior is parsing frontmatter/body, finding one valid IR block in mixed context, computing hashes, inferring metadata such as theme/archetype, collecting headings/component counts, and comparing expected vs actual snapshots. There is no code here for generating reports, rendering HTML, handling --plan/--generate/--review flags, ingesting URLs/files for document creation, bundling, or exporting images. This is a materially different primary purpose, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

This is a clear description/behavior mismatch. The declared purpose centers on creating or generating reports/documents/dashboards and mentions generation-oriented flags. The actual code does not generate reports at all; it only renders an existing HTML report into image files using Playwright screenshots. Moreover, the description explicitly states that exporting finished HTML to PPTX/PNG should use a different skill, making this code not just incomplete relative to the declaration but directly aligned with an excluded capability. No evidence appears for the declared generation, review, planning, theming, bundling, or source-ingestion behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description presents a content-generation skill whose core purpose is creating reports, dashboards, and research documents in Chinese/English with multiple generation-related modes. The code chunk instead implements a narrow preprocessing/guard validation utility for IR text before HTML generation. While such validation could be a supporting internal component of a report-generation system, this specific code does not itself perform the declared end-user behavior and lacks the advertised generation features and flags. Therefore the code chunk’s actual behavior is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description centers on creating/generating reports and dashboards, including multiple user-facing generation modes and export-related flags. The actual code chunk does not generate any content at all; it reads an existing HTML file and validates structural markers, theme fidelity, KPI numeric values, and embedded JSON metadata. This is a materially different primary purpose: quality assurance for final HTML output rather than report creation. While such validation could support a report-generation pipeline, this code chunk’s behavior is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill is for creating or generating reports, dashboards, business summaries, and research documents in Chinese/English. The supplied code does not implement document generation at all. Its primary purpose is to run late-context isolation evaluations over repo-contained cases, compare expected vs extracted snapshots, and report drift metrics. The CLI interface, inputs, outputs, and internal logic are all centered on evaluation/testing infrastructure, not report authoring or rendering. This is a clear material mismatch in primary purpose and capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents an end-user content creation skill for generating reports and dashboards. The actual code does not generate reports from user inputs, URLs, notes, or plans, and it does not implement the described flags such as --plan, --generate, --review, --themes, --from, --bundle, or --export-image. Instead, it is an internal evaluation utility for checking existing report artifacts against structural and rendering requirements. This is a materially different primary purpose, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill should be used to create or generate reports, dashboards, business reports, and research documents, including planning, rendering to HTML, review, themes, bundling, and export-related flags. The supplied code does none of that. Its primary purpose is to verify a repository release pipeline by invoking other scripts and checks: cleaning generated files, running pytest, running eval scripts, checking doc sync, and performing an image export smoke test on a fixture HTML file. While one step touches export-image functionality, it is only a QA smoke test and not user-facing report generation. Therefore the code’s actual behavior is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents an end-user content generation skill for producing reports, dashboards, and research documents. The actual code does not generate, plan, review, theme, bundle, or export reports. Instead, it is a test module that reads existing HTML files and asserts rendering and structural constraints related to charting and report layout. This is a materially different primary purpose: QA/regression validation of generated reports rather than report creation. While the tested artifacts are reports/dashboards, the code’s capability is not merely an implementation detail of generation; it is a separate testing function that is not represented in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a user-facing content generation skill for creating reports, dashboards, and research documents. The actual code does not implement generation, planning, review, theming, or export workflows. Instead, it is a frontend contract test focused narrowly on validating computed CSS styles in a local HTML fixture. This is a materially different primary purpose from the declared skill behavior, so the description does not accurately represent the supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description describes an end-user content generation skill focused on producing reports, dashboards, and research documents in Chinese and English. The actual code chunk is not implementing any such generation workflow. Instead, it is a test suite that reads local repository files and asserts that CSS, documentation, templates, and demo pages contain specific strings related to a corporate-blue theme and comparison-mode badge behavior. This is a materially different primary purpose: internal validation of styling/docs rather than report generation. There is no evidence of the declared flags, no rendering pipeline for user content, and no ingestion of notes/data/URLs. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a user-facing document/report generation skill, including planning, rendering HTML, refinement, theming, bundling, and export-related flags. The supplied code does not implement report creation or rendering behavior. Instead, it is test code for internal parsing and context-isolation behavior around IR extraction and snapshot comparison. This is a materially different primary purpose from the declared functionality. There is no evidence in this chunk of generating reports, dashboards, research documents, handling URLs/raw notes for authoring, or supporting the listed CLI-style flags. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill generates reports, dashboards, and research documents. The actual code does not generate any report content, render HTML, preview themes, process raw notes/data/URLs, or handle the described CLI flags as a user-facing report tool. Instead, it tests whether repository documentation files stay in sync with a skill contract by creating temp repos and invoking a doc-checking script. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill should generate reports, summaries, dashboards, or research documents from inputs and support generation-related flags. The supplied code does not implement any end-user report generation behavior. Instead, it is strictly a test file for evaluation contracts around report cases, including filesystem checks, schema validation, and subprocess execution of an eval runner. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description focuses on generating reports, dashboards, and research documents, and even explicitly excludes exporting finished HTML to PNG/PPTX in favor of another skill. The actual code is not about report generation at all; it is a test suite for an image export script that captures screenshots of HTML with specific rendering/export settings. This is a materially different primary purpose and directly aligns with the excluded export-image functionality rather than the declared report-creation behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description presents an end-user content generation skill for creating reports, dashboards, and research documents with multiple generation-related flags. The actual code chunk does not generate reports or render HTML; instead, it is a test module that exercises guard_validate behavior and enforces validation contracts. Its main role is QA/regression testing around validation semantics, file/context input handling, and reuse of shared validator logic. That is a materially different primary purpose from the declared report-generation functionality, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents an end-user content-generation skill whose main function is to create or render reports and related documents. In contrast, the supplied code is entirely a test suite for enforcing an HTML shell contract after generation. It uses pytest, reads fixture/template files, and checks for required IDs, CSS classes, JS bindings, metadata placeholders, watermark formatting, and forbidden legacy patterns. That is a materially different primary purpose: quality assurance and regression prevention rather than report generation. While these tests support a report-generation system, this chunk itself does not handle multilingual generation, inputs like URLs or notes, planning/review flows, bundling, or export rendering. Therefore the description does not accurately represent the actual behavior of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill generates reports, dashboards, and research documents, including planning, HTML rendering, review, theming, bundling, and image export. The actual code chunk does none of that. It is purely a test module that reads local documentation files and asserts the presence or absence of specific strings. Its primary purpose is documentation/contract validation for another skill or template, not content generation. There are no generation workflows, no handling of notes/data/URLs, no CLI flag implementation, and no report/dashboard rendering behavior in the provided code. This is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill should be used to create or generate reports, dashboards, business reports, and research documents, including planning, rendering, reviewing, theming, and file/URL-based generation workflows. However, the code shown is only a test file for validation logic around intermediate report blocks (KPI, timeline, chart, diagram). It asserts whether sample YAML/markdown-like blocks are valid, invalid_syntax, or invalid_semantics, and checks example files against those validators. This is materially different from a report-generation skill's primary purpose. While validation could be a supporting implementation detail inside a larger report creator, the supplied chunk itself exposes a distinct capability—schema/contract validation tests—and none of the core declared generation behaviors appear in the code. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a content-generation/reporting skill, but the supplied code is clearly test infrastructure for an evaluation pipeline. Its primary purpose is to verify that an evaluation script and manifest exist, that referenced context files resolve, and that running the eval script produces expected summary metrics. There is no evidence of report generation, dashboard creation, HTML rendering, review/theming, bundling, or export behavior. This is a material purpose mismatch, not merely an implementation detail.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/conftest.py:18