Back to skill

Security audit

academic-figures

Security checks for vulnerabilities and agentic risk

Overview

This is a local academic chart-generation skill with disclosed file output and environment checks; it has documentation and language rough edges, but no evidence of hidden exfiltration, persistence, or unsafe automatic execution.

Install only if you are comfortable running local Python scripts on your data and installing the listed scientific Python dependencies yourself. Expect Chinese/CJK-oriented prompts and examples, and review output paths before using batch or pipeline manifests because they can write multiple files. The optional setup script is local, but it may clear your matplotlib font cache.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (82)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code does not perform chart creation, export, theming, YAML pipelines, wizard selection, annotation, or watchdog retry logic itself. Instead, it is a support module for diagnostics and fallback error handling. While such diagnostics could plausibly be part of a larger figure tool, this chunk’s actual behavior is materially different from the declared primary purpose because it only analyzes exceptions and emits categorized troubleshooting messages. There is no evidence of network access or resource misuse, but the description does not accurately represent this code chunk’s function.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents a broad end-user skill with many high-level capabilities beyond plotting. The supplied code chunk is specifically a drawing/rendering module ('only does plotting' per its own header) and does implement meaningful scientific chart generation locally, so the general domain is aligned. However, the description overstates the implemented functionality in this chunk in several material ways. Many named capabilities are absent here: export pipeline, YAML orchestration, wizard mode, watchdog/retry logic, publication QA gates, and multiple claimed chart types. Because the code’s actual behavior is substantially narrower than the declared feature set, the description does not accurately represent what this supplied code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

The supplied code is clearly related to scientific figure generation and is consistent with the local/offline aspect: it performs plotting/statistical visualization and does not access network or external services. However, the declared description presents a much broader end-user skill with many features across chart coverage, export pipeline, UI/CLI helpers, and publication QA controls. This code chunk only covers a narrower extension module for several chart types and related validation/alt/caption support. Most headline capabilities in the description are not represented here. In addition, the module docstring says it depends only on numpy + matplotlib + af_v23_stats, but the actual code also imports SciPy inside several functions. So while the code is on-topic, the declared description materially overstates what this specific code chunk does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description is centered on figure creation and export: chart types, themes, journal presets, YAML orchestration, local rendering, annotation, watchdog behavior, and multi-format output. The supplied code chunk instead contains numerical/statistical routines and validation logic, with no plotting/rendering code, no file export, no themes/palettes, no journal presets, no wizard, no YAML pipeline handling, and no watchdog or PDF overlap/font-gate logic. Some implemented analyses conceptually support certain declared chart types (e.g., KM, ROC, funnel, Cox forest) by providing underlying statistics, so they may be supporting components of a larger figure system. However, this chunk on its face performs substantial undeclared analytical capabilities and does not itself implement the main declared behavior of producing publication-ready figures. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk implements only one narrow quality-check feature: auditing text sizes in a PDF using PyMuPDF. While the description does mention 'PDF text-overlap + font-size gates,' the overall declared purpose is a comprehensive publication-figure generation system. This code does not create figures, render charts, export TIFF/PNG/PDF, provide a Python API for plotting, or implement the many listed charting and workflow features. Therefore, for this chunk, the actual behavior is materially different from the declared primary purpose and represents only a small auxiliary auditing component of the described tool.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code is broadly aligned with the declared product category: it is indeed a local publication-oriented academic figure generator with themes, journal presets, pipeline/batch workflows, annotations, validation, and high-DPI export. However, the description presents a very specific feature inventory that is not accurately represented by this supplied code chunk. Several advertised features are only partially evidenced or delegated to external modules not shown here, especially the full set of claimed chart types and some quality-control/retry behaviors. There are also count mismatches: the code visibly defines more than 9 themes, and journal presets are dynamically loaded rather than cleanly matching the stated fixed set. In addition, the code has several substantial undeclared capabilities—Excel ingestion, batch processing, CJK/bilingual handling, accessibility artifact generation, diagnostics/doctor mode, and auto-suggestion/quick rendering. Because the evaluation criterion flags undeclared capabilities and materially inaccurate descriptions, this should be considered a mismatch rather than a perfect description-to-behavior match.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a figure-generation skill whose primary purpose is producing publication-ready scientific figures with extensive rendering/export/pipeline capabilities. The actual code chunk does not render figures at all; it generates prose legends from JSON metadata for different chart types. While this could be a supporting utility within a broader figure-making package, the description does not mention legend generation, and the code’s primary behavior is materially different from the declared purpose. There are no suspicious external resource accesses; it is local file I/O only. The mismatch is therefore due to behavior/purpose divergence, not permissions or triggers.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a full-featured scientific figure creation skill. However, the provided code chunk only performs installation/self-diagnostic tasks: checking Python version, testing imports for required libraries, printing manual pip install instructions, deleting the local matplotlib font cache, and invoking a font-detection helper script. While these actions are plausibly supportive of a figure tool and remain local, the code’s primary purpose in this chunk is environment preparation rather than generating publication-ready figures or exposing the listed chart/pipeline/export capabilities. Therefore, the code chunk does not accurately represent the declared functionality and should be marked as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code chunk is narrowly focused on one quality-check function: verifying text overlap in PDFs by rendering pages, extracting text spans with PyMuPDF, and confirming overlap via pixel-level distance measurements. This aligns only loosely with the single claimed feature about PDF text-overlap gates, but the declared description presents the skill primarily as a comprehensive scientific figure-generation system with many chart types, themes, export formats, and workflow features. None of those core generation/export/pipeline capabilities appear in this code chunk. The code is local-only and does not show suspicious external data access, but its actual behavior is materially narrower and different from the declared overall purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises and documents shell execution, environment setup, and file read/write behavior, but it does not declare an explicit tool scope such as allowed-tools or permissions. In an agent setting, missing scope boundaries increases the chance that a caller or orchestrator grants broader capabilities than intended, enabling command execution or file access beyond figure generation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The markdown states that over-limit inputs are rejected with a Chinese error message, which imposes a specific language for user-facing output. This is a natural-language locale policy concern because the skill does not indicate user choice or opt-in for that language behavior.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

These lines describe an expanded Chinese exception mapping, indicating the skill returns errors in Chinese. Because the file is bilingual elsewhere and does not state that language is selectable by the user, this suggests a locale preference may be enforced without explicit opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The phrase '中文提示' explicitly says user-facing error prompts are in Chinese. There is no indication here that users can choose another language or that the skill is limited to a Chinese-only environment, so this may violate language/locale policy.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The changelog explicitly says the exception diagnosis system uses a 'Chinese three-part diagnosis', which is a natural-language policy concern when presented as the default output behavior. There is no nearby indication that users can choose another language, so the file documents a forced locale behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file includes natural-language statements that validation messages are 'Chinese-first' and that CLI messaging was made 'fully Chinese'. That indicates a language/locale constraint presented as default behavior, without offering users a clear language choice or opt-in, which conflicts with the policy against forcing a specific language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The cheat sheet states lang=zh 中文标准措辞, which reads as a prescribed Chinese-language requirement rather than an optional locale choice. Under the language/locale policy rule, forcing a specific language without user opt-in is a natural-language policy violation unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file is primarily written in English, but lines L507-L546 switch to Chinese for the PRISMA and capability-reference sections. That creates a language/locale policy concern because the skill documentation effectively forces a different language for part of the experience without stating that this is optional or region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The FAQ says --caption will generate '中英双语' captions, which imposes a specific language/locale behavior rather than offering user choice. Under the policy, language constraints should be optional, user-selected, or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Line L004 says errors will be shown as '中文报错', which imposes a specific language behavior in the skill documentation. Under the policy, forcing a language without offering user choice or documenting a justified locale constraint is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown includes substantial operational guidance written in Chinese, including normative instructions for agent behavior, without stating that language selection depends on user preference. A skill that presents or enforces instructions in a specific language can violate language/locale policy when no opt-in or choice is offered.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file explicitly describes stderr messages as Chinese-only and later repeats that validation messages are fully in Chinese. This is a natural-language locale policy issue because it imposes a specific language on users without indicating any option to select another language.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.insecure_tls_verification

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
scripts/gen_figure.py:2917