Back to skill

Security audit

Large Model Visual Question Answering Skill | 大模型视觉问答技能

Security checks for vulnerabilities and agentic risk

Overview

This skill goes beyond simple visual Q&A by uploading media to cloud APIs and automatically managing a cloud-linked identity and history records.

Install only if you are comfortable with user-provided media and URLs being sent to the configured cloud service, with a local workspace database storing account/token data, and with automatic account-linked history queries. Avoid using it for sensitive personal, business, medical, or confidential images unless the publisher clarifies retention, identity, and deletion controls.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (55)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as generic visual question answering, but the associated behavior described by analysis suggests broader and different operations such as record listing, backend retrieval, local persistence, identity handling, and generic analysis workflows. This mismatch is dangerous because users and orchestrators may authorize the skill under false assumptions, causing unintended data access, uploads, or account-linked operations beyond simple image Q&A.

Content

No source excerpt is available for this finding.

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · SKILL.md (reported line 3)May include surrounding context.

md
---
name: "visual-qa-analysis"
description: "Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答"
version: "1.0.17"
license: "MIT-0"
---

# ❓ Large Model Visual Question Answering Skill | 大模型视觉问答技能
> **智能分析中枢** · 图片/视频智能分析 · 结构化报告 · 历史报告云端查询

---

## 🧭 技能概览 | Overview

| 模�

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation does not clearly warn users that local files, image URLs, and history queries are transmitted to external API/cloud services. In a vision-analysis skill that accepts attachments and automatically queries history, this omission undermines informed consent and can expose sensitive images, documents, metadata, and prior records to third-party processing.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documented API describes pet health analysis workflows, which materially conflicts with the skill’s declared purpose of visual question answering. This mismatch is dangerous because it can hide undisclosed data flows or capabilities, causing users and reviewers to grant access under false assumptions and potentially exposing sensitive pet/health-related data to an unexpected backend.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The core documented behavior focuses on starting health analysis jobs, retrieving results, querying historical reports, and exporting full reports rather than answering questions about image content. In a skill presented as generic visual Q&A, this indicates significant scope misrepresentation and raises the risk of unauthorized collection, retention, or export of sensitive analysis records beyond what users would reasonably expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest says this skill conducts open-ended question answering on image content, but the code explicitly requires a local or remote video input and sends it as videoUrl or uploaded file for analysis. Additional methods generate analysis reports and report-history listings, which is materially different from interactive image VQA.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The implementation materially diverges from the declared skill purpose: it performs video analysis and exposes a history-listing function instead of image-based visual question answering. This kind of capability mismatch is dangerous because users, orchestrators, or policy systems may grant the skill access based on the manifest description, while the code performs broader or different operations, including retrieval of prior analysis records via open_id.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file exposes a broad generic network/CRUD wrapper layer, including add, edit, delete, and arbitrary HTTP methods, which is materially broader than what an image VQA skill should need. In the context of a vision Q&A skill, this unnecessary capability expansion increases the attack surface and can enable unintended data access or state-changing backend actions if reachable by prompts or higher-level skill code.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This code derives, creates, and persists user identities outside the stated purpose of a visual Q&A skill. It can silently reuse a local API-key-derived identity or create a synthetic default account and store it, creating covert account linkage and enabling downstream authenticated network actions without clear user consent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The HTTP helper automatically performs phone-login/registration, fetches and caches tokens, and retries authorization, all unrelated to image question answering. This creates a hidden authenticated channel that can register accounts and transmit identifiers/tokens to external services, which is especially risky in a utility imported by many skills.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares no explicit tool scope even though the documentation indicates capabilities including shell execution, network access, filesystem reads/writes, and environment use. In an agent setting, missing permission boundaries increases the chance the skill can be invoked with broader powers than users or orchestrators expect, especially given it also handles local files and cloud queries.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The default trigger condition is ambiguous and may cause the skill to auto-activate whenever an image and question are present, even if another safer or more appropriate skill should handle the request. In practice this can lead to unnecessary uploads, shell/script execution, and cloud processing of user-provided files or URLs without sufficiently clear routing or consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The history-query trigger phrases are broad enough to activate on ordinary conversation, which can cause the agent to query cloud history unexpectedly. In this skill context, that creates privacy and data-minimization risk because historical records may be fetched or exposed without a clear, deliberate user request.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill exposes an analysis-history listing function that is not described by the manifest’s stated purpose of answering questions about images. Hidden or under-declared data access features increase the risk of unauthorized access to prior user activity, especially because the function takes an internal user identifier and may reveal historical analysis records.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.install_untrusted_source

Install source points to URL shortener or raw IP.

Warn
Code
suspicious.install_untrusted_source
Location
skills/smyx_common/scripts/config-dev.yaml:2