Back to skill

Security audit

alphaear-search

Security checks for vulnerabilities and agentic risk

Overview

The skill is a finance search/RAG tool, but it also performs under-disclosed external enrichment, sentiment analysis, model downloads, LLM-provider use, and local database writes.

Review this skill carefully before installing. It is not just a search helper: default structured search may contact Jina Reader, cache retrieved page content locally, score sentiment, load external model artifacts, and use configured LLM providers. Install only if those third-party services, environment credentials, local database writes, and model-download behavior are acceptable for your finance data and queries.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
scripts/sentiment_tools.py:48
Finding

Automatic Retrieval of Unpinned Model Artifacts

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/search_tools.py:262
Finding

Default Enrichment Discloses URLs and Content to Additional External Services

Content
View full analysis
List[Dict]: ``` The enrichment loop submits each result URL to the content extractor and may subsequently analyze the retrieved content: ```python if enrich and normalized_results: logger.info(f"🕸️ Enriching {len(normalized_results)} search results with Jina & Sentiment...") extractor = ContentExtractor() # Lazy load sentiment tool if not hasattr(self, 'sentiment_tool') or self.sentiment_tool is None: from .sentiment_tools import SentimentTools self.sentiment_tool = SentimentTools(self.db) for item in normalized_results: if item.get("url"): try: # If Jina Search already returned enough content, skip another fetch if skip_content_enrichment and item.get("content") and len(item.get("content", "")) > 100: full_content = item["content"] else: # Use Jina Reader to get full content full_content = extractor.extract_with_jina(item["url"], timeout=60) if full_content and len(full_content) > 100: item["content"] = full_content # Calculate sentiment using the title and content prefix text_to_analyze = f"{item['title']} {full_content[:500]}" sent_result = self.sentiment_tool.analyze_sentiment(text_to_analyze) score = sent_result.get('score', 0.0) item["sentiment_sc ...[truncated 3938 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (30)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Sentiment analysis with downloaded transformer models and external LLM calls, plus writing sentiment back to a news database, is substantially broader than the declared search/RAG scope. This increases supply-chain, data-egress, and integrity risks because the skill may download code/models, send data externally, and modify local datasets without clear disclosure.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file implements sentiment analysis and bulk database updates, which materially exceed the stated finance search/RAG behavior of the skill. This kind of hidden or undeclared capability increases attack surface, can surprise operators, and may enable unintended processing or modification of local data beyond user expectations.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares capabilities that imply network and environment access but does not define an explicit tool/permission scope. This creates ambiguity about what the skill is allowed to do and increases the risk of over-broad execution, unexpected external requests, or access to secrets through environment variables.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Broad invocation language like 'use when the user needs general finance info' lacks clear boundaries and exclusions, making accidental or overly frequent activation more likely. In combination with networked and possibly broader-than-declared behavior, vague triggers can cause the skill to run in contexts where sensitive data or unnecessary external access is involved.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Multiple docstrings and inline descriptions are exclusively in Chinese, with no indication that users may choose another language or locale. Under the stated policy, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code performs a safety-relevant model load and text encoding step via SentenceTransformer, which may download models or transmit text depending on backend configuration, but there is no confirmation prompt, user-visible disclosure, or explanatory comment about that behavior. The logging present is developer-oriented and does not clearly warn users that their query/document text may be processed by an external model stack.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest includes local finance retrieval, and this method specifically claims to load recent N days of history. However, the SQL query does not filter by date at all, so behavior does not match the described time-bounded retrieval semantics.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This is a direct contradiction between inline documentation and code behavior. The method documentation promises time-scoped history loading, while the implementation retrieves rows solely by descending publish_time and LIMIT.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest describes a skill for finance web searches and local document retrieval, but this file implements generic LLM tool-calling capability detection and model registry caching. Probing whether an LLM supports native function calling is not an obvious requirement of finance search or local context retrieval, making this capability unjustified by the stated skill purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The docstrings, tool description, and agent instruction are written only in Chinese, and the runtime test prompt also assumes Chinese input/output. This imposes a specific language/locale on use of the skill without any opt-in or documented justification, which matches the language-policy violation criteria.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes a skill for finance web searches and local document retrieval, but this module provisions connections to several third-party LLM providers by reading credentials from environment variables and configuring remote inference endpoints. Model-provider orchestration is not an obvious or declared capability required for search/RAG itself, especially across unrelated vendors like DeepSeek, DashScope, OpenRouter, ZAI, and UST.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/llm/factory.py (reported line 75)May include surrounding context.

python
return OpenAIChat(
            id=model_id,
            base_url="https://api.z.ai/api/paas/v4",
            api_key=api_key,
            timeout=60,
            role_map=role_map,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code sends user queries to external search providers and later sends result URLs to a content extraction service, but this file provides no explicit user warning, consent check, or data minimization control. For finance-related queries, this can leak sensitive research interests, company investigations, or proprietary local context to third parties.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The search tool goes beyond simple retrieval by fetching full page content and performing sentiment analysis on it, which is not disclosed by the stated skill purpose of web/local search and retrieval. This expands data processing scope, increases external data exposure, and can surprise users or operators who expect only search behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Importing and invoking a separate sentiment-analysis capability introduces an additional processing function unrelated to the declared search scope. In a finance context, this can materially influence outputs while silently processing retrieved content in ways users did not request or authorize.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Core descriptive text, prompts, and user-facing documentation in this file are written only in Chinese, with no indication that users can choose another language. This can violate language/locale policy when a skill imposes one language by default rather than offering choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a skill for finance web searches and local document-store retrieval, but this file implements sentiment analysis and also inspects UST_KEY_API from the environment to decide which remote model provider to use. Reading environment-based credentials/provider secrets is not an obvious requirement of search or local RAG retrieval and represents an extra capability outside the stated purpose.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code will automatically download a model from the network when it is not present locally. In a skill advertised only for finance search/RAG, undeclared outbound network access and execution of externally sourced model artifacts expand supply-chain and data-governance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The LLM path sends analyzed text to an external model provider without any disclosure, consent flow, or warning in the method contract. If the input contains proprietary, personal, or regulated finance data, this can cause unintended data exfiltration to third-party services.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.