Back to skill

Security audit

Smart Search

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent web-search skill, but it needs review because it can fetch arbitrary URLs and send broad user queries to external services without strong scoping or privacy safeguards.

Install only if you are comfortable with a skill that performs external web searches and page extraction. Avoid using it with confidential queries, private URLs, internal hostnames, customer data, or large batch files unless the runtime has network egress controls and you explicitly set reasonable timeouts.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ddgs_search.py:430
Finding

Arbitrary URL Extraction May Permit Server-Side Request Forgery

Content
View full analysis
0: if use_signal and threading.current_thread() is threading.main_thread(): signal.signal(signal.SIGALRM, _timeout_handler) signal.alarm(timeout) use_sigalrm = True else: timer = _DeadlineTimer(timeout) timer.start() start = time.time() last_error = None try: for attempt in range(max_retries): try: if timer: timer.check() with DDGS() as d: result = d.extract(url, fmt=fmt) elapsed = time.time() - start return result, elapsed except TimeoutError: raise TimeoutError(f"Extract timed out after {timeout}s") except Exception as e: last_error = e if attempt < max_retries - 1: wait = (2 ** attempt) + random.uniform(0, 1) print(f" Extract retry {attempt+1}/{max_retries}: {e}. " f"Waiting {wait:.1f}s...", file=sys.stderr) time.sleep(wait) raise last_error finally: if use_sigalrm: signal.alarm(0) if timer: timer.cancel() ``` The command-line argument and invocat ...[truncated 3399 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared purpose frames this as a general search skill, but the body also supports direct URL extraction, full-text retrieval from arbitrary remote pages, batch input from local files, and specialized media/book search. That mismatch can bypass user expectations and policy routing, causing the skill to handle more sensitive data flows and content classes than reviewers or users realize.

Content

No source excerpt is available for this finding.

Chaining Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 52)May include surrounding context.

python
# URLs only (one URL per line, for piping)
  python3.12 ddgs_search.py -q "query" --urls-only --fallback
  python3.12 ddgs_search.py -q "query" --urls-only --fallback | xargs -I{} curl {}

Search types: text, news, images, videos, books, all, all+ (default: text)
Backends: auto, all, bing, brave, duckduckgo, google, yandex, yahoo, wikipedia

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documents shell-based execution (python3.12 ...), file-based batch input (--batch-file queries.txt), and references content extraction flows, but it declares no explicit tool scope or permissions. This creates an authorization and review gap: an orchestrator may invoke broader capabilities than a user or policy expects, increasing the risk of unintended file or shell access.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger phrases are very broad (search, find info, look up, 搜索, 调研, 查一下) and likely to match many ordinary requests. Over-broad activation increases the chance the skill is selected in contexts where external queries, remote fetching, or extraction are not appropriate, potentially sending sensitive user content to third parties without clear intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill omits a clear disclosure that queries are sent to external search providers and that extract()/URL fetches contact remote sites directly. This is a real privacy and data-governance issue because users may paste sensitive names, internal project terms, or URLs, not realizing they will be transmitted off-platform.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The entire skill is written as Chinese-only guidance, with no indication that the user can choose another language or that the skill is intentionally limited to a Chinese-speaking context. Under the policy rule for language or locale constraints, this is a natural-language policy violation because it imposes a specific language without opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The entire skill document is written in Chinese and does not provide any language-selection option, bilingual fallback, or stated regional/compliance justification. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/search-apis.md (reported line 30)May include surrounding context.

用法: 在任何 URL 前加 https://r.jina.ai/

bash
curl -s "https://r.jina.ai/https://example.com/article"

核心特性:

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
83% confidence
Finding

The documented and implemented default behavior allows timeout=0, which disables the global timeout and can let network operations block indefinitely. In an automated agent context, this can tie up worker capacity, stall workflows, or be abused by slow/unresponsive endpoints to cause denial-of-service conditions.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 172)May include surrounding context.

python
def cancel(self):
        self._event.set()
        if self._thread:
            self._thread.join(timeout=0.1)


def search_with_retry(query, backend="auto", max_results=5, timelimit=None,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Search queries are sent to third-party backends, but the tool provides no explicit privacy warning or consent checkpoint before transmitting potentially sensitive user input. In an agent environment, users may assume local processing, so this can lead to inadvertent disclosure of confidential prompts, terms, or URLs to external services.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
81% confidence
Finding

search_with_retry() accepts timeout=None, enabling callers to perform unbounded network searches when no higher-level timeout is supplied. Repeated retries and backend fallback compound the resource-consumption risk in agent execution environments.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 177)May include surrounding context.

python
def search_with_retry(query, backend="auto", max_results=5, timelimit=None,
                      region=None, max_retries=3, fallback_backends=None,
                      timeout=None, use_signal=True):
    """Execute search with exponential backoff and optional fallback chain.

    Returns (results, backend_used, elapsed_seconds) or raises on total failure.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
81% confidence
Finding

search_by_type() also permits timeout=None, so type-specific searches can run without a hard deadline. Since this function may be invoked repeatedly across multiple backends, unbounded waits can accumulate into agent-level denial of service.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 281)May include surrounding context.

python
def search_by_type(query, search_type="text", backend="auto", max_results=5,
                   timelimit=None, region=None, max_retries=3,
                   fallback=None, timeout=None, use_signal=True,
                   force_lang=None):
    """Execute a type-specific search (text/news/images/videos/books) with optional fallback chain.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
82% confidence
Finding

search_all_types() fans out concurrent searches across several content types, but still allows timeout=None. Parallel unbounded operations increase the chance of thread exhaustion, long stalls, and amplified resource consumption if remote services hang.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 388)May include surrounding context.

python
def search_all_types(query, backend="auto", max_results=3, timelimit=None,
                     region=None, max_retries=2, fallback=None,
                     include_images=False, timeout=None, force_lang=None):
    """Search across text, news, and videos types at once. Returns dict keyed by type."""
    from concurrent.futures import ThreadPoolExecutor, as_completed

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
84% confidence
Finding

extract_from_url() permits timeout=None while fetching and parsing arbitrary remote content, which is riskier than simple search metadata retrieval. A malicious or slow target can keep extraction running indefinitely, consuming network and processing resources.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 425)May include surrounding context.

python
return results_by_type, total_elapsed, errors


def extract_from_url(url, fmt="text_markdown", max_retries=2, timeout=None, use_signal=True):
    """Extract full text from a URL using DDGS extract()."""
    try:
        from ddgs import DDGS

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill supports arbitrary URL extraction via --extract-url and also auto-extracts the first search result, causing the agent to fetch and process attacker-chosen external content beyond simple search metadata retrieval. In an agent setting, this expands exposure to untrusted remote data, can bypass intended scope limits for a 'search' skill, and may be abused for SSRF-like access depending on how the underlying DDGS library resolves URLs and what network the agent can reach.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

search_single() accepts and forwards timeout=None, so the main single-query path can execute without a deadline. Because this is the common entry point, the absence of a default bound materially increases availability risk for the skill.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 481)May include surrounding context.

python
def search_single(query, backend="auto", max_results=5, timelimit=None,
                  region=None, max_retries=3, fallback=False, output=None,
                  search_type="text", extract=False, extract_length=2000,
                  timeout=None, force_lang=None):
    """Execute a single search query, with type and extract support."""
    start_time = time.time()

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

The multi-type wrapper _search_multi_type() allows timeout=None and orchestrates parallel remote calls, so a single invocation may consume multiple threads indefinitely. This raises the blast radius of hung network operations within agent infrastructure.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 641)May include surrounding context.

python
def _search_multi_type(query, backend, max_results, timelimit,
                       region, max_retries, fallback, output,
                       include_images=False, timeout=None, force_lang=None):
    """Handle --type all / all+ search: text + news + videos (+ optional images) in parallel."""
    results_by_type, total_elapsed, errors = search_all_types(
        query, backend=backend, max_results=max_results,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Batch mode reads multiple queries from a file and sends them to external services without any added privacy notice, magnifying the volume of potential data disclosure. Because batch files may contain research notes, customer data, or internal terms, the absence of a warning or confirmation step increases the risk of accidental exfiltration.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
82% confidence
Finding

Batch search supports timeout=None while iterating over multiple external queries, making it possible for a single job to run for an unbounded period. In shared agent systems, this can monopolize execution slots and create cascading delays.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 710)May include surrounding context.

python
def search_batch(batch_file, backend="auto", max_results=5, timelimit=None,
                 region=None, max_retries=3, fallback=False, output=None,
                 timeout=None, force_lang=None):
    """Execute batch search from a file (one query per line)."""
    try:
        with open(batch_file, "r", encoding="utf-8") as f:

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The documentation states that the skill performs automatic language detection and chooses Chinese queries via Bing-first and English queries via Brave-first by default. This imposes locale-dependent behavior automatically rather than offering user choice up front, which can conflict with policies requiring language or locale choice unless clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The natural-language behavior auto-detects CJK content and otherwise treats queries as English, while the documented manual override only allows 'en' or 'cn'. This imposes a narrow language/locale policy that may route users into a language-specific backend strategy without an explicit opt-in or broader locale choice.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/ddgs_search.py (reported line 342)May include surrounding context.

python
if region:
                            kwargs["region"] = region

                        method = getattr(d, method_name)
                        results = list(method(query, max_results=max_results, **kwargs))

                        if results:

Static analysis

No suspicious patterns detected.