Back to skill

Security audit

Alibabacloud Data Agent Skill

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to be a real Alibaba Cloud data-analysis integration, but it needs Review because it handles cloud credentials, local files, reports, background monitoring, and custom network destinations without enough containment.

Review before installing. Use a tightly scoped Alibaba RAM policy, avoid broad DMS full-access credentials, do not set DATA_AGENT_ENDPOINT or ASYNC_TASK_PUSH_URL unless the destination is trusted, and avoid analyzing sensitive files or production databases until endpoint validation and path-safety issues are fixed. Disable or ignore HEARTBEAT-style autonomous notifications unless you explicitly want session progress and report content pushed to chat channels.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/data_agent/sse_client.py:303
Finding

Unrestricted Custom Endpoint Can Receive Signed Authentication Material and Sensitive Analysis Data

Content
View full analysis

Vulnerability Details

File Location: scripts/data_agent/sse_client.py:303-323
Related Configuration Source: scripts/data_agent/config.py:47-58, 94-103
Vulnerability Type: Unvalidated credential-bearing network endpoint
Risk Level: High

Vulnerable Code

python
# scripts/data_agent/config.py:94-103
return cls(
    api_key=os.environ.get("DATA_AGENT_API_KEY") or None,
    region=os.environ.get("DATA_AGENT_REGION", "cn-hangzhou"),
    endpoint=os.environ.get("DATA_AGENT_ENDPOINT"),
    timeout=int(os.environ.get("DATA_AGENT_TIMEOUT", "300")),
    max_retry=int(os.environ.get("DATA_AGENT_MAX_RETRY", "3")),
    poll_interval=int(os.environ.get("DATA_AGENT_POLL_INTERVAL", "2")),
    max_poll_count=int(os.environ.get("DATA_AGENT_MAX_POLL_COUNT", "60")),
    workspace_id=os.environ.get("DATA_AGENT_WORKSPACE_ID"),
)
python
# scripts/data_agent/sse_client.py:303-323
host = self._config.endpoint
headers = AliyunSignerV3.sign(
    access_key_id,
    access_key_secret,
    "POST",
    host,
    "GetChatContent",
    params,
    security_token=security_token if security_token else None,
)
headers["User-Agent"] = "AlibabaCloud-Agent-Skills/alibabacloud-data-agent-skill"

query_params = {"Action": "GetChatContent", "Version": "2025-04-14", **params}
query_string = "&".join(
    f"{quote(k, safe='')}={quote(str(v), safe='')}"
    for k, v in sorted(query_params.items())
)
url = f"{self._base_url}/?{query_string}"

with requests.post(
    url,
    stream=True,
    timeout=timeout,
    headers=headers,
) as response:

Technical Analysis

DATA_AGENT_ENDPOINT is accepted without validating its hostname, domain suffix, path, or expected Alibaba Cloud service identity. In AK/SK mode, the endpoint is used as the signing host and request destination.

The outbound headers contain an ACS authorization signature, the access key identifier, and ...[truncated 1973 chars]

Remediation
View remediation

Remediation Suggestions

  1. Parse configured endpoints with urllib.parse.urlsplit rather than treating them as arbitrary strings.
  2. Permit only documented Alibaba Cloud domains, such as exact service-specific hosts under aliyuncs.com.
  3. Validate that the endpoint corresponds to the configured region and expected service.
  4. Reject embedded user information, query strings, fragments, unexpected paths, IP literals, and non-HTTPS schemes.
  5. Require a separate, explicit development-only option before allowing custom endpoints.
  6. Do not send cloud signatures or STS tokens to a custom endpoint unless it has been independently authenticated.
  7. Revalidate endpoint configuration immediately before constructing each credential-bearing request.
  8. Recommend narrowly scoped custom RAM policies instead of broad DMS full-access policies.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/data_agent/file_manager.py:102
Finding

Unvalidated Upload Destination Can Exfiltrate User-Selected Local Files

Content
View full analysis

Vulnerability Details

File Location: scripts/data_agent/file_manager.py:102-146
Asynchronous Equivalent: scripts/data_agent/file_manager.py:401-449
Vulnerability Type: Unvalidated server-controlled file-upload destination
Risk Level: High

Vulnerable Code

python
data = signature_response.get("data", signature_response.get("Data", {}))
upload_host = data.get("uploadHost", data.get("UploadHost"))
upload_dir = data.get("uploadDir", data.get("UploadDir"))
policy = data.get("policy", data.get("Policy"))
oss_signature = data.get("ossSignature", data.get("OssSignature"))
oss_date = data.get("ossDate", data.get("OssDate"))
oss_security_token = data.get("ossSecurityToken", data.get("OssSecurityToken"))
oss_credential = data.get("ossCredential", data.get("OssCredential"))

if not all([upload_host, upload_dir, policy, oss_signature]):
    raise FileUploadError(
        f"Invalid signature response: missing required OSS upload parameters",
        file_path=file_path,
    )

file_key = f"{upload_dir}/{filename}"
upload_url = upload_host
upload_location = file_key

content_type = SUPPORTED_FILE_TYPES[suffix]
try:
    with open(file_path, "rb") as f:
        files = {
            "key": (None, file_key),
            "policy": (None, policy),
            "x-oss-signature": (None, oss_signature),
            "x-oss-date": (None, oss_date),
            "x-oss-security-token": (None, oss_security_token or ""),
            "x-oss-credential": (None, oss_credential),
            "x-oss-signature-version": (None, "OSS4-HMAC-SHA256"),
            "success_action_status": (None, "200"),
            "file": (filename, f, content_type),
        }
        response = requests.post(
            upload_url,
            files=files,
            timeout=timeout,
        )
        response.raise_for_status()

Technical Analysis

Uploading a selected file is necessa ...[truncated 1698 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require an https upload URL.
  2. Allow only expected Alibaba Cloud OSS hostname suffixes and verify that the hostname matches the configured region.
  3. Resolve the hostname and reject loopback, private, link-local, multicast, and reserved addresses.
  4. Disable redirects for upload requests. If redirects are operationally required, validate every redirect target before following it.
  5. Reject URLs containing user information, fragments, or unexpected ports.
  6. Bind the upload destination to trusted fields returned by an authenticated Alibaba Cloud endpoint.
  7. Apply the same validation to both synchronous and asynchronous implementations.
  8. Present the validated destination domain to the user before uploading sensitive data when a nonstandard endpoint is configured.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/data_agent/file_manager.py:183
Finding

Unvalidated Report Download URLs Enable Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/data_agent/file_manager.py:183-244
Invocation Site: scripts/cli/cmd_reports.py:48-63
Vulnerability Type: Server-side request forgery through remote report URLs
Risk Level: Medium

Vulnerable Code

python
return [
    FileInfo(
        file_id=f.get("FileId", ""),
        filename=f.get("FileName", ""),
        file_type=f.get("FileType", ""),
        size=f.get("FileSize", 0),
        download_url=f.get("DownloadLink") or f.get("DownloadUrl"),
    )
    for f in files
]
python
def download_from_url(
    self,
    download_url: str,
    save_path: str,
    timeout: int = 120,
) -> str:
    save = Path(save_path)
    save.parent.mkdir(parents=True, exist_ok=True)
    try:
        response = requests.get(download_url, timeout=timeout, stream=True)
        response.raise_for_status()
        with save.open("wb") as f:
            for chunk in response.iter_content(chunk_size=8192):
                f.write(chunk)
        return str(save.resolve())
    except requests.RequestException as e:
        raise FileDownloadError(f"Failed to download file: {e}")

Technical Analysis

Report URLs returned by the remote API are fetched directly. There is no validation of the scheme, hostname, port, resolved address, or redirect chain.

As a result, a malicious or compromised report response can direct the client to services reachable only from the user's machine or cloud environment. requests follows redirects by default, so validating only the initial URL would also be insufficient.

The same missing trust-boundary checks exist in the asynchronous downloader at scripts/data_agent/file_manager.py:513-533.

Attack Path

  1. A malicious or compromised API response returns a report whose DownloadLink targets localhost, a private-network service, or a cloud metadata address.
  2. The user runs the `r ...[truncated 943 chars]
Remediation
View remediation

Remediation Suggestions

  1. Permit only HTTPS report URLs on explicitly approved Alibaba Cloud OSS or report-hosting domains.
  2. Resolve destination addresses and reject loopback, private, link-local, multicast, unspecified, and reserved ranges for both IPv4 and IPv6.
  3. Set allow_redirects=False, or manually follow redirects only after validating each destination.
  4. Reject URLs with embedded credentials or unexpected ports.
  5. Enforce maximum response sizes and terminate oversized downloads.
  6. Set separate connection and read timeouts.
  7. Validate the asynchronous download implementation using the same centralized URL policy.
  8. Prefer opaque file identifiers and an authenticated download API over arbitrary URLs where supported.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/cli/cmd_reports.py:48
Finding

Remote Report Filenames Permit Path Traversal and Arbitrary File Overwrite

Content
View full analysis

Vulnerability Details

File Location: scripts/cli/cmd_reports.py:48-61
File-Write Sink: scripts/data_agent/file_manager.py:234-244
Vulnerability Type: Path traversal through an untrusted filename
Risk Level: High

Vulnerable Code

python
report_dir.mkdir(parents=True, exist_ok=True)
print(f"Downloading files to {report_dir.resolve()}...\n")

for category, rf in found_files:
    if not rf.download_url:
        print(f"  [{category}] {rf.filename or rf.file_id}: No download URL available")
        continue

    save_path = report_dir / (rf.filename or f"{rf.file_id}.bin")
    try:
        print(f"  Downloading [{category}] {rf.filename or rf.file_id}...")
        file_manager.download_from_url(rf.download_url, str(save_path))
        print(f"  ✅ Saved to {save_path.resolve()}")
        total_reports += 1
    except Exception as e:
        print(f"  ❌ Failed to download {rf.filename or rf.file_id} ({category}): {e}", file=sys.stderr)

The receiving method creates parent directories and writes the file:

python
save = Path(save_path)
save.parent.mkdir(parents=True, exist_ok=True)
response = requests.get(download_url, timeout=timeout, stream=True)
response.raise_for_status()
with save.open("wb") as f:
    for chunk in response.iter_content(chunk_size=8192):
        f.write(chunk)

Technical Analysis

rf.filename originates in the remote file-list response and is joined directly to the local report directory. The code does not reject absolute paths, path separators, or .. traversal components. It also does not verify that the resolved destination remains under the intended report directory.

Path joining does not provide containment protection. Traversal components can escape the report directory, and an absolute component can replace the preceding path on supported platforms. The downloader then creates parent directories and opens the target in ove ...[truncated 1085 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat remote filenames as display metadata, not trusted local paths.
  2. Reduce filenames to a safe basename and reject any input for which Path(filename).name != filename.
  3. Reject /, backslashes, drive prefixes, null characters, . components, and .. components.
  4. Generate local filenames from trusted file IDs where possible.
  5. Resolve both the report root and destination, then verify the destination is a strict descendant of the report root before creating directories or opening the file.
  6. Use exclusive file creation when overwriting is unnecessary.
  7. Apply an allowlist of safe filename characters and a conservative maximum length.
  8. Add regression tests for relative traversal, absolute paths, Windows drive paths, UNC paths, encoded separators, and symlink-based escapes.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/cli/worker_utils.py:49
Finding

Unvalidated Session Identifiers Are Used in Writable Filesystem Paths

Content
View full analysis

Vulnerability Details

File Location: scripts/cli/worker_utils.py:49-72
Additional Location: scripts/cli/cmd_reports.py:19-21
Vulnerability Type: Path traversal through a session identifier
Risk Level: Medium

Vulnerable Code

python
session_id = session.session_id
session_dir = Path(f"sessions/{session_id}")
session_dir.mkdir(parents=True, exist_ok=True)

existing_pid = check_worker_lock(session_dir)
if existing_pid:
    print(f"⚠️  A worker process (PID {existing_pid}) is already running for session {session_id}.", file=sys.stderr)
    print(f"   Check progress: cat sessions/{session_id}/progress.log", file=sys.stderr)
    print(f"   Current status: {(session_dir / 'status.txt').read_text().strip() if (session_dir / 'status.txt').exists() else 'unknown'}", file=sys.stderr)
    sys.exit(1)

input_data = vars(args).copy()
input_data.pop("func", None)
with open(session_dir / "input.json", "w", encoding="utf-8") as f:
    json.dump(input_data, f, ensure_ascii=False, indent=2)

with open(session_dir / "status.txt", "w", encoding="utf-8") as f:
    f.write("running")

Report handling similarly uses a CLI-provided identifier:

python
session_id = args.session_id
session_dir = Path(f"sessions/{session_id}")
report_dir = session_dir / "reports"

Technical Analysis

Session identifiers from API responses, environment variables, or command-line arguments are incorporated into paths without a format restriction or containment check.

A session identifier containing path separators or traversal components can cause directory creation and subsequent reads or writes outside the intended sessions root. Worker processing writes input.json, status.txt, lock files, and completion files beneath the constructed directory. Report handling also creates a reports subdirectory and writes downloaded content there.

Attack Path

  1. An attacker supplies a crafted se ...[truncated 1078 chars]
Remediation
View remediation

Remediation Suggestions

  1. Define and enforce the expected session-ID grammar, for example ^[A-Za-z0-9_-]{1,128}$.
  2. Reject empty identifiers, path separators, drive prefixes, . components, and .. components.
  3. Construct paths from a canonical session root and verify that the resolved path remains below that root.
  4. Centralize session-path creation in one validated helper and use it in worker, attach, report, streaming, and logging code.
  5. Do not trust session identifiers merely because they came from an authenticated API response.
  6. Refuse to operate through symbolic links inside the session hierarchy where practical.
  7. Add tests covering Unix traversal, Windows traversal, absolute paths, excessive lengths, Unicode separator lookalikes, and symlink escapes.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (103)

Tainted flow: 'req' from os.environ.get (line 96, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
97% confidence
Finding

The destination URL for notification delivery is taken directly from the ASYNC_TASK_PUSH_URL environment variable and used for an outbound HTTP request without validation or allowlisting. This enables server-side request forgery and unintended data exfiltration of session identifiers, message contents, and optional bearer tokens to attacker-controlled endpoints if the environment is influenced.

Content

Scanner excerpt · scripts/cli/notify.py (reported line 98)May include surrounding context.

python
req = Request(custom_url, data=data, headers=headers, method="POST")
        try:
            with urlopen(req, timeout=10) as response:
                success = response.status >= 200 and response.status < 300
                print(f"[Notify] HTTP push result: status={response.status}, success={success}", flush=True)
                return success

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description focuses on using the Alibaba Cloud Data Agent for Analytics to perform natural-language data analysis against enterprise databases, including resource discovery, starting analysis sessions, tracking progress, and fetching conclusions/reports. The supplied code chunk instead handles management/inspection of custom agents: it calls list_custom_agents and describe_custom_agent, prints agent metadata, and provides tips on using an agent elsewhere. There is no code here for database resource discovery, running analysis, querying data, monitoring jobs, or retrieving analysis outputs. This is a materially different primary purpose from the declared skill behavior, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description centers on an analytics agent capable of end-to-end natural language data analysis and reporting. The supplied code chunk is much narrower: it only exposes three discovery operations against DMS metadata (instances, databases, tables). While resource discovery is one declared feature, the primary advertised capabilities around analytics execution and report retrieval are absent from this code. Therefore the description materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description centers on enterprise database analytics via Alibaba Cloud Data Agent, especially discovery and querying of DMS-managed database resources. The supplied code chunk instead implements a dedicated file subcommand focused on analyzing uploaded files or existing file IDs. Its core behavior is validating local file paths, restricting formats to CSV/XLS/XLSX, uploading files, creating FILE data sources, asking preset file-analysis questions, streaming results, and optionally listing generated files. While it still uses the same Data Agent service and natural-language analysis sessions, the primary purpose and accessed resource type are materially different from the declared database-oriented description. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description centers on using Data Agent for natural language-driven analytics: discovering resources, launching analysis/query sessions, monitoring progress, and fetching insights/reports. The supplied code chunk does none of those. Instead, it performs a high-risk administrative import operation that adds DMS tables to the Data Center after confirmation. This is a materially different primary purpose and an undeclared write capability against managed data resources. While importing tables could be related to preparing data for later analysis, this specific command is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a full-featured Alibaba Cloud Data Agent CLI integration for enterprise database analytics. The supplied code chunk only implements local subprocess execution and post-hoc log file creation. It specifically writes progress.log and skips progress.jsonl creation despite the module docstring claiming both. There is no evidence of any Alibaba Cloud API/CLI invocation, database access, natural language processing, session management for analytics, resource discovery, report retrieval, or real-time tracking. This is a material description-behavior mismatch, not merely a supporting helper detail, because the code's actual primary purpose is generic logging around subprocess execution rather than data analysis functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about Alibaba Cloud's Data Agent for Analytics and enterprise database analysis workflows. The supplied code does not interact with Alibaba Cloud, DMS, databases, SQL, analytics jobs, or reports. Instead, it is solely a notification helper for OpenClaw/Clawdbot sessions, identifying session IDs and sending assistant messages through either a configured HTTP endpoint or a CLI. This is a materially different primary purpose and introduces undeclared messaging capabilities unrelated to the stated analytics functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description describes a database analytics skill with Alibaba Cloud CLI operations, enterprise data resource discovery, and analysis/report retrieval. The supplied code does none of that. It is a small utility module for preventing concurrent worker processes by storing and checking PIDs in a local lock file. While such locking could be a supporting implementation detail within a larger analytics tool, this code chunk itself does not implement the declared functionality and instead has a materially different immediate purpose. Therefore this chunk does not accurately represent the declared skill behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The description presents a concrete operational skill: invoke the Alibaba Cloud Data Agent via CLI for natural-language database analysis, including resource discovery, session initiation, progress tracking, and result retrieval. The supplied code chunk, however, is just an init.py file that aggregates and exports symbols from a Python SDK. While the exported names suggest the broader package may relate to Data Agent functionality, this snippet does not show a CLI interface or actual execution of those behaviors. This is a material description-versus-behavior mismatch for the provided code chunk, because the observable behavior here is package initialization/export, not the declared end-user analysis workflow.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a full-featured analytics skill that talks to Alibaba Cloud Data Agent for Analytics and supports resource discovery, analysis workflows, tracking, and report retrieval. The supplied code does not implement any of those functional behaviors. It only transforms object key casing for requests and responses and includes special-case casing preservation for file upload-related actions. While such an adapter could be a supporting utility inside a larger analytics integration, this code chunk by itself materially differs from the declared primary purpose and exposes no visible analytics or CLI behavior. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description promises a broader Data Agent for Analytics tool that can perform natural-language-driven analysis workflows end to end: discover resources, initiate query or deep analysis sessions, track progress, and fetch conclusions/reports. This code chunk only covers part of that scope: DMS metadata discovery via list_instances, search_database, and list_tables. Although the module docstring mentions askDatabase, the ask_database method is explicitly a placeholder and raises NotImplementedError, so the core NL2SQL/query execution capability is not actually present in this chunk. There is also no implementation for progress tracking, report retrieval, or analysis conclusion retrieval. No clearly undeclared malicious or unrelated behavior is present; the issue is that the declared description substantially overstates what this code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a user-facing analytics skill centered on Data Agent for Analytics: querying enterprise databases in natural language, running analyses, tracking sessions, and fetching reports. The supplied code chunk does none of that. It is a diagnostic/verification utility specifically for the dms-enterprise GetActiveRouteUnit API. Its only substantive behavior is building an OpenAPI client, calling GetActiveRouteUnit, parsing Route/data.RegionId, and reporting whether the tenant is on a default or non-default DMS unit. This is a materially different primary purpose and accesses different resources/semantics (routing metadata rather than analytics/query/reporting workflows). Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation does not clearly warn that user prompts, database contents, and uploaded file contents may be transmitted to Alibaba Cloud services for processing. In a data-analysis skill handling enterprise databases and local files, this omission is dangerous because users may unknowingly expose sensitive, regulated, or proprietary data to a third-party cloud service.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 75)May include surrounding context.

md
pip install -r scripts/requirements.txt

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 113)May include surrounding context.

md
- Or use local file `assets/example_game_data.csv` for file analysis experience

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill description presents a database analysis CLI workflow but does not disclose that the skill may proactively send session progress and report contents over external channels. This lack of transparency prevents informed consent and increases the chance that sensitive data is shared in ways the user did not expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This file implements OpenClaw/Clawdbot session discovery and message pushing, which is unrelated to the stated Alibaba Cloud data analytics purpose of the skill. The added capability creates an out-of-scope covert communication channel that can transmit data or operational status to external sessions, making the mismatch especially suspicious in this skill context.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/cli/worker_utils.py (reported line 76)May include surrounding context.

python
# Construct worker command
    cmd = [sys.executable] + [arg for arg in sys.argv if arg != "--async-run"]
    env = os.environ.copy()
    env["DATA_AGENT_ASYNC_WORKER"] = "1"
    env["DATA_AGENT_SESSION_ID"] = session_id
    env["DATA_AGENT_AGENT_ID"] = session.agent_id

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/cli/cmd_attach.py (reported line 153)May include surrounding context.

python
"""Create configuration from environment variables.

        Args:
            dotenv_path: Optional path to .env file to load.

        Returns:
            DataAgentConfig instance.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/cli/cmd_attach.py (reported line 194)May include surrounding context.

python
"""Create configuration from environment variables.

        Args:
            dotenv_path: Optional path to .env file to load.

        Returns:
            DataAgentConfig instance.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/data_agent/config.py (reported line 80)May include surrounding context.

python
"""Create configuration from environment variables.

        Args:
            dotenv_path: Optional path to .env file to load.

        Returns:
            DataAgentConfig instance.

Known Vulnerable Dependency: aiohttp==3.13.5 — 16 advisory(ies): CVE-2026-54279 (aiohttp: Host-Only Cookies Become Domain Cookies After CookieJar Persistence); CVE-2026-54273 (aiohttp: HTTP/1 Pipelined Requests Queue Without Limit); CVE-2026-54275 (aiohttp: TLS Server Hostname Override Is Ignored When Reusing HTTPS Connections) +13 more

High
Category
Supply Chain
Confidence
95% confidence
Finding

Pinning aiohttp to a version with multiple known advisories exposes any code using this dependency to documented security flaws, potentially including request handling, cookie isolation, connection reuse, or denial-of-service issues depending on how the library is used. In this skill's context, which interfaces with cloud data-analysis services and may handle authenticated API traffic, a vulnerable HTTP client is more concerning because it could affect confidentiality of session data, integrity of requests, or service availability.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares substantial capabilities including shell, network, environment access, and file read/write, but does not constrain them with an explicit permission or allowed-tools scope. In a skill that can access credentials and interact with external cloud services, missing scope boundaries increases the chance of overbroad execution, unintended side effects, and abuse if the skill is invoked in an unexpected context.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The activation guidance includes broad, common phrases like 'data analysis', 'database query', and 'data insights', which can cause the skill to trigger in many ordinary conversations. Because this skill can access local files, credentials, and remote cloud services, unintended invocation can lead to accidental data exposure or execution of higher-risk workflows than the user expected.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The HEARTBEAT trigger is underspecified, so the agent may run monitoring logic at unclear times or under broader conditions than intended. Ambiguous autonomous triggers are dangerous because they can cause repeated background access to session files and unintended disclosure or action without an active user request.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.