Back to skill

Security audit

GOG Weekly Sales Analytics

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it says, but it automatically publishes the whole working directory to ClawHub after handling API keys and generated reports, which creates a credible credential and data exposure risk.

Review this skill before installing. Use it only if you are comfortable sending GOG sales data to Gemini and Feishu, and do not run the current main.py with real credentials unless publication is removed or changed to a staged allowlist that excludes .env, generated data, logs, caches, and other private files. Publishing should be a separate explicit release action, not part of the recurring report workflow.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
main.py:47
Finding

Automatic publication of the project root may expose credentials and generated data

Content
View full analysis

Vulnerability Details

File Location: main.py:47-53
Related Location: SKILL.md:24-26
Vulnerability Type: Sensitive data exposure through overbroad publication scope
Risk Level: High

Vulnerable Code

main.py:47-53:

python
def publish_skill_to_clawhub():
    print("Publishing workflow as ClawHub skill...")
    result = subprocess.run([
        "clawhub", "publish", ".",
        "--slug", "gog-sales-analytics",
        "--name", "GOG Weekly Sales Analytics",
        "--version", "1.2.3",

The documented setup in SKILL.md:24-26 places credentials in the project root:

bash
cp .env.example .env
# Fill in API keys in .env
pip install -r requirements.txt

Technical Analysis

The workflow instructs users to create a credential-bearing .env file in the project root and later invokes clawhub publish ., passing that entire root directory as the publication source. No publication allowlist or ignore file exists in the audited project.

The same directory also receives runtime-generated files under data/, including scraped JSON and Gemini-generated Markdown reports. Consequently, publication safety depends on undocumented behavior of the external clawhub executable rather than an explicit security boundary enforced by this project.

Although publishing the reusable Skill is declared functionality, publishing the complete mutable working directory exceeds the minimum scope needed. Only static Skill source and configuration files need to be distributed.

Attack Path

  1. A user follows the documentation and creates .env in the project root.
  2. The user stores GEMINI_API_KEY, FEISHU_APP_ID, FEISHU_APP_SECRET, FEISHU_DRIVE_FOLDER_ID, and potentially CLAWHUB_API_TOKEN in that file.
  3. main.py generates scraped data and an analysis report beneath the same project directory.
  4. The workflow invokes clawhub publish ..
  5. If the publisher does n ...[truncated 1065 chars]
Remediation
View remediation

Remediation Suggestions

  1. Create a clean staging directory containing only explicitly approved static files, then publish that directory instead of ".".
  2. Add .env, .env.*, data/, generated reports, caches, logs, virtual environments, and credential files to the publisher's verified exclusion mechanism.
  3. Generate and inspect a publication manifest before upload. Abort if it contains secret-bearing or generated files.
  4. Add automated secret scanning before publication, checking both filenames and content patterns.
  5. Store credentials outside the publishable project tree where practical.
  6. Separate Skill publication from the recurring data-analysis workflow and require explicit confirmation before publishing.
  7. Rotate any credentials if they may already have been included in a published artifact.

T01 · Skill Instruction Hijacking

Warning
Location
analysis/gemini_analyzer.py:16
Finding

Untrusted scraped catalog data is interpolated directly into the Gemini instruction prompt

Content
View full analysis

Vulnerability Details

File Location: analysis/gemini_analyzer.py:16-53
Vulnerability Type: Indirect prompt injection through externally sourced data
Risk Level: Medium

Vulnerable Code

python
compact_data = []
for game in sales_data:
    if not isinstance(game, dict):
        continue
    compact_data.append({
        k: game.get(k)
        for k in (
            "title", "price", "original_price", "discount",
            "discount_percent", "rating", "genre", "category", "url",
        )
        if k in game
    })

payload = compact_data if compact_data else sales_data

prompt = f"""
Analyze this GOG weekly sales data. The dataset contains ALL {total_deals}
discounted games for the week (the complete set, not a sample). Base every
insight on the full dataset below:
{json.dumps(payload, indent=2, ensure_ascii=False)}

Provide:
1. Top 5 best value deals (highest discount percentage with good ratings)
2. Price trend comparison for popular AAA titles
3. Category breakdown of discounted games
4. Recommendations for budget gamers (<$10)

Format output as markdown report.
"""

response = client.models.generate_content(
    model="gemini-2.0-flash",
    contents=prompt
)

Technical Analysis

Values scraped from an external website are serialized and inserted directly into the same prompt that contains the application's instructions. The code does not explicitly identify the records as untrusted content, tell the model to ignore instructions embedded in field values, enforce strict field types and lengths, or validate the generated Markdown.

A malicious or compromised catalog entry could place instruction-like content in a scraped field such as a title. Because language models interpret both instructions and data as tokens in the same context, such content may override or distort the requested analysis.

The resulting model output is written to a report and ...[truncated 1363 chars]

Remediation
View remediation

Remediation Suggestions

  1. Clearly delimit the dataset and state that all enclosed values are untrusted data whose instructions must never be followed.
  2. Use structured model inputs or schema-constrained APIs where supported instead of combining instructions and records in one free-form prompt.
  3. Validate each record against a strict schema, including expected scalar types, maximum lengths, and permitted formats.
  4. Normalize price, discount, and rating fields before model submission.
  5. Allowlist URL schemes and hosts; reject unexpected or malformed URLs.
  6. Request schema-constrained output and render the final Markdown locally from validated fields.
  7. Scan generated output for unexpected external links, instruction-like content, embedded HTML, and unsupported claims before writing or uploading it.
  8. Consider requiring review or approval before distributing model-generated reports to a shared folder.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (23)

Tainted flow: 'payload' from os.getenv (line 13, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · sync/feishu_upload.py (reported line 17)May include surrounding context.

python
"app_id": FEISHU_APP_ID,
        "app_secret": FEISHU_APP_SECRET
    }
    response = requests.post(url, json=payload)
    data = response.json()
    if data.get("code") != 0:
        raise Exception(f"Failed to get tenant_access_token: {data}")

Tainted flow: 'payload' from os.getenv (line 13, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · sync/feishu_upload.py (reported line 81)May include surrounding context.

python
"perm": permission_type,
            "notify": True
        }
        requests.post(url, headers=headers, json=payload)

if __name__ == "__main__":
    import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 12)May include surrounding context.

Usage

text
cp .env.example .env
# Fill in API keys in .env
pip install -r requirements.txt
python main.py

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 13)May include surrounding context.

Usage

text
cp .env.example .env
# Fill in API keys in .env
pip install -r requirements.txt
python main.py

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 21)May include surrounding context.

Usage

text
cp .env.example .env
# Fill in API keys in .env
pip install -r requirements.txt
python main.py

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 22)May include surrounding context.

Usage

text
cp .env.example .env
# Fill in API keys in .env
pip install -r requirements.txt
python main.py

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill description emphasizes scraping, analysis, report generation, and Feishu sync, but the file also states that it publishes the workflow as a reusable skill on ClawHub. Undeclared distribution/deployment behavior is dangerous because users may provide secrets or run the skill expecting reporting only, while it also performs a broader external action that could expose content, metadata, or internal workflow details.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill's documented purpose is scraping, analysis, report generation, and syncing to Feishu, but the code also publishes itself to ClawHub. That hidden or unjustified capability is dangerous because it introduces a deployment/supply-chain action users would not reasonably expect from a sales-reporting workflow, enabling unauthorized propagation of the skill or its future modifications.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Embedding self-publication/deployment logic in a sales analytics workflow is unjustified and expands the trust boundary beyond data processing and file sync. In the context of agent skills, this is especially dangerous because it can turn routine execution into an unauthorized release action, increasing supply-chain and platform abuse risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README states that generated reports are synced to a shared Feishu Drive folder, but it does not warn users that workflow outputs will be transmitted to an external collaboration platform. Even if the intended data is only public GOG deal information, reports may include prompts, metadata, account identifiers, or later-added internal commentary, so silent external sharing increases the risk of unintended data disclosure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill declares broad operational behavior involving environment variables, filesystem access, network access, and shell execution, but does not define any explicit tool scope or permissions boundary. This creates unnecessary ambiguity for reviewers and operators and increases the risk of over-privileged execution, especially because the workflow also performs external syncing and publishing actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The function sends the full discounted-game dataset, including titles, URLs, pricing, ratings, and other fields, to an external Gemini API with no disclosure, consent flow, or configuration gate. Even if this dataset is not obviously sensitive, transmitting scraped or aggregated business data to a third-party model service can violate data-handling expectations, create privacy/compliance issues, and expose proprietary reporting inputs outside the local workflow.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · main.py (reported line 11)May include surrounding context.

python
output_file = f"./data/weekly_sales_{date_str}.json"
    os.makedirs("./data", exist_ok=True)
    
    result = subprocess.run([
        "web-scraper", "run", 
        "./scraper/gog_sales_scraper.yaml",
        "--output", output_file

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · main.py (reported line 24)May include surrounding context.

python
def run_analysis(sales_file):
    print("Running Gemini sales analysis...")
    result = subprocess.run([
        "python", "./analysis/gemini_analyzer.py", sales_file
    ], capture_output=True, text=True)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · main.py (reported line 35)May include surrounding context.

python
def sync_to_feishu(report_file):
    print("Syncing report to Feishu Drive...")
    result = subprocess.run([
        "python", "./sync/feishu_upload.py", report_file
    ], capture_output=True, text=True)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
95% confidence
Finding

This subprocess does not just run a helper; it publishes the current project as a ClawHub skill, creating an outbound deployment side effect unrelated to the stated reporting workflow. In an agent-skill context, self-publication materially increases supply-chain risk because executing the skill can propagate or redistribute code without a clearly justified business need or explicit operator approval.

Content

Scanner excerpt · main.py (reported line 46)May include surrounding context.

python
def publish_skill_to_clawhub():
    print("Publishing workflow as ClawHub skill...")
    result = subprocess.run([
        "clawhub", "publish", ".",
        "--slug", "gog-sales-analytics",
        "--name", "GOG Weekly Sales Analytics",

Known Vulnerable Dependency: requests==2.31.0 — 6 advisory(ies): CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi); CVE-2026-25645 (Requests has Insecure Temp File Reuse in its extract_zipped_paths() utility func) +3 more

Medium
Category
Supply Chain
Confidence
95% confidence
Finding

The skill pins requests==2.31.0, which is flagged by multiple published advisories, including issues involving credential leakage via .netrc handling and request/session security flaws. In this skill's context—scraping external GOG store data and likely making outbound HTTP requests—using a known-vulnerable HTTP client increases the risk of credential exposure or unsafe network behavior if a malicious URL, redirect, or crafted environment is encountered.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: python-dotenv==1.0.1 — 2 advisory(ies): CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)

Medium
Category
Supply Chain
Confidence
82% confidence
Finding

The skill pins python-dotenv==1.0.1, which is reported as affected by advisories involving unsafe file handling such as symlink following during .env modification. While dotenv is often used benignly for configuration loading, this workflow also syncs reports to external services and may run in automation contexts where filesystem manipulation or untrusted working directories could make such issues exploitable.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · sync/feishu_upload.py (reported line 17)May include surrounding context.

python
"app_id": FEISHU_APP_ID,
        "app_secret": FEISHU_APP_SECRET
    }
    response = requests.post(url, json=payload)
    data = response.json()
    if data.get("code") != 0:
        raise Exception(f"Failed to get tenant_access_token: {data}")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The function uploads any local file path it is given to Feishu without any in-code confirmation, path restrictions, or user-visible warning. In a recurring automation workflow, that increases the risk of accidental disclosure if the wrong file path is supplied or if upstream components can influence the path.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill includes the ability to grant file permissions to arbitrary Feishu users, which expands capability beyond simple report upload/sync. In an automation context, this can unintentionally expose uploaded reports to unintended recipients or be abused by upstream inputs to widen access without explicit operator awareness.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · sync/feishu_upload.py (reported line 81)May include surrounding context.

python
"perm": permission_type,
            "notify": True
        }
        requests.post(url, headers=headers, json=payload)

if __name__ == "__main__":
    import sys

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The script reads GEMINI_API_KEY from the environment to authenticate with an external service, but provides no docstring, comment, or user-facing explanation about its use. Under this rule, access to sensitive environment variables should have some form of disclosure unless it is already clearly documented elsewhere.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
sync/feishu_upload.py:29