Back to skill

Security audit

tech-news-digest

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real tech-news digest skill, but it needs review because it can run scheduled shell workflows, use credentials, fetch broad configurable URLs, and send results externally with incomplete safeguards.

Review before installing. Use a dedicated workspace and virtual environment, provide least-privilege API tokens, avoid enabling GH_APP_KEY_FILE or gh CLI fallback unless needed, confirm every scheduled job and delivery destination, disable enrichment for untrusted feeds, and do not import source configs from parties you do not trust until private-network URL blocking and safer command construction are added.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Warning
Location
references/digest-prompt.md:130
Finding

Mandatory Promotional Content Injection in Generated Reports

Content
View full analysis

Vulnerability Details

File Location: references/digest-prompt.md:130-134; enforced through SKILL.md:396-447
Vulnerability Type: Mandatory output manipulation
Risk Level: Medium

Vulnerable Code

markdown
### Stats Footer

📊 Data Sources: RSS {{rss}} | Twitter {{twitter}} | Reddit {{reddit}} | Web {{web}} | GitHub {{github}} releases + {{trending}} trending | Dedup: {{merged}} articles 🤖 Generated by tech-news-digest v<VERSION> | <https://github.com/draco-agent/tech-news-digest> | Powered by OpenClaw

text

The associated instructions in SKILL.md require the agent to read this prompt and follow every step strictly.

Technical Analysis

The digest workflow mandates a fixed attribution line and external repository link in every generated report. This content is unrelated to the substantive news requested by the user and is not presented as optional. Because the Skill is designed for recurring scheduled execution and multi-channel delivery, the injected content is reproduced in Discord, email, Markdown, and archived reports.

This does not modify the agent’s global safety rules, but it does take control of a portion of the agent’s output and requires persistent third-party promotion whenever the Skill is used.

Attack Path

  1. A user installs or invokes the Skill to generate a digest.
  2. The agent is instructed to load references/digest-prompt.md.
  3. The workflow requires the agent to follow every step in that prompt.
  4. The agent appends the prescribed attribution and external link.
  5. The branded content is delivered to configured recipients and stored in report archives.

Impact Assessment

The issue permits manipulation of report content, including recurring publication of a third-party link. It does not by itself provide filesystem, credential, or code-execution privileges. Its scope is the generated output and every configured delivery channel.

Remediation
View remediation

Remediation Suggestions

  • Make attribution explicitly optional and disabled by default.
  • Add a documented configuration field that allows users to choose or remove footer text.
  • Do not require strict reproduction of promotional URLs.
  • Clearly distinguish operational statistics requested by the user from project attribution.
  • Ensure scheduled workflows preserve the user’s selected attribution preference.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/enrich-articles.py:99
Finding

Server-Side Request Forgery Through Configurable Feeds and Article Enrichment

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch-rss.py:274-303 and scripts/enrich-articles.py:99-107
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

Vulnerable Code

From scripts/fetch-rss.py:

python
def fetch_feed_with_retry(source: Dict[str, Any], cutoff: datetime, no_cache: bool = False) -&gt; Dict[str, Any]:
    """Fetch RSS feed with retry mechanism and conditional requests."""
    source_id = source["id"]
    name = source["name"]
    url = source["url"]
    priority = source["priority"]
    topics = source["topics"]

    global _rss_cache, _rss_cache_dirty

    for attempt in range(RETRY_COUNT + 1):
        try:
            req_headers = {"User-Agent": "TechDigest/2.0"}

            # Add conditional headers from cache (thread-safe)
            with _rss_cache_lock:
                cache = _rss_cache if _rss_cache is not None else {}
                cache_entry = cache.get(url)
            now = time.time()
            ttl_seconds = RSS_CACHE_TTL_HOURS * 3600

            if cache_entry and not no_cache and (now - cache_entry.get("ts", 0)) &lt; ttl_seconds:
                if cache_entry.get("etag"):
                    req_headers["If-None-Match"] = cache_entry["etag"]
                if cache_entry.get("last_modified"):
                    req_headers["If-Modified-Since"] = cache_entry["last_modified"]

            req = Request(url, headers=req_headers)
            try:
                with urlopen(req, timeout=TIMEOUT) as resp:

From scripts/enrich-articles.py:

python
def fetch_full_text(url, max_chars=DEFAULT_MAX_CHARS):
    domain = get_domain(url)
    if domain in SKIP_DOMAINS:
        return {"text": "", "method": "skipped", "tokens": 0, "error": f"domain {domain} in skip list"}

    try:
        headers = {"Accept": "text/markdown, text/html;q=0.9", "User-Agent": USER_AGENT}
        req = Request(url
...[truncated 2222 chars]
Remediation
View remediation

Remediation Suggestions

  • Resolve every destination hostname before connecting and reject loopback, private, link-local, multicast, unspecified, and reserved addresses for both IPv4 and IPv6.
  • Revalidate the hostname and resolved address after every redirect.
  • Disable automatic redirects or implement a bounded redirect handler with destination validation.
  • Restrict allowed schemes to HTTPS where feasible and reject unusual ports.
  • Consider an explicit domain allowlist for feeds and article enrichment.
  • Protect against DNS rebinding by connecting only to a validated resolved address while preserving correct TLS hostname verification.
  • Apply strict response-size limits before reading response bodies.
  • Apply the same validation to initial feed URLs, resolved article links, and final response URLs.
  • Avoid placing content from untrusted or internal destinations into outbound reports.

T09 · Insecure Skill Coding Practices

Error
Location
references/digest-prompt.md:46
Finding

Shell Command Injection Through Unvalidated Prompt Placeholder Substitution

Content
View full analysis

Vulnerability Details

File Location: references/digest-prompt.md:46-52 and references/digest-prompt.md:141-148
Vulnerability Type: Shell command injection
Risk Level: High

Vulnerable Code

bash
python3 &lt;SKILL_DIR&gt;/scripts/run-pipeline.py \
  --defaults &lt;SKILL_DIR&gt;/config/defaults \
  --config &lt;WORKSPACE&gt;/config \
  --hours &lt;RSS_HOURS&gt; --freshness &lt;FRESHNESS&gt; \
  --archive-dir &lt;WORKSPACE&gt;/archive/tech-news-digest/ \
  --output /tmp/td-merged.json --verbose --force \
  $([ "&lt;ENRICH&gt;" = "true" ] &amp;&amp; echo "--enrich")
bash
python3 &lt;SKILL_DIR&gt;/scripts/send-email.py \
  --to '&lt;EMAIL&gt;' \
  --subject '&lt;SUBJECT&gt;' \
  --html /tmp/td-email.html \
  --attach /tmp/td-digest.pdf \
  --from '&lt;EMAIL_FROM&gt;'

Technical Analysis

The workflow instructs an agent to perform textual placeholder replacement inside shell commands. Several path placeholders are unquoted, including SKILL_DIR and WORKSPACE. Email-related values are enclosed in single quotes but are not validated or escaped, so a value containing a single quote can terminate the quoted argument and introduce shell syntax.

This contradicts the documentation’s assertion that no user input is interpolated into commands. Although the Python orchestration code uses argument arrays safely, the outer prompt still directs the agent to submit a shell command assembled from caller-supplied textual parameters.

Potential metacharacters include semicolons, command substitutions, redirection operators, and shell control operators. Whitespace in ordinary paths can also corrupt argument boundaries even without malicious intent.

Attack Path

  1. An attacker influences a cron placeholder, workspace path, Skill path, recipient address, sender value, or subject.
  2. The supplied value contains shell syntax, or contains a single quote followed by shell syntax for a q ...[truncated 956 chars]
Remediation
View remediation

Remediation Suggestions

  • Replace prompt-generated shell text with a structured launcher that invokes processes using argument arrays.
  • Store runtime parameters in a validated JSON configuration file and pass only that file’s path to the launcher.
  • Validate EMAIL and EMAIL_FROM with a strict email-address parser.
  • Restrict SUBJECT to printable text and reject control characters, newlines, and shell metacharacters.
  • Resolve and validate workspace and Skill paths in Python rather than interpolating them into shell commands.
  • If shell execution is unavoidable, use a well-tested shell-quoting function for every substituted value; do not rely on literal single quotes.
  • Replace the shell-based conditional ENRICH expression with explicit argument construction in Python.
  • Document which placeholders are trusted and reject values originating from fetched content.

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Third-Party Dependencies Without Integrity Verification

Content
View full analysis

Vulnerability Details

File Location: requirements.txt:1-8; installation documented in SKILL.md:325-329
Vulnerability Type: Dependency supply-chain weakness
Risk Level: Medium

Vulnerable Code

text
# Tech Digest Python Dependencies
# Install with: pip install -r requirements.txt

# RSS parsing (optional, will fallback to regex if not available)
feedparser&gt;=6.0.0

# JSON Schema validation (optional)
jsonschema&gt;=4.0.0

The Skill documentation instructs users to install these requirements:

bash
pip install -r requirements.txt

Technical Analysis

Both packages are specified using unrestricted lower bounds. Consequently, installation resolves to whichever compatible release is available at that time. No lock file, upper bound, exact version, or package hash is supplied.

The package names are established projects rather than obvious typosquatting attempts, and no evidence of a currently malicious dependency was identified. Nevertheless, the installation is not reproducible and does not guarantee that the installed artifacts are the versions reviewed by the Skill author.

A future compromised release, malicious package-index response, or incompatible update could be installed and execute code during installation or import.

Attack Path

  1. A user follows the documented pip install -r requirements.txt instruction.
  2. The package resolver selects the newest releases satisfying the lower bounds.
  3. A selected release is compromised or differs materially from the audited version.
  4. Package installation hooks or imported package code execute in the user’s Python environment.
  5. The compromised dependency obtains the privileges of the installing or executing user.

Impact Assessment

A compromised dependency could execute arbitrary Python code with the user’s privileges, access the active environment, modify files writable by that user, and perform outbound ne ...[truncated 269 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin dependencies to reviewed exact versions.
  • Generate and publish hashes for all permitted distributions and install with pip --require-hashes.
  • Maintain a lock file generated from a controlled dependency-resolution process.
  • Review and update pinned versions through a documented security-update procedure.
  • Recommend installation only in a dedicated, non-privileged virtual environment.
  • Separate optional dependencies into documented extras so users do not install them unnecessarily.
  • Add automated vulnerability and provenance checks for dependency updates.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (69)

Tainted flow: 'req' from os.getenv (line 380, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/fetch-web.py (reported line 83)May include surrounding context.

python
'X-Subscription-Token': api_key,
            'User-Agent': 'TechDigest/2.0'
        })
        with urlopen(req, timeout=TIMEOUT) as resp:
            limit_header = resp.headers.get('x-ratelimit-limit', '1')
            remaining = resp.headers.get('x-ratelimit-remaining', '')
            per_second = int(limit_header.split(',')[0].strip())

Tainted flow: 'req' from os.getenv (line 380, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/fetch-web.py (reported line 213)May include surrounding context.

python
'X-Subscription-Token': api_key,
            'User-Agent': 'TechDigest/2.0'
        })
        with urlopen(req, timeout=TIMEOUT) as resp:
            limit_header = resp.headers.get('x-ratelimit-limit', '1')
            remaining = resp.headers.get('x-ratelimit-remaining', '')
            per_second = int(limit_header.split(',')[0].strip())

Tainted flow: 'req' from os.getenv (line 380, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/fetch-web.py (reported line 384)May include surrounding context.

python
'X-Subscription-Token': api_key,
            'User-Agent': 'TechDigest/2.0'
        })
        with urlopen(req, timeout=TIMEOUT) as resp:
            limit_header = resp.headers.get('x-ratelimit-limit', '1')
            remaining = resp.headers.get('x-ratelimit-remaining', '')
            per_second = int(limit_header.split(',')[0].strip())

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 83)May include surrounding context.

  • WEB_SEARCH_BACKEND - Web search backend: auto|brave|tavily (optional, default: auto)
    • BRAVE_API_KEYS - Brave Search API keys, comma-separated for rotation (optional)
    • BRAVE_API_KEY - Single Brave key fallback (optional)
    • GITHUB_TOKEN - GitHub personal access token (optional, improves rate limits)
  1. Generate Digest:
    bash

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/fetch-github.py (reported line 428)May include surrounding context.

python
- `WEB_SEARCH_BACKEND` - Web search backend: auto|brave|tavily (optional, default: auto)
   - `BRAVE_API_KEYS` - Brave Search API keys, comma-separated for rotation (optional)
   - `BRAVE_API_KEY` - Single Brave key fallback (optional)
   - `GITHUB_TOKEN` - GitHub personal access token (optional, improves rate limits)

3. **Generate Digest**:
   ```bash

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · SKILL.md (reported line 453)May include surrounding context.

  • Anti-pattern: Do NOT copy pipeline steps into the cron prompt — it will drift out of sync

Multi-Channel Delivery Limitation

OpenClaw enforces cross-provider isolation: a single session can only send messages to one provider (e.g., Discord OR Telegram, not both). If you need to deliver digests to multiple platforms, create separate cron jobs for each provider:

text
# Job 1: Discord + Email

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file states that all prompts, templates, comments, and code must be in English. This is a natural-language policy constraint that forces a specific language choice and does not indicate user choice or opt-in, which matches the locale/language policy violation category.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README encourages broad natural-language installation and operation requests such as automatic setup, scheduling, and channel delivery without clear guardrails, confirmation steps, or exclusions. In an agent ecosystem, ambiguous trigger phrasing can cause overbroad actions, unintended configuration changes, or delivery to the wrong external destination if the agent interprets user intent too aggressively.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README states that the bot will handle installation, configuration, scheduling, and delivery through conversation, but it does not warn users that these actions may modify the system state and transmit content to external services. In an agent-driven environment, omission of these warnings increases the risk of users authorizing actions without understanding persistence, data flow, or channel exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README instructs users to provide multiple API keys and tokens for third-party services, but does not clearly disclose that the skill will consume those credentials to make outbound requests to external providers. In an agent-skill context, this is security-relevant because users may supply sensitive credentials without understanding which external services will receive data, what queries will be sent, or how broadly the skill can exercise those tokens.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documents capabilities including environment-variable access, file reads/writes, outbound network use, and shell execution, but it does not declare an explicit tool scope such as permissions or allowed-tools. That increases the chance the agent is invoked with broader authority than necessary, making any prompt-level misuse or future script change more dangerous.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest-style description broadly says the skill can 'Generate tech news digests' and lists capabilities, but it does not define specific activation phrases, scope boundaries, or exclusion conditions. In manifest files, this kind of open-ended description can overlap with generic requests about news summaries and increase the chance of unintended invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The scheduled prompt template sets 'LANGUAGE = English' as a fixed parameter, which imposes a specific language by default. The file does not indicate that this is user-selectable or limited to a justified region-specific use case, so it conflicts with the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The weekly scheduled prompt template also fixes 'LANGUAGE = English' rather than presenting language as a user choice. Because no opt-in or locale justification is provided, this is a natural-language policy issue under the language/locale rule.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 453)May include surrounding context.

  • Anti-pattern: Do NOT copy pipeline steps into the cron prompt — it will drift out of sync

Multi-Channel Delivery Limitation

OpenClaw enforces cross-provider isolation: a single session can only send messages to one provider (e.g., Discord OR Telegram, not both). If you need to deliver digests to multiple platforms, create separate cron jobs for each provider:

text
# Job 1: Discord + Email

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The network access section claims data is only sent to a fixed set of APIs, but other documented behavior includes full-article enrichment and Cloudflare Markdown for Agents, which implies requests to arbitrary article domains and additional services. In a skill that fetches external content, inaccurate endpoint disclosure hides the true egress surface and weakens review of data-handling risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The security documentation says no files outside the workspace are accessed, but the same file describes reading a GitHub App private key from an arbitrary path and using gh's credential store. This discrepancy can mislead reviewers and users about the real trust boundary, causing them to grant the skill more trust or access than intended.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a skill for generating tech news digests, with examples centered on LLM, AI-agent, and frontier-tech coverage. This source configuration repeatedly adds a separate 'crypto' topic across RSS, Twitter/X, GitHub, Reddit, and YouTube sources, expanding the skill into cryptocurrency/media monitoring rather than a clearly bounded tech-news digest.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This manifest enables multiple Chinese-language feeds and notes such as "中文科技媒体", "中文AI媒体", and "技术社区" as ordinary defaults, but the file does not indicate that users can choose their preferred language or opt in to non-default locales. Under the policy, forcing or silently defaulting to a specific language/locale without user choice can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest says the skill uses a six-source data collection model limited to RSS feeds, Twitter/X KOLs, GitHub releases, GitHub Trending, Reddit, and web search. This configuration adds a large number of additional sources beyond that model, including many explicit Reddit entries and YouTube channel feeds represented as RSS, which materially broadens the collection surface beyond the described scope.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest describes a tech news digest centered on six-source collection and digest generation, but this configuration explicitly defines a full 'Cryptocurrency' section covering Bitcoin, Ethereum, DeFi, NFTs, and crypto market developments. Crypto market news is not clearly implied by the stated purpose of a tech-news digest and materially broadens the content domain beyond the manifest's described scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The default language placeholder is set to Chinese, and the instructions later require writing the report in <LANGUAGE>. Because the file establishes a fixed default language without an explicit opt-in mechanism or justification for a locale restriction, it creates a natural-language policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The prompt instructs deletion of archive files older than 90 days without any safety guard, scope validation, or user-facing warning. In an agentic context, automated deletion can remove data a user expected to retain, especially if is misconfigured or shared, turning routine cleanup into unintended data loss.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This line requires the output to be written in <LANGUAGE>, and elsewhere the placeholder defaults that value to Chinese. Without an explicit user choice mechanism or documented justification, the skill effectively forces a locale, which falls under the language-policy violation category.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script defaults to writing output back to the input path when --output is omitted, which can silently overwrite the only copy of the source dataset. In automation or pipeline contexts, this increases the chance of destructive data loss or corruption if enrichment partially fails or produces unexpected output.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.