Back to skill

Security audit

Apify Google News Scraper

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent Apify-backed news scraper, but it should be reviewed because it sends queries to a third party and repeatedly places the Apify token in request URLs.

Review before installing. Use only non-sensitive search terms unless you accept Apify and the named actor handling them, prefer a narrowly scoped and revocable Apify token, and avoid copying or running the examples as written unless you change them to use an Authorization header instead of putting the token in the URL.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:48
Finding

Apify API Token Exposed Through URL Query Parameters

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The invocation phrases are broad enough to match common requests for news retrieval, which can cause the skill to activate in situations where a user did not clearly consent to using a third-party scraping service. In this skill, activation leads to sending user queries to Apify and retrieving full article content, so overbroad routing increases unintended data disclosure and use of external credentials.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill documentation does not clearly warn that user search terms and retrieval requests are sent to a third-party service using an API token, nor that the actor can fetch article text from external sites. This creates a consent and privacy risk because users may provide sensitive topics or assume the assistant is performing local/news-native retrieval instead of external transmission.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
83% confidence
Finding

The presence of the external API base URL indicates this skill depends on a third-party service for execution. By itself this is less severe than a concrete request, but in the surrounding context it confirms that user data and fetched content will leave the local environment.

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
import requests, os, time

TOKEN = os.environ["APIFY_API_TOKEN"]
BASE = "https://api.apify.com/v2"

# Step 1: Start the run
response = requests.post(

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This code transmits user-provided search queries and an API credential to an external Apify endpoint to start a scraping run. Although expected for the feature, it is security-relevant because it exports user intent to a third party and places the token in the URL query string, which may be logged by intermediaries, shells, or monitoring tools.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

md
BASE = "https://api.apify.com/v2"

# Step 1: Start the run
response = requests.post(
    f"{BASE}/acts/futurizerush~google-news-scraper/runs?token={TOKEN}",
    json={
        "searchQueries": ["AI"],

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This example again demonstrates transmission of search terms to Apify, here for region/language-specific news, with the API token embedded in the request URL. The context makes the transfer expected, but it still exposes user queries to a third party and risks credential leakage through URL logging.

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

Taiwan news in Chinese

python
requests.post(
    f"{BASE}/acts/futurizerush~google-news-scraper/runs?token={TOKEN}",
    json={
        "searchQueries": ["台灣"],

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This request sends multiple user search terms in one external API call, potentially broadening the amount of user interest/profile data disclosed to Apify. As with the other examples, the token is included in the URL, increasing the chance of credential exposure via logs or telemetry.

Content

Scanner excerpt · SKILL.md (reported line 107)May include surrounding context.

Multiple queries

python
requests.post(
    f"{BASE}/acts/futurizerush~google-news-scraper/runs?token={TOKEN}",
    json={
        "searchQueries": ["AI", "climate", "crypto"],

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

The bash example sends requests to Apify with the API token directly in the URL and may expose that token through shell history, process listings, CI logs, or proxy/server logs. It also transmits user queries externally, which is expected for the skill but still requires clear disclosure and safer credential handling.

Content

Scanner excerpt · SKILL.md (reported line 122)May include surrounding context.

bash
# Step 1: Start the run
RUN_RESPONSE=$(curl -s -X POST \
  "https://api.apify.com/v2/acts/futurizerush~google-news-scraper/runs?token=$APIFY_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchQueries": ["AI"], "region": "us", "language": "en", "dateFilter": "1d", "maxResults": 10}')

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

This curl invocation concretely transmits data and exposes the API token in the URL, which is a common secret-handling weakness. In a skill that may be copied into terminals or automation, this increases the risk that credentials and user search parameters are retained in logs or job traces.

Content

Scanner excerpt · SKILL.md (reported line 123)May include surrounding context.

bash
# Step 1: Start the run
RUN_RESPONSE=$(curl -s -X POST \
  "https://api.apify.com/v2/acts/futurizerush~google-news-scraper/runs?token=$APIFY_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchQueries": ["AI"], "region": "us", "language": "en", "dateFilter": "1d", "maxResults": 10}')

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

Polling the actor-run status with the token in the URL continues the credential exposure pattern across repeated requests, multiplying opportunities for token leakage in logs and observability tooling. While the endpoint itself is operationally necessary, the authentication method is unsafe and still reveals use of the third-party service.

Content

Scanner excerpt · SKILL.md (reported line 132)May include surrounding context.

md
# Step 2: Poll until done
while true; do
  STATUS=$(curl -s "https://api.apify.com/v2/actor-runs/$RUN_ID?token=$APIFY_API_TOKEN" \
    | jq -r '.data.status')
  [ "$STATUS" = "SUCCEEDED" ] && break
  [ "$STATUS" = "FAILED" ] || [ "$STATUS" = "ABORTED" ] && echo "Failed: $STATUS" && exit 1

External Transmission

Medium
Category
Data Exfiltration
Confidence
92% confidence
Finding

Fetching dataset items retrieves the scraped results from Apify and again places the API token in the URL. This is additionally sensitive because the returned data can include full article content and metadata, so both the credential and potentially sensitive output traverse third-party infrastructure.

Content

Scanner excerpt · SKILL.md (reported line 140)May include surrounding context.

done

Step 3: Fetch results

curl -s "https://api.apify.com/v2/datasets/$DATASET_ID/items?token=$APIFY_API_TOKEN" | jq '.'

text

## Output Format

Static analysis

No suspicious patterns detected.