Back to skill

Security audit

Google Maps Email Extractor (Apify)

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it advertises, but it needs review because it defaults to collecting personal contact data and handles the Apify token in riskier ways than necessary.

Install only if you are comfortable sending lead-search inputs and extracted contact data to Apify under your account. Prefer setting APIFY_TOKEN in the environment, avoid passing it on the command line, use a tight Apify token, set a budget, and set includePersonalData=false unless you have a lawful, consent-aware reason to collect person-like emails or personal profile URLs.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/google_maps_email_extractor_actor.py:238
Finding

Apify API Token Exposed in the Request URL

Content
View full analysis
dict[str, Any]: params: dict[str, str | int | float] = { "token": token, "timeout": timeout_sec, "clean": "true", } if budget_usd is not None: if budget_usd <= 0: raise SkillError("--budget-usd must be > 0.") params["maxTotalChargeUsd"] = budget_usd url = ( f"https://api.apify.com/v2/acts/{urllib.parse.quote(actor_id, safe='')}/run-sync-get-dataset-items" f"?{urllib.parse.urlencode(params)}" ) body = json.dumps(payload).encode("utf-8") req = urllib.request.Request(url=url, data=body, headers={"Content-Type": "application/json"}, method="POST") try: with urllib.request.urlopen(req, timeout=min(timeout_sec + 30, 3600)) as response: ``` ### Technical Analysis The Apify API token is added to the request URL as a `token` query parameter. TLS protects the URL while it is transmitted directly to Apify, and the destination is the documented official Apify endpoint. This is therefore not evidence of malicious exfiltration. However, URLs are commonly recorded by reverse proxies, HTTP access logs, debugging tools, monitoring agents, exception telemetry, and network security products. A query-string credential can consequently be disclosed to systems or personnel that do not otherwise need access to it. URL credentials may also be retained longer than authorization headers. The network request itself is necessary for the Skill's declared functionality, but placing the credential in the URL does not follow least-exposure practices. ### Attack Path 1. A user runs the Skill with a valid `APIFY_TOKEN`. 2. The run ...[truncated 756 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/google_maps_email_extractor_actor.py:46
Finding

Apify Token Accepted Through Process Command-Line Arguments

Content
View full analysis
str: token = explicit or os.getenv("APIFY_TOKEN", "") if not token: raise SkillError("Apify token missing. Pass --apify-token or set APIFY_TOKEN.") return token ``` ```python def add_common_args(parser: argparse.ArgumentParser) -> None: parser.add_argument("--apify-token", help="Apify API token (fallback: APIFY_TOKEN env)") parser.add_argument("--actor-id", default=DEFAULT_ACTOR_ID, help="Apify actor ID") parser.add_argument("--timeout-sec", type=int, default=DEFAULT_TIMEOUT_SEC, help="Run timeout in seconds") parser.add_argument("--budget-usd", type=float, help="Apify maxTotalChargeUsd run option") ``` ```python try: token = resolve_token(args.apify_token) ``` ### Technical Analysis The runner permits a secret to be supplied with `--apify-token`. Command-line arguments can be exposed through shell history, process inspection utilities, endpoint monitoring, operating-system auditing, CI/CD job logs, and orchestration metadata. The documentation recommends `APIFY_TOKEN`, which reduces the likelihood of this exposure, but the insecure command-line mechanism remains available and is explicitly mentioned in the error message. Environment variables are not universally secret, but they are generally less likely than command-line arguments to be stored in shell history or displayed in routine process listings. ### Attack Path 1. A user invokes the runner with `--apify-token apify_api_sensitive_value`. 2. The shell may save the entire command in its history file. 3. While the process runs, a local process monitor or sufficiently authorized local user reads its argument vector. 4. CI/CD, endpoint monitoring, or audit sof ...[truncated 469 chars]
Remediation
View remediation

other

Warning
Location
scripts/google_maps_email_extractor_actor.py:141
Finding

Personal Contact Data Collection Enabled by Default

Content
View full analysis
dict[str, Any]: payload["contactResultMode"] = args.result_mode payload["contactPagesLimit"] = positive_int(args.contact_pages, "contact-pages") payload["includePersonalData"] = not args.no_personal_data payload["website"] = "withWebsite" ``` The provided sample repeats the privacy-expansive default: ```json { "searchStringsArray": [ "wedding photographer" ], "locationQuery": "Austin, Texas, USA", "maxCrawledPlacesPerSearch": 25, "contactResultMode": "emailsOnly", "contactPagesLimit": 2, "includePersonalData": true, "website": "withWebsite", "skipClosedPlaces": true, "language": "en" } ``` ### Technical Analysis According to the project contract, `includePersonalData=true` permits collection of person-like email addresses and personal LinkedIn profile URLs found on public websites. Those fields are optional for the narrower purpose of obtaining generic business contact details. Both custom and quick-run paths enable this behavior without affirmative opt-in. The result is sent to the declared Apify actor and printed in the returned JSON; no unrelated recipient was identified. Nevertheless, enabling personal-data p ...[truncated 1177 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
README.md:76
Finding

Unpinned Third-Party Package Execution in Installation Instructions

Content
View full analysis
Remediation
View remediation
...`. 2. Use lockfiles and npm integrity hashes where the installation workflow supports them. 3. Document the expected package registry and warn users not to run the command with an untrusted custom registry. 4. Prefer an installation method that downloads reviewed Skill files without executing transient package code. 5. Periodically review and update the pinned version rather than automatically resolving the latest release. 6. Publish checksums or signed release artifacts for the Skill package. 7. In high-assurance environments, instruct users to inspect the resolved package and Skill contents before installation. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (21)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 60)May include surrounding context.

md
7. Run `scripts/google_maps_email_extractor_actor.py` or call the Apify API directly.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 96)May include surrounding context.

md
7. Run `scripts/google_maps_email_extractor_actor.py` or call the Apify API directly.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 107)May include surrounding context.

md
7. Run `scripts/google_maps_email_extractor_actor.py` or call the Apify API directly.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 117)May include surrounding context.

md
7. Run `scripts/google_maps_email_extractor_actor.py` or call the Apify API directly.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 126)May include surrounding context.

md
7. Run `scripts/google_maps_email_extractor_actor.py` or call the Apify API directly.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill explicitly requires an API token, references reading local files, and instructs direct calls to an external Apify actor, but it does not declare any tool scope or allowed-tools boundary. This creates an authorization and transparency gap: an agent/runtime may grant broader env, file, or network access than users expect, increasing the chance of over-privileged execution or misuse of credentials.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill promotes extraction of business email leads and even allows includePersonalData=true for person-like emails, yet it provides no privacy, consent, or legal-compliance warning. In context, this increases the risk of collecting and processing personal contact data for outreach in ways that may violate privacy law, anti-spam rules, or platform terms.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The default prompt auto-invokes a specific skill and external actor using a terse trigger-style instruction without stating when activation is appropriate, what inputs are required, or what safety checks should occur first. In a lead-generation context involving scraping public business emails, this increases the chance of unintended execution, over-collection, or use in contexts the user did not explicitly authorize.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The contract explicitly supports collecting person-like emails and personal LinkedIn URLs from public websites via includePersonalData without any accompanying privacy warning, consent guidance, or restrictions on lawful use. Even though the data is publicly accessible, documenting and operationalizing bulk collection of personal contact data increases privacy, compliance, and misuse risk for targeted outreach or profiling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The output contract describes extraction and return of emails, phone numbers, social profiles, and source URLs from public business websites, but provides no warning about privacy-sensitive handling, retention, or downstream use. In the context of a lead-generation skill, this omission makes bulk contact harvesting more likely to be used in ways that violate privacy expectations, platform rules, or anti-spam requirements.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script sends user-supplied search parameters and receives extracted contact data from a third-party service (Apify), while defaulting to includePersonalData=True. In a lead-generation skill that collects emails and related business contact details, this creates a real privacy and data-governance risk because users may not realize their queries and extracted data are being transmitted to and processed by an external provider.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 10)May include surrounding context.

md
params["maxTotalChargeUsd"] = budget_usd

    url = (
        f"https://api.apify.com/v2/acts/{urllib.parse.quote(actor_id, safe='')}/run-sync-get-dataset-items"
        f"?{urllib.parse.urlencode(params)}"
    )
    body = json.dumps(payload).encode("utf-8")

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/google_maps_email_extractor_actor.py (reported line 250)May include surrounding context.

python
params["maxTotalChargeUsd"] = budget_usd

    url = (
        f"https://api.apify.com/v2/acts/{urllib.parse.quote(actor_id, safe='')}/run-sync-get-dataset-items"
        f"?{urllib.parse.urlencode(params)}"
    )
    body = json.dumps(payload).encode("utf-8")

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · LICENSE (reported line 12)May include surrounding context.

text
permit persons to whom the Software is furnished to do so.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED,
INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A
PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT
HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION
OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The recommended inputs hard-code language: "en", and the document does not state that language is user-selectable or explain why English is required. Under the language/locale policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
72% confidence
Finding

The manifest explicitly instructs use of APIFY_TOKEN whenever the default prompt is followed, without indicating user consent, credential-scoping expectations, or whether invocation should be conditional. Although this does not directly expose the token, it normalizes automatic third-party credential use and can cause unnecessary external requests under account authority.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This manifest-like JSON file sets the language to "en" explicitly, which is a natural-language locale constraint. There is no indication in the file that users can opt in to this language choice or that the English-only setting is justified by a region-specific requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The payload normalization sets language to "en" by default, which imposes a specific locale behavior. The file does allow overriding language in some CLI paths, but the default still hard-codes English without documenting a policy reason or obtaining user preference.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.