Back to skill

Security audit

HasData

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly coherent for web-data retrieval, but it asks agents to install and use a broad scraping/enrichment CLI in ways that expose users to supply-chain, privacy, and session-cookie risks.

Install only if you trust the hasdata CLI and are comfortable with a third-party scraping service handling your queries and requested pages. Avoid the documented curl-to-sh installer unless you first inspect and verify what it will run. Do not provide session cookies or personal contact lists unless you have a clear lawful basis and understand that cookies can act like account credentials.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:19
Finding

Unverified Mutable Remote Installer Is Executed Directly by a Shell

Content
View full analysis
Remediation
View remediation
/install.sh" ``` 4. Publish and verify a cryptographic checksum or signature before execution. 5. Allow the user to inspect the downloaded script before running it. 6. Require explicit user approval before installing or executing new software. 7. Prefer a trusted package manager or signed release package where available. 8. Run installation with ordinary user privileges and never recommend `sudo` unless a specific, documented operation requires it. 9. Document the files and directories modified by the installer. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
references/web-scraping.md:69
Finding

Authenticated Session Cookies May Be Exposed Through CLI Arguments and Third-Party Scraping Infrastructure

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
package.sh:15
Finding

Caller-Controlled Packaging Output Path Can Delete an Existing File

Content
View full analysis
Remediation
View remediation
&2 exit 1 fi ``` 2. Add an explicit `--force` option if overwrite behavior is required. 3. Validate that the output filename ends in `.zip`. 4. Optionally constrain output to the project directory or a designated build directory. 5. Canonicalize and validate the parent directory before creating the archive. 6. Reject unexpected target types, including directories and symbolic links. 7. Use a securely created temporary archive and atomically rename it to the final destination after successful packaging. 8. Avoid deleting the destination until all validation is complete. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (16)

Chaining Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

The | sh construct is a classic command-chaining anti-pattern that turns a network response directly into shell commands, enabling immediate arbitrary code execution with the current user's privileges. In this skill context, the danger is elevated because an automated agent may follow setup instructions non-interactively, reducing the chance of human review and increasing the blast radius of a compromised install script.

Content

Scanner excerpt · README.md (reported line 14)May include surrounding context.

The skill itself is just instructions — it shells out to the hasdata binary, which the agent installs on first use:

sh
curl -sSL https://raw.githubusercontent.com/HasData/hasdata-cli/main/install.sh | sh

Configure

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says the skill fetches live web data from numerous sources via the hasdata CLI. However, the supplied code chunk does not perform any data fetching, scraping, searching, URL verification, rendering, monitoring, or API access. Its sole purpose is packaging the skill files into a distributable zip. This is a materially different primary purpose and behavior from the declared functionality, so it is a clear mismatch.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill presents privacy-sensitive people lookup and contact-data retrieval as standard workflows without any warnings, consent checks, or restrictions. In context, that omission is dangerous because the tool is designed for real-time web retrieval and can scale invasive collection far beyond incidental browsing.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The CSV lead-enrichment workflow explicitly supports fan-out collection of profile, role, employer, and email information at scale. Bulk enrichment dramatically amplifies privacy harm and abuse potential, enabling mass profiling, spam campaigns, targeted phishing, or creation of sensitive people datasets with minimal friction.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · references/all-commands.md (reported line 85)May include surrounding context.

md
| --- | --- |
| `configure` | Interactive setup; writes ~/.hasdata/config.yaml |
| `version` | Print version |
| `update` | Self-update from GitHub Releases |
| `completion {bash\|zsh\|fish\|powershell}` | Generate shell completion |

## Hidden / deprecated

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README instructs users or agents to execute a remote script directly via curl ... | sh, which gives unreviewed network-fetched content immediate shell execution. In an agent skill context, this is more dangerous because the skill is explicitly designed to trigger tool use and installation on first use, so a compromised upstream script, repo, or transport path could lead to arbitrary code execution on the host.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill description is extremely broad and encourages invocation for a wide range of loosely related requests. In an autonomous agent, ambiguous scope increases the chance the tool is used unnecessarily, including for sensitive scraping or personal-data retrieval when a narrower, less risky method would suffice.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Non-obvious triggers' section maps many everyday phrases to automatic use of this skill, including sensitive workflows like reputation checks, contact lookup, and enrichment. This broad trigger surface makes over-activation more likely and reduces opportunities for the agent to assess privacy, legality, and necessity before collecting data.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly guides users to perform people enrichment, reverse lookup, and discovery of personal identifiers such as employer, role, LinkedIn presence, email, and phone-linked identity. In an agent setting, this materially increases the chance of privacy-invasive data collection, doxxing, or targeted phishing workflows, especially because the guidance is framed as normal usage rather than exceptional or policy-gated behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation recommends searching for actual company or personal email addresses via quoted search-engine queries and even suggests pattern-guessing followed by verification. That facilitates harvesting contact data and can directly support spam, phishing, impersonation, or unwanted deanonymization of individuals.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation instructs users how to collect and reveal personal contact and identity data through plain-language enrichment workflows, including email and phone reverse lookups. This lowers the barrier to harmful targeting and makes misuse for phishing, social engineering, harassment, or deanonymization straightforward.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation explicitly recommends chaining place lookup with web scraping to extract emails from business websites, but provides no guardrails around privacy, consent, lawful basis, anti-spam, or terms-of-service considerations. In a lead-generation context, this materially enables bulk harvesting of contact data for unsolicited outreach or profiling, making misuse straightforward even if the emails are publicly exposed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The guidance explicitly suggests passing authentication cookies into scraping requests for cookie-walled or login-walled content. Even though it says to use this sparingly and with explicit user permission, it normalizes handling highly sensitive session material in a scraping workflow without strong safeguards, warnings about session-token exposure, or restrictions on exfiltration to third-party targets. In a web-scraping skill, this is more dangerous because the core function is making arbitrary outbound requests to user-specified URLs, which increases the risk of sending valid session cookies to unintended or malicious destinations.

Content

No source excerpt is available for this finding.

External Script Fetching

Low
Category
Supply Chain
Confidence
94% confidence
Finding

This finding is valid because the skill fetches an external script from GitHub at runtime and immediately uses it as part of installation guidance. External script fetching is risky in any documentation, but especially here because the skill states the agent installs the binary on first use, creating a path for automatic retrieval and potential execution of attacker-controlled content if the source is compromised.

Content

Scanner excerpt · README.md (reported line 14)May include surrounding context.

The skill itself is just instructions — it shells out to the hasdata binary, which the agent installs on first use:

sh
curl -sSL https://raw.githubusercontent.com/HasData/hasdata-cli/main/install.sh | sh

Configure

External Script Fetching

Low
Category
Supply Chain
Confidence
92% confidence
Finding

The skill recommends curl ... | sh, which executes a remote script directly from GitHub without integrity verification or pinning. If the upstream repository, network path, or referenced script is compromised, users could execute arbitrary code on their system.

Content

Scanner excerpt · SKILL.md (reported line 19)May include surrounding context.

md
## Prerequisites

- `command -v hasdata` — if missing, install with `curl -sSL https://raw.githubusercontent.com/HasData/hasdata-cli/main/install.sh | sh`.
- One-time setup: the user runs `hasdata configure`, pastes their API key, and it's saved to `~/.hasdata/config.yaml` (mode 0600). Every future call picks it up automatically.
- If a call fails with `no API key configured`, the user hasn't run `hasdata configure` yet — tell them to. **Never invent a key.**

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This is a markdown file, so SQP-2 applies to omissions in the skill description. The entries for configure and update describe writing ~/.hasdata/config.yaml and performing a self-update from GitHub Releases, but the file provides no caution about modifying the local system or fetching/installing remote updates.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.