Back to skill

Security audit

ResearchVault

Security checks for vulnerabilities and agentic risk

Overview

ResearchVault mostly matches its local research purpose, but concrete web-ingestion and local-file containment flaws make it worth careful review before installing.

Install only if you are comfortable running a local research portal/CLI that can fetch URLs, call configured search providers, and index selected local files. Keep the portal bound to localhost, protect .portal_auth, avoid exposing the port, do not enable private-network ingestion unless you intend to fetch internal resources, avoid ingesting untrusted URLs, avoid registering symlinked or /tmp artifacts, and prefer locked dependency installs before regular use.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/scuttle.py:76
Finding

SSRF Protection Can Be Bypassed Through DNS Rebinding

Content
View full analysis
None: parsed = urllib.parse.urlparse(url) if parsed.scheme not in {"http", "https"}: raise ScuttleError(f"Scheme '{parsed.scheme}' is not allowed. Only http/https.") hostname = parsed.hostname if not hostname: raise ScuttleError("URL must include a hostname.") # 1. Block obviously local hostnames if hostname.lower() in {"localhost", "metadata.google.internal"} or hostname.endswith(".local") or hostname.endswith(".localhost"): if not allow_private: raise ScuttleError(f"Blocked host: '{hostname}' is a local or internal address.") # 2. Resolve DNS and check all returned IPs try: addr_info = socket.getaddrinfo(hostname, parsed.port or (80 if parsed.scheme == "http" else 443)) except socket.gaierror as e: raise ScuttleError(f"DNS resolution failed for '{hostname}': {e}") for entry in addr_info: ip = entry[4][0] if not _is_safe_ip(ip, allow_private): raise ScuttleError(f"Blocked host: '{hostname}' resolves to a restricted IP: {ip}") ``` ### Technical Analysis The application validates a hostname by resolving it with `socket.getaddrinfo()`, but it does not bind the subsequent HTTP connection to one of the validated IP addresses. The Requests/urllib3 networking stack performs another DNS resolution when `super().request()` opens the connection. This creates a time-of-check/time-of-use gap. An attacker controlling DNS can return a ...[truncated 1698 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/core.py:439
Finding

Artifact Path Validation Can Be Bypassed with Symlinks and Weak Substring Checks

Content
View full analysis
str: # --- Security Hardening: Path Sanitization --- # Resolve and check against safe local directories abs_path = os.path.abspath(os.path.expanduser(path)) vault_root = os.path.abspath(os.path.expanduser("~/.researchvault")) # Check if path is within allowed boundaries is_safe = False for safe_root in [vault_root]: if abs_path.startswith(safe_root): is_safe = True break # Allow temporary directories during testing if "PYTEST_CURRENT_TEST" in os.environ or "TEMP" in abs_path or "tmp" in abs_path: is_safe = True if not is_safe: raise ValueError(f"Security violation: Artifact path must be within {vault_root}") # --------------------------------------------- artifact_id = f"art_{uuid.uuid4().hex[:10]}" now = datetime.now().isoformat() branch_id = resolve_branch_id(project_id, branch) conn = db.get_connection() c = conn.cursor() # Scrub sensitive data from metadata metadata = scrub_data(metadata) path = scrub_data(path) c.execute( """INSERT INTO artifacts (id, project_id, type, path, metadata, created_at, branch_id) VALUES (?, ?, ?, ?, ?, ?, ?)""", (artifact_id, project_id, type, path, json.dumps(metadata or {}), now, branch_id), ) ``` The registered path is subsequently opened during synthesis: ```python def _read_text_file(path: str, limit_chars: int = 20_000) -> str: try: with open(path, "r", ...[truncated 2990 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scuttle.py:83
Finding

Response Size Limit Is Not Enforced for Streaming or Misreported Responses

Content
View full analysis
self.scuttle_config.max_size_bytes: raise ScuttleError(f"Response too large: {cl} bytes") except (ValueError, TypeError): pass ``` Callers then buffer the response: ```python with SafeSession(config) as session: resp = session.get(url, headers=headers) resp.raise_for_status() soup = BeautifulSoup(resp.text, 'html.parser') ``` ```python with SafeSession(config) as session: resp = session.get(source, headers=headers) resp.raise_for_status() soup = BeautifulSoup(resp.text, 'html.parser') ``` ### Technical Analysis The nominal 10 MB limit is enforced only when the server supplies a valid `Content-Length` header greater than the configured maximum. HTTP responses may legitimately omit this header, use chunked transfer encoding, or report an inaccurate smaller value. Although the request is opened with `stream=True`, accessing `resp.text` causes Requests to consume and buffer the entire response body. No counter limits the bytes read from the stream. Consequently, the configured maximum is not an effective hard limit. ### Attack Path 1. An attacker hosts a URL that returns a very large or indefinitely streamed HTTP response. 2. The response omits `Content-Length`, uses chunked encoding, or reports a value below the configured maximum. 3. A user or authenticated portal caller submits the URL for ingestion. 4. `SafeSession` accepts the response because the header-based check does not reject it. 5. The connector accesses `resp.text`. 6. Requ ...[truncated 731 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
pyproject.toml:10
Finding

Python Dependencies Are Installed from Broad Unpinned Version Ranges

Content
View full analysis
=2.32.0", "beautifulsoup4>=4.12.0", "rich>=13.9.0", "pytest>=9.0.2", "pytest-mock>=3.15.1", "responses>=0.25.8", "fastapi>=0.100.0", "uvicorn[standard]>=0.23.0", ] ``` The installation instructions resolve these mutable ranges: ```bash python3 -m venv .venv source .venv/bin/activate pip install -e . ``` Portal startup also invokes dependency synchronization: ```bash echo "[1/4] Checking Python Dependencies..." UV_BIN="${UV_BIN:-$HOME/.local/bin/uv}" if [ ! -x "$UV_BIN" ]; then if command -v uv >/dev/null 2>&1; then UV_BIN="$(command -v uv)" else echo "Error: uv not found. Install uv or set UV_BIN." >&2 exit 1 fi fi echo "Using uv from $UV_BIN" "$UV_BIN" sync ``` ### Technical Analysis The audited directory does not contain a Python dependency lockfile. The use of broad `>=` constraints allows future releases to be selected without a source-code change or repeat audit. This makes installations non-reproducible and expands exposure to compromised, malicious, or unexpectedly incompatible package releases. The frontend includes a package lockfile with registry URLs and integrity hashes, which materially reduces the equivalent npm risk. The concern primarily applies to Python dependencies. Development and test packages such as `pytest`, `pytest-mock`, and `responses` are included in the main runtime dependency set, unnecessarily increasing the installed attack surface. ### Attack Path 1. A permitted dependency or one of its transitive dependencies publishes a compromised future release. 2. A user installs ResearchVault or starts the portal in an environment wher ...[truncated 1043 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (105)

Tainted flow: 'headers' from os.environ.get (line 282, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/core.py (reported line 288)May include surrounding context.

python
}
    params = {"q": query}
    
    response = requests.get(url, headers=headers, params=params)
    response.raise_for_status()
    return response.json()

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a substantive local-first research orchestration system with multiple capabilities. The supplied code chunk is merely a minimal stub that prints a greeting when executed. This is a materially different primary purpose and lacks the core described functionality, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code chunk is focused on authentication infrastructure, not research orchestration, synthesis, or background service management. It introduces a materially different capability: portal access control using a configured secret, HMAC-signed stateless sessions, cookie-based authentication, and expiration checks. These are substantive behaviors not represented in the declared description, so this is a description-behavior mismatch rather than a mere supporting implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description emphasizes a local-first research orchestration engine focused on state, synthesis, and optional background services. The supplied code instead implements the main entry point for a FastAPI-based portal backend: it configures CORS, registers auth/system/vault API routers, performs startup initialization by invoking a vault CLI command, and exposes a health endpoint. While the vault initialization could be loosely related to state management, the primary behavior here is operating a web API service with authentication and portal endpoints, which is not accurately represented by the declared description. This is a material purpose mismatch rather than a minor implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description focuses on research orchestration, state management, synthesis, and optional background services. The supplied code instead provides web authentication functionality: it accepts a login token, creates sessions, sets secure/HTTP-only cookies, revokes sessions on logout, and checks authentication status. Authentication may support a larger application, but it is a distinct capability not represented in the declared purpose. This is more than an implementation detail because the code exposes dedicated auth endpoints and manages user/session access control.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code does not implement or support a research orchestration engine, state management, synthesis, or background services. Instead, it is purely a static linting configuration for frontend source files. This is a materially different purpose from the declared description, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description describes a research orchestration engine with state management, synthesis, and optional background services. The supplied code does not implement or indicate any of those behaviors. Instead, it is a simple frontend build-tool configuration for PostCSS, specifying Tailwind CSS and Autoprefixer plugins. This is a materially different purpose and unrelated capability, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description describes an orchestration engine for research workflows and optional background services. The supplied code chunk does not implement any of that behavior; it only configures Tailwind CSS for a frontend, including theme colors, typography, and animations. This is a materially different purpose rather than a supporting implementation detail of orchestration logic, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description emphasizes a local-first orchestration engine concerned with state, synthesis, and optional background services. The supplied code instead primarily implements a remote content fetching/scraping module: it makes HTTP requests, resolves DNS, enforces SSRF protections, follows redirects, scrapes/parses Reddit/web/YouTube/Grokipedia content, and produces ArtifactDraft objects. This is a materially different primary purpose and introduces undeclared external network/resource access. The safety controls are implementation details, but the core behavior is still remote ingestion rather than local orchestration.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description is broad and centered on local-first research orchestration, state management, synthesis, and optional MCP/Watchdog background services. The provided code chunk instead performs a specific heartbeat/scuttling task: it retrieves a Moltbook feed item and records it as an event for a hard-coded project. That behavior is more akin to feed ingestion or telemetry logging than research orchestration or synthesis. While it could be considered a background service, the specific capability to fetch from Moltbook and log signals is not represented in the declared purpose, making the description materially incomplete relative to the code's actual function.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description focuses on a research orchestration engine handling state, synthesis, and optional background services. This code chunk does not implement orchestration logic; instead, its primary purpose is operational management of a local portal application. It starts and stops backend/frontend processes, performs dependency setup, manages PID files, checks ports, configures database/env settings, and handles portal authentication token creation and display. While this may support a broader local-first system, the code shown is materially centered on portal service lifecycle and web UI exposure, which is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description centers on research orchestration, state management, synthesis, and optional background services. The supplied code instead exercises a web-fetching component's SSRF defense behavior. That is a materially different functional area: network access control and URL retrieval safety rather than orchestration/state/synthesis. Even though this is only a test file, it still indicates the skill includes undeclared web-scuttling/network security capabilities not represented in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description focuses on a local-first research orchestration engine with state management, synthesis, and optional background services. The supplied code instead tests a connector for an external knowledge source (Grokipedia), including URL detection and HTTP-based fetching of article data. That is a materially different primary behavior than orchestration/state management, and it introduces undeclared external network/content-ingestion capability. Although this is only a test file, it clearly reflects functionality of the skill toward remote Grokipedia retrieval rather than the declared orchestration role.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description focuses on research orchestration, state/synthesis management, and optional background services. The supplied code instead tests web portal authentication mechanisms, including session cookies, expiry, login responses, and configuration of an authentication token. These are materially different capabilities and indicate a portal auth subsystem rather than a research orchestration engine. While auth could exist as supporting infrastructure in a larger system, this specific code chunk's primary purpose is not represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description emphasizes a local-first research orchestration engine concerned with state management, synthesis, and optional background services. The supplied code does not reflect orchestration, state management, synthesis, MCP, or Watchdog behavior. Instead, it tests a source-specific fetch/scuttle mechanism for moltbook:// content, including source attribution, confidence scoring, tagging, and artifact draft creation. That is a materially different primary purpose and introduces undeclared content-ingestion capabilities, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
88% confidence
Finding

The declared description emphasizes a local-first research orchestration engine concerned with state management, synthesis, and optional background services. The supplied code instead tests a specific external-content connector for YouTube, including URL recognition and fetching/parsing metadata from web content. That is a materially different functional area from orchestration and background service management. While such a connector could exist within a broader research system, this code chunk's behavior is not accurately represented by the declared description and introduces undeclared external web ingestion capabilities.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 107)May include surrounding context.

md
python scripts/vault.py init --objective "Analyze AI trends" --name "Trends-2026"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 112)May include surrounding context.

md
python scripts/vault.py init --objective "Analyze AI trends" --name "Trends-2026"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 117)May include surrounding context.

md
python scripts/vault.py init --objective "Analyze AI trends" --name "Trends-2026"

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · portal/backend/app/portal_state.py (reported line 55)May include surrounding context.

python
data = json.loads(path.read_text(encoding="utf-8"))
            return _coerce_state(data)
    except Exception:
        # Be resilient: a corrupt state file should never brick the Portal.
        return PortalState()

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · portal/backend/app/vault_exec.py (reported line 64)May include surrounding context.

python
root = _repo_root()

    # Ensure rich doesn't emit ANSI in captured output.
    env = dict(os.environ)
    env.setdefault("NO_COLOR", "1")
    env.setdefault("RICH_NO_COLOR", "1")
    env.setdefault("TERM", "dumb")

Known Vulnerable Dependency: brace-expansion==1.1.12 — 4 advisory(ies): CVE-2026-13149 (brace-expansion: DoS via exponential-time expansion of consecutive non-expanding); CVE-2026-33750 (brace-expansion: Zero-step sequence causes process hang and memory exhaustion); CVE-2026-14257 (brace-expansion: DoS via unbounded expansion length causing an out-of-memory pro) +1 more

High
Category
Supply Chain
Confidence
93% confidence
Finding

brace-expansion 1.1.12 is present transitively and the cited advisories describe denial-of-service conditions from pathological expansion patterns. This is a real vulnerability class, though in this file it is limited to development and tooling paths unless attacker-controlled glob patterns are fed into those tools.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: minimatch==3.1.2 — 3 advisory(ies): CVE-2026-27904 (minimatch ReDoS: nested *() extglobs generate catastrophically backtracking regu); CVE-2026-26996 (minimatch has a ReDoS via repeated wildcards with non-matching literal in patter); CVE-2026-27903 (minimatch has ReDoS: matchOne() combinatorial backtracking via multiple non-adja)

High
Category
Supply Chain
Confidence
93% confidence
Finding

minimatch 3.1.2 has reported ReDoS issues involving crafted glob patterns that trigger catastrophic backtracking or excessive computation. In this frontend lockfile it is a genuine dependency risk, but practical impact depends on whether untrusted patterns can reach lint/build tooling.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: brace-expansion==2.0.2 — 4 advisory(ies): CVE-2026-13149 (brace-expansion: DoS via exponential-time expansion of consecutive non-expanding); CVE-2026-33750 (brace-expansion: Zero-step sequence causes process hang and memory exhaustion); CVE-2026-14257 (brace-expansion: DoS via unbounded expansion length causing an out-of-memory pro) +1 more

High
Category
Supply Chain
Confidence
93% confidence
Finding

brace-expansion 2.0.2 carries similar denial-of-service risks from crafted expansion expressions. This is a genuine vulnerable package version in the lockfile, with impact mainly in build/lint tooling that accepts attacker-controlled patterns.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: browserslist==4.28.1 — 2 advisory(ies): CVE-2026-73088 (Browserslist: Uncaught crash / prototype write via untrusted browserslist-stats.); CVE-2026-73089 (Browserslist: Unbounded memory growth (no cache eviction) via distinct query res)

High
Category
Supply Chain
Confidence
87% confidence
Finding

browserslist 4.28.1 is reported vulnerable to crash/prototype-write behavior and unbounded memory growth under untrusted input scenarios. In a frontend toolchain this is a real issue, but it is primarily relevant when attacker-supplied browserslist queries or stats files are processed in CI or builds.

Content

No source excerpt is available for this finding.