Back to skill

Security audit

langgraph-for-agents

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent LangGraph/LangChain reference skill with expected LLM, search, RAG, and memory examples, but users should treat those examples as networked demo code.

Install/use is reasonable for LangGraph or LangChain development, but treat the reference Python files as live networked examples: prompts, documents, search queries, and memory contents may be sent to configured LLM, embedding, web, or Tavily providers if you run them. Avoid using secrets or private user data in these demos unless you add consent, redaction, provider controls, and per-user memory isolation.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (29)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description says the skill is for building agents with LangGraph/LangChain, but the provided code chunk is a RAG/retrieval example, not an agent implementation. It indexes documents from local and web sources, splits them, embeds them, stores them in an in-memory vector store, and retrieves similar documents for queries. This is a materially different primary purpose from agent construction, and it also accesses resources (local PDF and external web page) not implied by the description. Therefore the description does not accurately represent the code's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared purpose says the skill uses LangGraph/LangChain to build agents, but the provided code does not use LangGraph or LangChain at all. Instead, it directly calls the OpenAI-compatible chat completions API with a fixed prompt and prints detailed response fields, token usage, and raw JSON. This is a materially different primary purpose: an API response structure demo/debug script rather than agent-building with LangGraph/LangChain. No special triggers or permissions are declared, and none are obviously required beyond API access, but the main description-to-behavior mismatch is clear.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared purpose says the skill uses LangGraph/LangChain to build agents, but the supplied code only demonstrates a basic OpenAI chat completion streaming call. It creates an OpenAI client, submits fixed system/user messages, enables streaming, and prints token chunks. There is no LangGraph or LangChain import or usage, no agent construction, no workflow/graph behavior, and no trigger logic. This is a materially different primary purpose rather than a minor implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description says the skill is for building agents with LangGraph/LangChain, but the provided code chunk is a standalone reference example for OpenAI-compatible structured output techniques. It imports OpenAI and Pydantic, defines helper functions for logit-biased yes/no answering and schema-based math responses, and includes a simple test harness. There is no LangGraph or LangChain usage, no agent construction, and no related triggers or permissions. This is a material purpose mismatch rather than a minor implementation detail difference.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The inline comment at L06 states 'ignore env load and llm call', which suggests those actions are absent or intentionally omitted. However, the code immediately invokes the model at L09, contradicting the comment's stated intent and potentially misleading reviewers about externally impactful behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code invokes the Tavily search tool, which sends the tool call arguments to an external network service. While search output is printed afterward, there is no prior user-facing warning, confirmation, comment, or docstring disclosing that user-provided content may be transmitted off-system.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code sends state['topic'] to an external LLM via llm.invoke(messages) without any confirmation prompt, comment, docstring, or user-facing notice that user input is being transmitted off-process. Under the code-file criteria, network/HTTP calls that may transmit user data should have some visible disclosure unless clearly documented as part of the skill's stated purpose, which this file does not do.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This stage transmits state['topic'] and state['outline'] to the model using llm.invoke(messages), but the file contains no user-facing warning, confirmation, or explanatory comment about repeated external transmission of intermediate content. Because the code streams writing content through a remote model, some disclosure is expected for privacy-aware operation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The polishing step sends the full draft text to the LLM via llm.invoke(messages) with no user disclosure in comments, docstrings, or runtime messaging. This is a networked transmission of potentially sensitive generated or user-derived content, and the code provides no visible warning about that behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The comment at L012 minimizes or dismisses environment loading and model-calling behavior, yet the implementation is centered around invoking an LLM in both candidate_node and interviewer_node (L074, L085), and also imports dotenv support (L006). This is an intent/documentation mismatch because the comment suggests these behaviors are not relevant while the code's primary function depends on them.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code retrieves stored user memories and injects them directly into the prompt sent to the LLM, which can expose potentially sensitive user data to the model or downstream providers. In an agent-building skill, this pattern is more dangerous because long-term memory is a core feature and developers may reuse it broadly without adding consent, minimization, or redaction controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code sends the user's query directly to an external LLM service via ChatOpenAI without any disclosure, consent flow, or data-handling guardrails. In agent contexts, users may provide sensitive or proprietary information, so silent transmission to a third-party model provider can cause unintended data exposure or compliance issues.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The workflow sends user persona ground truth, generated conversations, and extracted persona data directly into LLM prompts, which may disclose sensitive profiling data to an external model provider. In a skill about building agents, this is more concerning because developers may reuse the pattern in production and process real user attributes without consent, minimization, or privacy controls.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The comment on L11 says to ignore env loading and LLM calls, suggesting those behaviors are absent or irrelevant. However, the module imports load_dotenv and ChatOpenAI, and the main block invokes a LangGraph workflow multiple times, which is inconsistent with the comment's stated intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The workflow uses an InMemorySaver checkpointer with a fixed thread_id, causing user-supplied conversation content to persist across invocations within the process. In an agent-building skill, this can expose prior prompts or sensitive data to later requests sharing the same runtime/thread context, especially if users are unaware that memory is retained.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code sends prompts to an external LLM provider without any user-facing disclosure, consent mechanism, or warning that data will leave the local environment. In an agent skill context, this pattern is risky because future reuse may substitute real user inputs, secrets, or proprietary code into the request, causing unintended third-party data exposure.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest describes a skill for using LangGraph/LangChain to build agents, but this file is focused on introspecting and printing detailed raw response internals such as metadata, tool call structures, fingerprints, and full JSON dumps. That diagnostic capability is not clearly necessary for the stated purpose and broadens the skill into model-response inspection/debugging.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The comment at L07 says to ignore env loading and LLM calls, yet the code below includes repeated model invocation via llm_with_tools.invoke(...) at L22 and L36. This is an active contradiction between the inline documentation and the actual behavior shown in the file.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The comment at L07 says to ignore environment loading and LLM calls, yet the file's core behavior is building and invoking an agent with an LLM model at L35-L49. This is a documentation-level contradiction because it downplays the primary operational behavior of the script rather than merely omitting detail.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest describes the skill as using LangGraph/LangChain to build agents, which suggests an implementation focused on agent construction patterns. This file goes beyond setup and actually configures a live Tavily web-search tool and streams execution of an agent answering a current-events query, introducing active external data access behavior not stated in the manifest description.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The TavilySearch tool adds external web-search capability. While agents can use tools, the manifest only states a generic purpose of building agents and does not specifically indicate networked search or retrieval, so this capability is broader than the stated purpose on its face.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The comment at L14 tells readers to disregard env loading and LLM invocation, yet the file imports load_dotenv and ChatOpenAI and the core node logic calls llm.invoke(...) at L36. This is an intent/documentation mismatch because the comment minimizes behavior that is actually central to what the code does.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code prints "===Testing memory===" before invoking the compiled workflow with only a new query, implying conversational memory is being tested. However, the graph contains only a single stateless node and no memory/checkpointing mechanism, so this invocation does not actually exercise persisted memory and the label is misleading about the code's behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The stated skill purpose is to use LangGraph/LangChain to build agents, which suggests framework-oriented agent construction examples. This file instead behaves as a specific opinion-synthesis content generation workflow, with hardcoded prompts and LLM invocations, which is a narrower and different end behavior than the manifest description implies.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.