Back to skill

Security audit

Code Mentor

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a legitimate programming tutor skill with disclosed optional local code tools, but users should understand that it saves progress locally and its test runner executes project code.

Before installing, be comfortable with the skill writing a local learning_log.md file after sessions. Do not paste secrets or proprietary code into saved progress notes. Use the optional test runner only on projects you trust, ideally in an isolated environment, because running tests can execute local code and dependencies.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
scripts/run_tests.py:187
Finding

Unpinned Dependencies and Implicit Package Retrieval

Content
View full analysis
=2.15.0 # For testing (run_tests.py) pytest>=7.2.0 # For better output formatting colorama>=0.4.6 ``` `README.md:332-337`: ```bash pip install -r requirements.txt ``` ```bash npm install --save-dev jest ``` ### Technical Analysis The Python dependencies use open-ended minimum version constraints rather than exact, reviewed versions. No lockfile or package hashes are present in the audited project. Consequently, installation at different times can resolve to different package versions whose code was not included in this audit. The JavaScript test runner invokes `npx jest`. Depending on the npm environment and local package availability, `npx` may resolve or retrieve the Jest package before executing it. This creates a supply-chain execution channel in which package code and installation lifecycle behavior can run with the privileges of the user invoking the script. The subprocess call uses an argument list and does not enable a shell, so the audited code does not expose a direct shell-command injection vulnerability through `self.target`. The risk instead arises from trusting dynamically resolved third-party packages. ### Attack Path 1. An attacker compromises an allowed package release, its publishing account, or the package-resolution infrastructure. 2. A user follows the documented installation instructions, causing the open-ended Python constraints to resolve to the affected release; alternatively, the user invokes ...[truncated 1213 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (31)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad AI programming tutor whose main function is teaching and interactive guidance across many educational use cases. The supplied code instead implements a local static analysis CLI tool for code review-style checks and basic metrics. While there is partial overlap with 'review their code' and limited debugging/best-practice guidance, the overwhelming majority of declared capabilities are absent: no conversational interface, no lesson generation, no algorithm or design-pattern instruction, no mentoring, no interview prep, and no homework assistance. The code's actual primary purpose is materially different from the declared purpose, so this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad AI programming tutor with many educational and mentoring capabilities across Python and JavaScript. The supplied code instead implements a narrow static analysis tool for Python that estimates algorithmic complexity from AST patterns. While complexity analysis could be a small supporting feature within a programming tutor, this code chunk does not implement the described interactive teaching, debugging, mentoring, or multi-language support. The primary purpose is materially different and significantly narrower than declared, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description emphasizes an educational AI tutor for programming instruction, code explanation, interview prep, and mentoring. The actual code does not provide tutoring, explanation, interactive lessons, code review, algorithm guidance, or mentoring behavior. Instead, its primary purpose is operational: run tests on local targets using external tools and summarize outcomes. While test execution could loosely support debugging workflows, that is only incidental and far narrower than the declared tutoring function. This is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 312)May include surrounding context.

md
- GET /tasks - List all tasks
  - GET /tasks/:id - Get specific task
  - PUT /tasks/:id - Update task
  - DELETE /tasks/:id - Delete task

- Project structure:
  /src

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/languages/javascript-reference.md (reported line 112)May include surrounding context.

md
// Modification
person.city = 'SF';     // Update
person.email = 'a@example.com';  // Add
delete person.age;      // Remove

// Methods
Object.keys(person);    // ['name', 'age', 'city']

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README claims the skill automatically saves persistent user progress after each session, which introduces stateful data retention beyond a simple tutoring interaction. In an educational skill, users may paste code, errors, notes, or other sensitive content, so undocumented or under-specified persistence can create privacy and data handling risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Automatic progress tracking is described without a clear warning or consent flow for persistent storage of session-derived data. Because users are encouraged to submit code and problem details, the saved log may contain sensitive proprietary or personal information, making silent retention a meaningful privacy risk.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Cross-session logging of learning progress in plain language can capture user-provided code snippets, debugging context, and other freeform content that may include secrets or sensitive business logic. In the context of a programming tutor, this is more dangerous because users commonly paste real source code and error output into the tool.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The debugging section states 'I will NEVER directly point to the bug or give you the answer,' and later says it will resist saying where the bug is. But other sections explicitly promise a 'Refactored Version' after discussion and a 'Final' full solution if the user is stuck, which contradicts the absolute wording of the earlier guidance.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to persistently write user learning history to a local file after each session. Storing session topics, goals, and performance without explicit opt-in creates a privacy risk, can retain sensitive user data longer than expected, and turns a tutoring interaction into silent stateful data collection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill mandates persistent logging of learning progress to disk without warning the user that their interaction data will be stored. This undermines informed consent and can expose personal study history, code-related notes, or other sensitive content to anyone with filesystem access or later tooling that reads the log.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · references/algorithms/common-patterns.md (reported line 112)May include surrounding context.

md
function lengthOfLongestSubstring(s) {
    const seen = new Set();
    let left = 0;
    let maxLength = 0;

    for (let right = 0; right < s.length; right++) {
        // Shrink window until no duplicates

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/languages/javascript-reference.md (reported line 399)May include surrounding context.

md
.finally(() => console.log('Done'));

// Promise chaining
fetch('https://api.example.com/data')
    .then(response => response.json())
    .then(data => console.log(data))
    .catch(error => console.error(error));

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/languages/javascript-reference.md (reported line 409)May include surrounding context.

md
.finally(() => console.log('Done'));

// Promise chaining
fetch('https://api.example.com/data')
    .then(response => response.json())
    .then(data => console.log(data))
    .catch(error => console.error(error));

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring claims support for Python, JavaScript, Java, and C++, but the actual analyzer only instantiates analyzers for Python and JavaScript. For Java, C++, and C files, the code exits with an unsupported-language error instead of performing analysis, so the documented capability does not match real behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill's test runner materially expands capability from tutoring and code guidance into executing arbitrary local code across Python and JavaScript frameworks. That broader execution surface is dangerous in an agent setting because users may reasonably expect analysis/help, while the implementation can actually run untrusted repository code with side effects.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
93% confidence
Finding

This subprocess invocation runs pytest against a user-supplied target, which causes arbitrary test code in the repository to execute with the agent's privileges. Although shell injection is mitigated by passing an argument list, the core risk remains arbitrary code execution because Python test files and pytest hooks/conftest modules can run attacker-controlled code during collection and execution.

Content

Scanner excerpt · scripts/run_tests.py (reported line 74)May include surrounding context.

python
'--tb=short'
            ]

            process = subprocess.run(
                cmd,
                capture_output=True,
                text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code executes pytest without informing the user that running tests will execute arbitrary code from the target project, including setup hooks and imports. The missing warning increases the chance of unsafe execution in a context where the skill is presented primarily as an educational mentor rather than a code-execution tool.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
92% confidence
Finding

This subprocess call launches unittest discovery on a user-controlled directory target, which can import and execute arbitrary Python modules during test discovery and runtime. In an agent skill context, that means a crafted workspace can trigger code execution beyond the tutoring purpose and potentially access local files, tokens, or network resources available to the agent.

Content

Scanner excerpt · scripts/run_tests.py (reported line 138)May include surrounding context.

python
'-v'
            ]

            process = subprocess.run(
                cmd,
                capture_output=True,
                text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Launching Jest through npx without a warning is especially risky because it may execute untrusted project dependencies and test code, not just a trusted system binary. In a mentoring skill, users may not anticipate that a simple test command can trigger arbitrary JavaScript execution with access to the agent's environment.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
96% confidence
Finding

This subprocess call invokes Jest through npx on a target chosen from the local project, which can execute arbitrary JavaScript test code as well as project-resolved tooling and dependencies. Using npx increases risk because it may rely on project-defined packages and scripts, enabling untrusted repository content to run code in the agent environment.

Content

Scanner excerpt · scripts/run_tests.py (reported line 189)May include surrounding context.

python
try:
            cmd = ['npx', 'jest', self.target, '--verbose']

            process = subprocess.run(
                cmd,
                capture_output=True,
                text=True,

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The debugging section states the mentor never gives direct answers, presenting that as a core behavior. Later, the progressive hints section includes a final 'Solution: Full solution with explanation,' which contradicts the absolute claim and creates intent-documentation divergence.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This markdown file includes example fetch(...) calls to external endpoints, which demonstrate transmitting requests over the network. The surrounding documentation does not include any warning that these examples contact external services or may send user/system data, which is the type of disclosure SQP-2 asks for in markdown files.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

This markdown file includes examples for opening files in write and append modes and writing JSON data, which can modify or overwrite user files. The section does not include any warning or note about these effects, so readers are not alerted to the data-impacting behavior.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/analyze_code.py:197