Back to skill

Security audit

Buck Mason Stylist

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly disclosed Buck Mason shopping/lookbook assistant, but it needs Review because it handles account tokens, public voting data, payments, and untrusted lookbook inputs in ways that can affect privacy or local security.

Before installing, decide whether you are comfortable with public-by-link lookbook hosting, unauthenticated voting endpoints, and local persistence of shopping/account state. Prefer the order-code flow over saved JWTs, use --no-voting or add authentication for private feedback, do not run the lookbook builder on untrusted config/picks JSON, and require explicit confirmation before any payment, return, email access, or public deploy.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/build-html-lookbook.py:110
Finding

Arbitrary Remote Resource Fetching Enables Server-Side Request Forgery

Content
View full analysis
0: return tmp = dst.with_suffix(".tmp") subprocess.run(["curl", "-sSL", "-o", str(tmp), src_url], check=True) img = Image.open(tmp).convert("RGB") h = round(w * 4 / 3) ratio = max(w / img.width, h / img.height) new = img.resize((round(img.width * ratio), round(img.height * ratio)), Image.LANCZOS) left = (new.width - w) // 2 top = (new.height - h) // 2 new.crop((left, top, left + w, top + h)).save(dst, "JPEG", quality=85, optimize=True) tmp.unlink() ``` ```python src_url = ((p0.get("try_on") or {}).get("url") or (p0.get("hero") or {}).get("url") or p0.get("image_url")) if not src_url: print(f"warn: no hero image for {look_id}", file=sys.stderr) continue dst = OUT / f"{look_id}.jpg" if not dst.exists(): tmp = dst.with_suffix(".tmp") subprocess.run(["curl", "-sSL", "-o", str(tmp), src_url], check=True) web_jpeg(tmp, dst, max_w=1200) tmp.unlink() ``` ### Technical Analysis The builder accepts image URLs from the caller-controlled picks JSON and passes them directly to `curl`. It does not validate: - The URL scheme - The destination hostname - The resolved IP address - Redirect destinations - Whether the address is loopback, link-local, private, or otherwise reserved - Response size or content type The use of a subprocess argument array prevents ordinary shell metacharacter injection, but it does not prevent SSRF. The `-L` option also causes redirects to be followed without validating the final destination. Because the downloaded response is subsequently processed by Pillow and may be incorporated into a generated deployment, the behavior creates ...[truncated 1432 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/build-html-lookbook.py:123
Finding

Unvalidated Look Identifiers Permit Filesystem Path Traversal

Content
View full analysis
filename relative to OUT for look in CFG["looks"]: look_id = look["id"] if args.no_tryon: pieces = [p for p in PICKS if p["look"] == look_id] if not pieces: continue p0 = pieces[0] src_url = ((p0.get("try_on") or {}).get("url") or (p0.get("hero") or {}).get("url") or p0.get("image_url")) if not src_url: print(f"warn: no hero image for {look_id}", file=sys.stderr) continue dst = OUT / f"{look_id}.jpg" if not dst.exists(): tmp = dst.with_suffix(".tmp") subprocess.run(["curl", "-sSL", "-o", str(tmp), src_url], check=True) web_jpeg(tmp, dst, max_w=1200) tmp.unlink() look_hero[look_id] = dst.name else: src = args.look_images / f"{look_id}.png" if not src.exists(): print(f"error: missing try-on image: {src}", file=sys.stderr) sys.exit(2) dst = OUT / f"{look_id}.jpg" web_jpeg(src, dst, max_w=1200) look_hero[look_id] = dst.name # Per-piece thumbnails for p in PICKS: src_url = ((p.get("try_on") or {}).get("url") or (p.get("hero") or {}).get("url") or p.get("image_url")) if not src_url: continue p["thumb_path"] = f"thumb-{p['id']}.jpg" thumb(src_url, OUT / p["thumb_path"], w=240) ``` ### Technical Analysis The values in `CFG["looks"][].id` and `PICKS[].id` are interpolated into filesystem paths without validation. The code does not reject path separators or traversal components and does not verify that resolved paths remain under the intended output or look-image directories. T ...[truncated 2011 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:218
Finding

Account-Wide Session JWT Is Persisted in a Plaintext Profile File

Content
View full analysis
` (raw, no `Bearer` prefix). Skip to step 2. ``` ```text The customer clicks the link in their inbox, OR the agent reads the email itself and extracts the token, OR the customer pastes the URL/token back into the chat. Then `POST /api/login_via_token` with `{ token: "" }` returns a JWT. Save it to `profile.md → jwt` so the next session starts at path (a). ``` ### Technical Analysis The Skill explicitly directs the agent to save an account-wide bearer credential in `profile.md`. The same documentation describes `profile.md` as a plain-text file kept in persistent agent memory or a workspace. No controls are specified for: - Encryption at rest - Restrictive file permissions - Secret-manager storage - Token expiry checks - Token rotation or revocation - Redaction from model context, logs, backups, or synchronization - Automatic deletion after use A bearer JWT generally grants access based on possession. Any process, Skill, backup service, workspace collaborator, or logging system that obtains the file may be able to replay the token. The repository does recommend the lower-privilege order-code path for ordinary tracking requests. Persisting the account-wide JWT nevertheless exceeds the minimum privilege necessary for those common operations. ### Attack Path 1. The customer requests account-wide order access. 2. The magic-link flow returns a JWT after login. 3. Following `SKILL.md`, the agent writes the JWT into `profile.md`. 4. The plaintext profile remains in persistent workspace storage across sessions. 5. Another local process, Skill, workspace collaborator, backup system, or compromised i ...[truncated 717 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
templates/voting/functions-api-vote.js:16
Finding

Default Voting API Permits Unauthenticated Abuse and Public Disclosure of Participant Metadata

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (142)

Harmful Content Injection

Critical
Category
Prompt Injection
Confidence
95% confidence
Finding

This content may contain harmful instructions that could cause physical harm if followed. CRITICAL: Review carefully before use.

Content

Scanner excerpt · README.md (reported line 1)May include surrounding context.

md
# Buck Mason Stylist Skill

A personal-shopping skill for [Buck Mason](https://www.buckmason.com), built for Claude Code, Codex, ChatGPT custom GPTs, and any agent that loads `SKILL.md`–style skills. Talks to the pima.io MCP at `pima.io/mcp/buckmason/*`.

## What it does

Tainted flow: 'req' from os.environ.get (line 158, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/verify-face.py (reported line 168)May include surrounding context.

python
)

try:
    with urllib.request.urlopen(req, timeout=60) as r:
        resp = json.loads(r.read())
except urllib.error.HTTPError as e:
    body = e.read().decode("utf-8", errors="replace")[:400]

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · PUBLISHING.md (reported line 77)May include surrounding context.

}

text

2. **Update SKILL.md frontmatter** to include the ClawHub-required fields:
```yaml
---
name: buck-mason-stylist

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a broad end-user shopping assistant with discovery, recommendation, stock-checking, persistence, and checkout capabilities. The supplied code chunk instead only builds a deployable HTML lookbook from already-prepared inputs. It consumes config/picks files, validates lookbook/image provenance, downloads and converts imagery, creates thumbnails/Open Graph assets, and emits an HTML page with client-side selection UI and a copyable 'send to my stylist' handoff message. While it is related to the 'AI try-on lookbooks' portion of the description, it does not itself generate try-on images, fetch live inventory/store availability, perform wardrobe analysis, decide outfits, persist customer sizes, or run checkout links. The primary purpose of this code is materially narrower and different from the declared end-to-end shopping skill.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description centers on shopping-assistant behaviors: stock checking, wardrobe analysis, outfit recommendations, AI try-on, checkout, and persistent sizing reuse. The supplied code does none of those. Its primary purpose is infrastructure/deployment for a lookbook website, including Cloudflare Pages publishing and optional voting functionality. While 'AI try-on lookbooks' in the description loosely relates to lookbooks, this script is specifically a hosting/deployment utility rather than a customer-facing shopping or styling function. The Cloudflare deployment, project management, KV vote storage, and public URL publication are undeclared capabilities and materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad interactive personal shopping skill with multiple customer-facing capabilities: online and nearby-store stock checks, wardrobe gap analysis, outfit suggestions, AI try-on lookbooks, and checkout. The supplied code instead implements a narrow backend data-prep task for a weekly newsletter/lookbook workflow. Its primary function is to fetch Buck Mason product data, match items to stored sizes, avoid recently proposed SKUs using a local wishlist/history file, and emit candidate products. That is materially different from the declared primary purpose. While it loosely relates to shopping and lookbooks, several advertised capabilities are absent, and the weekly newsletter candidate-discovery behavior is itself undeclared. Therefore this code chunk does not accurately represent the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a retail personal shopping assistant focused on inventory checks, wardrobe analysis, outfit recommendations, try-on lookbooks, checkout, and remembering sizes. The supplied code does none of those things. Instead, it is a deployment/post-processing utility that parses an existing lookbook HTML file, extracts look/item metadata, injects a thumbs up/down voting interface, and adds client-side code to POST collected feedback to an API. This is a materially different primary purpose and introduces undeclared capabilities around UI modification and feedback collection. No shopping, stock-checking, checkout, or size-memory behavior appears in this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk does not implement or expose the declared shopping-related behavior. It is merely an init.py docstring describing shared helpers and import patterns for script modules. Since the actual code shown has a materially different purpose (library/package scaffolding) and lacks any evidence of the declared end-user capabilities, the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code chunk is a profile parser, not a shopping workflow implementation. It processes markdown text and optionally reads a local file path to extract customer profile fields, including sizes, reference photo file paths, and preferred Link payment method mappings. While reusing customer sizes is loosely related to the declared skill, the chunk does not implement the declared primary capabilities such as stock checks, outfit suggestions, wardrobe gap analysis, AI try-on lookbooks, or checkout link generation. The file-reading wrapper and payment-method parsing are concrete behaviors not reflected in the declared description. Because the actual code's purpose is materially narrower and somewhat different from the declared end-user functionality, this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description overlaps partially with the code: both concern Buck Mason personal styling, event-aware suggestions, AI try-on lookbooks, and reuse of profile sizes/preferences. However, several declared core capabilities are not represented in this code chunk. There is no checkout flow or link-cli usage, no wardrobe gap analysis, and no nearby-store stock lookup. The script mainly orchestrates candidate discovery, curation, build/deploy/validation of lookbooks, plus premium-image gating and face verification. It also adds operational capabilities like deployment, validation, voting KV integration, and wishlist append/logging that are not described. Because the declared description emphasizes broader shopping and checkout functions while the actual code is primarily a headless lookbook pipeline, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents an end-user shopping assistant for Buck Mason with inventory lookup, styling intelligence, AI try-on, and checkout. The supplied code instead is a QA/validation utility for a generated lookbook/deployment. It inspects index.html, checks for Buck Mason product links, prices, stock labels, AI disclosure text, OG metadata, og.jpg, and deployed URL accessibility via curl. This is materially different from the declared primary purpose and includes undeclared behavior related to website/deployment validation and HTTP requests. Although the script is loosely related to lookbooks and Buck Mason branding, it does not implement the shopping assistant capabilities described.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a personal shopping assistant focused on stock checking, outfit recommendations, sizing reuse, AI try-on lookbooks, and checkout. The supplied code does none of those things. Instead, it implements a feedback/voting API for a lookbook, persisting votes and comments in KV and recording requester metadata. This is a materially different primary purpose and includes undeclared data collection/storage behavior, so it is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about a personal shopping assistant for apparel retail features such as stock checking, outfit recommendations, AI try-on, checkout, and size reuse. The supplied code instead implements a lookbook voting results endpoint: it lists vote keys from a KV namespace, loads and parses vote objects, aggregates up/down counts for looks and items, collects voter/comment metadata, and returns both tally and raw votes. This is a materially different purpose and introduces capabilities not described in the skill, especially retrieval of voting data and exposure via an unauthenticated endpoint. Therefore the description does not accurately represent the code's behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a customer-facing retail shopping assistant with inventory, styling, and checkout features. The supplied code chunk does not implement any of those behaviors or interact with shopping, inventory, customer profiles, stores, AI try-on, or checkout. Instead, it is solely test infrastructure that sets up sys.path for importing helper modules during pytest runs. This is a materially different primary purpose from the declared skill behavior, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a consumer-facing shopping/styling skill with inventory lookup, outfit generation, try-on lookbooks, and checkout. The supplied code does none of those things. It is strictly a test suite for parsing profile.md content into structured fields. While some parsed fields (sizes, reference photos, payment preferences) could support a shopping assistant, this code chunk’s primary purpose is unrelated implementation/testing logic and does not implement the declared shopping capabilities. Therefore the description does not accurately represent the behavior of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a retail/personal-shopping assistant focused on stock checking, wardrobe analysis, outfit suggestions, AI try-on, and checkout. The actual code does none of those things. Instead, it tests a deterministic script that classifies calendar events and decides whether they merit actions like 'premium', 'editorial', or 'skip' based on event characteristics. This is a materially different primary purpose. While event-aware outfit suggestions are mentioned in the description, this code is narrowly about event scoring logic and hard-vetoing certain event types, not shopping workflows, stock access, size reuse, or checkout. Therefore the supplied code chunk does not accurately represent the declared skill description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents an end-user shopping assistant with multiple commerce and personalization capabilities. The supplied code chunk does not implement any of those behaviors. Instead, it is strictly an end-to-end test suite for a validate-lookbook script, focused on local deployment/content validation rules for an HTML lookbook page and associated image assets. While the test fixtures mention Buck Mason products and AI-generated try-on previews, that is only fixture content used to test validation gates, not actual shopping, inventory, recommendation, or checkout functionality. This is a material mismatch in primary purpose and capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for a retail personal shopping assistant focused on apparel inventory, outfit recommendations, sizing reuse, and checkout. The supplied code chunk is unrelated: it is a test suite for a script that verifies facial similarity between generated and reference images using mocked OpenAI responses and threshold logic. There is no evidence of shopping, stock lookup, wardrobe analysis, store proximity, checkout, or customer size persistence. This is a clear material mismatch in primary purpose and capabilities.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 45)May include surrounding context.

md
> **What's loaded by agents at runtime, vs repo-only.** Follow refs from this `SKILL.md` only. Files like `README.md`, `PUBLISHING.md`, `SECURITY.md`, and `CLAU

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 281)May include surrounding context.

md
> **What's loaded by agents at runtime, vs repo-only.** Follow refs from this `SKILL.md` only. Files like `README.md`, `PUBLISHING.md`, `SECURITY.md`, and `CLAU

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 294)May include surrounding context.

md
convention. **Hard rule**: never share images/picks/configs across lookbooks. `scripts/build-html-lookbook.py` enforces the marker check.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 299)May include surrounding context.

md
convention. **Hard rule**: never share images/picks/configs across lookbooks. `scripts/build-html-lookbook.py` enforces the marker check.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 307)May include surrounding context.

md
- `templates/voting/functions-api-vote.js`, `functions-api-votes.js`, `wrangler.toml.example` — Cloudflare Pages Functions + binding template for voting. `scrip

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 307)May include surrounding context.

md
- `templates/voting/functions-api-vote.js`, `functions-api-votes.js`, `wrangler.toml.example` — Cloudflare Pages Functions + binding template for voting. `scrip

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file documents returns-management and order-tracking endpoints for a skill described primarily as shopping/styling. That broadens the accessible customer-data and action surface beyond the declared need, enabling access to shipment status, return flows, and other order operations that can expose sensitive customer information or trigger customer-impacting actions if the agent is over-permissioned.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.