Back to skill

Security audit

architect-review

Security checks for vulnerabilities and agentic risk

Overview

This architecture-review skill is purpose-aligned and disclosed, with only local project reading and a confirmed local report write as notable behaviors.

Install only if you are comfortable with the agent reading your project structure and relevant local files for an architecture review. Expect a local `.arch-review/` report file to be created only after confirmation, and review the generated report before sharing it because it may contain internal architecture details and file paths.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · examples.md (reported line 97)May include surrounding context.

md
## Security — 9/10 ✅

### Strengths
- JWT-based auth with short-lived access tokens + refresh token rotation
- All API inputs validated with zod schemas at controller boundary
- Secrets loaded from environment via vault integration, not in code
- Rate limiting on auth endpoints (5 attempts / minute)

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Lines L118-L120 explicitly instruct that the main agent performs the review directly and 'Do NOT delegate to subagents via Task tool.' But lines L128-L130 immediately define an optional parallel mode that does use the Task tool and subagents. These instructions actively contradict each other rather than merely omitting detail.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The AskQuestion title, prompt, and follow-up text are hard-coded in Chinese, which imposes a specific language on the user regardless of their preferences. The file does not indicate that language should match the user's locale or be selectable.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs writing a report into the project tree, but the skill description does not prominently warn users that invoking the skill may modify the repository by creating .arch-review/ and saving files. Hidden or insufficiently disclosed write behavior can surprise users, pollute repos, and be abused to persist artifacts in sensitive workspaces.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language trigger phrase is broad enough to activate on ordinary conversation such as 'help me do an architecture review,' which can cause the agent to enter a tool-using workflow unexpectedly. In this skill, unexpected activation is more concerning because the workflow includes repository scanning, interactive triage, and eventual file writing, so accidental invocation can lead to unintended project actions.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · dimensions.md (reported line 67)May include surrounding context.

md
- Prop drilling through 5+ layers without intermediate state management
- Shadow state that drifts from the canonical source
- Fire-and-forget mutations with no confirmation of success
- Implicit dependencies on execution order

---

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The manifest description presents the skill as an architecture review that yields structured scoring and recommendations, and the core principles state it 'produces a report.' However, the file also defines a multi-step interactive triage process using AskQuestion prompts before finalizing output, which is broader behavior than a straightforward report-generation flow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The text says the skill can be customized to control report 'language,' but it does not state that language selection should follow the user's preference or require opt-in. This can create a natural-language policy issue if the skill ends up forcing a language or locale implicitly.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.