Back to skill

Security audit

Low-Visibility Fix

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it claims, but it needs review because its page scoping and generated handoff report can exceed the intended target boundary.

Install only if you are comfortable reviewing its generated report before handing it to another agent. Use tightly controlled target paths and page names, avoid running it on untrusted projects with arbitrary --pages or --selector values, and treat report.md as untrusted rendered evidence rather than executable instructions.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/scope.py:35
Finding

Explicit Page Resolution Allows Filesystem Path Traversal

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/emit_docs.py:145
Finding

Untrusted Audit Metadata Can Inject Content into Agent-Facing Markdown

Content
View full analysis
未找到的页面: {', '.join(sc['missing'])}", ""] lines.append("## Findings(按严重度)") if not doc_set["findings"]: lines.append("无 — 该范围内未发现低能见度问题。") else: for sev in ("critical", "major", "minor"): fs = [f for f in doc_set["findings"] if f["severity"] == sev] if not fs: continue lines.append(f"### {sev} ({len(fs)})") for f in fs: m = (f" — measured {f['measured']} / need {f['threshold']}" if "measured" in f and "threshold" in f else "") lines.append(f"- `{f['rule']}` @ {f['file']} {f['location']}{m}") lines += ["", "## 修改建议(交给实现 agent,本技能不直接改文件)"] if not doc_set["recommendations"]: lines.append("无。") else: rmap = {r["finding_id"]: r for r in doc_set["recommendations"]} for f in doc_set["findings"]: r = rmap.get(f["id"]) if r: lines.append(f"- [{f['id']}] {r['recommendation']} ({r.get('snippet_ref', '')})") if doc_set["needs_judgment"]: lines += ["", "## 待判定(需视觉/浏览器二次确认)"] for n in doc_set["needs_judgment"]: ...[truncated 2556 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code is not the described low-visibility auditing skill itself; it is an evaluation/check script for that analyzer. Its primary purpose is regression testing: enumerate fixture files, invoke scripts/analyze.py, parse stdout JSON, compare to expected golden files, and exit nonzero on drift. While this is related to the broader project, it does not implement the user-facing behavior described, such as auditing specific UI pages/components, performing a bounded visual/browser pass, or generating structured fix-plan documents. No undeclared sensitive permissions appear, but the code’s actual purpose is materially different from the declared skill behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose centers on auditing existing mobile UIs for low-visibility field usability issues and producing implementer-ready handoff documents. The actual code contains only generic JSON Schema validation logic: type checks, required/properties handling, array/item validation, enum and numeric bounds, local $ref resolution, and a helper to load JSON files and validate them. None of this implements UI inspection, browser/visual passes, page/component scoping, severity ranking, or fix-plan document generation. While schema validation could be a supporting utility inside a larger skill, the supplied chunk by itself reflects a materially different primary purpose and unrelated capability set.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a full audit skill that inspects an existing UI under low-visibility conditions and produces structured remediation documents. The supplied code instead implements a narrow safety policy: it checks whether a proposed contrast ratio for a critical control falls below a configured field threshold and returns allow/refuse guidance. While this relates to low-visibility contrast, it is only one small policy component and does not match the described primary purpose or capabilities of a scoped UI auditing/document-generation skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The declared behavior in the file broadens the skill from analysis/document generation into direct remediation, contradicting the higher-level description that it never edits the target. This scope expansion is dangerous because orchestration layers, users, or other agents may rely on the YAML metadata when deciding permissions and execution paths, leading to unauthorized or unsafe file modifications.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The UI-facing metadata promises the skill will 'audit & fix' and the default prompt says 'audit and fix this UI', which expands user expectations beyond the stated handoff-only behavior. In an agent ecosystem, this kind of mismatch can cause downstream systems or users to authorize unintended modification workflows, undermining safety boundaries and making the skill more likely to be invoked for direct changes.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/fix-snippets.html (reported line 7)May include surrounding context.

html
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Low-visibility fix snippets (reference, not a page)</title>
<!--
  Compliant component patterns to graft into a UI during Step 5.
  All values trace to references/design-tokens.json (field tier).
  These are copy-paste starting points, not a finished screen.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/handoff-doc.template.md (reported line 3)May include surrounding context.

md
# 低能见度审计交接文档 — <target>

<!--
  Shape of the report.md that scripts/audit.py emits beside audit.json.
  Rendered automatically by scripts/emit_docs.render_markdown — this template
  documents the contract for an implementer agent reading the output.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/analyze.py (reported line 268)May include surrounding context.

python
rules.append((spec, simple, decls))
    for k in list(varmap):
        varmap[k] = resolve_value(varmap[k], varmap)
    return rules, varmap


# --- HTML tree ---------------------------------------------------------------

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
81% confidence
Finding

The skill advertises capabilities that involve reading files, writing output, invoking Python scripts, and potentially rendering pages, but it does not declare an explicit tool/permission boundary. In agent environments, missing scope declarations can allow broader-than-necessary access and make it harder to enforce least privilege, especially since the skill text references shell execution, file writes, and network-adjacent rendering behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Enabling implicit invocation without narrow trigger conditions or exclusions increases the chance that this skill is auto-selected in contexts where it should not run. Because the skill metadata already overstates its authority to 'fix' UIs, broad implicit invocation makes that mismatch more dangerous by increasing accidental activation and scope confusion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This file sets lang="zh" and all visible button/status labels are Chinese-only (确认, 正常, 报警, 停止设备). The policy for natural-language violations applies to all file types, and there is no indication here that the locale is optional, user-selectable, or constrained to a documented region-specific use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The template headings and field labels are written entirely in Chinese, which implies the generated handoff document will be produced in a fixed language. The file does not mention any user opt-in, language selection mechanism, or region-specific justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · evals/check.py (reported line 23)May include surrounding context.

python
def run_analyzer(html_path):
    proc = subprocess.run([sys.executable, ANALYZE, html_path],
                          capture_output=True, text=True)
    if proc.returncode not in (0, 1):  # 0 clean, 1 findings present
        raise RuntimeError(f"analyzer crashed on {html_path}:\n{proc.stderr}")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The evaluation cases define activation largely through broad natural-language requests, which can cause the skill to trigger outside its intended scope or in adjacent workflows. In this skill, over-broad activation is risky because it can make the agent perform audits when the user actually wants direct modification, new design, or unrelated accessibility work, increasing the chance of incorrect tool use and policy boundary violations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Natural-language policy review applies to all file types. Every user task field is named 'task_zh' and contains Chinese-only prompts, with no indication that the skill supports other languages or that Chinese is a justified region-specific requirement, which can amount to a forced language/locale constraint without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document declares a Chinese language context and all visible user-facing text is presented only in Chinese, with no indication that the user can select another language or that the skill is region-specific. This creates a natural-language policy concern because the interface enforces a specific locale without opt-in.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · evals/run_all.py (reported line 65)May include surrounding context.

python
def run_audit(args):
    p = subprocess.run([sys.executable, AUDIT, *args], capture_output=True, text=True)
    return p.returncode, p.stdout, p.stderr

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · evals/run_all.py (reported line 129)May include surrounding context.

python
fails = []
    for g in sorted(glob.glob(os.path.join(FIX, "*.expected.json"))):
        html = g[: -len(".expected.json")] + ".html"
        p = subprocess.run([sys.executable, os.path.join(SCRIPTS, "analyze.py"), html],
                           capture_output=True, text=True)
        if p.returncode not in (0, 1):
            fails.append(f"{os.path.basename(html)} crashed")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The render_markdown function emits all user-facing report headings and status text in Chinese, regardless of user preference or runtime configuration. This is a natural-language locale policy concern because the file hard-codes a specific language with no opt-in, selection mechanism, or justification that the skill is region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The HTML declares lang="zh" and all visible UI text is in Chinese, which imposes a specific language/locale without any opt-in or explanation. This matches the policy category for language or locale constraints that are not documented as optional or region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document declares lang="zh" and the visible content is entirely in Chinese, which imposes a specific language/locale on users. The file does not provide any opt-in, alternative language option, or justification that this is a region-specific artifact, so it matches the language/locale policy violation category.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The file sets lang="zh", which imposes a specific language/locale in the document. Under the policy, locale constraints should either offer user choice or be clearly documented as justified for a region-specific purpose, neither of which appears in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document sets lang="zh", which explicitly fixes the content locale to Chinese. Under the policy rule, forcing a specific language without user opt-in or documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The HTML declares lang="zh" and the visible warning text is only in Chinese, which imposes a specific language on users without offering an opt-in or alternative. This matches a natural-language policy concern because the file presents a fixed locale with no documented justification for being region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document sets lang="zh", which imposes a specific language/locale on users. Under the policy rule, locale constraints should either be optional for the user or clearly justified as region-specific; no such opt-in or justification appears in this file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.