Back to skill

Security audit

边看边记-视频要点自动归纳

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent subtitle-summary purpose, but it asks agents to use authenticated browser sessions in overbroad ways, including token inspection, API interception, and saving broad course data locally.

Review this skill before installing. Use it only in a contained browser/session and expect it may capture more than subtitles from authenticated learning pages. Avoid using it with accounts or course pages containing sensitive personal, business, or restricted material unless the token handling and raw JSON persistence are removed or tightly minimized.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:16
Finding

Skill Instructions Attempt to Override Security Controls

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/zhixueyun_extractor.py:134
Finding

Raw Bearer Token Is Read Despite Being Unnecessary for Response Interception

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/zhixueyun_extractor.py:153
Finding

Authenticated Page and API Data Are Overcollected and Persisted in Plaintext

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
scripts/zhixueyun_extractor.py:12
Finding

Playwright and Browser Components Are Installed Without Version or Integrity Pinning

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

This mismatch is security-relevant because the skill claims to limit itself to exposed subtitles/transcripts, yet the described behavior expands to intercepting backend APIs, extracting full DOM text, reading localStorage token information, and writing raw JSON. That broadens collection beyond user-visible transcript data and can expose authenticated content or token-derived session data contrary to user expectations.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

This mismatch is security-relevant because the skill claims to limit itself to exposed subtitles/transcripts, yet the described behavior expands to intercepting backend APIs, extracting full DOM text, reading localStorage token information, and writing raw JSON. That broadens collection beyond user-visible transcript data and can expose authenticated content or token-derived session data contrary to user expectations.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 30)May include surrounding context.

md
在线 URL 的站点族规则与具体 API 策略见 `references/online-subtitle-strategy.md`,仅在输入为 URL 或浏览器页面时加载。

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Guidance to intercept platform responses and read bearer tokens from localStorage is highly sensitive because tokens are authentication artifacts that can enable replay, broader data access, or account compromise if mishandled. In a skill whose stated purpose is subtitle summarization, accessing token material is unnecessary and materially increases the risk of unauthorized access and sensitive-data exposure.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The combined workflow collects login-token material, intercepts authenticated responses, and persists captured course/session data locally. In a subtitle summarization skill, this is more dangerous because the requested capability is narrow, but the implementation grants broad access to private account-scoped information and creates durable copies of that data.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instructions normalize extracting authentication tokens from localStorage and using them in Authorization headers, without any warning about credential sensitivity. This is dangerous because tokens are bearer secrets: disclosure or misuse can directly grant access to the user's authenticated platform account and related data.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/extract_subtitle.py (reported line 366)May include surrounding context.

python
output_dir.mkdir(parents=True, exist_ok=True)

    now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
    # Never persist the raw source URL to disk — it may carry access tokens,
    # session IDs, or other credentials in query parameters. The sanitized
    # form is used for metadata, report file names, and rendered titles.
    display_source = sanitize_url(source) if source_is_url else source

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares broad capabilities involving network access and file creation but does not explicitly scope or constrain tool permissions. In practice, this makes review and enforcement harder and increases the chance the skill can access external resources or write artifacts in ways users and policy layers did not clearly approve.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description says to use the skill whenever a user asks to open, analyze, or summarize video links, subtitle files, browser-accessible course pages, or generate study notes from video content. This trigger scope is broad and lacks negative examples or tighter invocation constraints, increasing the risk of unintended activation for generic video-analysis requests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description requires output as a Markdown knowledge report but the title and surrounding description are written as a Chinese-only skill experience, and no user language choice is offered anywhere in the file. This can violate language/locale policy because it implicitly fixes the interaction/output language without user opt-in.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The instruction that all learning platforms should be treated as non-blockable and that authenticated subtitle/API data may be intercepted after login undermines normal trust boundaries around authenticated sessions. Even with user login, automating collection from authenticated APIs can capture content beyond what is minimally necessary and bypass internal platform restrictions or user expectations about session use.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Directing the workflow to capture data from authenticated sessions via API interception and browser-exposed transcript sources expands the attack surface from visible content extraction to session-bound backend harvesting. In the context of a learning platform, this can expose protected educational materials or account-scoped data, especially if the implementation stores raw responses.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Using platform-provided AI knowledge-point summaries as an input source exceeds the stated subtitle/transcript-only boundary and can silently ingest derived or non-user-visible course data. This increases data-scope creep and may expose content the user did not intend to extract under the narrower subtitle-only promise.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document promotes intercepting backend API responses during page load as the core strategy, rather than limiting itself to subtitles/transcripts visibly exposed to the user. This creates a capability to harvest non-visible platform data and bypasses the stated scope boundary of simple subtitle extraction.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Polling localStorage for login tokens accesses credential-like session material unrelated to summarizing subtitles. In the context of a browser automation skill, this creates a direct path to session abuse, token leakage, and unauthorized expansion of access to private course data.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The guide explicitly instructs reading an authentication token from localStorage and reusing it for direct API calls. That expands the skill from extracting browser-visible subtitles into credential-backed API access, which can expose account-scoped data and enable access patterns beyond what the user intended to reveal in the page UI.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The guide instructs writing page text and intercepted API payloads to a local JSON file without minimization or a warning about sensitive educational or account-linked data. Local persistence increases the chance of unintended retention, later reuse, or exposure of private learning records and metadata.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The template hard-codes the report title and section headings in Chinese, which imposes a specific language on generated output. The file does not indicate that this is optional, user-selected, or justified as a region-specific skill, so it conflicts with the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file mandates Chinese phrasing for the estimate message templates, including the fallback message when duration is unknown. This imposes a specific language on users without opt-in, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Polling localStorage for a token to detect login status is not justified by the stated skill purpose and unnecessarily accesses credential-related data. In a browser automation tool operating after manual login, this creates needless handling of session artifacts that could be repurposed or accidentally exposed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script inspects authentication-related localStorage content without clearly warning the user that browser credential storage will be read. In a tool whose declared purpose is subtitle extraction, this undermines transparency and increases the sensitivity of the operation beyond what a user would reasonably expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The extractor collects and persists much more than subtitle/transcript content: full page body text, video metadata, and multiple intercepted API payloads. In the context of a login-gated learning platform, this broad collection violates data minimization and can expose protected course content or unrelated personal/course data if the output files are accessed or reused.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code reads login token data from localStorage even though token access is unnecessary for subtitle extraction. Even though it only persists derived fields, inspecting credential-related browser storage expands the script's access to authentication context and creates avoidable exposure around user session data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script writes full extracted data, including page text and intercepted API responses from a login-gated course platform, to local disk without an upfront warning or explicit user confirmation. This creates a confidentiality risk because sensitive educational content or user-associated data may be retained longer than expected and exposed through local file access, backups, or sharing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

A language-specific skill description can create a locale policy issue when it effectively forces one language without user opt-in or justification. Here, all instructional content is presented only in Chinese, and the file does not state that the skill is intentionally limited to Chinese-speaking users or provide an alternative language path.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.