Back to skill

Security audit

x-list-digest

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed X-list collection and reporting workflow with local archives, optional Feishu publishing, and opt-in scheduling controls.

Install only if you are comfortable letting the agent use an already logged-in browser to read authorized X lists and write local archives. Configure your own list URLs, data directory, timezone, and Feishu profile/root folder before use; enable Feishu publishing or scheduling only after reviewing those destinations and permissions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (26)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a broader workflow skill: collecting authorized X content over time, filtering archived posts, classifying viewpoints, producing reports/statistics, and optionally archiving. The supplied code chunk does not implement those user-facing capabilities. Instead, it validates that an externally supplied analysis exactly matches raw tweet IDs, enforces schema/value constraints, normalizes ticker symbols, checks thread linkage rules, and writes merged output. Although an aggregate() helper exists for counts by category/author/ticker, it is not used in the main execution path and does not itself generate the described daily/weekly reports. Therefore the code's actual purpose is a narrower internal data-validation/enrichment step, materially different from the declared end-to-end skill description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about collecting and analyzing X/Twitter content and producing reports, with optional Feishu archiving and no default scheduling. The supplied code does none of that. Instead, it checks whether a local Chrome debug endpoint is available, inspects running processes and TCP listeners, validates that the browser uses a specific user-data directory, and may launch the Hermes-authorized browser profile. This is a materially different primary purpose and involves undeclared local process, filesystem, and browser-control capabilities. While such functionality could theoretically support later web collection, this chunk itself is not merely a minor implementation detail of the declared behavior; it is an unrelated browser setup/verification component not represented in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a higher-level workflow focused on user-authorized list collection over a fixed window plus analysis/reporting and optional archiving. This code chunk instead performs low-level scraping of whatever tweets are currently visible on an X page and reports browser/page state. While tweet collection is related, the implementation shown is materially broader/different than 'authorized X lists' and includes undeclared environmental/automation data (captcha/login/rate-limit, viewport, scroll position). It also omits the key declared behaviors of time-window filtering, archived-post skipping, viewpoint classification, reporting, Feishu archiving, and scheduling controls. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a higher-level analytics and optional archiving workflow around X list collection and reporting. The supplied code chunk instead only normalizes existing records: it reads input, filters by time window, validates tweet IDs, deduplicates by ID, sorts, and writes output. While the time-window aspect loosely aligns with part of the description, the primary purpose and major claimed capabilities are absent. This is therefore a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description centers on collecting authorized X lists, classifying viewpoints, producing reports, and optionally archiving via Feishu. This code does not perform collection, filtering, classification, or report generation. Instead, it implements a post-publication verification/finalization step for a Feishu daily document: validating a Feishu doc URL, comparing fetched remote content to a local report, enforcing Sunday weekly-statistics presence, checking source tweet IDs in the published document, consulting a local SQLite digest database, and updating database/receipt records with the Feishu URL. Those are materially different behaviors and include undeclared state/database modification capabilities tied to publication tracking.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a production-style workflow for collecting and analyzing authorized X list content on user request, with optional archiving and consent-based scheduling controls. The actual code chunk is materially different: it is a standalone offline demo that imports test components, constructs synthetic fixture data, enriches and renders a report, writes local files, and persists to a local database. The docstring directly says it never fetches X or publishes messages. While there is some overlap in report-generation themes, the core capability and resource access differ significantly from the declared purpose, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description is for a content-ingestion/reporting skill operating on X lists with optional Feishu archiving and consent-based scheduling. The supplied code instead implements a 'doctor' utility that reports structural readiness of the local installation. Its primary behavior is to inspect configuration and filesystem paths, detect dependencies, locate the lark-cli executable, and emit command arguments for external native checks. While Feishu and browser/Hermes are tangentially related to the broader skill, this code chunk does not perform the declared core functions at all. That is a material description-versus-behavior mismatch in primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The supplied code chunk is a collection engine, not an analysis/reporting/archive workflow. Its primary behavior is to drive a native browser session against an X list URL, collect snapshots, optionally open detail tabs for expandable items, enforce collection budgets, and support authorized resumption of interrupted runs. That is consistent with only one fragment of the declaration: collecting list content within a window and excluding some prior history. However, the declared description emphasizes additional higher-level functions—classifying market viewpoints, producing daily and weekly reports, optional Feishu archiving, and consent-gated scheduling—that are absent from this code. Conversely, the code has substantial undeclared operational capabilities around browser session control, tab navigation, and resumable budgeted collection. Because the actual behavior materially differs from the declared end-to-end purpose and includes undeclared collection-control capabilities while omitting most declared outputs, this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description focuses on X/Twitter list collection, market-view classification, report/statistics generation, and optional Feishu archiving under explicit user request. This code does not implement any X collection, content classification, time-window handling, archived-post skipping, daily/weekly reporting, or scheduling behavior. Instead, it performs a distinct Feishu publication finalization workflow, including URL validation, remote content verification, integrity checking via hashes, database updates, and receipt/state persistence. While optional Feishu archiving is mentioned in the description, this chunk's primary purpose is a narrow publication verification/recording step that is not accurately represented by the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
87% confidence
Finding

The code broadly aligns with collecting/processing X-list data over a fixed window and producing a report, but several declared specifics are not represented in this chunk. The primary behavior here is orchestration of digest runs, local persistence, and markdown report generation. The description promises skip-archived handling, market-viewpoint classification, weekly off-topic statistics, and optional Feishu CLI archiving; none of those are concretely implemented in this code chunk, aside from a generic 'enrich' call and an 'feishu' output enum. There is also explicit persistence into local files and a SQLite-style digest database via persist(), which is a meaningful behavior not mentioned in the declared description. No scheduling behavior is present, so there is no mismatch on that point. Overall, the description overstates some capabilities and omits local persistence behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about collecting and analyzing X/Twitter list content and producing reports with optional archival. The actual code does nothing related to X lists, post filtering, market viewpoint classification, reporting, Feishu archiving, or scheduling consent. Instead, it tests infrastructure for verifying a local browser process and its debugging port using process inspection and socket ownership checks. This is a materially different primary purpose and capability set, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The provided code chunk does not implement the declared reporting workflow. Instead, it is a test for a diagnostic script's unconfigured output structure. While test code may be ancillary to a larger project, this chunk's actual behavior is materially different from the declared purpose: it validates diagnostic JSON fields and browser/login access status, not data collection, classification, report generation, archiving, or scheduling consent behavior. Therefore this code chunk is not accurately represented by the description.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

md
and other hosts must use only their supplied browser tools and documentation. `scripts/collect.js` is a read-only DOM snapshot helper, usable only if the host

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 20)May include surrounding context.

Use a trusted Agent Skills installer or place this directory in your host's skill discovery path. Review files before enabling execution. See host adapters.

After authorizing Python dependencies, create a package-local virtualenv:

sh
python3 -m venv .venv

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/category-taxonomy.md (reported line 7)May include surrounding context.

md
Choose the core thesis, not isolated keywords. Keep cross-market details in tickers/subcategory. Unknown tickers are empty arrays; uppercase words and project names do not automatically imply tickers. Sentiment is bullish/bearish/neutral/mixed. Actionable requires explicit entry/exit conditions. Confidence high/med/low measures extraction reliability, not expected return.

Only off_topic counts toward weekly author off-topic rate; other is not automatically off-topic. Fewer than three archive days means insufficient samples. Never automatically remove list members. Consensus requires cited posts from two independent authors, and each author has one sentiment vote per ticker.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring says the script does not inspect credentials, yet the code loads config, reads Feishu profile and root folder token settings, and builds native CLI authentication/status commands. Even if it does not execute them here, this is still credential/auth-related inspection logic that contradicts the stated documentation.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The docstring promises 'never use the default identity', but the auth_status_argv assembled at L30 contains no '--as user' flag, unlike the folder read command at L32. The code relies on profile selection and environment unsetting rather than explicitly enforcing identity in all generated commands, which contradicts the stronger documentation claim.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The timestamp formatter defaults to Asia/Bangkok when no timezone is provided. This imposes a specific locale choice in generated output without any visible user opt-in or explanation that the skill is region-specific, which matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The config sets timezone: Asia/Bangkok as the default locale-related behavior, with no accompanying note that users may change it or that the setting is region-specific by design. This can violate language/locale policy expectations when a skill forces a locale without user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The comment says to add an "authorized list URL" before running, but it does not define what URLs are valid, what list source is expected, or any constraints or exclusions. In a manifest/config file, this lack of specificity can lead to ambiguous setup and unintended behavior when the skill is configured.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The config sets timezone: Asia/Bangkok, which imposes a locale-specific default in a natural-language-adjacent configuration value. The provided file does not indicate that users can choose or override this locale, nor that the region-specific setting is required for a documented purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The skill manifest describes producing daily writing reports and weekly off-topic statistics, and this file even defines an aggregate() helper for summary statistics, but the main execution path only writes the enriched per-post dataset. This is a semantic mismatch between the claimed reporting behavior and what this script actually does when invoked.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This script writes collected identifiers and source run paths to history-excluded.json, and elsewhere persists snapshots, merged raw records, event logs, and metadata. The file contains no confirmation prompt, print/log disclosure during execution, or inline warning comments describing that user data is being written to disk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The write() function creates parent directories and writes JSON data to the supplied path, which modifies the filesystem. There is no confirmation prompt, logging, docstring, or comment disclosing this behavior in the code shown.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script writes the bundled report to the user-supplied output path and then writes a JSON manifest alongside it, but there is no confirmation prompt, warning message, or comment/docstring disclosing that existing files may be overwritten. Because these are filesystem modifications and the output path is externally provided, the behavior lacks visible user disclosure in the code.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_browser_preflight.py:8