Back to skill

Security audit

Roundtable Forge

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed roundtable-discussion workflow that creates local discussion artifacts and uses optional subagents in a purpose-aligned way.

Install this if you want structured multi-agent roundtable discussions and are comfortable with local Memory JSON and rendered Markdown artifacts being created for continuation. Avoid placing sensitive private data in roundtable prompts or Memory files, and expect Chinese-default wording unless you adapt the skill or explicitly request another language.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (51)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The documented purpose says this is a roundtable-discussion skill, but the instructions also direct local file operations, validation, rendering, and script execution. This mismatch is dangerous because operators and routing systems may grant trust or invoke the skill under assumptions that do not match its actual behavior, hiding data-access and execution side effects behind an innocuous description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The documented purpose says this is a roundtable-discussion skill, but the instructions also direct local file operations, validation, rendering, and script execution. This mismatch is dangerous because operators and routing systems may grant trust or invoke the skill under assumptions that do not match its actual behavior, hiding data-access and execution side effects behind an innocuous description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The documented purpose says this is a roundtable-discussion skill, but the instructions also direct local file operations, validation, rendering, and script execution. This mismatch is dangerous because operators and routing systems may grant trust or invoke the skill under assumptions that do not match its actual behavior, hiding data-access and execution side effects behind an innocuous description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The documented purpose says this is a roundtable-discussion skill, but the instructions also direct local file operations, validation, rendering, and script execution. This mismatch is dangerous because operators and routing systems may grant trust or invoke the skill under assumptions that do not match its actual behavior, hiding data-access and execution side effects behind an innocuous description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The documented purpose says this is a roundtable-discussion skill, but the instructions also direct local file operations, validation, rendering, and script execution. This mismatch is dangerous because operators and routing systems may grant trust or invoke the skill under assumptions that do not match its actual behavior, hiding data-access and execution side effects behind an innocuous description.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill instructs the host to read and write files and invoke shell-based linter/renderer scripts, but it declares no explicit tool scope or permission boundaries. That creates a capability/contract gap: a caller may treat the skill as harmless orchestration while it actually requires filesystem and subprocess access, increasing the chance of over-broad execution in a privileged environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
71% confidence
Finding

The manifest description presents invocation examples primarily in Chinese and frames the skill's activation cues around Chinese phrases, without stating that users may choose their preferred language. This can be read as a locale/language default baked into the skill rather than an opt-in or explicitly multilingual behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger cues include broad phrases such as requests for deeper analysis or other perspectives that commonly occur in ordinary conversations. Over-broad routing can unintentionally activate a powerful skill that performs file and shell-related operations, creating unnecessary exposure and surprising side effects in contexts where the user did not request this capability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger phrases on lines 8 to 10 are ambiguous and lack clear activation boundaries, so common requests like asking for multiple perspectives or a roundtable-style discussion may unintentionally route into this skill. Because the skill is designed to create multiple character lanes, shared memory, and possibly subagent execution, misrouting has higher security and reliability impact than a simple formatting skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrase at line 13 is broad enough to match ordinary user requests for discussion or analysis, which can cause the skill to activate when the user did not explicitly intend to invoke this agent behavior. In a multi-agent orchestration skill, accidental activation can change system behavior, introduce unnecessary tool/subagent use, and increase the chance of prompt-scope confusion or policy bypass through unintended role simulation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This JSON file contains user-facing natural-language content entirely in Chinese across the purpose, category labels, rationale text, mnemonic, and usage guidance. Under the language/locale policy, forcing a specific language without user opt-in or a documented regional justification is a policy concern.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The six-hats semantic routing test explicitly expects activation from broad phrases like '全面分析', '权衡', '风险和价值', and '结构化讨论'. Those cues are common in ordinary analytical requests, so an implementation optimized to satisfy this eval may over-trigger a specialized discussion structure when the user did not ask for it, causing misrouting and unintended workflow activation. In this skill context, that matters because routing controls memory creation, multi-agent orchestration, and output shaping, so false activation expands processing scope beyond user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The sample JSON hard-codes Chinese text for fields like title, label, summary, and rationale, which communicates a specific language expectation in a protocol document. There is no surrounding instruction that this is merely illustrative or that language should follow user preference, so it can conflict with organizational language/locale policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file includes extensive mandatory Chinese-language instructions and required Chinese phrasing, such as the Chinese-only sections for 融思者 and 主播 and required spoken expressions in podcast mode. Because the file does not indicate that the skill is region-specific or that language should be selected based on user preference, it appears to impose a locale/language policy without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The title and the entire protocol are written as a Chinese-only instruction set for the Conductor, with no indication that language should follow the user's preference or that other locales are supported. This creates a natural-language policy concern because it effectively fixes the interaction language without explicit user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document’s natural-language instructions, prompts, and examples are all Chinese, including user-facing invitation phrases. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale constraint is explicitly documented and justified, which is not present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The conductor script is written as a mandatory opening line in Chinese, which indicates a fixed language choice for user-facing interaction. The file does not offer a language option or explain that the protocol is intended only for a Chinese-language context, so this conflicts with the language/locale policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The revision prompt is specified only in Chinese, and the surrounding protocol presents these prompts as normative behavior rather than optional examples. Without opt-in or a documented locale scope, this imposes a specific language on interactions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The protocol uses Chinese participant labels and Chinese sample prompts as the default interaction pattern, which creates a language constraint across core workflow elements. Because the document does not state that the skill is China-specific or allow user language selection, this is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The template mandates appending an exact disclaimer, and the only provided text is in Chinese. This imposes a specific language on all outputs without indicating user choice, opt-in, or a documented locale-specific justification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list includes very broad phrases such as '全面分析', '从多个维度', and '结构化讨论', which are common in ordinary user requests and can cause the skill to select a specialized discussion structure when the user did not actually request one. In this skill, that can silently alter reasoning flow, output length, anonymity behavior, or orchestration mode, creating prompt-routing confusion and making the agent easier to steer through ambiguous phrasing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The only concrete speaking examples are written entirely in Chinese, which can function as a natural-language instruction signal for the skill's outputs. The file does not offer a language choice or explain that Chinese is required for a region-specific use case, so this creates a locale-policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The glossary states that the production-quality podcast spec is modeled on Chinese-market reference podcasts, which creates a locale-specific framing in the skill's normative vocabulary. Because the file does not indicate that this locale is optional, user-selectable, or justified as a region-specific skill, it may violate the language/locale policy requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The sample values and field constraints require summary and key_takeaways lengths in Chinese characters (字), and the consumption example also uses Chinese phrasing. This implies a language-specific output policy, but the document does not offer user opt-in or explain a region-specific need, which violates the language/locale policy rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The contract requires rendered output sections to use Chinese labels such as "议题背景", "与会角色", "讨论过程", "合成", and "如何继续". This imposes a specific language on all renderers and consumers without any stated user opt-in or justification for a locale-specific requirement.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/run_eval_fixture.py:24