Back to skill

Security audit

tiangong-skill

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent skill-building toolkit with disclosed local output creation and validation scripts; the main caution is broad auto-trigger wording, not hidden or destructive behavior.

Install this if you want a workflow for creating and validating agent/skill definitions. Before use, be aware it may create files under output/ and can use web search for persona research only when authorized; review any generated SKILL.md before installing or using it. Consider narrowing the trigger words if you do not want ordinary creation requests to route into this skill.

Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill declares no required permissions, yet the documented workflow references local file access and validation of multiple repository files such as references/, examples/, and generated outputs. This creates a permission/behavior mismatch that can cause the agent to read local content without an explicit, user-visible declaration of that capability, increasing the chance of unintended data exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The manifest describes an agent-design assistant, but the instructions also include repository auditing, structural validation, conflict detection, and local path-restricted verification behavior. This hidden expansion of scope is dangerous because users and orchestrators may invoke the skill expecting content design while it performs broader analysis of local files and project structure.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill instructs the agent to perform web_search/web_fetch for persona research even though that external network access is not clearly disclosed in the manifest. Undeclared external retrieval expands the trust boundary and may cause unanticipated data exfiltration, prompt injection exposure from remote content, or policy bypass through fetched materials.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The workflow directs the agent to write generated deliverables into an output/ directory, but disk-writing behavior is not reflected in the skill's stated role or permission metadata. Undisclosed file creation/modification can overwrite artifacts, persist sensitive content locally, or surprise downstream systems that expected a read-only design assistant.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The job-oriented pipeline also mandates writing a final SKILL.md file to disk, reinforcing that the skill performs persistence beyond advisory design. Persisting generated content without clearly declared write capability increases risk of unauthorized file changes and accidental storage of sensitive prompts or user data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The quickstart states the workflow will 'automatically execute' J1→J5 and elsewhere defines J5 as writing a SKILL.md into the output directory, but it does not clearly warn users about filesystem-modifying behavior before they follow the instructions. In an agent context, implicit write actions can surprise users, overwrite local artifacts, or normalize unsafe autonomous file operations without explicit consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

Broad trigger phrases like '做一个' and '帮我写个' are likely to match many unrelated conversations, causing the skill to activate outside its intended context. Over-triggering is dangerous because it can route benign requests into a more privileged or more complex workflow that reads files, writes outputs, or performs external research unexpectedly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

Restricting disclaimers to only the first activation can weaken ongoing transparency in a roleplay/persona system, especially when later turns may be detached from the original context or shared independently. In a skill designed to emulate real people and maintain character, users may forget they are interacting with a distilled persona, increasing the risk of misplaced authority or confusion.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The template explicitly instructs the agent to remain in first-person character and to give a disclaimer only once, then stop repeating it. That can mislead users in longer conversations by making synthetic persona output look increasingly authentic and reducing ongoing transparency that the system is roleplaying a reconstructed public figure rather than providing verified, attributable statements.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger phrases are extremely broad and include common conversational text such as '做一个', '帮我写个', and generic English equivalents. This can cause the skill to activate in unrelated conversations, leading to unintended takeover of the interaction flow, prompt/context pollution, or accidental use of a powerful agent-design skill when the user did not explicitly request it.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document states that mentioning any one intent word will automatically activate the skill, which creates ambiguous activation semantics. In practice, mere discussion about agents or examples could trigger the skill without clear user consent, increasing the risk of unintended behavior and reducing the reliability of routing and safety boundaries.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.