Back to skill

Security audit

guionizacion-humanizada

Security checks for vulnerabilities and agentic risk

Overview

This skill is a real text-humanization workflow, but it asks the agent to collect and persist Instagram-based voice profiles and frames the output as hiding AI authorship.

Install only if you are comfortable with the agent building a persistent voice profile from your or a client's content. Prefer giving a small set of writing samples yourself, confirm consent for any client account, avoid using it to hide required AI disclosure, and review or delete `cerebro/voice-guide.md` when it is no longer needed.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The activation criteria are extremely broad and include an open-ended 'any variation' clause, which can trigger the skill for ordinary editing requests. In context, that matters because activation may initiate Instagram analysis, scraping, or persistent voice-profile creation without users realizing the skill has escalated from editing into data collection.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly aims to make text 'not seem written by AI,' which promotes concealment of AI assistance rather than transparent editing. In many contexts this can facilitate deception, policy evasion, academic/professional misrepresentation, or bypass of AI-origin scrutiny, and the rest of the skill operationalizes that concealment by imitating a real person’s voice.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill’s stated purpose is text humanization, but it expands into collecting third-party social media content and persisting a voice profile in cerebro/voice-guide.md. That creates unnecessary data collection and retention risk, especially if the analyzed account belongs to a client or another person who did not meaningfully consent to long-term profiling.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Using an APIFY-backed actor to fetch public Instagram captions adds scraping/collection capability beyond what is necessary to rewrite text naturally. Even if content is public, automated harvesting for persistent stylistic profiling increases privacy, compliance, and misuse risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill promises style adaptation, but the workflow forces an Instagram-analysis phase whenever no voice guide exists. This hidden prerequisite changes the trust boundary from simple rewriting to external data gathering, which can surprise users and cause collection of more personal content than needed for the task.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The instructions direct the agent to browse profiles, compare engagement, open multiple posts, and read author comments. That is broader than needed for text editing and effectively builds behavioral and stylistic dossiers from social content, increasing privacy exposure and the chance of collecting unrelated sensitive information.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.