Back to skill

Security audit

Content Scorer

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent marketing-copy scorer, but users should understand that normal scoring can send pasted copy to Anthropic if the local backend is unavailable or Haiku is selected.

Install only if you are comfortable with pasted marketing copy being processed by a local LLM server or by Anthropic Claude when Haiku is used or local auto-detection fails. Use --demo or --compliance-only for no remote call, set LLM_BACKEND=local for sensitive drafts, and treat the TRIBE v2/fMRI calibration claims as unverified unless the publisher provides the referenced calibration artifacts.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The public description understates actual behavior: the skill performs remote LLM analysis, hook comparison, and compliance scanning beyond the narrow scoring description. Description-behavior mismatches are dangerous because they defeat informed consent, interfere with security review, and can cause sensitive copy or internal business text to be sent to third-party services unexpectedly.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The advertised calibration suggests a concrete scoring methodology or model lineage tied to TRIBE v2 and fMRI brain-response data. In code, scoring is produced by prompting either a local OpenAI-compatible model or Anthropic Haiku and combining outputs with hard-coded weights; there is no dataset, model artifact, calibration routine, or reference to TRIBE v2 beyond the manifest claim.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill invokes networked and environment-backed capabilities via Anthropic API access but does not declare any explicit tool scope or permissions. That creates a transparency and governance gap: users and orchestration systems may treat the skill as lower-risk than it is, while pasted content can still be transmitted externally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill invites users to paste marketing copy for premium scoring without a clear warning that the content will be transmitted to Anthropic. If users submit unreleased campaigns, customer data, or proprietary sales material, this can produce unintended third-party disclosure and policy noncompliance.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
79% confidence
Finding

The file repeatedly endorses and rewards 'embedded commands,' 'assume-the-close,' urgency, and other NLP persuasion tactics in natural-language instructions and examples. While not a static security issue, this is a semantic policy concern because the skill encourages covert influence techniques as a default behavior rather than transparently framing them as optional or constrained.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest says the skill scores marketing copy for resonance, hook strength, NLP technique usage, and conversion readiness, returning a 0-100 score with breakdown and rewrite suggestions. The code also exposes a forbidden-words compliance checker, a standalone hook comparison mode, and a full rewrite generator as first-class CLI capabilities, which materially expand the skill beyond the described scoring scope.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script description emphasizes local MLX scoring by default, but the implementation can fall back to Anthropic and transmit user-provided marketing copy off-host. This creates a confidentiality and consent risk because users may supply proprietary or regulated content believing processing is local-only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

User copy is sent to the Claude backend without a runtime warning or consent prompt when auto-fallback occurs. If users input confidential campaign drafts, client data, or sensitive business text, the skill may silently exfiltrate that content to a third-party service.

Content

No source excerpt is available for this finding.

Internal Network Request

Medium
Category
Server-Side Request Forgery
Confidence
70% confidence
Finding

Code issues a request to a loopback, link-local, or private-range host. This can reach internal services not meant to be exposed and is a common SSRF pivot.

Content

Scanner excerpt · score_content.py (reported line 114)May include surrounding context.

python
try:
            import openai
            import urllib.request
            urllib.request.urlopen("http://localhost:8800/health", timeout=1)
            return openai.OpenAI(base_url="http://localhost:8800/v1", api_key="local"), "local"
        except Exception:
            import anthropic

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest markets the skill as a calibrated scoring engine, but actual premium behavior depends on a third-party hosted LLM and an API key. This is a trust and disclosure issue: users may assume deterministic local analysis when their content is actually processed by Anthropic, creating unexpected confidentiality and compliance exposure.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest and main description present this as a scoring/rewrite skill operating on provided copy, with implementation centered on score_content.py. However, the documentation explicitly directs recalibration via vault/learnings/2026-03-27-tribe-v2-colab-spec-task47.md, indicating dependence on or reference to external project data beyond the described scoring workflow.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.