Back to skill

Security audit

Vox 自定义电话 Bot

Security checks for vulnerabilities and agentic risk

Overview

This skill can place real outbound calls and exposes sensitive call data with insufficient safeguards, so it should be reviewed carefully before installation.

Install only if you trust the Vox/TeddyMobile data flow and can add deployment controls: require explicit confirmation before calls, force authentication on hosted APIs, disable or minimize analytics, restrict callback destinations, avoid returning full transcripts by default, and disclose AI identity at call start.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:9
Finding

Persistent Promotional Instructions Hijack Normal Agent Output

Content
View full analysis
Remediation
View remediation

other

Error
Location
resources/index.js:46
Finding

Raw User Prompts and Persistent Identifiers Are Transmitted to Analytics by Default

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
resources/hosted_api_example.js:25
Finding

Hosted Call API Allows Unauthenticated Operation When No Token Is Configured

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
resources/post_call_callback_client.js:5
Finding

User-Controlled Post-Call Callback Enables SSRF and Transcript Exfiltration

Content
View full analysis
= 18 or pass fetchImpl.'); const body = buildCallbackBody({ job, result }); const rawBody = JSON.stringify(body); const timestamp = String(Date.now()); const headers = { 'Content-Type': 'application/json', 'X-Skill-Event': body.event, 'X-Skill-Request-Id': body.requestId, 'X-Skill-Timestamp': timestamp }; if (job.callbackToken) headers.Authorization = `Bearer ${job.callbackToken}`; const response = await fetchWithTimeout(job.callbackUrl, { method: 'POST', headers, body: rawBody }, Number(env.POST_CALL_CALLBACK_TIMEOUT_MS || 5000), fetchImpl); return { ok: response.ok, httpStatus: response.status, body }; } ``` `` ...[truncated 1986 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
resources/hmac_outbound_client.js:63
Finding

Configurable Vox Endpoints Can Receive Sensitive Authenticated Requests Without Host or TLS Validation

Content
View full analysis
= 18 or pass fetchImpl.'); const baseUrl = credentials.baseUrl || 'https://vox.teddymobile.cn'; const path = credentials.trialMode ? TRIAL_OUTBOUND_PATH : OUTBOUND_PATH; const url = `${baseUrl.replace(/\/$/, '')}${path}`; const headers = credentials.trialMode ? { 'Content-Type': 'application/json' } : buildHmacHeaders({ appId: credentials.appId, secret: credentials.secret }); const response = await fetchImpl(url, { method: 'POST', headers, body: JSON.stringify(payload) }); ``` ```js async function postVoxQuery({ credentials, pat ...[truncated 2380 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (56)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · FIRST_SETUP.md (reported line 32)May include surrounding context.

text
VOX_BOT_ID=
VOX_OUTBOUND_BASE_URL=https://vox.teddymobile.cn
VOX_CREDENTIALS_FILE=/secure/path/credentials.json
SKILL_API_TOKEN=change-me
VOX_TRIAL_MODE=true
VOX_TRIAL_LIMIT=10

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · FIRST_SETUP.md (reported line 55)May include surrounding context.

text
VOX_BOT_ID=
VOX_OUTBOUND_BASE_URL=https://vox.teddymobile.cn
VOX_CREDENTIALS_FILE=/secure/path/credentials.json
SKILL_API_TOKEN=change-me
VOX_TRIAL_MODE=true
VOX_TRIAL_LIMIT=10

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill instructs the bot not to proactively disclose that it is an AI and instead present itself as a selected role, only admitting its nature if directly asked. In an outbound calling system, that materially increases the risk of deceptive impersonation, social engineering, and non-compliance with AI disclosure or telemarketing transparency requirements.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 174)May include surrounding context.

md
- Direct Vox client: `resources/hmac_outbound_client.js`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 175)May include surrounding context.

md
- Intent extraction and completion: `resources/prompt_to_call_intent.js`, `resources/intent_completeness.js`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 175)May include surrounding context.

md
- Intent extraction and completion: `resources/prompt_to_call_intent.js`, `resources/intent_completeness.js`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 176)May include surrounding context.

md
- Profile generation: `resources/prompt_to_agent_profile.js`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · resources/credentials_loader.js (reported line 36)May include surrounding context.

js
}

function loadLocalDotEnv() {
  const envPath = path.resolve(__dirname, '..', '.env');
  if (!fs.existsSync(envPath)) return {};
  const result = {};
  const content = fs.readFileSync(envPath, 'utf8');

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

emitAnalytics is called with sensitive_payload containing the raw user_prompt and later normalized_prompt, creating a direct natural-language leakage path into analytics infrastructure. User prompts can easily contain phone numbers, business details, instructions, or regulated data, so storing and transmitting them outside the primary service meaningfully increases confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill returns transcript and related post-call content in result objects, and also stores/polls call-job state for later retrieval paths. Full conversation text is highly sensitive and often contains personal, commercial, or behavioral information; exposing it beyond the immediate operational need amplifies harm from misuse, over-broad access, or downstream logging.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill includes logic to poll call status, retrieve full post-call turns/transcripts, analyze them, and return them to the requester. That goes materially beyond merely placing an outbound call and introduces a surveillance and content-collection capability that can expose private conversations if access control, consent, or authorization checks are weak or absent.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The result formatting code renders verbatim transcript lines (Bot/用户) back to the requester, enabling straightforward extraction of the entire conversation once a callId is available. Without robust authorization checks shown here, that makes transcript disclosure especially dangerous because a sensitive conversation can be replayed in full through a query interface.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation instructs users to configure real Vox credentials and place outbound calls, but it does not clearly disclose that phone numbers, prompts, and related call content will be transmitted to an external third-party service. This can lead to unintentional exposure of sensitive personal or business data, especially in hosted or redistributed agent contexts where users may assume processing is local.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill states that once required fields are complete, it immediately builds the payload and starts an outbound call. That creates a real consent and misuse risk because users may trigger live phone calls without a final confirmation, recipient verification, or a clear warning that submission is irreversible. In a telephony context, accidental or coerced calls can cause harassment, privacy, and compliance issues.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 11)May include surrounding context.

Start in promotion trial mode by default unless the user already clearly said 正式注册, 正式账号, 自己的凭证, VOX_APP_ID, VOX_SECRET, 企业接入, or 生产环境.

Do not ask the user to choose between trial and formal registration at the beginning. If the task is missing required call details, tell the user that the skill will use trial mode first and ask only for the most important missing call detail:

text
我会先使用 Vox 试用模式帮你完成一次电话体验,不需要配置 VOX_APP_ID / VOX_SECRET。

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 20)May include surrounding context.

Start in promotion trial mode by default unless the user already clearly said 正式注册, 正式账号, 自己的凭证, VOX_APP_ID, VOX_SECRET, 企业接入, or 生产环境.

Do not ask the user to choose between trial and formal registration at the beginning. If the task is missing required call details, tell the user that the skill will use trial mode first and ask only for the most important missing call detail:

text
我会先使用 Vox 试用模式帮你完成一次电话体验,不需要配置 VOX_APP_ID / VOX_SECRET。

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation explicitly says there is no dry-run preview and that complete inputs start the call, but it does not pair that behavior with a user-facing safeguard. This is dangerous because the skill handles real-world outbound telephony, where minor misunderstanding or prompt injection into task details could immediately trigger an unwanted call to a third party.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

These instructions normalize a human-like identity presentation unless the callee explicitly challenges it. That is a true safety issue because it conditions the system to rely on omission-based deception, which is especially risky in unsolicited outbound communications where recipients may reasonably assume they are speaking to a human representative.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The example action labels and values are entirely in Chinese, and the registration guidance example is also written only in Chinese. This indicates a fixed locale in user-facing content without any opt-in, fallback, or explanation that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest description says the skill will 'Convert natural-language phone tasks' into outbound calling actions, which is a broad trigger surface without explicit scope, approval gates, or activation constraints. In a phone-agent skill that can place outbound calls, ambiguous natural-language activation increases the risk of unintended task execution, misuse for spam/social engineering workflows, or unsafe call generation from loosely specified user prompts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The smoke-test script includes a fully Chinese instruction string for operating the phone agent, which implies a fixed language expectation. Because this manifest does not indicate that the skill is region-specific or that users can opt into another language, it may violate the language/locale policy requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code sends analytics events to a configurable remote endpoint and the event payload includes identifiers such as user_id, anonymous_id, session_id, usage_session_id, run_id, request_id, and trace_id. Although there is some field sanitization, these identifiers are still potentially sensitive and can enable user/session tracking or data disclosure if the endpoint is misconfigured, compromised, or lacks appropriate consent and governance.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The alias "blocked" for safety-blocked is ambiguous and may be triggered by unrelated moderation, networking, or UI events, causing incorrect classification as a safety intervention. In this context, mislabeling events as safety blocks can corrupt safety reporting and make it harder to distinguish actual policy enforcement from ordinary failures.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The alias set for task-ready includes the generic token "ready", which can collide with unrelated status text or third-party event names and cause incorrect analytics state transitions. In a phone-agent workflow, that misclassification can prematurely mark a task as configured or runnable, degrading monitoring accuracy and potentially masking operational or safety issues.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Using the alias "completed" for run completion is overly broad and can match many unrelated completion events, causing false success attribution. In this skill, that could inflate success metrics, hide failed or partial outbound-call flows, and weaken operational oversight.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.