Back to skill

Security audit

xinyi-drink

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent coffee-order helper, but it handles phone-number-linked order and reward data with review-worthy privacy and authorization risk.

Install only if you trust the Xinyi backend and are comfortable providing your own mini-program-bound phone number for reward and order-history features. Avoid using it on shared machines unless you clear the local cache afterward, and do not use it to look up another person's phone number or orders.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/query_orders.py:227
Finding

Phone Numbers Are Exposed in GET URLs and Debug Logs

Content
View full analysis

Vulnerability Details

File Locations:

  • scripts/query_orders.py:227-238
  • scripts/recommend_drink.py:177-184
  • scripts/claim_reward.py:149-160

Vulnerability Type: Sensitive information exposure through URL query parameters and debug logging
Risk Level: Medium

Vulnerable Code

scripts/query_orders.py:227-238:

python
query = {
    "mobile": resolved_mobile,
    "status": args.status,
}
orders_url = build_url(
    config["apiBaseUrl"].rstrip("/"),
    "/skill/xinyi/orders",
    query,
)
debug_log(args.debug, f"fetching orders from {orders_url}")

try:
    orders_response = fetch_json(orders_url, config["timeoutSeconds"])

scripts/recommend_drink.py:177-184:

python
context_url = build_url(
    base_url,
    "/skill/xinyi/context",
    {"mobile": resolved_mobile},
)
debug_log(args.debug, f"fetching context from {context_url}")
try:
    context_response = fetch_json(context_url, timeout)

scripts/claim_reward.py:149-160:

python
context_url = build_url(
    base_url,
    "/skill/xinyi/context",
    {"mobile": resolved_mobile},
)
try:
    debug_log(args.debug, f"fetching context from {context_url}")
    context_response = fetch_json(
        context_url,
        config["timeoutSeconds"],
    )

Technical Analysis

The order and personalized-context endpoints place the user's phone number in the URL query string. The resulting complete URL is also written to standard error when debug mode is enabled.

HTTPS protects the URL while it is in transit, but it does not prevent the URL from being recorded at either endpoint. Query strings are frequently retained in:

  • Reverse-proxy and web-server access logs
  • API gateway and load-balancer logs
  • Monitoring, tracing, and error-reporting systems
  • Debug output captured by an Agent runtime
  • Shell transcripts and continuous-integration logs
  • Network-security ...[truncated 1632 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace phone-number-bearing GET requests with POST requests whose identifiers are carried in a JSON body.

  2. Require HTTPS for every endpoint that receives personal data. Validate the configured URL scheme and reject plaintext HTTP except for explicitly enabled local testing.

  3. Never log complete URLs containing personal identifiers. Log only the endpoint path or a sanitized URL:

    python
    debug_log(args.debug, "fetching orders from /skill/xinyi/orders")
    
  4. If correlation is necessary, use a short-lived request identifier rather than the phone number. Do not log a reversible encoding of the number.

  5. Configure backend servers, proxies, gateways, monitoring platforms, and tracing systems to suppress request bodies and sensitive query parameters.

  6. Add automated tests asserting that debug output never contains the supplied or saved phone number.

  7. Review existing logs for retained phone numbers and remove or restrict them according to the applicable privacy and retention requirements.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/query_orders.py:191
Finding

Order Lookup Relies on a Caller-Supplied Phone Number Without Client-Visible Ownership Authentication

Content
View full analysis

Vulnerability Details

File Locations:

  • scripts/query_orders.py:191-238
  • scripts/skill_http.py:40-46
  • references/privacy-boundaries.md:40-41

Vulnerability Type: Potential insecure direct object reference and missing account-ownership enforcement
Risk Level: High

Vulnerable Code

scripts/query_orders.py:191-238:

python
def main() -> int:
    parser = argparse.ArgumentParser(description="按手机号查询新一好喝用户订单")
    parser.add_argument("--mobile", help="用户在【新一咖啡】微信小程序绑定的手机号")
    parser.add_argument(
        "--use-saved-mobile",
        action="store_true",
        help="内部参数:仅订单摘要或口味偏好等明确个性化场景复用本地手机号",
    )
    parser.add_argument(
        "--status",
        type=int,
        choices=(2, 4),
        help="可选订单状态分组;不传查询全部订单,2=正在进行中订单,4=历史订单/已完成订单",
    )
    parser.add_argument("--query", help="用户原始订单问题")
    parser.add_argument("--debug", action="store_true", help="输出调试信息到 stderr")
    args = parser.parse_args()

    resolved_mobile = args.mobile
    if not resolved_mobile and args.use_saved_mobile:
        resolved_mobile = load_mobile()

    if not resolved_mobile:
        sys.stdout.write(render_missing_mobile())
        return 0

    config = load_config()
    query = {
        "mobile": resolved_mobile,
        "status": args.status,
    }
    orders_url = build_url(
        config["apiBaseUrl"].rstrip("/"),
        "/skill/xinyi/orders",
        query,
    )
    debug_log(args.debug, f"fetching orders from {orders_url}")

    try:
        orders_response = fetch_json(orders_url, config["timeoutSeconds"])

scripts/skill_http.py:40-46:

python
def fetch_json(url: str, timeout: int) -> dict:
    request = urllib.request.Request(
        url,
        headers={"Accept": "application/json"},
        method="GET",
    )
    return request_json(request, timeout)

Technical Analysis

The order-query executa ...[truncated 2543 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require authentication before processing order queries. Use a server-issued, short-lived session or access token derived from the official account-login flow.
  2. Bind the authenticated principal to the phone number server-side. The backend should derive the account identifier from the authenticated session rather than trusting a caller-supplied mobile value.
  3. If phone verification is required, use a challenge-response or one-time-code flow and issue a scoped token only after successful verification. The Skill should not collect or persist the one-time code itself.
  4. Remove the phone number as a freely selectable order identifier where possible. A preferred request would contain only an authenticated token and an optional order-status filter.
  5. Enforce authorization on every order request at the backend, regardless of Agent instructions or client behavior.
  6. Return a generic authorization error for mismatched or unverified identities and apply rate limiting, abuse detection, and audit logging to deter enumeration.
  7. Add integration tests proving that:
    • An unauthenticated request is rejected.
    • A token for one account cannot query another account.
    • Modifying the supplied phone number does not alter the authenticated account scope.
    • Saved local state alone is insufficient authorization.
  8. Continue treating the natural-language third-party-phone policy as user guidance, but do not rely on it as the access-control mechanism.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (57)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill stores a user's mobile number and activity participation state in a local JSON file, which is more privacy-sensitive than the short capability description suggests. Because this cache can later be reused or displayed, unintended disclosure is possible on shared machines or permissive agent profiles, especially since the data is directly tied to order history and account-linked actions.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares executable installation and runtime behavior with network, shell, and local file access, but does not define an explicit permission or allowed-tools boundary. That weakens the trust model for reviewers and hosts because the skill can access sensitive capabilities without a clearly scoped declaration, increasing the chance of overbroad execution in permissive agent environments.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Broad trigger keywords like everyday beverage or fatigue phrases can cause the skill to activate in contexts where the user did not intend to engage a phone-number-handling, networked commerce skill. Misrouting increases privacy risk because the agent may steer users into workflows involving mobile numbers, order history, or local cache access without sufficiently specific intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

该 shell 脚本的用户可见提示、用法说明和安装完成提示均为中文,例如参数错误提示和 --help 文本。根据规则,强制特定语言且没有用户选择或合理地域限定,属于自然语言层面的语言/locale 政策问题。

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · install.sh (reported line 136)May include surrounding context.

sh
exit 4
fi

mkdir -p "$(dirname "$DEST")"
if [ -e "$DEST" ]; then
  BACKUP="$DEST.backup.$(date +%Y%m%d%H%M%S)"
  if [ -e "$BACKUP" ]; then

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · install.sh (reported line 136)May include surrounding context.

sh
exit 4
fi

mkdir -p "$(dirname "$DEST")"
if [ -e "$DEST" ]; then
  BACKUP="$DEST.backup.$(date +%YmdHMS)"
  if [ -e "$BACKUP" ]; then

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The documentation instructs users to execute npx clawhub@latest install xinyi-drink, which fetches and runs the latest published package version at install time. Because the version is not pinned, a compromised upstream package, malicious publish, or breaking update could cause arbitrary code execution on the user's machine during installation.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly asks the user to provide the mobile phone number bound to the mini-program, but the response guidance does not tell the user why this personal data is needed, how it will be used, or that account/gift/order information may be queried with it. This creates a privacy and social-engineering risk because users may disclose sensitive identifiers without informed consent, and a leaked or misused phone number could enable unauthorized lookups or profiling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This guidance invites the user to submit their bound phone number to retrieve or analyze order history, but it omits any warning that personal account and purchase data will be accessed. In the context of a consumer ordering skill, that omission is more dangerous because order history can reveal behavioral patterns and personal preferences, making uninformed disclosure of the identifier materially sensitive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This instruction normalizes collecting a bound phone number for activity participation and order analysis without warning the user that the number is sensitive personal information or explaining the verification boundary. In this skill context, the phone number can unlock account-linked order data, so collecting it casually in conversation increases the risk of privacy violations, misdelivery of benefits, or unauthorized exposure if the chat is seen by others or handled insecurely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The guidance explicitly tells the agent to ask the user to send a phone number tied to a mini-program account, but it provides no adjacent notice about why the data is needed, how it will be used, or how it will be protected. Because the same number is used for identity linkage, activity participation, and potential access to order history, this creates unnecessary privacy risk and raises the chance of oversharing sensitive personal data in chat.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The default script text directly asks the user to send their bound phone number, with no privacy, security, or anti-impersonation warning. Since this is canned copy likely to be reused broadly, it operationalizes unsafe collection of personal data and may train users to disclose account-linked identifiers in an untrusted conversational channel.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file hardcodes all user-facing labels, instructions, and response requirements in Chinese, including locale-specific wording and stylistic constraints, with no indication that the user can choose another language. This can violate language/locale policy when the skill is expected to adapt to user preference unless the restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code explicitly prompts the user to send the mobile number bound to the mini-program and later reuses that value in rendered context and claim flows, but there is no user-facing privacy notice, purpose limitation, or guidance on how the number will be used and retained. In a skill that links account identity, rewards, and order history, collecting a phone number without clear disclosure increases privacy and account-linkage risk if the response is surfaced to users or operators.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill flow directs collection of a user's phone number to query reward eligibility and then ties that identity to order-history access, expanding sensitive-data use beyond a single minimal transaction. In this context, the phone number is an account identifier, so encouraging the model to solicit and reuse it raises risks of over-collection, unauthorized account lookup, and privacy leakage if the wrong number is provided or echoed back.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code builds response sections that surface sensitive context including full mobile number, whether it came from local cache, activity participation state, order history, purchased goods, visited stores, and later even nickname. Rendering and summarizing this data for conversational use increases the chance of exposing personal information to unintended viewers, logs, downstream models, or support staff, especially because some of it is not necessary for answering every user request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file contains multiple user-facing strings exclusively in Chinese for summaries, errors, and prompts, with no indication that the user can choose another language. This is a natural-language policy issue because the skill appears to enforce a locale/language without documented opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code sends the user's mobile number as part of a network request to an orders endpoint, but this file contains no user-facing consent, notice, or minimization controls before transmitting that personal data. In a skill that handles order history for a specific user, a phone number is a sensitive identifier; undisclosed transmission increases privacy and compliance risk and could expose user data if the downstream service or logs are mishandled.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.