Back to skill

Security audit

Study Abroad Assistant - 留学助理

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a disclosed study-abroad assistant, but it can send stored API credentials and applicant data to an arbitrary engine URL set by the environment.

Review before installing. The default behavior is understandable for a cloud-backed admissions assistant, but only use the default compliancehub.cn engine or a local loopback engine you control. Do not run it with a custom STUDY_ENGINE_URL unless you are willing to send your API key, anon_id, profile, essay text, and application data to that endpoint. Avoid installing on shared machines unless ~/.study-abroad/api_key permissions and backups are acceptable.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/api.py:82
Finding

Unrestricted Custom Engine URL Can Exfiltrate Stored API Credentials and Applicant Data

Content
View full analysis

Vulnerability Details

File Location: scripts/api.py:16-18, 62-68, 82-92
Vulnerability Type: Server destination injection leading to credential and sensitive-data disclosure
Risk Level: High

Vulnerable Code

python
ENGINE_URL = os.environ.get("STUDY_ENGINE_URL", "https://compliancehub.cn")
API_PREFIX = "/api/study"
CONFIG_DIR = Path(os.environ.get("STUDY_CONFIG_DIR", Path.home() / ".study-abroad"))
python
def api_key() -> str | None:
    """Registered user API key: prefer STUDY_API_KEY, otherwise use KEY_FILE."""
    k = os.environ.get("STUDY_API_KEY") or ""
    if k:
        return k.strip()
    if KEY_FILE.exists():
        k = KEY_FILE.read_text().strip()
        return k or None
    return None
python
def request(method: str, path: str, params=None, body=None) -> dict:
    """Unified request returning parsed JSON response data."""
    url = ENGINE_URL.rstrip("/") + API_PREFIX + path
    headers = {"x-anon-id": anon_id(), "Content-Type": "application/json"}
    if api_key():
        headers["x-api-key"] = api_key()
    payload = json.dumps(body, ensure_ascii=False).encode() if body is not None else None

    if _http():
        httpx = _http()
        try:
            r = httpx.request(method, url, params=params, headers=headers,
                              content=payload, timeout=20, trust_env=False)

Technical Analysis

The STUDY_ENGINE_URL environment variable controls the destination of all API requests without hostname validation, scheme validation, or an enforced origin allowlist. At the same time, request() automatically attaches the registered user's API key and anonymous identifier to every request.

Consequently, an attacker who can influence the process environment or command invocation can redirect requests to an attacker-controlled HTTP or HTTPS server. The client will then transmit the x-api-key header, x-anon-id h ...[truncated 2241 chars]

Remediation
View remediation

Remediation Suggestions

  1. Enforce an explicit destination allowlist. Production requests should accept only the exact trusted HTTPS origin:

    python
    from urllib.parse import urlparse
    
    ALLOWED_HTTPS_ORIGINS = {"https://compliancehub.cn"}
    ALLOWED_LOCAL_HOSTS = {"127.0.0.1", "localhost", "::1"}
    
    def validate_engine_url(value: str) -> str:
        parsed = urlparse(value)
        origin = f"{parsed.scheme}://{parsed.netloc}"
    
        if origin in ALLOWED_HTTPS_ORIGINS:
            return value.rstrip("/")
    
        if parsed.scheme == "http" and parsed.hostname in ALLOWED_LOCAL_HOSTS:
            return value.rstrip("/")
    
        raise ValueError("Untrusted STUDY_ENGINE_URL")
    
  2. Restrict plaintext HTTP to loopback development endpoints. Require HTTPS for every non-loopback destination.

  3. Do not automatically attach credentials to custom endpoints. If custom remote engines must be supported, require a separate endpoint-specific credential and explicit user confirmation.

  4. Validate the final request origin immediately before adding x-api-key. Credential attachment should depend on the validated origin rather than merely on the existence of an API key.

  5. Disable redirects or independently validate every redirect destination before forwarding authentication headers.

  6. Warn users prominently when a non-default engine is configured, displaying the exact destination before any profile or document data is transmitted.

  7. Add automated security tests confirming that:

    • Arbitrary external HTTP and HTTPS origins are rejected.
    • HTTP is accepted only for loopback development.
    • Credentials are never sent to untrusted origins.
    • Redirects cannot move authenticated requests to another origin.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Presenting backend packaging and deployment tooling as an end-user study-abroad assistant is dangerous because it can cause an agent to run internal maintenance workflows in an inappropriate context. Such hidden operational behavior may generate files, invoke local scripts, or manipulate engine configuration beyond what a normal user request should trigger.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Presenting backend packaging and deployment tooling as an end-user study-abroad assistant is dangerous because it can cause an agent to run internal maintenance workflows in an inappropriate context. Such hidden operational behavior may generate files, invoke local scripts, or manipulate engine configuration beyond what a normal user request should trigger.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Presenting backend packaging and deployment tooling as an end-user study-abroad assistant is dangerous because it can cause an agent to run internal maintenance workflows in an inappropriate context. Such hidden operational behavior may generate files, invoke local scripts, or manipulate engine configuration beyond what a normal user request should trigger.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file contains user-facing instructions and examples solely in Chinese, but does not state that the skill is intended only for Chinese-speaking users or provide an opt-in language choice. That creates a natural-language policy issue because it effectively forces one language/locale without user selection.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill advertises capabilities that imply file, environment, network, and shell access, but it does not declare any explicit tool restrictions or permission boundaries. In an agent environment, undeclared broad capability surfaces increase the chance of over-privileged execution, unintended data access, or outbound transmission of sensitive user information to local files or remote services.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The natural-language description presents the skill entirely as a Chinese-language assistant and does not state that users may choose another language. Because the trigger list includes English phrases like "study abroad" and "grad school", the skill may be invoked by non-Chinese-speaking users without any opt-in to the enforced locale.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger list includes broad terms like general application-related phrases that may match many ordinary conversations. Over-broad activation increases the likelihood that the skill runs in contexts where users did not intend to share educational background, drafts, or planning data, which is more concerning here because the skill also claims remote engine transmission and local persistence.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 131)May include surrounding context.

md
- **不构成法律/签证/移民意见**:重大事项指向官方渠道
- **数据准确性**:知识库中 deadline/学费/录取率/GRE 要求等字段部分为待核实状态,申请前务必以各校官网为准;本 skill 不因数据误差承担录取相关责任
- **用户责任**:所有提交材料与套磁内容由用户本人核实、定稿并提交;因内容真实性/时效性导致的后果由用户承担
- **隐私(数据流向)**:画像/选校/文书等输入经 HTTPS 发送到云端引擎(`compliancehub.cn/api/study`)做确定性计算;本地仅落盘 `~/.study-abroad/` 的 anon_id(匿名额度/进度延续)与注册后的 API Key(保存时 chmod 600);本 skill 不收集账号密码;完全离线场景请勿使用云端引擎
- **降级模式**:引擎不可达时结果基于通用知识,仅供参考
- **教育咨询而非承诺**:本 skill 提供的是信息整理与流程建议,不构成对申请结果的任何保证

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The title and surrounding natural-language content present the skill backlog entirely in Chinese, with no indication that users can choose another language or locale. Under the policy rule for language/locale constraints, this can be a violation when a specific language is imposed without documented opt-in or justification.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · TODO.md (reported line 14)May include surrounding context.

md
## P1(功能补全)
- [ ] **skill 端选校支持多 discipline**:`--discipline` 目前单值,导致 `data` 类项目(如 UW MS Data Science)在 cs 方向用户下被过滤 → 改 nargs+ 或引擎支持逗号分隔
- [ ] **教授库实名核实**:占位姓名 → 按 faculty 页录入真实教授(Top 校优先,见 PROFESSOR_CANDIDATES 候选池),套磁才有真实落点
- [ ] **方案 B:真实 LLM 接入**(DeepSeek key 即可,服务器已验证 api.deepseek.com 可达):LLM_BASE_URL=https://api.deepseek.com/v1、LLM_MODEL=deepseek-chat、compose 注入重启
- [ ] 网站加一页 `/study.html`(产品说明 + 触发引导 + 注册入口)

## P2(扩展位)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest sets the skill language to "zh-CN", which can constitute a language/locale restriction when no alternative or user opt-in is described. The policy for this audit flags natural-language locale constraints unless the skill offers a choice or clearly justifies the restriction as region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The client automatically sends a persistent anon_id and, when present, an API key to a remote service on every request, but the code shown does not provide any explicit runtime notice or consent flow to the user before doing so. In this skill context, the service handles study-abroad planning/profile data, so correlating a stable identifier with profile contents and account credentials can expose personal data or enable unexpected tracking if users are unaware of the transmission.

Content

No source excerpt is available for this finding.

Dynamic import via __import__()

Medium
Category
Dangerous Code Execution
Confidence
75% confidence
Finding

Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

Content

Scanner excerpt · scripts/calibrate_ucla.py (reported line 108)May include surrounding context.

python
"仅作画像权重初版参考,必须用含真实录取结果+科研/实习细节的内部数据复核后才可采纳。"
        ),
    }
    print(__import__("json").dumps(out, ensure_ascii=False, indent=2))


if __name__ == "__main__":

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The module docstring and interactive prompts present the skill entirely in Chinese, which imposes a specific language on users without any opt-in or alternative locale path. Under the policy, forced language or locale without user choice is a natural-language policy concern.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/export_adoption.py (reported line 28)May include surrounding context.

python
def run_calibrate(script):
    r = subprocess.run([sys.executable, os.path.join(SCRIPTS, script)],
                       capture_output=True, text=True, cwd=BASE)
    if r.returncode != 0:
        raise RuntimeError(f"{script} 运行失败: {r.stderr[:500]}")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The register flow persists an API key locally via api.save_api_key(key) and then prints the storage path, but it does not warn the user about local credential persistence, expected file permissions, or risks on shared machines. In a CLI skill that may be used on personal, lab, or shared academic systems, silent credential storage increases the chance of unintended disclosure through weak filesystem permissions, backups, or other local users.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This Python file contains natural-language output and metadata entirely in Chinese, starting with the module docstring identifying the skill as a Chinese-output renderer. The policy requires flagging language or locale constraints when the skill forces a specific language without offering the user a choice or documenting a justified region-specific scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The only natural-language content in this file is written in Chinese, which imposes a specific language for readers without indicating any user choice or opt-in. Under the language/locale policy rule, this is a natural-language policy issue because the file does not offer an alternative language or document a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The top-level docstring and surrounding comments present operational guidance exclusively in Chinese, with no indication that users can choose another language. Under the language/locale policy, forcing a single language without opt-in can be a natural-language policy concern unless clearly justified.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/guide.py (reported line 201)May include surrounding context.

python
else:  # update
            fields = {}
            for k in ("gpa", "toefl", "gre", "targetDegree"):
                v = getattr(args, k, None)
                if v is not None:
                    fields[k] = v
            if args.direction:

Static analysis

Detected: suspicious.install_untrusted_source

Install source points to URL shortener or raw IP.

Warn
Code
suspicious.install_untrusted_source
Location
package.json:20