Back to skill

Security audit

智能批改作业

Security checks for vulnerabilities and agentic risk

Overview

The skill’s thesis-review purpose is understandable, but it stores sensitive document text in temporary files and its advertised evidence guardrail can validate evidence against the wrong source document.

Install only if you are comfortable with the skill reading all supported files in the chosen archive or directory and writing extracted plaintext to local temporary storage. Use it in a restricted workspace, review the files to be processed first, clean /tmp/auto_grading after use, and treat the evidence guardrail as imperfect because findings may be attributed to the wrong document when filenames or evidence overlap.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/harness.py:58
Finding

Evidence Verification Is Not Bound to the Claimed Source Document

Content
View full analysis
= 0: preview = content[max(0, idx - 10):idx + len(evidence) + 30] return {"verified": True, "found_in": txt_file, "match_preview": preview.replace('\n', ' ')[:80]} # 降级:模糊匹配前30字符 idx = content.find(query) ``` ```python # scripts/batch_extract.py:99-102 txt_name = Path(f).stem + '.txt' txt_path = os.path.join(args.output, txt_name) with open(txt_path, 'w', encoding='utf-8') as wf: wf.write(r.get('text', '')) ``` ### Technical Analysis `verify_evidence()` retrieves the issue's claimed source filename into `filename`, but never uses that variable when selecting the text to search. Instead, it searches every `.txt` file in the extraction directory and accepts the first exact match or a match of only the first 30 characters. Consequently, evidence from one document can validate a finding attributed to an entirely diffe ...[truncated 1793 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/batch_extract.py:20
Finding

Temporary Plaintext Files Are Not Reliably Deleted

Content
View full analysis
100: return {'text': text, 'method': 'textutil'} # 2. antiword r = subprocess.run(['antiword', path], capture_output=True, text=True, timeout=30) ``` ### Technical Analysis The temporary file is created with `delete=False`, but deletion occurs only inside the successful `textutil` return-code branch and after the temporary output has been read. The temporary file is not reliably removed when: - `textutil` returns a nonzero status; - `subprocess.run()` raises an exception, including timeout or missing-executable errors; - opening or reading the converted file fails; - an exception occurs before `os.unlink()` executes. Because document conversion can place student identity, academic work, or other document contents into this temporary file, an interrupted or failed conversion can leave plaintext data behind. No `finally` block guarantees cleanup. ### Attack Path 1. A crafted or malformed `.doc` file is submitted for extraction. 2. `NamedTemporaryFile(delete=False)` creates a persistent temporary path. 3. The external `textutil` conversion fails, times out, or produces output before returning a nonzero status. 4. Execution leaves the successful cleanup path or raises an exception before `os.unlink()` runs. 5. The temporary plaintext file remains in the system temporary directory. 6. A local process or user with ...[truncated 664 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:135
Finding

Third-Party Dependencies Are Installed Without Version or Integrity Pinning

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents an automated review/grading system for theses and related archive documents. The supplied code does not implement review, grading, validation, or evidence-based issue detection. Its primary function is limited to extracting text from supported document types and saving the extracted text plus basic metadata (images/pages/slides/domain code counts) into output files and a manifest. While batch reading of doc/docx/pdf/pptx is consistent with part of the description, the core promised behavior—multi-dimensional assessment, cross-document checks, PPT evaluation, bibliography standard validation, and safety guardrails—is absent. Therefore the description materially overstates what the code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

代码与声明存在明显的描述-行为不一致。声明描述的是一个完整的毕业设计/论文自动化审查系统,重点在多文档解析、内容评分与格式规范检查;而给出的代码只是一段 harness.py,作用是对外部审查结果进行结构化校验和证据回溯,并辅以 doc/docx 版本冲突判断。虽然其中的“证据可溯源、排除文件混淆、宁可漏报不可误报”等护栏理念与声明中的安全护栏部分一致,但这仅覆盖了声明中的一小部分辅助能力,无法支撑其主要宣称功能。故应判定为 mismatch。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documents capabilities equivalent to file read, file write, and shell execution but does not declare any explicit tool scope or allowed-tools boundary. In practice this creates an over-privileged, weakly constrained workflow that could read arbitrary local content, write extracted sensitive student material to disk, and invoke system utilities without a clear policy gate.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Allowing upload of thesis/archive files or a directory to serve as an ambiguous activation condition weakens user-intent verification. Given the skill's documented ability to enumerate and extract from multiple file types, this ambiguity makes accidental data processing more dangerous in context.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Allowing upload of thesis/archive files or a directory to serve as an ambiguous activation condition weakens user-intent verification. Given the skill's documented ability to enumerate and extract from multiple file types, this ambiguity makes accidental data processing more dangerous in context.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The workflow invokes shell/system utilities such as text extraction helpers for document processing, which expands the attack surface beyond normal grading logic. Shell-based processing of untrusted user-supplied documents or paths increases the risk of command misuse, unsafe argument handling, and unintended access to the host environment.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill is framed as document review and scoring, but this code shells out to host-installed programs to process .doc files. Spawning subprocesses is not an obvious or declared requirement from the manifest itself and expands the runtime capability beyond straightforward document parsing libraries.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/batch_extract.py (reported line 20)May include surrounding context.

python
import subprocess, tempfile
    tmp = tempfile.NamedTemporaryFile(suffix='.txt', delete=False)
    tmp.close()
    r = subprocess.run(['textutil', '-convert', 'txt', '-output', tmp.name, path],
                       capture_output=True, timeout=30)
    if r.returncode == 0:
        with open(tmp.name, 'r', encoding='utf-8', errors='ignore') as f:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/batch_extract.py (reported line 29)May include surrounding context.

python
if len(text) > 100:
            return {'text': text, 'method': 'textutil'}
    # 2. antiword
    r = subprocess.run(['antiword', path], capture_output=True, text=True, timeout=30)
    if r.returncode == 0 and len(r.stdout) > 100:
        return {'text': r.stdout, 'method': 'antiword'}
    # 3. olefile 二进制提取

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script persists full extracted document text and a manifest containing source file paths to disk under /tmp/auto_grading by default. For a grading/review skill that may process sensitive student papers or thesis archives, this creates unintended data retention and metadata exposure, especially on shared systems where temporary directories may be accessible or insufficiently monitored.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script writes extracted contents plus a manifest of filenames and paths to disk without any consent, warning, or retention control. In the context of academic submissions, this can expose personally identifiable information, unpublished research, and directory metadata beyond what is necessary for immediate grading.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This Python file contains user-facing natural-language text entirely in Chinese, including the module description and operational labels, with no indication that language is configurable or optional. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The natural-language description and operating instructions are presented only in Chinese, with no indication that users may choose another language or locale. Under the stated policy, forcing a specific language without opt-in can be a language/locale policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill writes extracted document text to /tmp, which can expose sensitive academic content, student identifiers, and thesis material to other processes or later reuse if cleanup is weak. Because the inputs are likely confidential educational records, unnecessary persistence materially increases privacy and data-leakage risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This markdown guidance uses mandatory language such as “强制项”, “必须”, and “禁止” to enforce Chinese GB/T 7714 formatting conventions, including language-specific punctuation and English-name formatting. Under the policy rule, forcing a specific language/locale convention without opt-in or an explicit region-specific justification can be a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The module docstring presents the skill description only in Chinese, and the command-line help text is also localized in Chinese. This can violate language/locale policy when the skill forces a specific language without user opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.