Back to skill

Security audit

Paper Polisher Pro — AI Detector & Academic Polishing

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed local academic-writing analysis toolkit, with scanner concerns mostly explained by local evaluation and maintenance scripts rather than hidden or harmful behavior.

Install only if you are comfortable with a local academic-writing detector that reads your drafts and can write reports or corrected files where you request. Treat AI-score and de-AI wording as self-review aids, not as a way to hide AI use; follow your institution’s disclosure rules. Optional model setup writes to ~/.cache/paper-polisher and should use only the author-signed files accepted by pp_setup.py.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (87)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk does not implement academic-writing self-checking, polish guidance, compliance label checking, paragraph attribution, journal precheck, or batch rewriting guidance. Instead, it is a standalone attack-data generator used to create adversarial evaluation samples from a corpus. It loads synonym resources, rewrites AI texts via synonym replacement and sentence swapping, mixes human and AI paragraphs, and exports the results as attack-labeled JSONL records. While this may support evaluation of an AI-text detector, its direct purpose and behavior are materially different from the declared end-user skill description, and the CLI triggers (--corpus, --out, --n-per-model) also do not match the declared user-facing functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code chunk does not implement the declared end-user skill behavior. Instead of analyzing academic writing, scoring AI-likeness, giving polish guidance, checking AIGC labeling compliance, or performing paragraph-level attribution on provided documents, it generates a controlled evaluation benchmark from internal corpora. Its primary function is to sample human and AI paragraphs, synthesize mixed documents under different patterns and AI fractions, attach ground-truth paragraph labels, and save the benchmark for later evaluation. While this may support development or validation of a paragraph-attribution engine mentioned in the description, this chunk’s actual behavior is a separate internal benchmarking capability that is not reflected in the declared purpose. Therefore, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code does not implement academic writing quality analysis, polishing guidance, AI-generated text detection, attribution, compliance label checks, journal precheck, or batch rewriting guidance. Instead, its sole purpose is data hygiene auditing: loading training and evaluation corpora, hashing normalized text prefixes, detecting exact overlaps, summarizing hit counts by corpus/train set, and failing the process if any leakage is found. This is a materially different primary purpose and introduces undeclared capabilities related to training/evaluation contamination auditing. The code is local and credential-free, which is consistent with part of the description, but that does not resolve the strong functionality mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a user-facing academic writing self-check/polishing/detection tool with optional batch guidance and attribution features. The supplied code chunk does not implement any of those functions. Instead, it is a dataset preparation utility for evaluation: it loads two named benchmark datasets from local directories, normalizes and filters entries, flips labels for one source, samples records by strata, and writes a consolidated JSONL corpus. This is a materially different primary purpose. While such a script could support development/evaluation of the broader tool, the code chunk itself is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a user-facing local academic-writing audit/polish skill with many analysis and compliance features, plus batch rewriting guidance. The supplied code chunk does not implement those functions. Instead, it is an internal regression/evaluation framework for the detector engine: it reads corpora/attacks JSONL files, runs ai_detector on them, computes benchmark metrics, evaluates split/sample settings, and saves results. While this supports the broader product claims about detector quality, this chunk’s primary purpose is materially different from the declared skill behavior and includes undeclared evaluation capabilities. Therefore this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code substantially matches the detection-related parts of the description: bilingual local AI-rate scoring, paragraph-level attribution, journal precheck, optional supervised layer integration, model fingerprint attribution, and --batch directory processing are all present. It is also fully local and requires no credentials/network access. However, the declared description presents a broader skill that includes polishing guidance, style/terminology/translation-smell review, metaphor audit, and batch rewriting guidance. None of those capabilities are implemented in this code chunk. The code is narrowly an AI-writing detection/reporting tool, not a writing polisher or rewriting advisor. It also does not perform AIGC compliance label checking itself; it merely mentions another script for that purpose. Therefore the description materially overstates this code chunk’s implemented capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The supplied code chunk is narrowly focused on aggregating outputs from four local detection/check scripts into a composite AI-risk verdict for one file. That aligns with part of the description about local AI-rate self-checking and terminology/translation-smell/style analysis. However, the declared purpose describes a much broader multifunction tool, including compliance labeling, paragraph attribution, journal precheck, batch thesis-scale processing, rewriting guidance, metaphor audit, supervised ONNX inference, and LLM-family fingerprint attribution. None of those major capabilities appear in this code chunk. The code also only accepts a single file plus --json, directly contradicting the declared --batch DIR trigger. This is therefore a material description-versus-behavior mismatch, even though some core AI-risk scoring claims partially match.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a broad academic-writing and AIGC-analysis system with many specialized detection, attribution, compliance, and batch-processing features. The supplied code chunk instead implements only a narrow n-gram overlap/repetition utility for comparing two files or checking internal repetition within one file. While such a utility could be a supporting component of a larger rewriting/polishing toolkit, this chunk by itself does not substantiate the declared primary purpose or most of the claimed capabilities. Its triggers and resource usage are limited to local file input/output and CLI arguments, which are consistent with being local-only, but the actual behavior is materially narrower and different from the declared functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents an end-user academic-writing analysis and guidance skill. The supplied code chunk instead implements an internal maintenance/calibration script for the detector’s pattern database. Its primary function is to evaluate regex/string patterns against a labeled corpus, drop or demote patterns based on false-positive and lift thresholds, and optionally persist the recalibrated pattern set to ai_patterns_zh.json. While this could be a supporting component of a broader AI-detection system, it does not itself perform the user-facing functions emphasized in the description, and it introduces an undeclared file-modifying calibration capability. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The description presents a comprehensive academic-writing audit and compliance tool with many features beyond AI-rate scoring. The actual code chunk is much narrower: it only computes four simple heuristic predictability/repetition metrics and a combined AI-like score for zh/en text. The code is consistent with a small part of the claimed 'AI-rate self-check' functionality, but it does not substantiate most of the declared capabilities, especially compliance labeling, attribution, journal checks, batch processing, fingerprinting, or any polishing guidance. There are no undeclared sensitive behaviors or external access; the mismatch is that the declared scope and sophistication are materially broader than what this code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The supplied code implements only a narrower 'quality report' component. Its primary behavior is reporting AI-trace score, structure completeness, and basic readability statistics for one input file, with optional before/after comparison. That partially aligns with the declared 'quality report' and local bilingual self-check themes, but the description makes many concrete claims that are not represented in this code chunk: batch processing, compliance labeling checks, attribution, journal precheck, rewrite/polish guidance, model fingerprinting, and advanced multilayer/ONNX/freshness features. Since the declared description substantially overstates what this code actually does, this chunk is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code does not implement the declared skill's primary purpose. Instead of a comprehensive academic-writing/AIGC analysis system, it only checks terminology normalization against a local JSON database and optionally writes a corrected output file. Its scope is specifically medical term standardization, with single-file CLI input and optional JSON output/auto-fix. While this is loosely related to the declared mention of terminology/polish guidance and local processing, the major advertised capabilities—AI-rate analysis, compliance labeling, attribution, batch mode, multilayer engine, and model fingerprinting—are absent from this code chunk. Additionally, the code includes an undeclared auto-fix file-rewriting capability. Therefore this chunk materially differs from the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code’s primary purpose is much narrower than the declared description. It implements a local rule-based Chinese translation-smell checker with specific heuristics: repeated 的, excessive 被, abstract nominalization, connector overuse, redundant verbs, and a small seed lexicon of suspected literal translations. It cross-checks against a local terminology.json file and a tiny hardcoded approximation of translationese patterns, then prints/report JSON. That fits only a small subset of the declared 'polish guidance/style/translation-smell' functionality. Major advertised capabilities—AI-rate scoring, multilayer fusion engine, optional supervised ONNX layer, AIGC labeling compliance checks, paragraph attribution, journal precheck, LLM source fingerprinting, CN/EN bilingual support, and batch rewriting guidance—are absent from this code chunk. There is also a notable behavior discrepancy: when given a directory it recursively scans only *.ts files, which does not align with the broad academic document/thesis-scale batch description. Therefore this chunk does not accurately represent the declared skill as stated.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · eval/release_smoke.py (reported line 54)May include surrounding context.

python
def run_py(script, args, env_extra=None, timeout=180):
    env = dict(os.environ)
    if env_extra:
        env.update(env_extra)
    p = subprocess.run([sys.executable, str(SCRIPTS / script)] + args,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · eval/release_smoke.py (reported line 487)May include surrounding context.

python
_p = subprocess.run([sys.executable, str(SCRIPTS / "ai_detector.py"), _hu_file_hold,
                         "--profile", "journal", "--format", "json"],
                        capture_output=True, text=True, encoding="utf-8", timeout=180,
                        env={**os.environ, "PP_ORT_THREADS": "8"})
    try:
        _jd = json.loads(_p.stdout)
        _ok_j = _jd.get("journal_precheck") is not None and _jd.get("risk_bands") is not None

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · eval/release_smoke.py (reported line 509)May include surrounding context.

python
_p = subprocess.run([sys.executable, str(SCRIPTS / "ai_detector.py"), _hu_file_hold,
                         "--profile", "journal", "--format", "json"],
                        capture_output=True, text=True, encoding="utf-8", timeout=180,
                        env={**os.environ, "PP_ORT_THREADS": "8"})
    try:
        _jd = json.loads(_p.stdout)
        _ok_j = _jd.get("journal_precheck") is not None and _jd.get("risk_bands") is not None

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · eval/release_smoke.py (reported line 524)May include surrounding context.

python
_p = subprocess.run([sys.executable, str(SCRIPTS / "ai_detector.py"), _hu_file_hold,
                         "--profile", "journal", "--format", "json"],
                        capture_output=True, text=True, encoding="utf-8", timeout=180,
                        env={**os.environ, "PP_ORT_THREADS": "8"})
    try:
        _jd = json.loads(_p.stdout)
        _ok_j = _jd.get("journal_precheck") is not None and _jd.get("risk_bands") is not None

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description and user-facing instructions are predominantly in Chinese, including the title, feature list, setup guidance, and compliance notes. The file does not indicate that the tool is China-specific or offer an alternative language, which creates a language/locale policy concern under the rule for forced language without user opt-in.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises broad local capabilities including file reads/writes, shell execution, environment-variable use, and references to optional model setup, but it does not declare any explicit tool scope or permission boundary. That omission makes it harder for a host platform or reviewer to constrain what the skill may access, increasing the risk of over-privileged execution if the implementation later performs unintended filesystem, shell, or network actions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The documented fingerprint-mining and freshness workflow expands the skill from end-user writing self-check into model-output sampling and detector-maintenance activity. That broader data-collection and corpus-building behavior increases the chance of processing third-party or sensitive text outside user expectations, and it weakens the claim of a narrowly scoped local writing assistant.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrases include broad academic-assistance and rewriting-style requests such as checking AI rate, polishing papers, and rewriting-related language. Broad routing terms can cause the skill to activate in contexts where users are seeking help to disguise AI authorship, making misuse more likely even if the docs disclaim that intent.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The related-tool reference to a 'de-AI rewriting companion' conflicts with the skill's integrity-focused framing and can normalize or encourage use for evasion of AI-detection or disclosure rules. In this context, that contradiction is dangerous because the skill already offers detection and polishing features that could be repurposed to iteratively lower detectability.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

From L06740 onward, many records place English disease names in the cn field while en is left empty, effectively embedding an undocumented language shift inside a file otherwise structured as Chinese/English terminology pairs. This creates a locale-policy issue because the file starts as Chinese-first terminology data but later forces English labels without any explicit opt-in or justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This file is explicitly an "对抗攻击样本生成器" for generating paraphrase and mixed attack corpora to red-team a detector, rather than performing paper polishing, attribution, journal precheck, or compliance checks for user documents. While internal evaluation tooling can exist in a repository, this capability is not justified by the manifest’s claimed user-facing purpose and represents a materially different function.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
eval/run_eval.py:225