Back to skill

Security audit

论衡 — 严肃长文流水线

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a disclosed research-and-writing pipeline, but it contains conflicting tool-scope instructions that may lead agents to use external search services beyond the declared tool and consent boundaries.

Install only if you are comfortable with a large multi-agent workflow that writes many project files and may send research queries to external services. Before use, restrict the host tool policy to the declared tools, disable or explicitly approve any multi_search/exa/consensus/AI4Scholar/Firecrawl integrations, and avoid providing confidential topics or documents unless you choose the no-external-services path.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (18)

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The T7.5 section explicitly says its input boundary forbids reading final/, but the pseudocode immediately validates final/交付说明.md and checks isolation using final/定稿.md and final/交付说明.md. This contradiction defeats the boundary control the section claims to enforce, so an agent can justify accessing future-stage artifacts or fail unpredictably depending on which paragraph it follows.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agents/00-主控-扩展职责.md (reported line 146)May include surrounding context.

md
| T2.5 完整性门 | M-Integrity-1 + M-Form-6 + M-Exist-3 | 数据卡完整 + 信任级别齐全 | status.md「M 门执行记录」 |
| T7.5 完整性门 | M-Integrity-2 + M-Form-8 + M-Exist-2 | 审计完整 + 三角验证覆盖 + 证据包齐全 | status.md「M 门执行记录」 |
| v1→v2 修订后 | **章节级 M 门**(按 `## ` 拆章节识别变更点)| 修订说明含「章节级变更检测」段 | 修订说明头部 |
| T8 终检 | M-Form 8 项 + M-Exist 3 项 = 11 项全过 | 四档判定 = 「通过」 | `final/M-Gate-Report-v2.2.12.json` | <!-- 注:`v2.2.12` 是 M-Gate 算法自有版本号(非 SKILL 版本),保持不变 -->

T8 终检前**分片必读** [`_shared/真源/M-Gate-核心.md`](../_shared/真源/M-Gate-核心.md) 的 **M-Form / M-Exist / M-Integrity 伪代码段**(附录部分按需 → [`_shared/真源/M-Gate-Algorithm-appendix.md`](../_shared/真源/M-Gate-Algorithm-appendix.md)),按伪代码执行,产出 `final/M-Gate-Report-v2.2.12.json`(标准 JSON 格式)。

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agents/00-主控-扩展职责.md (reported line 146)May include surrounding context.

md
| T2.5 完整性门 | M-Integrity-1 + M-Form-6 + M-Exist-3 | 数据卡完整 + 信任级别齐全 | status.md「M 门执行记录」 |
| T7.5 完整性门 | M-Integrity-2 + M-Form-8 + M-Exist-2 | 审计完整 + 三角验证覆盖 + 证据包齐全 | status.md「M 门执行记录」 |
| v1→v2 修订后 | **章节级 M 门**(按 `## ` 拆章节识别变更点)| 修订说明含「章节级变更检测」段 | 修订说明头部 |
| T8 终检 | M-Form 8 项 + M-Exist 3 项 = 11 项全过 | 四档判定 = 「通过」 | `final/M-Gate-Report-v2.2.12.json` | <!-- 注:`v2.2.12` 是 M-Gate 算法自有版本号(非 SKILL 版本),保持不变 -->

T8 终检前**分片必读** [`_shared/真源/M-Gate-核心.md`](../_shared/真源/M-Gate-核心.md) 的 **M-Form / M-Exist / M-Integrity 伪代码段**(附录部分按需 → [`_shared/真源/M-Gate-Algorithm-appendix.md`](../_shared/真源/M-Gate-Algorithm-appendix.md)),按伪代码执行,产出 `final/M-Gate-Report-v2.2.12.json`(标准 JSON 格式)。

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/dispatch/G14-中文AI痕迹检测器.md (reported line 11)May include surrounding context.

md
> 公共工具白名单 / 零 exec / 叶子纪律 / 降级自报 / token 统计:见 [`_shared/真源/dispatch-header.md`](../_shared/真源/dispatch-header.md)

<!-- generated: dispatch-contract (do not edit by hand) -->
node: g14_style_gate
phase: Phase 4.4 前置(G14 风格闸)
default: required

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/dispatch/G14-中文AI痕迹检测器.md (reported line 11)May include surrounding context.

md
> 公共工具白名单 / 零 exec / 叶子纪律 / 降级自报 / token 统计:见 [`_shared/真源/dispatch-header.md`](../_shared/真源/dispatch-header.md)

<!-- generated: dispatch-contract (do not edit by hand) -->
node: g14_style_gate
phase: Phase 4.4 前置(G14 风格闸)
default: required

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file declares an execution model where M-Gate reads only final/定稿.md and final/证据包/ as inputs, but later sections require decisions based on many other files such as status.md, audits/, 01-任务简报.md, data/数据卡.md, and drafts/. This creates a confused-deputy/specification-drift condition: an implementing agent may either under-read required evidence and falsely pass a gate, or over-read unintended files and violate the documented trust boundary.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document defines 🟢🟡🔴 earlier as source trust levels for data cards, but later reuses the same symbols for timeliness on case cards. This semantic collision can cause agents or auditors to misclassify evidence quality, weakening gate checks and allowing lower-trust sources to be mistaken for merely stale-or-fresh material.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The comments and condition state that if the target language includes Chinese, G14 must run, while purely foreign-language content is marked not applicable. This is a language-based policy decision that appears mandatory rather than user-selectable, and the file does not indicate user opt-in or a documented justification for the locale-specific constraint.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This YAML manifest states the node is 'default: triggered' and continues even when the opt-out condition is undecidable, but it does not clearly define concrete trigger conditions or exclusion cases beyond an opt-out key. That makes invocation broad and potentially unintended, which matches the vague-trigger category for manifest files.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This YAML manifest marks the node as default: triggered but does not define what event, command, or context causes the skill to trigger. Because manifest files are in scope for vague-trigger checks, the lack of specificity creates ambiguity about when the node should run versus remain inactive.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/_shared/真源/中文数据源集成.md (reported line 72)May include surrounding context.

md
- **API**:完全免费、**无需 API Key**、REST 接口
- **中文文献覆盖**:收录大量中文学术期刊/学位论文(通过 Crossref/DOI 及中文出版商数据源)
- **能力**:按标题/作者/机构/主题/引用检索,支持全文搜索、过滤、分页
- **用法**:`GET https://api.openalex.org/works?search=人工智能&per-page=20`
- **适合**:文献计量、开题调研、引用分析

#### Crossref(强烈推荐)

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/_shared/真源/中文数据源集成.md (reported line 79)May include surrounding context.

md
- **性质**:DOI 注册机构([api.crossref.org](https://api.crossref.org)),全球文献元数据
- **API**:免费、无 Key、REST 接口
- **能力**:按 DOI/标题/作者检索论文元数据
- **用法**:`GET https://api.crossref.org/works?query=深度学习&rows=20`
- **适合**:查 DOI、找英文文献元数据、引用格式

> **以上两者默认关闭**,须主人在 Phase 0 **显式勾选**才调用;未勾选或选「④全部拒绝」则不调用。

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file states that the LLM does not directly connect to provider billing/quota APIs and that mapping is inferred from read-only host metadata. However, later sections add a Phase 0 '顶配档探活门' with up to three active liveness probes to verify whether a top-tier model can be called, which goes beyond purely declarative metadata-based mapping.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Line L65 says the system no longer performs model availability prechecks, with only a brief exception note. But the earlier detailed section L53-L61 defines a concrete pre-dispatch liveness gate with retries, intervals, and recheck timing, which is functionally an availability precheck. This is an internal documentation contradiction about what the system actually does.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The injected adaptive tool-priority list directs the agent to consider tools such as search_google_scholar, exa_search, consensus_search, and AI4Scholar_search, while the earlier section explicitly states that the only authoritative source of allowed tools is SKILL.md metadata.subagent_tiers.research and that undeclared tools must be rejected. This contradiction weakens the trust boundary around tool authorization: if runtime capability checks are missing, bypassed, or inconsistently enforced, the agent may attempt to use undeclared external tools, causing policy drift, unintended data egress, and inconsistent operator consent handling.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

This YAML manifest uses a single high-level condition, 'target_language_includes_zh', to decide whether the gate must run, but it does not define concrete matching criteria or exclusions in this file. Without more specific trigger scope or negative examples, mixed-language or incidental-Chinese cases could be ambiguously classified and cause unintended invocation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.