Back to skill

Security audit

secretary-core-秘书核心模块

Security checks for vulnerabilities and agentic risk

Overview

This skill looks like a local assistant prototype, but its documentation asks for messaging tokens and promises live calendar, reminder, and platform actions that the reviewed code does not actually implement or safely scope.

Review carefully before installing. Do not rely on this skill's confirmations for real meetings, reminders, or notifications unless a verified integration is present. Avoid placing real bot tokens in plaintext config files, prefer pinned reviewed releases, and require explicit confirmation plus privacy controls before enabling memory, habit learning, calendar changes, or message sending.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:260
Finding

Bot Credentials Are Stored in a Plaintext Configuration File

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:260-268
Vulnerability Type: Plaintext storage of authentication credentials
Risk Level: Medium

Vulnerable Code

yaml
# Location: ~/.secretary/config.yaml
platform: feishu
tokens:
  feishu: "xxx"
  dingtalk: "xxx"
  wechat: "xxx"

Technical Analysis

The documented configuration places Feishu, DingTalk, and WeChat bot tokens directly in a plaintext YAML file under the user's home directory. The instructions do not require restrictive file permissions, encryption, secret-manager integration, or controls preventing the file from being included in source repositories and backups.

These tokens are bearer credentials. Anyone who obtains the configuration file may be able to authenticate as the associated bot without knowing an additional password. Although the repository only contains placeholder values, following the documented configuration process would result in real credentials being stored in this format.

Attack Path

  1. A user follows the documented setup process and writes real bot tokens to ~/.secretary/config.yaml.
  2. The file is created with permissive permissions, copied into a backup, included in a support archive, or accidentally committed to a repository.
  3. An attacker with access to that file extracts the bearer tokens.
  4. The attacker reuses the tokens against the relevant messaging-platform APIs.
  5. The attacker performs operations permitted to the compromised bot until the credentials are revoked.

Impact Assessment

Exploitation requires read access to the configuration file or a copy of it. It does not directly provide operating-system privilege escalation. The resulting platform access is limited to the permissions assigned to each compromised bot token, but may include sending messages, accessing bot-visible resources, impersonating trusted automation, or triggering workflows connected to the bot.

Remediation
View remediation

Remediation Suggestions

  • Store credentials in an operating-system keychain, managed secret store, or deployment-platform secret mechanism.
  • Prefer references to environment variables over literal token values in YAML.
  • If a local configuration file is unavoidable, create it with mode 0600 and verify ownership before reading it.
  • Maintain a separate example file containing placeholders only.
  • Add the real configuration path to version-control ignore rules.
  • Prevent credentials from appearing in logs, exception messages, generated reports, and support archives.
  • Document credential rotation and revocation procedures.
  • Grant every bot only the minimum platform permissions necessary for its intended functions.

T09 · Insecure Skill Coding Practices

Warning
Location
CONTEXT_MANAGER.md:27
Finding

Permanent Conversation Memory Is Designed Without Data-Protection Controls

Content
View full analysis

Vulnerability Details

File Location: CONTEXT_MANAGER.md:27-35
Vulnerability Type: Unprotected persistent storage of potentially sensitive conversation data
Risk Level: Medium

Vulnerable Code

python
class LongTermMemory:
    storage = "SQLite"
    
    def store(self, memory):
        # Important information is stored permanently
        db.execute(
            "INSERT INTO memories VALUES (?, ?, ?)",
            (memory.content, memory.timestamp, memory.tags)
        )
    
    def retrieve(self, query):
        # Semantic search
        return db.search(query)

Technical Analysis

The long-term memory design permanently stores complete memory content, timestamps, and tags in SQLite. It does not define encryption at rest, database permissions, user isolation, sensitive-data filtering, consent, expiration, deletion, or retention limits.

Parameterized insertion reduces SQL-injection risk, but it does not protect the confidentiality of stored records. Assistant conversations may contain personal data, customer information, contact details, schedules, confidential business material, or credentials pasted by users.

The permanent-storage design also conflicts with the privacy claims in README.md:183-187, which state that there is no data collection and that user habit data is encrypted. No encryption implementation is present in the supplied artifact.

This is a documented design rather than an active implementation in the bundled Python files. The risk becomes exploitable if the described persistent-memory module is implemented or deployed without additional controls.

Attack Path

  1. Persistent memory is implemented according to the documented design.
  2. A user submits sensitive information during normal assistant interactions.
  3. The application writes the content permanently to the SQLite database.
  4. The database is exposed through weak filesystem permissions, host ...[truncated 620 chars]
Remediation
View remediation

Remediation Suggestions

  • Keep conversation context in bounded process memory by default.
  • Require explicit user consent before enabling persistent memory.
  • Apply a documented retention period rather than permanent storage.
  • Encrypt records at rest using keys stored separately from the database.
  • Restrict the database and parent directory to the service account that requires access.
  • Isolate records by user or tenant and enforce authorization on every retrieval.
  • Detect and redact credentials, authentication tokens, and other high-risk secrets before storage.
  • Provide user-accessible export, selective deletion, and complete deletion controls.
  • Record only the minimum information necessary for the feature.
  • Update privacy documentation so that it accurately describes collection, retention, encryption, and deletion behavior.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:374
Finding

Installation Instructions Execute Unpinned Code Outside the Audited Artifact

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:374-384
Additional Locations: README.md:75-86, requirements.txt:1
Vulnerability Type: Unverified and unpinned software supply chain
Risk Level: Medium

Vulnerable Code

bash
# Install from the package registry
pip install secretary-core

# Install from a mutable source repository
git clone https://github.com/pengong101/secretary-core
cd secretary-core
pip install -e .
text
numpy>=1.20.0

Technical Analysis

The supplied artifact does not contain the documented secretary_core.py implementation or setup.py, even though the installation and usage documentation refers to them. Consequently, the recommended package-registry and repository installation paths retrieve and execute content that was not included in this audit.

The Git installation instructions use a mutable branch rather than a reviewed commit or signed release. The package-registry command does not pin a reviewed version or verify artifact hashes. The NumPy dependency specifies only a lower bound, permitting future releases to be selected automatically.

No evidence of dependency confusion, typosquatting, or a currently malicious upstream package was found. The risk is that mutable or compromised upstream content could differ from the reviewed artifact and that Python build or installation hooks may execute with the installing user's privileges.

Attack Path

  1. A user follows the documented pip install or git clone instructions.
  2. The installer retrieves the current upstream package, repository branch, build dependencies, and transitive dependencies.
  3. An upstream account, package release, repository branch, or dependency is compromised or replaced after this audit.
  4. Python installation hooks or imported runtime code execute the modified content.
  5. The malicious code operates with the filesystem, network, and process privileges of the us ...[truncated 523 chars]
Remediation
View remediation

Remediation Suggestions

  • Include the actual documented implementation and packaging files in the reviewed artifact.
  • Pin package installations to reviewed, immutable versions.
  • Pin Git installations to a specific commit digest rather than a mutable branch.
  • Publish and verify cryptographic hashes for release artifacts.
  • Use a lock file with exact direct and transitive dependency versions.
  • Use hash-checked installation, such as pip --require-hashes, where practical.
  • Add an upper bound or exact reviewed version for NumPy and test upgrades before release.
  • Verify repository ownership, package-registry ownership, release signatures, and source provenance.
  • Avoid installing as an administrator and perform builds in an isolated environment.
  • Ensure the documentation and registry metadata identify the same version and implementation.

other

Warning
Location
secretary_efficiency_v1.py:193
Finding

Assistant Reports External Actions as Completed Without Executing Them

Content
View full analysis

Vulnerability Details

File Location: secretary_efficiency_v1.py:193-201
Additional Locations: secretary_v3.0.0.py:421-426, INTENT_UNDERSTANDING.md:15-16
Vulnerability Type: Misleading action-status spoofing
Risk Level: Medium

Vulnerable Code

python
def _concise_response(self, text: str, intent: IntentResult) -> str:
    """Generate a concise response"""
    if intent.intent_type == IntentType.MEETING:
        return f"Okay, the meeting has been scheduled. {self._format_entities(intent.entities)}"
    elif intent.intent_type == IntentType.REMINDER:
        return f"Okay, the reminder has been set. {self._format_entities(intent.entities)}"
    elif intent.intent_type == IntentType.EMAIL:
        return "Okay, please provide the email content."
    else:
        return "Okay, it has been executed."

The later implementation similarly reports immediate processing without invoking an integration:

python
if intent == IntentType.COMMAND:
    return {
        'content': f"Okay, I will process this immediately: {message}",
        'action_required': True,
        'style': style
    }

The design document explicitly prescribes this behavior:

text
Response strategy: Execute immediately and confirm the result.

Technical Analysis

The bundled implementations classify text, retain in-memory context, and generate response strings. They contain no calendar, reminder, email, or messaging API invocation that could perform the actions being reported.

Despite this, the response generator states that meetings were scheduled, reminders were set, and generic operations were executed. The action_required flag in the later implementation confirms that an action remains outstanding, but the user-facing text implies active or completed processing.

This creates a spoofed success state: the interface presents a legitimate-looking completion message without ...[truncated 1216 chars]

Remediation
View remediation

Remediation Suggestions

  • Never state that an external action succeeded unless a tool or API returned a verified success result.
  • Clearly distinguish among proposed, awaiting confirmation, in progress, completed, and failed states.
  • For the current local-only implementation, use wording such as “I identified this as a meeting request, but no calendar integration is configured.”
  • Require explicit confirmation before consequential operations where appropriate.
  • Invoke only authenticated, least-privileged calendar, messaging, or reminder integrations.
  • Validate tool responses and include a verifiable event ID, message ID, reminder ID, or timestamp in success confirmations.
  • Propagate API failures to the user instead of replacing them with generic success text.
  • Add automated tests asserting that completion language cannot be emitted without a successful tool result.
  • Maintain an action audit log containing the requested operation, confirmation state, tool result, and final status while excluding secrets.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (43)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

该描述显著夸大了代码能力。代码的主要功能是一个轻量级会话秘书原型:关键词匹配意图识别、基础实体抽取、简单双风格回复、有限上下文和用户习惯学习。虽然与“助理/秘书”方向相关,但声明中的多个核心卖点未在代码中体现,尤其是20轮上下文、情感识别、主动提醒机制以及飞书/钉钉/企业微信集成,均缺失或仅停留在文本层面的模拟。因此描述与实际行为存在明显不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

声明将该技能描述为“核心智能助理”,重点能力包括长上下文对话、情感识别、主动提醒、日程管理和企业通讯平台集成。但实际代码片段的可见功能较窄:它在父类基础上新增实体识别模式、时间表达式提取,以及根据意图和实体生成最多 3 条“预判行动”建议,并把这些建议附加到响应文本中。这里的“主动预判”更像文本建议,不等同于真正创建提醒、管理日程或调用外部系统。代码中也没有看到飞书、钉钉、企业微信 API 调用、网络访问、消息推送、日历操作、持久化会话管理或情感分析逻辑。虽然某些能力可能隐藏在未提供的父类 EfficientSecretary 中,但就该代码块本身而言,已展示的行为与声明的核心能力存在明显落差,因此应判定为描述与实际行为不完全匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

声明描述的是一个通用智能助理核心能力集合,重点在多轮上下文、情感识别、提醒、日程管理和企业 IM 集成。但代码并未展示这些核心能力:没有会话记忆窗口或 20 轮上下文机制,没有任何情感分析逻辑,没有提醒发送、定时任务或外部消息接口,也没有飞书/钉钉/企业微信 API 调用。相反,代码的主要功能是从文本中抽取工作需求、按优先级维护工作计划、接收 intelligence_report 生成额外任务,以及输出报告和流程优化建议。这些功能更像“工作规划/任务管理秘书”,与声明的主要能力集存在实质偏差,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

该技能描述与代码行为存在明显不一致。代码的主要功能是对输入文本做情感、意图和关系的简单规则匹配,并生成固定模板回复,这与“智能助理核心技能”的部分表述(如情感识别)相符。但描述中几个关键能力未被实现:一是“20轮对话上下文”与实际deque(maxlen=10)不符;二是不存在主动提醒或日程管理相关数据结构、调度机制或接口;三是没有任何飞书、钉钉、企业微信集成代码。整体上属于功能夸大,描述覆盖的关键能力明显超出实际实现。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

该代码与“智能助理核心技能”总体方向一致,确实实现了 20 轮上下文和基础情感识别,也有一定的任务/计划数据结构支持。但声明中的关键能力存在明显夸大或缺失:首先,‘集成飞书/钉钉/企业微信’完全没有落地,代码没有网络访问、Webhook、SDK、认证或任何平台适配逻辑;其次,‘主动提醒’只体现在基于关键词和时间的建议文案,不会真正创建提醒或主动触达用户;再次,‘日程管理’实现较弱,仅是内存中的工作计划对象操作,不足以支撑通常理解的完整日程管理。另一方面,代码实际还实现了用户习惯学习,这是描述中未提及的附加能力。综合判断,描述未准确反映代码实际能力,属于实质性不匹配。

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a secretary core skill with reminder, calendar/schedule management, and enterprise platform integration capabilities. The actual code only performs local emotion/intent/relationship heuristics, stores recent utterances, and returns text responses; it contains no reminder scheduling, calendar operations, or integrations with 飞书/钉钉/企业微信.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The design explicitly describes permanent storage of user memories in SQLite with no mention of consent, retention limits, deletion controls, encryption, or access restrictions. In an assistant skill that handles conversational context and personal preferences, indefinite storage increases privacy risk, enables over-collection of sensitive data, and raises the impact of compromise or misuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module describes behavioral preference learning plus scenario and relationship inference from language patterns and history, but provides no transparency, consent, or guardrails around profiling. This is dangerous because it can silently infer sensitive attributes or social context, influence assistant behavior in opaque ways, and create privacy, bias, and manipulation risks if the inferences are wrong or abused.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

命令型意图使用“帮我、立即、现在、去”等高频日常词作为触发特征,缺少约束条件、置信度阈值和二次确认机制,容易把普通对话误判为可执行指令。在该技能具备日程管理和外部集成能力的上下文中,这种过度分类可能导致未经充分确认的实际操作。

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The ambiguous-intent examples include very short and underspecified phrases like "那个..." and "你知道的", but the document does not define when such inputs should remain unclassified versus trigger clarification logic. Without explicit scope or thresholds, ordinary incomplete utterances may be swept into this category too broadly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

示例展示系统可直接“已为您预定明天 14:00 的会议室”,并进一步询问是否通知参会人员,但未说明权限校验、冲突检测、审计记录或对外通知前的确认流程。对于接入飞书/钉钉/企业微信的秘书技能,这类默认自动化会修改日程和资源状态,可能造成误预定、越权操作或错误通知。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README advertises proactive actions such as preparing materials, auto-classifying emails, showing client information, and sending reminders, but it does not define user consent, approval prompts, scope limits, or authorization boundaries. In an assistant integrated with enterprise messaging and scheduling systems, this can lead to unintended actions or exposure of sensitive business data based only on inferred intent or context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The examples normalize automatic meeting booking and attendee notification from short or ambiguous user input without any safety warning or confirmation step. In a secretary/automation skill tied to workplace systems, this increases the risk of mistaken bookings, unauthorized notifications, privacy leaks, or abuse through prompt injection or misunderstood context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The release notes describe 7-day habit learning, interaction recording, active-hours analysis, and preference modeling, but provide no notice about what behavioral data is stored, how long it is retained, or whether the user consents to this profiling. In an assistant skill that integrates workplace platforms, this can silently create persistent behavioral profiles from routine conversations and metadata, increasing privacy and surveillance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The predictive suggestions section includes actions such as sending reminders, booking meeting rooms, helping reply to emails, and potentially canceling meetings, yet the document does not state that these are gated behind confirmation or consent. In a secretary-style assistant connected to enterprise messaging and scheduling systems, proactive or mis-triggered actions can affect communications, calendars, and operations on the user's behalf.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The known-issues section says the skill currently supports only Chinese. A forced language constraint is a natural-language policy concern unless the tool offers language choice or clearly documents a justified region-specific limitation, which this file does not.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill advertises proactive reminders and external message delivery across enterprise messaging platforms without warning that user content, schedules, and recipient data may be transmitted to third-party services. This can cause unintended disclosure of sensitive business or personal information, especially in assistant workflows that act on conversational context automatically.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill claims calendar create, edit, and delete capabilities but does not warn users about the risk of modifying live schedule data, attendee lists, locations, and reminders. In an assistant context, silent or poorly understood write actions can lead to data integrity issues, missed meetings, or unauthorized changes to business workflows.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest description is written as a Chinese-only skill description ('智能助理核心技能'), which indicates a language-specific presentation without any stated user choice or opt-in. Under the policy rule, forcing a specific language or locale without explicit user selection can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill’s title, comments, trigger patterns, and all user-facing responses are written in Chinese, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-only region or audience. This is a natural-language locale policy concern because it imposes a specific language by default rather than offering choice or justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The stated purpose includes 情感识别, yet intent understanding is implemented only via keyword scoring, simple context boosting, and entity extraction. There is no emotion taxonomy, sentiment detection, or affect-driven response behavior anywhere in the file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill appends raw user text to a habit-learning buffer and also learns contacts from extracted entities without any consent, notice, minimization, or retention control. In a secretary skill, users are likely to provide sensitive scheduling, contact, and workplace information, so silent collection increases privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill tells users that meetings and reminders have been arranged or set, but the code only generates reassuring text and never calls any calendar, reminder, or messaging backend. In an assistant context, this can mislead users into relying on actions that were never performed, causing missed meetings, missed reminders, or operational disruption.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code stores raw user inputs, extracted entities, and timestamps in conversation history, while also retaining a separate habit-learning buffer. In a secretary assistant context, this naturally captures potentially sensitive personal and business data such as names, meeting times, locations, and tasks, which could later be exposed through debugging, logs, memory inspection, or downstream processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The class docstring labels the skill in Chinese, and the surrounding natural-language strings and examples indicate the skill is designed to operate only in Chinese. Under the policy rule, forcing a specific language without opt-in or a documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.