Back to skill

Security audit

Test Template Starter Pack

Security checks for vulnerabilities and agentic risk

Overview

This looks like a local agent-template/test package, but it overstates production readiness and its test runner can write and delete fixed database files outside the skill directory.

Review before installing or running. Treat this as a prototype/test template, not a production Telegram agent. Run tests only in an isolated workspace, set FUTURE_WORKSPACE to a disposable directory, and add proper secret handling, webhook verification, access controls, logging limits, and data-retention rules before connecting real CRM, payment, tax, legal, healthcare, or customer data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
_test-template.py:37
Finding

Predictable Test Database Paths Can Overwrite and Delete Existing Files

Content
View full analysis

Vulnerability Details

File Location: _test-template.py, lines 37-39, 425-449, 463-494, 506-528, and 535-555
Vulnerability Type: Unsafe use and deletion of predictable filesystem paths
Risk Level: Medium

The test suite creates SQLite databases under a shared external workspace using fixed filenames. After each test, it unconditionally deletes the selected file if it exists.

Vulnerable Code

Workspace and data-directory selection:

python
WORKSPACE = Path(os.environ.get("FUTURE_WORKSPACE", "/root/.openclaw/workspace-future"))
DATA_DIR = WORKSPACE / "data"
DATA_DIR.mkdir(exist_ok=True)

Tier-gating test:

python
def test_tier_gate(r: TestResult):
    """Test tier gating logic."""
    db_path = DATA_DIR / "test-template-tier.db"
    conn = init_db(db_path)
    cur = conn.cursor()

    # Test operations omitted

    conn.close()
    if db_path.exists():
        db_path.unlink()

Analytics test:

python
def test_analytics(r: TestResult):
    """Test analytics engine."""
    db_path = DATA_DIR / "test-template-analytics.db"
    conn = init_db(db_path)
    cur = conn.cursor()

    # Test operations omitted

    conn.close()
    if db_path.exists():
        db_path.unlink()

Persistence test:

python
def test_db_persistence(r: TestResult):
    """Test database operations."""
    db_path = DATA_DIR / "test-template-persist.db"
    conn = init_db(db_path)
    cur = conn.cursor()

    # Test operations omitted

    conn.close()
    if db_path.exists():
        db_path.unlink()

Follow-up test:

python
def test_follow_up_pipeline(r: TestResult):
    """Test follow-up sequence scheduling."""
    db_path = DATA_DIR / "test-template-followup.db"
    conn = init_db(db_path)
    cur = conn.cursor()

    # Test operations omitted

    conn.close()
    if db_path.exists():
        db_path.unlink()

Tech

...[truncated 2275 chars]

Remediation
View remediation

Remediation Suggestions

  1. Create a unique temporary directory for every test run:

    python
    import tempfile
    from pathlib import Path
    
    with tempfile.TemporaryDirectory(prefix="test-template-") as temp_dir:
        db_path = Path(temp_dir) / "tier.db"
        conn = init_db(db_path)
        try:
            # Run the test
            pass
        finally:
            conn.close()
    
  2. Prefer an in-memory SQLite database when filesystem persistence is not specifically under test:

    python
    conn = sqlite3.connect(":memory:")
    
  3. For persistence tests, use tempfile.NamedTemporaryFile() or a unique UUID-based path inside a private temporary directory.

  4. Never unlink a path merely because it exists. Track whether the current process created the file, and only remove artifacts owned by the current test run.

  5. Create temporary directories with restrictive permissions and avoid shared or privileged workspace locations.

  6. Use try/finally or context managers so database connections and temporary resources are safely cleaned up even when assertions or database operations fail.

  7. Avoid the hard-coded /root/.openclaw/workspace-future default for tests. Test artifacts should not be written to an application data directory or require elevated privileges.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The description overstates the code’s purpose and capabilities. The supplied code is a reusable local test template for prototyping or validating a generic vertical-agent architecture, not a production-ready Telegram agent builder. It does include supporting pieces aligned with the description—intent classification, scoring, routing, follow-up scheduling, analytics, and persistence—but only as simple local template logic with SQLite and tests. There is no Telegram API usage, no real CRM integration, no deployment automation, and no evidence of actual production agent operation in this chunk. Therefore the declared description does not accurately represent what this code actually does.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The guide explicitly tells users to replace mock data with real CRM, calendar, payment, and webhook integrations, but provides no guidance on protecting secrets, validating webhook authenticity, minimizing collection of personal data, or securing storage and transport. In the context of Telegram/CRM/payment-connected agents handling customer records, this omission can lead builders to deploy insecure integrations that expose sensitive user data or credentials.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest advertises a production-ready Telegram vertical AI agent template with CRM, routing, follow-up, and analytics. In this file, the implemented behavior is limited to local SQLite-backed test utilities, keyword intent classification, scoring, routing, analytics, and a CLI test runner; there is no Telegram API interaction, no external CRM integration, and no production deployment logic.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The intent patterns include extremely generic phrases such as "help", "explain", and "what is", which can cause the agent to activate or route into FAQ handling on ordinary conversation that was not meant for this skill. In a Telegram CRM/lead-management agent, this increases the chance of unintended invocation, misclassification, noisy automation, and accidental processing of user or lead data in the wrong workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The brief is explicitly framed as a Russian assistant and all sample templates, reminders, and user-facing phrases are written in Russian, with no indication that users can choose another language or locale. This is a natural-language policy concern because the skill appears to enforce a specific language experience without documenting user choice or opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The FAQ intent includes very broad trigger terms such as "как", "что", and "help", which are common in ordinary conversation and can cause misclassification of unrelated user messages. In a tax assistant, this can route users into the wrong workflow, produce incorrect tax guidance, or bypass more specific intent handling, increasing the chance of unsafe or misleading automation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The guide presents intent matching examples that assume only Russian and English inputs, but it does not state that language support is configurable or user-selectable. This can amount to a language/locale policy issue because the skill behavior appears biased toward fixed languages without explicit opt-in or choice.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The banner says 'CORE ENGINE — Do not modify per vertical', which implies these components are fixed. However, several class docstrings and comments immediately below state that methods such as classification, scoring, entity extraction, and urgency/value logic should be overridden per vertical, contradicting the stated immutability of the core engine.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.