Back to skill

Security audit

Dev Factory

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent automated builder skill, but it grants remote/model-driven code, shell, credential, and publishing capabilities without enough containment or approval controls.

Only install or run this in a disposable container or VM with a minimal environment, no host credential files, restricted network egress, private repositories by default, and manual review before any shell command, Notion-driven task, source upload to GLM/Claude, GitHub push, or release creation. Use least-privilege Notion and GitHub tokens and avoid running the daemon against task queues writable by untrusted users.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
builder/symphony/glm5_agent.py:231
Finding

Unrestricted Model-Controlled Shell Execution and Workspace Escape

Content
View full analysis
str: """Run a command in the workspace Args: workspace: Workspace directory command: Command to run Returns: Command output """ import subprocess try: result = subprocess.run( command, shell=True, cwd=str(workspace), capture_output=True, text=True, timeout=120 ) output ...[truncated 2430 chars]
Remediation
View remediation

other

Error
Location
builder/symphony/notion_tracker.py:190
Finding

Indirect Prompt Injection from Notion Tasks into a Bash-Capable Coding Agent

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
builder/correction/fixer.py:27
Finding

Error-Controlled Arbitrary File Modification Outside the Project

Content
View full analysis
Tuple[Optional[str], Optional[int]]: """에러 위치 추출 (우선순위순 패턴 리스트)""" patterns = [ # 표준 CPython: File "path", line N [, in func] r'\s*File "([^"]+)", line (\d+)', # 짧은 형식: path:N: message r'([\w./\\-]+\.py):(\d+):', ] for pattern in patterns: match = re.search(pattern, output) if match: return match.group(1), int(match.group(2)) return None, None ``` The fixer accepts the extracted path without enforcing project containment: ```python file_path = error.get('file_path') line_number = error.get('line_number') if not file_path: logger.warning("No file path in error, cannot auto-fix") return False file_path = Path(file_path) if not file_path.exists(): logger.warning("File %s not found", file_path) return False # 규칙 기반 수정 시도 if self._fix_by_rules(error, file_path, line_number): return True ``` A rule-based correction writes directly to that path: ```python content = file_path.read_text() lines = content.splitlines() if line_number > len(lines): return False line = lines[line_number - 1] new_line = re.sub( rf"(\w+)\['{re.escape(key)}'\]", rf"\1.get('{key}', '')", line ) if new_line != line: lines[line_number - 1] = new_line file_path.write_text('\n'.join(lines)) logger.info("Fixed KeyError: changed ['%s'] to .get('%s', '') at line %d", key, key, line_number) return True ``` The Claude fallback deliberately retains an external absolute path: ```python try: rel_path = file_path.relative_to(project_path) except ValueError: rel_pat ...[truncated 2381 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
builder/symphony/workspace.py:21
Finding

Configurable Workspace Lifecycle Hooks Execute Arbitrary Shell Strings

Content
View full analysis
bool: """Run lifecycle hooks Args: workspace: Workspace to run hooks in hook: Hook type commands: List of commands to run Returns: True if all hooks succeeded, False otherwise """ if not commands: return True all_success = True for cmd in commands: try: logger.info("Running %s hook: %s", hook.value, cmd) result = subprocess.run( cmd, shell=True, cwd=str(workspace.path), capture_output=True, text=True, timeout=60 ) if result.returncode != 0: logger.warning( "Hook %s failed: %s\nstderr: %s", hook.value, cmd, result.stderr ) all_success = False else ...[truncated 1845 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
check_status.py:63
Finding

Status and Health Scripts Read the Entire Shared Credential File

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (258)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The description advertises automatic repository creation and publishing but does not prominently warn users that running the skill can create external artifacts and push generated code to GitHub. For an autonomous builder agent, omission of this warning creates a meaningful risk of unintended disclosure, reputational damage, and release of unsafe or sensitive code.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 341)May include surrounding context.

md
- Python 3.11+
- ChatDev 2.0
- GLM-5 API
- GitHub Personal Access Token
- Notion API

## 라이선스

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
94% confidence
Finding

The workflow executes a shell command with forceful recursive deletion as an automated hook. In agentic or orchestrated environments, shell-based cleanup is especially risky because incorrect working-directory assumptions, variable expansion in future edits, or malicious symlink placement inside the workspace can turn a routine cleanup into deletion of arbitrary filesystem content.

Content

Scanner excerpt · WORKFLOW.md (reported line 34)May include surrounding context.

md
after_create:
    - command: "python -m venv .venv"
  before_run:
    - command: "rm -rf dist/ build/"
  after_run:
    - command: "pytest --cov=. --cov-report=xml"
---

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · acp_self_correction.py (reported line 165)May include surrounding context.

python
Return the fixed code or explain what needs to be changed.
"""
        return prompt
    
    def _apply_fix_locally(self, error: Dict) -> Dict:
        """로컬에서 직접 수정 적용"""

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · builder/orchestrator.py (reported line 303)May include surrounding context.

python
Return the fixed code or explain what needs to be changed.
"""
        return prompt
    
    def _apply_fix_locally(self, error: Dict) -> Dict:
        """로컬에서 직접 수정 적용"""

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · builder/symphony/glm5_agent.py (reported line 349)May include surrounding context.

python
Return the fixed code or explain what needs to be changed.
"""
        return prompt
    
    def _apply_fix_locally(self, error: Dict) -> Dict:
        """로컬에서 직접 수정 적용"""

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · builder/symphony/glm5_agent.py (reported line 356)May include surrounding context.

python
Return the fixed code or explain what needs to be changed.
"""
        return prompt
    
    def _apply_fix_locally(self, error: Dict) -> Dict:
        """로컬에서 직접 수정 적용"""

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill executes project tests in a subprocess, which is a powerful code-execution capability not constrained by sandboxing or user confirmation. In this context the project under test may be adversarial, so test execution becomes a direct path to arbitrary code execution under the user's privileges.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
98% confidence
Finding

The subprocess environment is built from os.environ, exposing all parent-process environment variables to untrusted test code. If tests are malicious, they can read API keys, tokens, cloud credentials, and other secrets from the inherited environment.

Content

Scanner excerpt · acp_self_correction.py (reported line 277)May include surrounding context.

python
capture_output=True,
                text=True,
                timeout=30,
                env={**os.environ, 'PYTHONPATH': 'src'}
            )
            
            return {

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · builder/correction/base.py (reported line 45)May include surrounding context.

python
capture_output=True,
                text=True,
                timeout=timeout,
                env={**os.environ, 'PYTHONPATH': 'src'}
            )

            return {

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · hybrid_acp_correction.py (reported line 354)May include surrounding context.

python
capture_output=True,
                text=True,
                timeout=timeout,
                env={**os.environ, 'PYTHONPATH': 'src'}
            )

            return {

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · self_correction_engine.py (reported line 146)May include surrounding context.

python
capture_output=True,
                text=True,
                timeout=timeout,
                env={**os.environ, 'PYTHONPATH': 'src'}
            )

            return {

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This code repair component includes an undocumented network-exfiltration capability by transmitting full source code and error context to a third-party GLM API. In the context of an automated builder/fixer, that is especially dangerous because it can leak sensitive intellectual property or secrets from arbitrary repositories under repair.

Content

No source excerpt is available for this finding.

Tainted flow: 'req' from pathlib.Path.read_text (line 225, file read) → urllib.request.urlopen (network output)

High
Category
Data Flow
Confidence
98% confidence
Finding

The fixer sends full file contents and error context to an external API over the network. Because source files may contain proprietary code, secrets, credentials, or embedded prompt-injection content, this is a clear exfiltration channel and expands trust to a remote service without user approval.

Content

Scanner excerpt · builder/correction/fixer.py (reported line 234)May include surrounding context.

python
}
            )

            with urllib.request.urlopen(req, timeout=30) as response:
                result = json.loads(response.read().decode())
                content = result['choices'][0]['message']['content']

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The fixer delegates repair to an external CLI agent and explicitly authorizes Bash, extending its capabilities beyond straightforward code editing. Given that the prompt is built from untrusted repository/error content, this creates a powerful execution surface for prompt-injection-driven shell actions and unsafe repository mutation.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
92% confidence
Finding

The subprocess environment is populated by copying the entire parent process environment into the external Claude CLI invocation. That exposes all available environment variables, including tokens and credentials, to a third-party tool process that also has Bash capability, increasing the blast radius if the tool is compromised or manipulated.

Content

Scanner excerpt · builder/correction/fixer.py (reported line 288)May include surrounding context.

python
'claude', '-p', prompt,
                '--allowedTools', 'Edit,Write,Bash'
            ], cwd=str(project_path), capture_output=True, text=True,
               timeout=120, env={**os.environ,
                'CLAUDE_OUTPUT_DIR': str(project_path)})

            if result.returncode == 0:

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

The skill automatically reads a Notion API key from a workspace .env file without requiring explicit caller-provided credentials. In an agent context, implicit credential discovery increases the blast radius of the component: code that gains access to this class can silently reuse stored secrets to access or modify external Notion data.

Content

Scanner excerpt · builder/integration/notion_sync.py (reported line 27)May include surrounding context.

python
}

    def _load_token(self) -> Optional[str]:
        """.env 파일에서 Notion API 토큰 로드"""
        import os
        workspace = os.getenv("OPENCLAW_WORKSPACE", str(Path.home() / ".openclaw" / "workspace"))
        env_file = Path(workspace) / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

The code targets a predictable secret location under the workspace and parses .env contents for NOTION_API_KEY. In a multi-skill or agent environment, predictable secret harvesting patterns are risky because they normalize secret access from local storage and can enable unauthorized use of credentials if the component is invoked unexpectedly.

Content

Scanner excerpt · builder/integration/notion_sync.py (reported line 30)May include surrounding context.

python
""".env 파일에서 Notion API 토큰 로드"""
        import os
        workspace = os.getenv("OPENCLAW_WORKSPACE", str(Path.home() / ".openclaw" / "workspace"))
        env_file = Path(workspace) / ".env"
        if env_file.exists():
            for line in env_file.read_text().splitlines():
                if line.startswith('NOTION_API_KEY='):

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
80% confidence
Finding

Passing the full parent process environment into a tool-capable external agent exposes all inherited secrets and configuration to that subprocess. In this context, the Claude CLI is also granted Bash access, increasing the risk that prompt-influenced behavior can read and exfiltrate API keys, tokens, or internal settings from environment variables.

Content

Scanner excerpt · builder/orchestrator.py (reported line 157)May include surrounding context.

python
'--output-format', 'json',
                '--allowedTools', 'Edit,Write,Bash'
            ], cwd=str(project_path), capture_output=True, text=True,
               timeout=self.claude_timeout, env={**os.environ,
                'CLAUDE_OUTPUT_DIR': str(project_path)})

            if result.returncode == 0:

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
80% confidence
Finding

The fix path also forwards the complete environment to the external CLI, again exposing inherited secrets to an agent operating on untrusted prompts and project state. Because this flow is triggered during error handling, attackers may have additional opportunities to shape inputs that induce secret access or disclosure.

Content

Scanner excerpt · builder/orchestrator.py (reported line 238)May include surrounding context.

python
result = subprocess.run([
                'claude', '-p', prompt
            ], cwd=str(project_path), capture_output=True, text=True,
               timeout=120, env={**os.environ,
                'CLAUDE_OUTPUT_DIR': str(project_path)})

            if result.returncode == 0:

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
99% confidence
Finding

Merging os.environ into the child process environment explicitly propagates all current environment variables to code executed from the project under test. If that project or its tests are malicious, they can harvest CI secrets, service credentials, and tokens directly from the environment. In a pipeline that automatically builds and tests discovered/generated projects, this is a concrete secret-exposure path, not just a theoretical concern.

Content

Scanner excerpt · builder/pipeline.py (reported line 355)May include surrounding context.

python
capture_output=True,
                text=True,
                timeout=30,
                env={**os.environ, 'PYTHONPATH': 'src'}
            )

            return {

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Shell commands emitted by the model are executed automatically with no user-facing disclosure, review, or confirmation step. In this context, the model is remote and can generate arbitrary commands, so silent execution substantially raises the risk of destructive actions, dependency poisoning, credential access, or data exfiltration.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

The run_command tool accepts an unconstrained command parameter from model output and executes it with shell=True. This is a classic tool-parameter-abuse issue: an adversarial or compromised model can supply payloads that run arbitrary OS commands, pivot outside intended build actions, or exfiltrate data.

Content

Scanner excerpt · builder/symphony/glm5_agent.py (reported line 309)May include surrounding context.

python
import subprocess

        try:
            result = subprocess.run(
                command,
                shell=True,
                cwd=str(workspace),

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · builder/symphony/orchestrator.py (reported line 67)May include surrounding context.

python
"python -m venv .venv",
                ],
                WorkspaceHook.BEFORE_RUN: [
                    "rm -rf dist/ build/",
                ],
                WorkspaceHook.AFTER_RUN: [
                    "pytest --cov=. --cov-report=xml || true",

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

Using subprocess.run on a free-form command string with shell=True is a classic tool-parameter abuse pattern. Any untrusted data that reaches the hook command can inject shell metacharacters or entirely different commands, resulting in arbitrary execution, data exfiltration, or destructive filesystem actions.

Content

Scanner excerpt · builder/symphony/workspace.py (reported line 397)May include surrounding context.

python
try:
                logger.info("Running %s hook: %s", hook.value, cmd)

                result = subprocess.run(
                    cmd,
                    shell=True,
                    cwd=str(workspace.path),

Static analysis

No suspicious patterns detected.