Back to skill

Security audit

Sentinel - AI Agent State Guardian

Security checks for vulnerabilities and agentic risk

Overview

Sentinel looks like a legitimate local backup and restore skill, but its file restore safeguards are too weak for a tool that can overwrite workspace state.

Install only if you are comfortable granting a local tool authority to read, back up, and overwrite the configured workspace. Keep it unprivileged, avoid running it as a system service at first, disable automatic restore until tested, keep backups outside shared locations, and do not monitor paths containing untrusted symlinks or traversal-prone patterns.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
sentinel_restore.py:284
Finding

Workspace Boundary Bypass Through Path Traversal and Symlink Following

Content
View full analysis

Vulnerability Details

File Location: sentinel.py:249-263, sentinel.py:169-185, sentinel_restore.py:42-59, sentinel_restore.py:107-121, and sentinel_restore.py:284-290
Vulnerability Type: Path traversal and unsafe symlink handling
Risk Level: High

Vulnerable Code

python
# sentinel.py:249-263
def scan_critical_files(self) -> List[Path]:
    """Discover all critical files to monitor."""
    critical_files = []
    
    for pattern in self.config.CRITICAL_FILES:
        if '*' in pattern or '?' in pattern:
            # Glob pattern
            try:
                matches = list(self.workspace.glob(pattern))
                critical_files.extend(matches)
            except Exception as e:
                self.logger.warning(f"Failed to glob pattern {pattern}: {e}")
        else:
            # Exact path
            filepath = self.workspace / pattern
            if filepath.exists():
                critical_files.append(filepath)
python
# sentinel.py:169-185
def backup_file(self, filepath: Path, workspace_root: Path) -> Optional[Path]:
    """Create timestamped backup of a file."""
    try:
        # Create backup path preserving directory structure
        rel_path = filepath.relative_to(workspace_root)
        timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
        backup_path = self.backup_root / timestamp / rel_path
        
        # Create parent directories
        backup_path.parent.mkdir(parents=True, exist_ok=True)
        
        # Copy file
        shutil.copy2(filepath, backup_path)
        self.logger.debug(f"Backed up: {filepath} -> {backup_path}")
        return backup_path
    except Exception as e:
        self.logger.error(f"Backup failed for {filepath}: {e}")
        return None
python
# sentinel_restore.py:284-290
filepath = tool.workspace / args.file

if not filepath.is_relative_to(tool.wo
...[truncated 3977 chars]
Remediation
View remediation

Remediation Suggestions

  1. Resolve the workspace once and resolve every candidate source and destination before use:
    python
    workspace = Path(config_obj.WORKSPACE_ROOT).resolve(strict=True)
    candidate = (workspace / user_path).resolve(strict=False)
    if candidate != workspace and workspace not in candidate.parents:
        raise ValueError("Path escapes workspace")
    
  2. Apply equivalent canonical containment checks to backup sources using the resolved backup root.
  3. Explicitly reject symbolic links in monitored files and restore destinations when links are not required. Use lstat() and inspect every path component where appropriate.
  4. Revalidate containment and symlink status immediately before opening or copying a file to reduce time-of-check/time-of-use exposure.
  5. Avoid writing through an existing destination path that may be a symlink. On supported platforms, use safe descriptor-based operations with no-follow semantics.
  6. Validate CRITICAL_FILES during initialization and reject absolute paths, traversal components, and paths that resolve outside the workspace.
  7. Run Sentinel under a dedicated, least-privileged operating-system account with access only to the intended workspace and backup directory.
  8. Add regression tests covering ../ traversal, absolute paths, source symlinks, destination symlinks, nested symlink components, and backup-directory escape attempts.

T09 · Insecure Skill Coding Practices

Warning
Location
sentinel.py:258
Finding

Integrity Monitoring Fails Open for Deleted, Permission-Changed, and Skipped Files

Content
View full analysis

Vulnerability Details

File Location: sentinel.py:142-150, sentinel.py:215-240, sentinel.py:258-263, sentinel.py:278-281, and sentinel.py:370-379
Vulnerability Type: Fail-open integrity monitoring and incomplete state validation
Risk Level: Medium

Vulnerable Code

python
# sentinel.py:142-150
def update_file(self, filepath: Path, file_hash: str, size: int, mtime: float):
    """Update manifest entry for a file."""
    key = str(filepath)
    self.manifest[key] = {
        'hash': file_hash,
        'size': size,
        'mtime': mtime,
        'last_checked': time.time()
    }
python
# sentinel.py:258-263
else:
    # Exact path
    filepath = self.workspace / pattern
    if filepath.exists():
        critical_files.append(filepath)

return list(set(critical_files))  # Remove duplicates
python
# sentinel.py:278-281
# Skip if file is in use
if ProcessDetector.is_file_in_use(filepath):
    self.logger.debug(f"Skipping {filepath}: file in use")
    return True
python
# sentinel.py:370-379
for filepath in critical_files:
    if self.monitor_file(filepath):
        success_count += 1
    else:
        failure_count += 1

# Save manifest
if self.manifest.save():
    self.logger.debug("State manifest saved")
else:
    self.logger.error("Failed to save state manifest")

The Unix lock probe also opens files in read/write mode:

python
# sentinel.py:110-121
else:
    # On Unix, check if we can get exclusive access
    try:
        with open(filepath, 'r+b') as f:
            import fcntl
            fcntl.flock(f.fileno(), fcntl.LOCK_EX | fcntl.LOCK_NB)
            fcntl.flock(f.fileno(), fcntl.LOCK_UN)
        return False
    except (IOError, OSError, ImportError):
        # If fcntl not available or file is locked, assume in use
        return True

Technical Analysis

Exact ...[truncated 2875 chars]

Remediation
View remediation

Remediation Suggestions

  1. Reconcile the current scan with all prior manifest entries during every cycle. If a previously tracked file is absent, report an explicit deletion violation and apply the configured recovery policy.
  2. Retain configured exact paths in the scan plan even when they do not exist so missing-file checks can execute.
  3. Store and compare permission mode, ownership, file type, and other required metadata using lstat() where supported.
  4. Honor SKIP_FILES_IN_USE. When a check is skipped, return an indeterminate or failed status rather than success.
  5. Distinguish lock contention from permission denial, missing files, unsupported locking, and other I/O failures. Do not map all exceptions to “file in use.”
  6. Avoid requiring write access merely to inspect locking. If advisory lock checking remains necessary, document its limitations and use a read-only-compatible approach where possible.
  7. Emit warning or error alerts for files that remain unchecked across one or more cycles.
  8. Add tests for deleted exact paths, deleted glob matches previously present in the manifest, read-only files, permission-only changes, inaccessible files, and actively locked files.

T09 · Insecure Skill Coding Practices

Warning
Location
sentinel.py:297
Finding

Changed Content Is Promoted and Backed Up Before Corruption Recovery

Content
View full analysis

Vulnerability Details

File Location: sentinel.py:297-320 and sentinel.py:326-340
Vulnerability Type: Backup poisoning and improper trusted-baseline management
Risk Level: Medium

Vulnerable Code

python
# sentinel.py:297-320
else:
    # Check for changes
    if current_hash != previous_state['hash']:
        self.logger.info(f"Change detected: {filepath}")
        
        # Create backup before updating manifest
        backup_path = self.backup_manager.backup_file(filepath, self.workspace)
        if backup_path:
            self.logger.info(f"Backup created: {backup_path}")
        
        # Update manifest
        self.manifest.update_file(filepath, current_hash, stat.st_size, stat.st_mtime)
        
        self.logger.alert('INFO', f"File changed: {filepath}")

# Check integrity
is_ok, issues = self.integrity_checker.check_file(filepath, previous_state or {})
if not is_ok:
    self.logger.warning(f"Integrity issues in {filepath}: {', '.join(issues)}")
    self.logger.alert('WARNING', f"Integrity issue: {filepath} - {', '.join(issues)}")
    
    # Auto-restore if configured and file is corrupted
    if self.config.AUTO_RESTORE_ON_CORRUPTION and 'empty' in ' '.join(issues).lower():
        self.attempt_restore(filepath)
python
# sentinel.py:326-340
# Find latest backup
backup_path = self.backup_manager.get_latest_backup(filepath, self.workspace)
if backup_path is None:
    self.logger.error(f"No backup found for {filepath}")
    self.logger.alert('ERROR', f"Auto-restore failed: no backup for {filepath}")
    return False

# Verify backup is not empty
if backup_path.stat().st_size == 0:
    self.logger.error(f"Backup is empty: {backup_path}")
    return False

# Restore
shutil.copy2(backup_path, filepath)
self.logger.info(f"Restored from backup: {backup_path} -> {filepath}")
self.logger.alert('INFO', f"Auto-restored {filepath} 
...[truncated 2933 chars]
Remediation
View remediation

Remediation Suggestions

  1. Separate the trusted baseline from the current observed state. Never overwrite the trusted hash merely because a change was observed.
  2. Validate a changed file according to policy before accepting it. Require explicit administrative approval for unexpected modifications when semantic validation is unavailable.
  3. Keep a known-good backup corresponding to the trusted manifest state. Record and verify its cryptographic hash.
  4. Do not create the recovery candidate from content already classified as corrupted. Store changed content in a separate quarantine or untrusted snapshot area if preservation is required.
  5. During restoration, iterate from newest to oldest and select the newest backup that passes size, hash, readability, and any format/schema validation.
  6. After restoration, hash the restored destination and verify that it matches the selected trusted backup before updating state.
  7. Use atomic writes for manifest updates and restored files so interruption cannot leave partially updated trust state.
  8. Define explicit policies for authorized changes, unexpected non-empty changes, empty files, missing files, and backup validation failures.
  9. Add regression tests proving that truncation restores an older valid backup and that unauthorized non-empty content is not silently promoted into the trusted baseline.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

A second description-behavior mismatch is present around claims of automated backup, integrity monitoring, and recovery, while the documented operations require explicit invocation and user interaction for some restore actions. In security-sensitive backup tooling, overstating automation can produce dangerous assumptions about resilience and tamper detection.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A second description-behavior mismatch is present around claims of automated backup, integrity monitoring, and recovery, while the documented operations require explicit invocation and user interaction for some restore actions. In security-sensitive backup tooling, overstating automation can produce dangerous assumptions about resilience and tamper detection.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · config_example.py (reported line 72)May include surrounding context.

python
# ============================================================================

# File where Sentinel stores its state manifest (file hashes, last check times)
# This file is critical for change detection — DO NOT DELETE
STATE_FILE = "/path/to/sentinel_state.json"

# ============================================================================

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 193)May include surrounding context.

macOS (launchd):

xml
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 194)May include surrounding context.

macOS (launchd):

xml
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 209)May include surrounding context.

macOS (launchd):

xml
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 193)May include surrounding context.

macOS (launchd):

xml
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises and demonstrates filesystem read/write operations but does not declare any explicit tool scope such as permissions or allowed-tools. This creates an authorization ambiguity where an agent may invoke file-capable behavior without clear least-privilege boundaries, increasing the risk of unintended workspace modification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill describes restore and quarantine behaviors that can overwrite, move, or revert workspace files but does not prominently warn about the risk of unintended data loss or rollback of legitimate changes. In a state-management skill, such silent modification capability is especially dangerous because users may enable it expecting protection while actually introducing destructive automated behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The example configuration enables automatic restoration on corruption without a clear warning that recent legitimate user changes may be reverted. Because users often copy configuration examples verbatim, this can normalize unsafe defaults and lead to unexpected loss of work.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The integrity violation response section lists auto-restore and quarantine as response options but omits any warning that those actions alter file locations or contents. In the context of an AI agent workspace, automated mutation of state files can corrupt workflows, hide malicious tampering, or destroy forensic evidence if performed prematurely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The auto-restore logic copies a backup over the current workspace file without any confirmation, approval gate, or atomic safety check at the point of overwrite. In an agent workspace, this can destroy legitimate recent changes or allow rollback to attacker-influenced backup content if the backup set or corruption signal is manipulated, making integrity recovery itself a destructive action.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest intentionally records detailed workspace metadata including relative file paths, sizes, modification times, permissions, and sometimes file hashes, then writes it to disk as JSON. Even without file contents, this metadata can reveal sensitive project structure, secret-bearing filenames, operational timelines, and integrity fingerprints; if the manifest is stored insecurely or shared, it can aid reconnaissance or leak private information.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · sentinel_restore.py (reported line 282)May include surrounding context.

python
parser.add_argument('--latest', action='store_true', help='Restore latest backup')
    parser.add_argument('--interactive', action='store_true', help='Interactive backup selection')
    parser.add_argument('--list', action='store_true', help='List all backups')
    parser.add_argument('--auto', action='store_true', help='Skip confirmation prompts')
    
    args = parser.parse_args()

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · README.md (reported line 260)May include surrounding context.

md
**USE AT YOUR OWN RISK.**

- The author(s) are NOT liable for any damages, losses, or consequences arising from 
  the use or misuse of this software — including but not limited to financial loss, 
  data loss, security breaches, business interruption, or any indirect/consequential damages.
- This software does NOT constitute financial, legal, trading, or professional advice.
- Users are solely responsible for evaluating whether this software is suitable for

Static analysis

No suspicious patterns detected.