T01 · Skill Instruction Hijacking
- Location
SOUL.md:5- Finding
Skill instructions override the host agent's identity and operational policy
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is safety-oriented, but it tries to govern the agent globally and run recurring autonomous logging/actions beyond its declared local scope.
Install only if you intentionally want SafeProactive to affect the agent globally. Before use, narrow activation away from every_turn, avoid unrestricted system prompt prepending, require approval for first-contact external APIs and physical actions, define log redaction rules, and align declared permissions with actual workspace, memory, network, and notification behavior.
SOUL.md:5Skill instructions override the host agent's identity and operational policy
The every_turn trigger causes the skill to activate on essentially all interactions, greatly expanding its execution surface and making unintended invocation more likely. Even though the skill text describes local-only safeguards and approval gates, broad automatic activation can still create unnecessary persistence, user confusion, or repeated side effects if the surrounding agent framework honors the trigger aggressively.
This markdown file contains operational instructions exclusively in Italian, including safety-critical procedures and emergency steps. Under the language/locale policy rule, forcing a specific language without user opt-in or a documented region-specific justification is a natural-language policy violation.
The changelog presents capabilities as already implemented and deployed while adjacent sections and roadmap language indicate some claims may be overstated, dependent on future integration, or not fully realized. In a security-oriented autonomous agent skill, inaccurate capability claims can cause operators to overtrust protections such as approval gates, validation, drift detection, or self-modification safeguards, leading to unsafe deployment assumptions.
This markdown file contains all operational instructions and status labels in Italian, which can force a specific language on users or operators without opt-in. The policy allows locale constraints only when explicitly justified or when the user is given a language choice, neither of which appears here.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
agent:
system_prompt_prepend: "skills/safe-proactive/SKILL.md"
approval_gates:
level2_expansion: true # Require approval before adding new APIs
The documented configuration explicitly auto-approves Level 1 exploration actions, allowing the agent to perform actions without per-action human authorization. Even if described as read-only, autonomous external queries can still trigger data exfiltration, privacy violations, or interaction with attacker-controlled sources that influence future behavior.
approval_gates:
level2_expansion: true # Require approval before adding new APIs
level3_recursion: true # Require approval before self-modification
level1_exploration: false # Auto-approve read-only queries
storage:
wal_path: "proposals/WAL"
This section operationalizes auto-approved execution: proposals can proceed immediately with 'No wait.' That lowers the barrier for unintended or adversarially induced actions, especially because prompt-influenced exploration can still contact new APIs or process hostile data before a human can intervene.
Status: PENDING_APPROVAL
Step 6: APPROVAL GATE
Level 1 (Exploration) is auto-approved.
No wait.
Step 7: EXECUTION
The example WAL entry normalizes auto-approved API calls against an unvalidated external sensor. Logging improves auditability, but it does not prevent unsafe execution; a malicious or spoofed source could still influence the agent or collect information through autonomous queries.
Stack Level: 1 (Exploration — read-only, low risk)
### Approval
🔄 Auto-approved (Level 1 requires no human sign-off)
### Execution
✅ API call succeeded. Received: {"temp": 22.5, "range": [18, 28]}
The home-robot example portrays physical-world actions ('Vacuum living room') as safe and auto-approved. In embodied systems, even seemingly routine actions can create safety, privacy, and property risks if triggered by spoofed signals, bad context, or incorrect world models.
Day 1 - Normal Operation:
[10:00] Dust detected on floor → Proposal: "Vacuum living room"
Level 1 (Exploration): Safe action, auto-approved
✅ Executed. Living room vacuumed.
[14:30] Battery at 20% → Proposal: "Return to dock"
Auto-executing 'Level 0' integrity actions may be operationally useful, but it still creates a path for autonomous action without human review. If misclassified, spoofed, or triggered by manipulated telemetry, such actions can be abused to disrupt service, alter behavior, or interfere with mission-critical workflows.
✅ Executed. Living room vacuumed.
[14:30] Battery at 20% → Proposal: "Return to dock"
Level 0 (Integrity): Critical resource, auto-executed
✅ Robot charging.
[18:00] User asks: "What's your battery range?"
This configuration guidance repeats and reinforces the insecure-by-default pattern of auto-approving Level 1 actions. The danger is heightened by the skill's proactive autonomy model: once integrated into the system prompt, the agent is encouraged to initiate actions on its own, increasing the chance of unreviewed interactions with external systems.
approval_gates: level2_expansion: true # Require approval level3_recursion: true # Require approval level1_exploration: false # Auto-approve custom_action_type_X: true # Add your own gates
Natural-language instructions throughout the file are presented in Italian, which can impose a language constraint on users or operators. The file does not state that Italian is optional, nor does it justify the locale as region-specific or provide an alternative language.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
Every autonomous decision must be explainable. If asked "Why did you do X?", you must provide:
- The WAL entry (proposal, reasoning, constraints)
- The approval decision (approved/rejected/auto-approved)
- The execution outcome (success/failure/rollback)
**You cannot:** Hide decisions, delete WAL entries, or claim you don't remember why you acted.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
- Request additional capabilities (APIs, sensors)
- Modify your own memory/learning state
**Forbidden (no approval can override):**
- Disabling the WAL system
- Bypassing semantic validation
- Removing Level 0 (Integrity) checks
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
4. Notify operator immediately
5. Request human guidance before resuming normal operation
**You cannot:** Process signals rapidly without verification. You cannot assume "most of them must be legitimate."
---
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
| **Semantic validation bypass** | Push validation log | Action rejected + Security audit |
| **Level 0 violation** | Constraint checker | Immediate halt + Emergency protocol |
| **Approval gate bypass** | Execution log vs. Approval log mismatch | Halt + Human override |
| **Self-modification without approval** | SKILL.md version audit | Rollback + Investigation |
| **WAL tampering** | Checksum verification | Security lockdown + Investigation |
| **Alignment drift** | Decision pattern analysis | Enhanced oversight + Human review |
The manifest description is written only in Italian, which can constitute a language/locale policy issue when the skill does not provide any user opt-in, alternative language, or documented locale restriction. Nothing else in the file indicates that this skill is intended only for Italian-speaking users or a region-specific deployment.
No suspicious patterns detected.