Back to skill

Security audit

Pharmaclaw Tox Agent

Security checks for vulnerabilities and agentic risk

Overview

This looks like a local chemistry screening skill rather than malware, but its drug-safety labels are incomplete and should only be treated as rough research heuristics.

Install only if you need a lightweight local RDKit descriptor helper. Do not rely on its Low/Medium/High output for clinical, regulatory, synthesis-gating, or safety-critical decisions without expert review and more complete validated toxicology models.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

other

Warning
Location
scripts/tox_agent.py:49
Finding

Risk Classification Ignores Documented QED and Veber Indicators

Content
View full analysis

Vulnerability Details

File Location: scripts/tox_agent.py:49
Vulnerability Type: other: Unsafe risk-classification logic
Risk Level: Medium

Vulnerable Code

python
'risk': 'Low' if lipinski_viol == 0 and pains_count == 0 else 'Medium/High',

Technical Analysis

The risk decision considers only Lipinski violations and PAINS matches. It does not consider the calculated QED score or Veber violations.

This behavior contradicts the documented classification in SKILL.md, which states that:

  • A Low classification requires no Lipinski violations, no PAINS alerts, and QED greater than 0.5.
  • Low QED should produce a Medium classification.
  • Multiple Veber violations can produce a High classification.
  • PAINS hits can produce a High classification.

The implementation can consequently label a compound as Low risk even when its QED is at or below 0.5 or it has multiple Veber violations. It also combines Medium and High into the non-specific value Medium/High, rather than emitting the documented distinct classifications.

These descriptor-based rules are drug-likeness heuristics rather than comprehensive toxicology predictions. Presenting their result as a general safety-risk classification may further encourage downstream consumers to place more confidence in the output than the underlying checks support.

Attack Path

  1. An attacker or upstream system supplies a syntactically valid SMILES value.
  2. The molecule is selected so it has no Lipinski violations and does not match either simplified PAINS pattern.
  3. The molecule nevertheless has a QED score at or below 0.5, one or more significant Veber violations, or both.
  4. The implementation calculates those indicators but omits them from the final risk decision.
  5. The result is returned with "risk": "Low".
  6. A downstream synthesis, screening, or derivative-selection component trusts the favorable classification and advances the co ...[truncated 1055 chars]
Remediation
View remediation

Remediation Suggestions

Replace the binary expression with explicit, documented Low, Medium, and High classification rules. At minimum:

  1. Require all documented Low-risk conditions, including QED greater than 0.5 and acceptable Veber results.
  2. Return a distinct Medium or High value instead of the ambiguous Medium/High value.
  3. Define precedence when multiple indicators produce different classifications, with the most severe applicable classification taking priority.
  4. Add unit tests covering QED boundary values, each Lipinski threshold, PAINS hits, both Veber violations, and combinations of these conditions.
  5. Keep SKILL.md and the implementation synchronized through tests that validate the documented output contract.
  6. Clearly label the result as heuristic drug-likeness and assay-interference screening, not a comprehensive toxicology determination.
  7. Require expert review and validated toxicology models before using the result to make safety-critical decisions.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains very broad terms such as "risk," "safety," "toxicology," and "QED," which can cause the skill to activate in unrelated contexts. Because this skill performs pharma/chemistry safety analysis and chains into other agents, unintended invocation could expose sensitive workflow context, generate misleading safety guidance, or trigger downstream actions when the user did not intend to use this skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The implementation materially underdelivers on the safety-analysis claims in the skill metadata by computing only a limited set of descriptors and collapsing risk into 'Low' or 'Medium/High' rather than discrete classes. In a pharmacology/toxicology context, this can mislead downstream agents or users into overtrusting incomplete safety assessments, causing unsafe prioritization or decision-making.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script emits toxicity and drug-safety classifications without any warning that results are heuristic, non-clinical, and unsuitable as definitive medical or regulatory guidance. Because this skill is specifically framed for pharma safety profiling, users or chained agents may treat the output as authoritative, increasing the risk of harmful real-world decisions based on incomplete computational heuristics.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

Labeling PAINS detection as a 'mock' while presenting the skill as a real safety profiler indicates that alerting is based on an incomplete hard-coded substructure list rather than robust PAINS screening. This creates a false sense of chemical-risk coverage and may miss problematic motifs that users assume are being checked.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.