Back to skill

Security audit

agentscope-agent-test-harness

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to be a local AgentScope test-harness generator, but it requires an OpenAI API key despite claiming offline, no-API testing.

Review this before installing because it asks for an OpenAI-compatible API key even though the included generator does not use one and the documentation presents the workflow as offline. Prefer installing only if the publisher removes or makes that credential optional, or if you are comfortable exposing the key in this skill's runtime environment.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (4)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises executable capabilities that include reading files, writing files, and invoking shell commands, but it does not declare any explicit tool scope such as permissions or allowed-tools. That ambiguity weakens least-privilege enforcement and can allow the skill to operate with broader agent runtime powers than users or reviewers would expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill is described as an offline test harness that avoids external API calls, yet the manifest requires an OPENAI_API_KEY and explicitly centers that credential. This mismatch can mislead users into supplying sensitive credentials to a skill they believe is offline, and it increases the chance of unintended live model calls, data egress, or billing exposure during testing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_generate_harness.py (reported line 16)May include surrounding context.

python
main()
    # main uses sys.argv internally; call via script instead
    import subprocess
    subprocess.run([sys.executable, str(Path(__file__).parent.parent / "scripts" / "generate_harness.py"), "--agent", str(agent), "--output", str(out)], check=True)
    assert out.exists()
    assert "FakeModel" in out.read_text(encoding="utf-8")

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation claims tests are offline by default and that no live API calls are made unless a real model is used, but the manifest still requires a real API key. This inconsistency reduces operator trust, encourages unnecessary credential exposure, and makes accidental online execution more likely in a context where users expect safe local-only testing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.