Back to skill

Security audit

ia-test-audit

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed test-audit workflow with local helper scripts, and I did not find hidden, destructive, or unrelated behavior.

Install only if you are comfortable letting the agent read in-scope repository tests and source files and run local test/audit commands. Use the documented scratch-copy approach for mutation or reach probes, and avoid live services, paid APIs, or shared databases unless you separately authorize and isolate them.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code’s actual behavior is narrowly focused on identifying likely duplicate or structurally similar tests. It tokenizes/parses test files in several languages, normalizes actions and assertions, computes signatures, and emits findings such as REDUNDANT, SUBSUMED, FOLD, and PARAMETRIZE. This is related to test maintenance and test quality, but it is not the declared capability of auditing whether tests detect regressions in the sense described. The declared description promises checks for mocked-away subjects, weak or circular assertions, indiscriminate fixtures, swallowed failures, and tests missing from gates; none of those are implemented here. The primary purpose is materially different: duplicate-test detection and consolidation suggestions, not regression-detection audit. No concerning undeclared external-resource access is present beyond reading matched files and reporting results, but the purpose mismatch is substantial.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description promises a comprehensive test-quality audit focused on whether tests detect regressions, including several classes of issues such as mocked-away subjects, weak/circular assertions, undiscriminating fixtures, swallowed failures, and missing gates. The supplied code does not implement that broad behavior. Instead, it specifically extracts probable guard/refusal strings from source code and compares them against string literals found in tests to identify messages that may be unasserted. That is related to one narrow aspect of regression-detection quality—whether negative-path messages are textually pinned in tests—but it does not cover most of the declared audit areas. So the code’s actual primary purpose is materially narrower and different from the declared capability set.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

There is a material mismatch. The declared description promises a broad semantic audit of test quality and regression-detection strength, including mocked-away subjects, weak/circular assertions, fixtures, swallowed failures, and missing gate coverage. The supplied code does only a narrow report-level census of JUnit XML outputs: it aggregates skip reasons, counts failures, notes tests with assertions="0" when the report provides that field, and compares reported identifiers against test files on disk to find files never mentioned by reports. Only the 'tests missing from gates' part is substantially represented. The rest of the claimed capabilities would require inspecting test source code or richer metadata, which this script does not do.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description promises a broad audit of test quality issues such as mocked-away subjects, circular assertions, swallowed failures, fixture quality, and gate coverage. The supplied code does not implement those analyses. Instead, it provides a specialized pytest probe for subprocess-driven tests, capturing nonzero subprocess outcomes and stderr snippets to help determine which guard/refusal a test actually exercised. While this may support some form of test auditing, its primary behavior is materially narrower and different from the declared purpose. There are no obvious dangerous extra permissions beyond writing its output file, but the main functionality is mismatched.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 103)May include surrounding context.

md
| [duplicate_tests.py](./scripts/duplicate_tests.py) `<test globs> --json` | Structural duplicate/consolidation candidates, not an effectiveness score. |

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SPEC.md (reported line 16)May include surrounding context.

md
Out of scope:
- Acting as the runtime instructions themselves (those live in `SKILL.md`).
- Trigger phrasings already covered by adjacent `ia-*` skills (`validate-plugin` flags >70% description overlap as DUPLICATE_TRIGGER).
- <!-- to fill in: domain-specific exclusions when the skill drifts -->

## Trigger Context

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/stacks.md (reported line 10)May include surrounding context.

md
Tool names and flags below were current when written. Confirm each with `--help` or the
tool's docs before running it, and prefer the repository's pinned wrapper over a global
binary. Install audit-only tools ad hoc, for example with `uvx`, `npx`, or a scratch
Composer project. Never add them as project dependencies without approval.

Each section covers three things:
- **junit**: how to get the report `junit_census.py` reads. Produce it with the gate's

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/duplicate_tests.py (reported line 390)May include surrounding context.

python
node = stack.pop()
        yield node
        for name in reversed(node._fields):
            value = getattr(node, name, None)
            if isinstance(value, list):
                stack.extend(v for v in reversed(value) if isinstance(v, ast.AST))
            elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context):

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/duplicate_tests.py (reported line 839)May include surrounding context.

python
node = stack.pop()
        yield node
        for name in reversed(node._fields):
            value = getattr(node, name, None)
            if isinstance(value, list):
                stack.extend(v for v in reversed(value) if isinstance(v, ast.AST))
            elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context):

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/duplicate_tests.py (reported line 489)May include surrounding context.

python
out.extend(fold(stmt.body if value else stmt.orelse))
                continue
        for name in ("body", "orelse", "finalbody"):
            if isinstance(getattr(stmt, name, None), list):
                setattr(stmt, name, fold(getattr(stmt, name)))
        if isinstance(stmt, ast.Try | ast.TryStar):
            for handler in stmt.handlers:

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/duplicate_tests.py (reported line 490)May include surrounding context.

python
continue
        for name in ("body", "orelse", "finalbody"):
            if isinstance(getattr(stmt, name, None), list):
                setattr(stmt, name, fold(getattr(stmt, name)))
        if isinstance(stmt, ast.Try | ast.TryStar):
            for handler in stmt.handlers:
                handler.body = fold(handler.body)

Static analysis

No suspicious patterns detected.