Back to skill

Security audit

CI Whisperer

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its CI-triage purpose, but it merits Review because its helper can print raw GitHub Actions failure logs from authenticated repositories before redaction.

Install only if you are comfortable letting the agent use your authenticated GitHub CLI to read Actions runs and logs for the target repository. Keep the default read-only mode, enable CI_WHISPERER_WRITE=1 only for sessions where you want PR creation, and treat helper output as raw CI logs that may need manual redaction before sharing.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ci_autopsy.py:45
Finding

Raw GitHub Actions Logs Are Exposed Without Secret Redaction

Content
View full analysis

Vulnerability Details

File Location: scripts/ci_autopsy.py, lines 45–53
Vulnerability Type: Sensitive information exposure through unredacted CI logs
Risk Level: Medium

Vulnerable Code

python
def cmd_failed_logs(args: argparse.Namespace) -> None:
    out = run([
        gh_bin(), "run", "view", str(args.run_id),
        "--repo", args.repo,
        "--log-failed",
    ])
    print(out)

Technical Analysis

The failed-logs command retrieves complete failed-job logs through the authenticated GitHub CLI and writes them directly to standard output. No redaction, structured filtering, output-size restriction, or manual review boundary is applied before print(out).

This implementation conflicts with the security requirements in SKILL.md, which state that tokens must never be printed and secrets in logs must be redacted before quotation. GitHub masks registered secrets in many circumstances, but that protection is not comprehensive. Logs may still contain unregistered credentials, transformed secrets, authorization headers, signed URLs, private keys, personal data, or values printed by compromised dependencies.

Repository and run identifiers are passed to subprocess.run as separate arguments rather than through a shell, so this code does not establish command injection. The flaw is specifically the unrestricted disclosure of remotely retrieved log content.

Attack Path

  1. A workflow step, malicious contributor, or compromised dependency writes a sensitive value into a failing job's output.
  2. The value is not recognized or masked by GitHub, such as when it is transformed, dynamically generated, or not registered as a repository secret.
  3. A user or agent invokes:
    bash
    python3 scripts/ci_autopsy.py failed-logs --repo owner/repo --run-id 123
    
  4. The script uses the locally authenticated gh client to retrieve the failed-job logs.
  5. The complete response ...[truncated 996 chars]
Remediation
View remediation

Remediation Suggestions

  1. Sanitize all retrieved logs before printing them. Redact common GitHub, cloud-provider, bearer-token, authorization-header, URL-credential, private-key, and password patterns.
  2. Support user-configured secret values and replace exact matches with a fixed marker such as [REDACTED].
  3. Extract only short excerpts surrounding relevant errors instead of emitting complete failed-job logs.
  4. Add strict output-length and line-count limits, with an explicit opt-in mechanism for reviewing additional content.
  5. Warn users that heuristic redaction cannot guarantee removal of every secret and require confirmation before exposing raw logs.
  6. Keep raw log data out of exception messages, persistent files, telemetry, and agent transcripts wherever possible.
  7. Add automated tests using representative GitHub tokens, cloud credentials, bearer headers, signed URLs, private-key blocks, transformed secrets, and ordinary non-secret text to detect both redaction failures and excessive false positives.
  8. Continue using minimally scoped GitHub credentials. Read-only analysis should use only repository and Actions read permissions; write permissions should be enabled solely for explicitly approved PR operations.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (4)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code does interact with GitHub Actions runs using the GitHub CLI, which partially aligns with the declared purpose of fetching metadata/logs safely. However, the primary declared purpose is higher-level failure analysis and proposing fixes, including concise root-cause reporting and possibly automated PR creation. The actual code only retrieves and prints run data/logs and lists recent runs. It lacks any logic for diagnosis, summarization, fix generation, or PR automation. Additionally, it exposes a listing capability not described. This is a material description-behavior mismatch because the implemented behavior is a narrower data-fetching utility rather than an analyzer/remediator.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill instructs use of shell-accessible tooling (gh) and references bundled scripts, but the manifest does not declare any explicit tool scope such as allowed tools or permissions. That creates an authorization and review gap: an agent may invoke shell commands with broader capabilities than users or platform policy expect, including repository-affecting operations in PR mode.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest description contains broad trigger phrases like 'CI is failing' and 'why did this workflow fail,' which can cause the skill to activate on ordinary conversation without the user intentionally requesting repository/log access. In context, that matters because the skill is designed to query GitHub run metadata and logs via gh, potentially pulling sensitive build output or private repository details based on an ambiguous request.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/ci_autopsy.py (reported line 35)May include surrounding context.

python
def run(cmd: list[str]) -> str:
    p = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True)
    if p.returncode != 0:
        raise SystemExit(p.stdout.strip())
    return p.stdout

Static analysis

No suspicious patterns detected.