Back to skill

Security audit

oss-workflow

Security checks for vulnerabilities and agentic risk

Overview

The skill is a GitHub workflow router, but it includes a broad shell hook and under-scoped checks around GitHub write operations that warrant Review before installation.

Install only if you want a strict GitHub workflow system that creates local audit checklists and may use a shell hook to block or allow GitHub commands. Review the guard setup before enabling profile-level hooks, and do not rely on its current checks as sufficient authorization for merges, closes, edits, comments, or issue operations.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
references/guard/guard.py:53
Finding

Release authorization can be satisfied by stale or unrelated checklist evidence

Content
View full analysis

Vulnerability Details

File Location: references/guard/guard.py:53-80, 164-190
Related Workflow Location: references/checklists/s2.md:21-22, 85-87
Vulnerability Type: Improper authorization validation
Risk Level: High

Vulnerable Code

python
R14_LINE = re.compile(r"(?:R14|放行|release)", re.IGNORECASE)
CHECKED = re.compile(r"\[[xX]\]")
EVIDENCE = re.compile(
    r"(?:\bm\d{4,}\b|https?://|[“”「」『』]|“|\d{1,2}:\d{2}:\d{2})"
)
python
def r14_state(paths):
    """返回 (有勾, 有带证据的勾)。"""
    checked = evidenced = False
    for path in paths:
        try:
            with open(path, encoding="utf-8", errors="replace") as fh:
                for line in fh:
                    if R14_LINE.search(line) and CHECKED.search(line):
                        checked = True
                        if EVIDENCE.search(line):
                            evidenced = True
        except OSError:
            continue
    return checked, evidenced
python
paths = find_checklists(project)

if not paths:
    return deny(
        # ...
        payload,
    )

if is_a:
    checked, evidenced = r14_state(paths)
    if not checked:
        return deny(
            # ...
            payload,
        )
    if not evidenced:
        return deny(
            # ...
            payload,
        )

The S2 checklist contains two distinct R14 gates:

markdown
- [ ] 🛑 R14 放行原句获得
  → skill: oss-contribution/references/rules-and-arbitration.md §R14
markdown
- [ ] 🛑 R14 放行(推送前再次确认)
  → skill: oss-contribution/references/rules-and-arbitration.md §R14
- [ ] fork → push branch → gh pr create(或直推如有权限)

Technical Analysis

The guard searches every .oss-task/*/CHECKLIST.md file and accepts any checked line containing the broad terms R14, 放行, or release. It then treats a URL, timestamp, quotation mark, or mNNNN token on that line as sufficient authorization evidence.

The authorization record is not bound to:

  • The active task or che ...[truncated 1890 chars]
Remediation
View remediation

Remediation Suggestions

  1. Assign every checklist a unique task identifier and require the hook payload or active-task state to identify exactly one checklist.
  2. Replace keyword matching with a structured, exact gate identifier such as P9_PRE_PUSH_APPROVAL.
  3. Store authorization in a machine-readable record containing:
    • Task identifier.
    • Repository and remote URL.
    • Branch and target branch.
    • Permitted operation.
    • Referenced user-message identifier.
    • Approval timestamp.
  4. Validate that the approval applies to the exact git push or gh pr create target parsed from the command.
  5. Require the dedicated pre-push gate; do not allow the preliminary S2 R14 item to authorize publication.
  6. Reject multiple active checklists rather than accepting authorization from any matching file.
  7. Enforce an appropriate freshness or single-use policy and invalidate approval when the target, branch, or relevant diff changes.
  8. Add regression tests covering stale checklists, unrelated tasks, mismatched repositories, the preliminary S2 gate, and reused approval evidence.

T09 · Insecure Skill Coding Practices

Error
Location
references/guard/guard.py:57
Finding

Destructive GitHub operations are authorized by checklist existence alone

Content
View full analysis

Vulnerability Details

File Location: references/guard/guard.py:57-61, 164-179, 181-213
Supporting Documentation: references/guard/LIMITATIONS.md:19-23, 52-54
Vulnerability Type: Missing authorization check for remote state-changing operations
Risk Level: High

Vulnerable Code

python
CLASS_B = re.compile(
    r"\b(?:"
    r"gh\s+pr\s+(?:comment|edit|review|merge|close)"
    r"|gh\s+issue\s+(?:comment|create|close|edit)"
    r")\b"
)
python
paths = find_checklists(project)

if not paths:
    return deny(
        "已拦截 GitHub 写操作:{op}\n"
        # ...
        payload,
    )

if is_a:
    checked, evidenced = r14_state(paths)
    if not checked:
        return deny(
            # ...
            payload,
        )
    if not evidenced:
        return deny(
            # ...
            payload,
        )
    missing = missing_slots(paths, project)
    if missing:
        return deny(
            # ...
            payload,
        )

return 0

The documented limitation confirms the behavior:

markdown
- **B 类**(gh pr comment/edit/review/merge/close、gh issue 写):仅存在性
  (评论预览等门禁仍由 oss-workflow s2.md 文本规则承担)。

Technical Analysis

The guard groups comments and reviews together with materially destructive actions such as:

  • gh pr merge
  • gh pr close
  • gh pr edit
  • gh issue close
  • gh issue edit
  • gh issue create

For all class-B operations, the only enforced condition is that at least one checklist exists somewhere under .oss-task. The R14 approval check and evidence validation are executed only for class-A commands.

A checklist's existence demonstrates that a workflow file was created; it does not demonstrate that the user authorized a specific merge, closure, edit, comment, or issue operation. The guard therefore converts unrelated local workflow state into authorization for remote GitHub mutations.

Textual instructions elsewhere in a workflow do not provide an equivalent security boundary because ...[truncated 1457 chars]

Remediation
View remediation

Remediation Suggestions

  1. Split class-B operations into separate risk groups:
    • Lower-risk publication: comments and reviews.
    • Destructive state changes: merge, close, and edit.
    • Resource creation: issue creation.
  2. Require explicit, structured, task-specific user authorization for every destructive state change.
  3. Bind approval to the repository, pull-request or issue number, exact operation, and relevant command arguments.
  4. Require fresh confirmation for merge, close, and consequential edits; do not rely on checklist existence.
  5. Apply a preview-and-confirm gate to externally published comments, reviews, and issue bodies where applicable.
  6. Parse GitHub CLI arguments structurally rather than relying solely on broad regular-expression matching.
  7. Use deny-by-default behavior for recognized remote mutations when approval state is missing, ambiguous, stale, or malformed.
  8. Add tests proving that an empty, incomplete, stale, or unrelated checklist cannot authorize any class-B operation.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

A second independent mismatch report indicates the skill may actually run automated guard-policy tests and validate blocking behavior for git push, gh pr, and gh issue comment flows rather than simply dispatching other skills. Hidden testing or enforcement behavior in a skill advertised as an entry router creates a trust-boundary problem and can lead to unreviewed command interception, denial of expected workflow actions, or confusion during high-impact GitHub operations.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A second independent mismatch report indicates the skill may actually run automated guard-policy tests and validate blocking behavior for git push, gh pr, and gh issue comment flows rather than simply dispatching other skills. Hidden testing or enforcement behavior in a skill advertised as an entry router creates a trust-boundary problem and can lead to unreviewed command interception, denial of expected workflow actions, or confusion during high-impact GitHub operations.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
83% confidence
Finding

The skill declares required binaries and clearly orchestrates file reads, environment checks, and shell use, but it does not declare an explicit tool scope such as allowed-tools or permissions. In an agentic environment, this can cause over-broad capability inheritance and make it easier for the skill to perform unintended filesystem, environment, or command actions without transparent review.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger definition is extremely broad, covering common words like PR, issue, push, commit, README, docs, and fork. In practice this can cause unintended invocation on benign requests, forcing users into an unexpected workflow, increasing the chance of overreach into unrelated tasks, and creating opportunities for accidental execution of GitHub-oriented procedures when not needed.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The scenario mapping includes underspecified triggers such as 'fix', '新项目', and other broad phrases that are common in many non-GitHub contexts. This weak specificity increases false activation risk and can route unrelated work into a workflow with checklist, shell, and repository expectations that do not fit the user’s actual intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
77% confidence
Finding

The dependency section specifies '非中文/非自有仓' handling and a Chinese writing chain 'zh-tech-writing', and the document elsewhere frames Chinese-specific tooling as part of the workflow. This suggests a locale/language preference embedded in the skill without an explicit user choice or opt-in path.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The title explicitly frames the checklist in Chinese, and multiple instructions require producing repository analysis artifacts in that language context without any opt-in or alternative locale. This creates a natural-language policy concern because the file appears to force a specific language rather than allowing user preference.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The checklist hard-codes a language-selection rule based on inferred 'Chinese signals' and requires the PR body, title, and commit message to follow that determination. This can override the user's actual preference or project norms, causing unintended external communications and reducing operator control over content sent to third-party platforms like GitHub.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The title "S3 Checklist: 文档维护" and the checklist content are entirely in Chinese, and L24 explicitly includes a "正文语言判定" gate, which suggests enforced language handling. Because the file does not offer a language/locale choice or explain that this checklist is only for a China-specific workflow, it may violate the policy against forcing a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The checklist requires a multilingual file strategy tied to a specific rule set ('主语言与命名按 I2') without any visible user consent, exception path, or justification in this file. In an agent workflow, this can silently steer repository language, naming, or documentation decisions away from user intent and create unauthorized changes to project structure or content policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module docstring and all user-facing denial messages are written only in Chinese, including the reasons and remediation text returned to the model/user. This imposes a specific language on interactions without offering any language choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · references/guard/guard_v2_test.py (reported line 21)May include surrounding context.

python
payload = {"tool_input": {"command": command}, "cwd": project}
    if bridge:
        payload["hook_event_name"] = "PreToolUse"
    p = subprocess.run([sys.executable, GUARD], input=json.dumps(payload).encode("utf-8"),
                       capture_output=True)
    return p.returncode, p.stderr.decode("utf-8", "replace"), p.stdout.decode("utf-8", "replace")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The hook matcher applies to very generic shell tool names ("pwsh|Bash|bash|Shell"), so nearly any shell invocation in the agent environment will trigger execution of guard.py. In this skill, which is the mandatory entrypoint for many GitHub-related workflows, that broad interception increases the blast radius: a compromised, buggy, or overly permissive guard script could monitor, block, or alter a wide range of unrelated shell activity rather than only narrowly scoped repository operations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.