Back to skill

Security audit

C++ 算法竞赛自动化测试数据生成与校验框架

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent contest-problem generation purpose, but it asks users to run unverified Docker code and untrusted C++ with broad workspace access.

Install only if you are comfortable running a third-party Docker image and generated C++ code. Use a clean, dedicated workspace with no secrets, verify or rebuild the Docker image yourself where possible, avoid mounting unrelated project files, and prefer adding Docker isolation such as no network, non-root user, read-only inputs, and a separate writable output directory.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Error
Location
SKILL.md:24
Finding

Unverified Third-Party Docker Image Is Loaded and Executed

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:81
Finding

Untrusted Generated Programs Execute with Read/Write Access to the Entire Workspace

Content
View full analysis
4 3 3 docker run --rm -v "$PWD:/data" -w /data cpp-sandbox python3 scripts/generate.py 4 3 3 ``` Supporting execution behavior from `scripts/generate.py`: ```python for name, src, exe in [("生成器(gen)", gen_cpp, gen_exe), ("校验器(valid)", valid_cpp, valid_exe), ("标程(std)", std_cpp, std_exe)]: ok, err = compile_cpp(src, exe, include_path=tmpdir) if not ok: print(json.dumps({"status": "error", "message": f"{name} 编译失败:\n{err}"})) return ``` ```python with open(in_file, 'w') as fin: cmd_gen = [gen_exe, str(subtask_id), str(tc)] res = subprocess.run(cmd_gen, stdout=fin, stderr=subprocess.PIPE, **run_kwargs) with open(in_file, 'r') as fin: res = subprocess.run([valid_exe], stdin=fin, capture_output=True, **run_kwargs) with open(in_file, 'r') as fin, open(out_file, 'w') as fout: res = subprocess.run([std_exe], stdin=fin, stdout=fout, stderr=subprocess.PIPE, **run_kwargs) ``` The non-English strings in the original source are diagnostic labels only; they do not alter the security behavior. ### Technical Analysis The Docker command mounts the entire current working directory at `/data` with read/write permissions. The script then compiles and executes the generator, validator, and standard solution. These source files include generated code and user-provided code and therefore must be treated as untrusted. The container command does not include ...[truncated 2511 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (32)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The script compiles and then later executes user-provided C++ programs, a powerful arbitrary native code execution capability not disclosed in the manifest. In this skill context, that is a major security boundary violation because users may expect data generation, not host-level execution of untrusted binaries.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
97% confidence
Finding

The skill instructs the agent to read files, write files, and execute shell/Docker commands, but it declares no explicit tool scope or permission boundaries. That creates an authorization gap where a host may allow broader filesystem and command access than the user expects, increasing the chance of unintended command execution or workspace modification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill description and instructions are written to enforce Chinese prompts and outputs, including fixed Chinese user-facing messages and formatting requirements, but do not offer any user-selectable language or locale option. This can violate organizational language/locale policy when a skill forces a specific language without explicit opt-in.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill requires orchestrating host and container shell commands, environment inspection, Docker image checks, repeated rebuild loops, and command-based troubleshooting. That materially expands the trust boundary from content generation into privileged system interaction, creating risk of unintended command execution, environment probing, and file modification well beyond the user's likely expectation for a problem-generation skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill directs creating and overwriting local workspace files without any explicit user-facing warning or confirmation. Silent file modification can destroy user work, replace trusted files, or prepare inputs for later execution steps without informed consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs running Docker with the current directory mounted into the container and executing a build script, but it does not clearly warn the user about command execution, container access to workspace files, or the resulting trust implications. In this context, the mounted workspace plus script execution could expose or modify many files, making the omission dangerous.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
79% confidence
Finding

The Docker invocation relies on a mutable local image tag (cpp-sandbox) rather than a pinned immutable digest, so the executed environment can be replaced silently. If an attacker or misconfiguration changes that image, the skill would run arbitrary code against the mounted workspace.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
79% confidence
Finding

This Docker command again uses an unpinned mutable image reference, allowing environment substitution without changing the skill text. Because the current directory is mounted into the container, a compromised or unexpected image could read, alter, or exfiltrate workspace contents.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The title and core directive are written exclusively in Chinese and include mandatory wording such as '必须严格遵守', but the document does not offer any user language/locale choice. Under the policy, forcing a specific language without opt-in is a natural-language policy violation unless the locale constraint is clearly justified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/generate.py (reported line 21)May include surrounding context.

python
def compile_cpp(src_path, output_exe, include_path="."):
    cmd = ["g++", "-O2", "-std=c++14", "-I", include_path, src_path, "-o", output_exe]
    try:
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
        if result.returncode != 0:
            return False, result.stderr
        return True, ""

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

This executes a compiled validator binary derived from user-provided C++ code. Although a timeout and memory limit are attempted, the script still runs untrusted native code on the host, which can abuse CPU, fork/exec behavior, filesystem access, or exploit compiler/runtime/kernel bugs; the skill context makes this more dangerous because arbitrary generators/validators/solutions are exactly what this tool processes.

Content

Scanner excerpt · scripts/generate.py (reported line 53)May include surrounding context.

python
# 1. 喂给 valid 校验输入合法性
        with open(in_file, 'r') as fin:
            v_res = subprocess.run([valid_exe], stdin=fin, capture_output=True, **run_kwargs)
            if v_res.returncode != 0:
                raise Exception(f"AI 编造的样例输入 {i} 格式不合法,未通过校验器:\n{v_res.stderr or v_res.stdout}")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

This runs the compiled standard-solution binary built from untrusted user input. Native code execution is the core danger here: the program can perform unauthorized local actions within the agent environment despite the light timeout/resource controls.

Content

Scanner excerpt · scripts/generate.py (reported line 59)May include surrounding context.

python
# 2. 喂给 std 计算真实输出
        with open(in_file, 'r') as fin:
            s_res = subprocess.run([std_exe], stdin=fin, capture_output=True, **run_kwargs)
            if s_res.returncode != 0:
                raise Exception(f"标程运行 AI 样例 {i} 时崩溃:\n{s_res.stderr}")

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script writes meta.json into the current workspace root for an external consumer, which is broader than generating the requested package artifacts. In an agent environment, writing integration metadata outside the package directory can overwrite expected files, interfere with other tools, or create a covert side effect channel.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script automatically performs multiple executions of untrusted user-supplied binaries without any user-facing warning or confirmation. Lack of disclosure and gating increases the chance that dangerous code runs in contexts where the operator did not understand or approve the risk.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

This is the main execution of the user-supplied generator binary to produce test data. Because it executes attacker-controlled native code, it can misuse local resources or interact with the host filesystem/environment; the skill's purpose directly amplifies the risk because such code is expected and repeatedly run.

Content

Scanner excerpt · scripts/generate.py (reported line 200)May include surrounding context.

python
# a. 运行 gen 生成 .in
                    with open(in_file, 'w') as fin:
                        cmd_gen = [gen_exe, str(subtask_id), str(tc)]
                        res = subprocess.run(cmd_gen, stdout=fin, stderr=subprocess.PIPE, 
                                             **run_kwargs)
                        if res.returncode != 0:
                            print(json.dumps({"status": "error", "message": f"生成器在 Subtask {subtask_id} Case {tc} 崩溃:\n{res.stderr}"}))

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

This executes the user-provided validator binary against generated inputs. A validator is just arbitrary native code here, so without strong isolation it can perform unauthorized actions on the agent host despite the timeout and memory settings.

Content

Scanner excerpt · scripts/generate.py (reported line 208)May include surrounding context.

python
# b. 运行 valid 校验 .in
                    with open(in_file, 'r') as fin:
                        res = subprocess.run([valid_exe], stdin=fin, capture_output=True, 
                                             **run_kwargs)
                        if res.returncode != 0:
                            print(json.dumps({"status": "error", "message": f"数据校验失败 (Subtask {subtask_id} Case {tc}):\n{res.stderr or res.stdout}"}))

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

This executes the user-supplied standard solution binary to generate expected outputs for official test cases. The skill context makes this especially dangerous because execution is automatic and repeated, creating a reliable path for running arbitrary native code under the agent's privileges.

Content

Scanner excerpt · scripts/generate.py (reported line 216)May include surrounding context.

python
# c. 运行 std 生成 .out
                    with open(in_file, 'r') as fin, open(out_file, 'w') as fout:
                        res = subprocess.run([std_exe], stdin=fin, stdout=fout, stderr=subprocess.PIPE, 
                                             **run_kwargs)
                        if res.returncode != 0:
                            print(json.dumps({"status": "error", "message": f"标程运行失败 (Subtask {subtask_id} Case {tc}):\n{res.stderr}"}))

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

This invokes the untrusted generator again to create student-download sample inputs. Repeated execution increases exposure, and the generated binary still has host-level access unless isolated; in this skill, running attacker-controlled native generators is central and therefore high risk.

Content

Scanner excerpt · scripts/generate.py (reported line 239)May include surrounding context.

python
# 为防止随机种子与线上数据冲突,给 tc 传参加一个偏移量(如 100)
                with open(s_in, 'w') as fin:
                    res = subprocess.run([gen_exe, str(s_id), str(file_idx + 100)], stdout=fin, stderr=subprocess.PIPE, **run_kwargs)
                    if res.returncode != 0:
                        raise Exception(f"生成线下学生数据(输入)失败: {res.stderr}")
                with open(s_in, 'r') as fin, open(s_out, 'w') as fout:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

This runs the untrusted compiled standard solution again for student sample output generation. Each additional execution path broadens the chance of abuse or sandbox escape if isolation is weak or absent.

Content

Scanner excerpt · scripts/generate.py (reported line 243)May include surrounding context.

python
if res.returncode != 0:
                        raise Exception(f"生成线下学生数据(输入)失败: {res.stderr}")
                with open(s_in, 'r') as fin, open(s_out, 'w') as fout:
                    res = subprocess.run([std_exe], stdin=fin, stdout=fout, stderr=subprocess.PIPE, **run_kwargs)
                    if res.returncode != 0:
                        raise Exception(f"生成线下学生数据(输出)失败: {res.stderr}")

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script conditionally deletes the source directory if its basename is problem_temp. This is a destructive side effect beyond the stated packaging role, and in a shared workspace it can erase user data or outputs unexpectedly if path assumptions are wrong or attacker-influenced.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code opens and writes to filenames taken directly from command-line/user-controlled parameters such as test overview logs, markup files, and test case files, using fopen(..., "wb") without any confirmation, path restriction, or overwrite protection. In the context of an agent skill that generates validators/tests automatically, this can overwrite arbitrary files reachable by the process if an attacker can influence those arguments or generated invocations.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · scripts/testlib.h (reported line 3355)May include surrounding context.

text
quit(_fail, #readMany ": size should be non-negative.");                \
    if (size > 100000000)                                                       \
        quit(_fail, #readMany ": size should be at most 100000000.");           \
                                                                                \
    std::vector<typeName> result(size);                                         \
    readManyIteration = indexBase;                                              \
                                                                                \

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · scripts/testlib.h (reported line 3358)May include surrounding context.

text
quit(_fail, #readMany ": size should be non-negative.");                \
    if (size > 100000000)                                                       \
        quit(_fail, #readMany ": size should be at most 100000000.");           \
                                                                                \
    std::vector<typeName> result(size);                                         \
    readManyIteration = indexBase;                                              \
                                                                                \

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · scripts/testlib.h (reported line 3366)May include surrounding context.

text
quit(_fail, #readMany ": size should be non-negative.");                \
    if (size > 100000000)                                                       \
        quit(_fail, #readMany ": size should be at most 100000000.");           \
                                                                                \
    std::vector<typeName> result(size);                                         \
    readManyIteration = indexBase;                                              \
                                                                                \

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest says the skill generates problem assets from the original problem, but the instructions additionally require reading local workspace files such as references/backgrounds.md, references/testlib-manual.md, and template directories. While related to generation quality, this broader file-reading capability is not stated in the manifest description and goes beyond the minimal purpose presented to the user.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.