T03 · Remote Payload Retrieval and Execution
- Location
scripts/advanced_ml_skill.py:661- Finding
Unrestricted Execution of Remotely Generated Python Code
- Content
View full analysis
Vulnerability Details
File Location:
scripts/advanced_ml_skill.py, lines 382–396 and 657–674
Vulnerability Type: Remote payload retrieval followed by unrestricted dynamic execution
Risk Level: CriticalVulnerable Code:
python if not self.openai_client: return fallback_code prompt = self._build_code_prompt(algorithm_name) try: response = self.openai_client.completions.create( model="code-davinci-002", prompt=prompt, max_tokens=900, temperature=0.2, ) candidate = self._strip_code_fence(response.choices[0].text) if "def train_and_evaluate" not in candidate: return fallback_code return candidate + "\n" except Exception: return fallback_codepython error: Optional[str] = None accuracy: Optional[float] = None namespace: Dict[str, Any] = {} try: exec(code, namespace) # noqa: S102 train_fn = namespace.get("train_and_evaluate") if not callable(train_fn): raise ValueError("生成代码中未定义 train_and_evaluate 函数。") score = train_fn( data_bundle["x_train_processed"], data_bundle["x_test_processed"], data_bundle["y_train"], data_bundle["y_test"], self.random_state, data_bundle["n_classes"], ) accuracy = round(float(score), 4)Technical Analysis
When an OpenAI client is configured, the application requests executable Python source code from a remote AI service. The only validation applied to the response is a substring check for
def train_and_evaluate. This does not establish that the response contains only a safe model implementation.Python code may contain executable top-level statements before or after the expected function. It can also import unrestricted modules and perform arbitrary actions from inside the function. The complete response is passed to
execwith normal built-ins and the full per ...[truncated 1978 chars]- Remediation
View remediation
Remediation Suggestions
- Remove dynamic execution of remotely generated code. Map supported algorithm names to reviewed local estimator implementations instead.
- If code generation is retained as a user-facing feature, treat generated source strictly as display-only content and do not execute it.
- If execution is an unavoidable product requirement, run generated code in a disposable external sandbox or container with:
- No host filesystem mounts.
- No application secrets or API keys.
- Outbound networking disabled.
- A read-only base filesystem.
- An unprivileged user and dropped Linux capabilities.
- Strict CPU, memory, process, and execution-time limits.
- Destruction of the environment after every execution.
- Use a narrow, structured model specification rather than generated source code—for example, permit the remote service to select from an allowlisted algorithm and validated numeric hyperparameters.
- Do not rely on substring checks, regular expressions, restricted
globals, or AST filtering as the sole security boundary. Python cannot be safely sandboxed inside the same trusted process using these techniques alone. - Add tests that reject generated content containing top-level statements, imports, filesystem access, network access, dynamic evaluation, or subprocess calls, even if isolated execution remains as defense in depth.
