T09 · Insecure Skill Coding Practices
- Location
scripts/on_serverless/emr_serverless_submit_cli.py:29- Finding
Long-Lived Cloud Credentials Embedded in Remote Spark Job Configuration
- Content
View full analysis
Dict[str, Any]: ak = os.getenv("VOLCENGINE_AK") sk = os.getenv("VOLCENGINE_SK") if ak: conf["serverless.spark.access.key"] = ak if sk: conf["serverless.spark.secret.key"] = sk return conf ``` The function is invoked automatically for Spark Jar and PySpark jobs: ```python def _cmd_jar(args: argparse.Namespace) -> str: conf = _merge_conf({}, _parse_json(args.conf)) conf = _add_spark_credential(conf) ``` ```python def _cmd_pyspark(args: argparse.Namespace) -> str: conf = _merge_conf({}, _parse_json(args.conf)) conf = _add_spark_credential(conf) ``` This behavior is also explicitly documented in: - `references/emr_serverless/job_instance/emr_serverless_job_instance_guide.md:77` - `references/emr_serverless/job_instance/emr_serverless_job_instance_guide.md:102` ### Technical Analysis The submission CLI reads the caller's long-lived Volcengine access key and secret key directly from environment variables and inserts both values into the Spark job configuration. This configuration is then included in the remote task submitted through the Serverless SDK. The SDK already receives credentials through `build_serverless_client()`. Automatically copying those credentials into remote job configuration expands their exposure beyond the local authenticated client. Depending on EMR and Spark visibility controls, the values may become accessible through: - Job-definition or job-detail APIs. - Spark runtime configuration interfaces. - Driver or executor logs. - Control-plane request records. - Job history and debugging interfaces. - User-provided Jar or PySpark code running inside the job. The behavior ...[truncated 1571 chars]- Remediation
View remediation
