T09 · Insecure Skill Coding Practices
- Location
scripts/pdf2md_marker.py:29- Finding
Hardcoded Volcengine API Credential
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This PDF extraction skill is mostly purpose-aligned, but its optional LLM mode can send PDF content to remote services and includes a hardcoded fallback API key.
Review this skill before installing. Use local-only conversion for sensitive PDFs unless you intentionally want LLM enhancement, understand which remote endpoint will receive document-derived content, and provide your own scoped API key. The embedded Volcengine key should be removed or rotated by the publisher, and dependency setup should ideally be pinned or run in an isolated environment.
scripts/pdf2md_marker.py:29Hardcoded Volcengine API Credential
README.md:15Unpinned Third-Party Dependency Installation
The embedded fallback API key creates a path where PDF contents can be sent to an external LLM service using credentials bundled in the script, rather than credentials deliberately provided by the user. Because this tool handles document extraction, the context makes accidental disclosure of confidential PDFs especially dangerous.
A hardcoded fallback API key allows the script to authenticate to an external LLM service even when the operator did not intentionally configure credentials. If LLM mode is enabled, sensitive PDF content may be sent to a remote endpoint under an embedded key with inadequate notice, creating both data leakage and credential misuse risk.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
conda create -n pdf-extractor python=3.10 -y
The README advertises an optional LLM enhancement flag but does not disclose that document content may be transmitted to an external service for processing. Because academic PDFs can contain unpublished research, personal data, or confidential material, this omission can lead users to exfiltrate sensitive content without informed consent.
The skill says it is self-contained and should not perform extra installs or downloads, yet later states that Marker models are downloaded automatically. This creates a misleading trust boundary: using the skill may trigger unexpected network access and retrieval of remote model artifacts, which can violate offline-use assumptions, enterprise policy, or reproducibility expectations.
The documented --ark-code-latest and custom OpenAI-compatible parameters imply that PDF content may be transmitted to an external service, but the skill does not warn users about this data flow. Because PDFs may contain sensitive unpublished research, proprietary data, or personal information, omission of this disclosure can lead to unintended exfiltration and policy violations.
This code sends the user prompt and any attached image data to an external OpenAI-compatible HTTP service at a configured base URL, which constitutes data exfiltration to another service boundary. Even though the default endpoint is localhost, the code provides no consent, classification, or safeguards around potentially sensitive prompt/image contents, so deployments that point this URL to a remote proxy could unintentionally disclose confidential data.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
openai_base_url: Annotated[
str, "The base url to use for OpenAI-like models. No trailing slash."
] = "https://api.openai.com/v1"
openai_model: Annotated[str, "The model name to use for OpenAI-like model."] = (
"gpt-4o-mini"
)
The script adds optional outbound LLM/network capability to a PDF-to-Markdown converter and pairs it with a built-in credential fallback, which can cause document contents to be transmitted to a third-party service during conversion. In this context the skill processes potentially sensitive PDFs, so enabling remote processing without strong gating, explicit consent, or secret management materially increases data-exfiltration risk.
The function assigns a default language list of ['Chinese', 'English'] when the caller does not specify languages. This imposes a locale/language assumption by default rather than offering a neutral default or explicit opt-in, which matches the policy concern about forcing a specific language setting.