T08 · Insecure Dependencies
- Location
src/transformers-embedding.js:24- Finding
Unverified Model Downloads from a Hardcoded Third-Party Mirror
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a coherent RAG document-search skill, but it needs review because it can persist indexed content and use under-disclosed external embedding/model sources with supply-chain and prompt-injection risks.
Review this before installing in sensitive environments. Avoid indexing secrets, credentials, personal data, or proprietary corpora unless you intend them to be stored locally and possibly embedded. Use the default local/simple path for sensitive data, require explicit approval before enabling OpenAI embeddings, pin and verify model artifacts, avoid the hardcoded mirror or make it opt-in, and update vulnerable transitive dependencies. Treat retrieved document text as untrusted evidence, not instructions for an agent to follow.
src/transformers-embedding.js:24Unverified Model Downloads from a Hardcoded Third-Party Mirror
src/rag2.js:207Retrieved Documents Are Inserted Verbatim into Downstream LLM Prompts
src/retriever.js:45Invalid Chunking Configuration Can Cause a Non-Terminating Loop
This content may contain harmful instructions that could cause physical harm if followed. CRITICAL: Review carefully before use.
"disgrace": 29591,
"infused": 29592,
"pudding": 29593,
"stalks": 29594,
"##urbed": 29595,
"arsenic": 29596,
"leases": 29597,
"##hyl": 29598,
"##rrard": 29599,
"collarbone": 29600,
"##waite": 29601,
This content may contain harmful instructions that could cause physical harm if followed. CRITICAL: Review carefully before use.
"disgrace": 29591,
"infused": 29592,
"pudding": 29593,
"stalks": 29594,
"##urbed": 29595,
"arsenic": 29596,
"leases": 29597,
"##hyl": 29598,
"##rrard": 29599,
"collarbone": 29600,
"##waite": 29601,
protobufjs 7.5.4 is reported with multiple serious advisories including denial of service and code-generation-related injection issues. Even though this is transitive through onnxruntime-web, bundling a dependency with numerous critical flaws materially increases attack surface, especially in a retrieval/ML skill that may ingest external model artifacts or serialized data formats.
tar 7.5.12 has multiple critical advisories including decompression and parsing denial-of-service conditions. This matters because onnxruntime-node has an install script and depends on tar, so the package introduces risk during installation and possibly runtime archive handling; supply-chain and CI environments are especially exposed if attacker-controlled archives can be processed.
Referenced artifact was not completely inspected
node src/rag-skill.js init my_docs
Referenced artifact was not completely inspected
node src/rag-skill.js init my_docs
Referenced artifact was not completely inspected
node src/rag-skill.js init my_docs
Referenced artifact was not completely inspected
node src/rag-skill.js init my_docs
Referenced artifact was not completely inspected
import { RAGRetriever } from './src/retriever.js';
Referenced artifact was not completely inspected
node test/retriever.test.js
form-data 4.0.5 is flagged for CRLF injection via multipart field names/values. This can be dangerous if the skill constructs multipart requests from untrusted input, because attacker-controlled fields may smuggle unintended headers or manipulate downstream HTTP request bodies; the presence of the OpenAI client in the dependency tree makes outbound multipart-capable HTTP behavior plausible.
sharp 0.34.5 is flagged for inherited native-library vulnerabilities in image parsing libraries such as libvips/libheif. This is particularly relevant in ML/retrieval contexts because image preprocessing is common; if the skill ever accepts attacker-supplied images, native parser flaws can lead to crashes, denial of service, or potentially memory-safety exploitation in the underlying library stack.
The entire skill design document is written in Chinese and does not indicate that users may choose another language or that the language restriction is required for a region-specific purpose. Under the policy, forcing a specific language without opt-in is a natural-language policy violation.
This markdown file documents commands and API examples for adding arbitrary documents into a local RAG database, including examples that read files and store their contents for later retrieval. The README does not warn users that ingested content may contain sensitive or private data that will be persisted and surfaced in retrieval results, which is a user-data/privacy relevant behavior for markdown descriptions.
The natural-language note says the current system '只支持英文分词' (only supports English tokenization), which imposes a language constraint. The file does not present this as a user-selectable option or clearly justify it as a narrowly scoped region/language-specific tool, so it conflicts with the language/locale policy criteria.
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).
nquest": 9187,
"dubbed": 9188,
"##tv": 9189,
"halt": 9190,
"brave": 9191,
"brunswick": 9192,
"doi": 9193,
"arched": 9194,
"curtis": 9195,
"divorced": 9196,
"predominantly": 9197,
"somerset": 9198,
"streams": 9199,
"ugly": 9200,
"zoo": 9201,
"horrible": 9202,
"curved": 9203,
"buenos": 9204,
"fierce": 9205,
"dictionary": 9206,
"vector": 9207,
"theological": 9208,
"unions": 9209,
"handful": 9210,
"stability": 9211,
"chan": 9212,
"punjab": 9213,
"segments": 9214,
"##lly": 9215,
"altar": 9216,
"ignoring": 9217,
"gesture": 9218,
"monsters": 9219,
"pastor": 9220,
"##stone": 9221,
"thighs": 9222,
"unexpected": 9223,
"operators": 9224,
"abruptly": 9225,
"coin": 9226,
"compiled": 9227,
"associates": 9228,
"improving": 9229,
"migration": 9230,
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).
nquest": 9187,
"dubbed": 9188,
"##tv": 9189,
"halt": 9190,
"brave": 9191,
"brunswick": 9192,
"doi": 9193,
"arched": 9194,
"curtis": 9195,
"divorced": 9196,
"predominantly": 9197,
"somerset": 9198,
"streams": 9199,
"ugly": 9200,
"zoo": 9201,
"horrible": 9202,
"curved": 9203,
"buenos": 9204,
"fierce": 9205,
"dictionary": 9206,
"vector": 9207,
"theological": 9208,
"unions": 9209,
"handful": 9210,
"stability": 9211,
"chan": 9212,
"punjab": 9213,
"segments": 9214,
"##lly": 9215,
"altar": 9216,
"ignoring": 9217,
"gesture": 9218,
"monsters": 9219,
"pastor": 9220,
"##stone": 9221,
"thighs": 9222,
"unexpected": 9223,
"operators": 9224,
"abruptly": 9225,
"coin": 9226,
"compiled": 9227,
"associates": 9228,
"improving": 9229,
"migration": 9230,
The file's human-readable comments and docstrings consistently use Chinese, which effectively forces a specific language for operators or maintainers reading the skill. Under the stated policy, language-specific natural-language instructions should not be imposed without explicit opt-in or justification.
The hardcoded OpenAI endpoint confirms that embedding requests leave the local environment for a third-party service. The danger is contextual rather than inherently malicious: in a data-processing skill, remote transmission without strong safeguards can violate privacy expectations or compliance requirements.
return this.cache[cacheKey];
}
const response = await fetch('https://api.openai.com/v1/embeddings', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
The hardcoded OpenAI endpoint confirms that embedding requests leave the local environment for a third-party service. The danger is contextual rather than inherently malicious: in a data-processing skill, remote transmission without strong safeguards can violate privacy expectations or compliance requirements.
return this.cache[cacheKey];
}
const response = await fetch('https://api.openai.com/v1/embeddings', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
This code sends arbitrary input text to the OpenAI Embeddings API, which is an external third party, without any built-in disclosure, consent check, or data-classification guardrail. If callers pass sensitive documents, secrets, or personal data, the module will exfiltrate that content off-host by design, creating a real privacy and data-handling risk in a RAG context.
The batch path uses the same external endpoint, so it carries the same trust-boundary risk with potentially greater volume. If exploited through misuse or misconfiguration, many records can be sent offsite before operators realize remote processing is occurring.
for (let i = 0; i < toProcess.length; i += this.batchSize) {
const batch = toProcess.slice(i, i + this.batchSize);
const response = await fetch('https://api.openai.com/v1/embeddings', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
The batch path uses the same external endpoint, so it carries the same trust-boundary risk with potentially greater volume. If exploited through misuse or misconfiguration, many records can be sent offsite before operators realize remote processing is occurring.
for (let i = 0; i < toProcess.length; i += this.batchSize) {
const batch = toProcess.slice(i, i + this.batchSize);
const response = await fetch('https://api.openai.com/v1/embeddings', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
The batch embedding path transmits all uncached texts to an external API with no warning or consent mechanism. In practice this can leak large volumes of internal corpus data at once, making the privacy impact higher than the single-text path because many documents may be uploaded in one operation.
This code sends full document content to an external embedding provider via embeddingProvider.embed(content) without any visible consent, minimization, or trust-boundary checks in this module. If documents contain sensitive or proprietary data, indexing them can disclose that data to a third party or another service boundary.
Detected: suspicious.env_credential_access