T08 · Insecure Dependencies
- Location
README.md:8- Finding
Unpinned Third-Party Dependency Installation
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a coherent iFLYTEK transcription skill, but users should know it uploads selected audio to a third-party API and currently works as MP3-only despite broader claims.
Install only if you are comfortable sending the selected MP3 audio and related metadata to iFLYTEK/XFYun using your XFEI credentials. Avoid confidential, regulated, or non-consented recordings unless your policy allows that provider, and use a virtual environment with pinned dependencies where possible.
README.md:8Unpinned Third-Party Dependency Installation
The code generally matches the stated primary purpose of transcribing audio with the iFLYTEK/XFYun speed transcription service, including upload, task creation, polling, and result parsing. However, there are material description-behavior mismatches. Most importantly, the transcribe() method explicitly rejects any file whose extension is not .mp3, contradicting the declared support for WAV/PCM/MP3. The create_task() method also hardcodes format='audio/mp3' and encoding='lame', reinforcing that only MP3 is actually handled by this implementation. Additionally, the description claims automatic language detection, but the code exposes --language with a default of zh_cn and never performs any detection logic. These are not minor implementation details; they change what users can actually do with the skill. There is no evidence of unrelated malicious or undeclared exfiltration behavior beyond the expected API calls for transcription.
The README explicitly states that audio is sent to third-party iFLYTEK endpoints, but it does not provide a clear user-facing privacy or data-transmission warning before use. Because audio may contain sensitive personal, medical, legal, or business information, users could unknowingly transmit regulated or confidential data to an external service.
The skill documents use of environment variables, network access, and output file writing, but it does not declare explicit tool permissions or allowed-tools scope. This weakens least-privilege controls and makes it harder for a host system or reviewer to understand and constrain what the skill can do, increasing the risk of unintended file writes or outbound data transfer.
The skill handles potentially sensitive audio such as meetings, interviews, medical consultations, and legal proceedings, yet it does not clearly warn users that recordings are uploaded to a third-party service. This creates a significant privacy and compliance risk because users may disclose confidential or regulated data without informed consent.
The skill uploads user-provided audio to an external third-party transcription service, but the code provides no explicit notice at the point of use that local audio will leave the system. In an agent-skill context, this matters because audio may contain sensitive conversations, personal data, or regulated content, and users may reasonably assume local-only processing unless clearly told otherwise.
The code sets default transcription parameters to zh_cn and mandarin, and the CLI repeats zh_cn as the default. This enforces a specific language/locale behavior by default rather than asking the user to choose or clearly opt in, which matches the language/locale policy violation rule.
The manifest describes a transcription skill for WAV, PCM, and MP3 inputs, but the transcribe workflow explicitly rejects any file whose extension is not .mp3. This is a direct behavior mismatch at the skill level because users invoking the skill for WAV or PCM transcription would be told the format is unsupported.
The command-line description says 'Transcribe audio files' and the positional argument is documented generically as a 'Path to audio file', which implies normal support for common audio formats. However, the core transcribe method rejects every non-MP3 file, so the inline documentation overstates what the program actually accepts.
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
if args.output:
output_path = Path(args.output)
if args.output_format == "json":
output_path.write_text(
json.dumps(result, ensure_ascii=False, indent=2),
encoding='utf-8'
)
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
encoding='utf-8'
)
else:
output_path.write_text(text, encoding='utf-8')
print(f"\nSaved to: {args.output}")
except Exception as e:
The file includes substantial Chinese-only user-facing guidance such as the error table's friendly prompts and the FAQ section, while the rest of the skill is primarily in English. This creates a language-policy concern because the skill imposes a specific language for part of the experience without explicitly asking for user preference or stating that the skill is intended for Chinese-speaking users.
The script writes transcription text or JSON results to a user-specified path, which is a file write operation covered by the warning rule. Although the save action is user-requested via --output, there is no prior comment, help text, or disclosure that the tool will create or overwrite a file at the provided path.
No suspicious patterns detected.