T09 · Insecure Skill Coding Practices
- Location
main.py:63- Finding
OS Command Injection Through User-Controlled CLI Arguments
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a real-estate scraper that openly teaches users how to bypass anti-bot and CAPTCHA protections while saving reusable browser sessions.
Install only if you have explicit permission to collect data from the target sites and are prepared to manage saved browser sessions as secrets. Avoid using real account sessions, do not upload CAPTCHA images to third-party solvers, pin dependencies before installation, and do not pass untrusted CLI values to main.py.
main.py:63OS Command Injection Through User-Controlled CLI Arguments
scripts/bypass_anjuke.sh:155Reusable Browser Sessions Are Persisted Without Access Protection
docs/captcha_strategies.md:107CAPTCHA Images and API Credentials May Be Sent to an Unverified Third Party
SKILL.md:57Unpinned Dependencies and Global Package Installation Create Supply-Chain Exposure
The description claims a comprehensive crawler supporting 安居客、贝壳找房、链家、搜房网 with anti-crawling bypass and data extraction. However, the provided code chunk is mainly a CLI launcher that delegates work to other scripts via subprocess. In the visible code, Python-mode scraping exists only for 安居客 and 贝壳; 链家 is supported only through a shell script in agent-browser mode; 搜房网 is listed as a choice but not implemented in either Python or agent-browser mode. Because a declared target site is unimplemented and another is only partially supported, the description overstates actual behavior of this code chunk. Also, the batch/config orchestration behavior is an undeclared additional capability, though relatively related. Overall, this is a description-behavior mismatch.
该代码与声明存在明显范围不一致。声明描述的是一个覆盖多个房产平台的综合抓取技能,但提供的代码只包含 anjuke_crawler.py,且逻辑全部围绕安居客展开,没有贝壳、链家、搜房网相关请求、解析、路由或适配器实现。因此其主要能力被明显高估。另一个偏差是,声明提到“包含反爬虫绕过策略”,而实际代码仅使用常规请求头、随机睡眠和注释性建议,属于基础限频规避措施,不足以支撑广义的反爬虫绕过能力表述。整体上,声明的功能范围和实现内容不匹配。
The declared description presents a broad multi-site real-estate scraping skill covering Anjuke, Beike, Lianjia, and Fang, with anti-crawling bypass and data extraction. The provided code chunk is much narrower: it is a Beike-specific test script using agent-browser to mimic a mobile browser, inspect the page, and check for CAPTCHA/anti-bot behavior. While this does align partially with the anti-bot bypass aspect and touches light content capture, it does not implement the claimed comprehensive multi-site crawler behavior or structured extraction functionality. Therefore the description overstates the scope and functionality of this code chunk.
The declared description presents a broad real-estate scraping skill covering several platforms and including both anti-crawling bypass and data extraction. The supplied code chunk only implements a Lianjia-specific browser automation flow to reduce captcha triggers and establish a verified session. While anti-bot bypass is consistent with part of the description, the primary behavior here is much narrower than declared: there is no scraping logic, parsing, or extraction of property data, and no support for Anjuke, Beike, or SouFun in this chunk. Therefore the description materially overstates the functionality represented by the provided code.
The guidance to use real cookies, referers, and manually complete verification to restore an authenticated session is a direct workaround for verification and anti-bot mechanisms. In this context, it can facilitate unauthorized continued access under a validated human session, making the skill substantially more dangerous than a normal scraper.
This document explicitly provides operational guidance for bypassing CAPTCHA, rotating proxies, spoofing device fingerprints, reusing verified sessions, and reducing detection by anti-bot systems. In the context of a real-estate crawler skill, these are not ordinary extraction instructions; they materially enable evasion of access controls and abusive scraping against third-party sites.
The file has undeclared shell-execution capability via subprocess and bash/python invocations, which materially expands what the skill can do beyond ordinary data extraction. Missing permission declaration weakens governance and user awareness, making dangerous execution paths easier to ship unnoticed.
Tool parameter abuse is present because attacker-controlled parameters are interpolated into shell commands and executed with shell=True. Since this skill accepts several CLI inputs and then passes them into python3/bash command strings, a crafted value can escape the intended parameter boundary and run arbitrary OS commands.
import subprocess
try:
result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
if result.returncode == 0:
print(f"命令执行成功")
The code issues live requests to external real-estate sites while the static finding indicates network capability is not covered by declared permissions. In an agent skill, undeclared external communication is dangerous because it can exfiltrate data, interact with third parties unexpectedly, or bypass operator review.
The code issues live requests to external real-estate sites while the static finding indicates network capability is not covered by declared permissions. In an agent skill, undeclared external communication is dangerous because it can exfiltrate data, interact with third parties unexpectedly, or bypass operator review.
The file gives direct operational instructions for bypassing CAPTCHA and avoiding detection, including 'simulate normal users,' use proxy IPs, complete verification manually, then save and reuse the verified session. In this skill context, that is especially dangerous because the skill is a multi-site crawler, so the content is tailored to defeating platform protections to enable continued scraping rather than legitimate access.
The README advertises exporting scraped real-estate data to JSON, CSV, HTML, and charts without warning that scraped content may include personal or sensitive information and will be written to local files. Because this skill targets large-scale collection from real-estate platforms, users may inadvertently store regulated or private data in insecure locations or redistribute it without safeguards.
The README explicitly instructs users to save a verified browser session after completing CAPTCHA, which can preserve authentication cookies, tracking identifiers, and other session artifacts. In the context of an anti-crawler skill, this increases the risk of mishandling sensitive session data and can enable unauthorized reuse if the saved state is stored insecurely or shared.
The skill explicitly instructs users to evade anti-bot defenses using fingerprint spoofing, proxies, cookie/session manipulation, and CAPTCHA workarounds, but does not present a clear upfront warning about the legal, policy, and abuse risks. In this context, the omission is significant because the skill's core purpose is not neutral scraping but operational bypass of third-party protections.
The skill encourages exporting scraped data and saving browser/session state without warning that these artifacts may include personal data, account tokens, cookies, or other sensitive information. In a crawler that manipulates authenticated or semi-authenticated sessions, this can lead to accidental retention, leakage, or reuse of sensitive data.
The skill’s read_when triggers are broad and operationally oriented, including requests to scrape real-estate websites and bypass anti-bot mechanisms, without narrowing scope, authorization, or target constraints. This increases the chance the skill is invoked in inappropriate contexts and can enable unauthorized scraping behavior more readily, especially given the explicitly declared browser automation capability.
The configuration explicitly documents anti-crawler evasion techniques such as simulating human behavior, setting cookies, and introducing random delays. In the context of a real-estate crawler skill, this goes beyond ordinary scraping configuration and meaningfully facilitates bypassing site protections and access controls.
The natural-language tips directly encourage bypassing anti-crawler controls without any policy guardrails or user-choice mechanism. Because the skill is specifically a crawler for third-party real-estate sites, this guidance makes operational misuse more likely and lowers the barrier to violating site restrictions.
Recommending proxies, session management, and mobile-device simulation is explicit advice for avoiding crawler detection. This meaningfully increases the capability of the skill to evade defensive controls on third-party websites.
The proxy list enables proxy rotation infrastructure that can be used to distribute requests and reduce detection by target sites. In this skill's context, that materially supports anti-detection behavior rather than being necessary for basic configuration alone.
This section instructs users to save and reload authenticated browser session state and cookies after passing anti-bot checks, but does not warn that these artifacts may contain active authentication tokens, tracking identifiers, and other sensitive state. If the saved session files are reused, shared, or stored insecurely, they can enable unauthorized access, account/session hijacking, or unintended persistence of a verified identity across scraping runs.
The proxy rotation guidance encourages routing scraping traffic through third-party proxies without warning that proxies can observe destination URLs, request metadata, and potentially collected data, while also exposing the operator's network behavior and identifiers. In the context of anti-crawler evasion, this increases the chance that users will send sensitive traffic through untrusted infrastructure and mishandle legal, privacy, or data-exposure risks.
The skill metadata frames the capability as property-site crawling and extraction, but this file broadens scope into anti-detection behavior such as CAPTCHA handling, referer spoofing, browser fingerprint tuning, and alternate-page strategies to get around defenses. That mismatch increases risk because users may deploy the skill as a seemingly normal crawler while it contains hidden or understated evasion functionality.
The CAPTCHA-solving example uploads an image to an external recognition service, but the document does not warn that challenge data will leave the local environment or discuss legal, privacy, and confidentiality implications. This creates data-governance and third-party exposure risk, especially if screenshots or challenge pages contain identifiers, cookies, or other sensitive context.
The example performs an outbound POST to a third-party CAPTCHA-solving endpoint and sends image data externally. In this skill's context, the transmission supports CAPTCHA circumvention and may expose challenge content or related sensitive material to an unvetted external service.
import requests
def solve_captcha(image_path):
response = requests.post(
"https://captcha.service.com/api/solve",
files={"image": open(image_path, "rb")},
headers={"Authorization": "Bearer YOUR_API_KEY"}
No suspicious patterns detected.