T09 · Insecure Skill Coding Practices
- Location
scripts/grading_pipeline.py:806- Finding
Batch Size Limit Is Documented but Not Enforced
- Content
View full analysis
Vulnerability Details
File Location:
scripts/grading_pipeline.py, lines 806–821
Vulnerability Type: Unbounded input processing and resource exhaustion
Risk Level: MediumVulnerable Code
python # Collect patent numbers patent_numbers = [] if args.patents: patent_numbers.extend(args.patents) if args.patents_file: with open(args.patents_file, "r", encoding="utf-8") as f: for line in f: line = line.strip() if line and not line.startswith("#"): patent_numbers.extend( [p.strip() for p in line.replace(",", " ").split() if p.strip()] ) if not patent_numbers: parser.print_help() sys.exit(0) patent_numbers = list(dict.fromkeys(patent_numbers))The skill documentation states that each batch is limited to 50 patents, but the CLI does not enforce that restriction. Every supplied identifier is loaded into memory, deduplicated, processed, retained in the
resultslist, and subsequently written into a Word or Excel document.Technical Analysis
An attacker or untrusted user can provide a
--patents-filecontaining an arbitrarily large number of entries. The program reads the entire list and performs grading for every unique value. The resulting memory, CPU, document-generation, and disk usage scale with the number of entries.The risk is amplified because
run_grading()accumulates every generated result in memory before creating the output document. Word and Excel generation also creates a large in-memory document structure. This can exhaust the agent worker's memory or consume excessive CPU and disk space.The limit described in
SKILL.mdis therefore only advisory and does not establish a security boundary.Attack Path
- An attacker submits or references a patent list containing hundreds of thousands or millions of unique identifiers.
- The agent invokes the script with
--patents-fileor passes the identifiers through `--p ...[truncated 849 chars]
- Remediation
View remediation
Remediation Suggestions
Enforce the documented limit after combining and deduplicating all input sources:
python MAX_PATENTS = 50 patent_numbers = list(dict.fromkeys(patent_numbers)) if len(patent_numbers) > MAX_PATENTS: parser.error( f"A maximum of {MAX_PATENTS} unique patent numbers is allowed per batch." )Additional hardening should include:
- Limit the size of
--patents-filebefore reading it. - Reject excessively long patent identifiers and validate their format.
- Stream input records rather than loading an unbounded file into memory.
- Apply output-size, execution-time, memory, and disk quotas.
- Enforce the same limit inside
run_grading()so callers cannot bypass the CLI validation. - Rate-limit or budget MCP requests independently of the number of supplied identifiers.
- Add tests for 0, 1, 50, and 51 identifiers, as well as oversized input files.
- Limit the size of
