T09 · Insecure Skill Coding Practices
- Location
scanner.py:12- Finding
Unbounded Graph Expansion and Quadratic Memory Allocation
- Content
View full analysis
Vulnerability Details
File Location:
scanner.py:12-17, 66-69, 109
Vulnerability Type: Uncontrolled resource consumption
Risk Level: MediumVulnerable Code
python def list_files_recursive(directory, extensions=('.txt', '.def', '.erl', '.ex', '.md')): """Yield all files with given extensions under directory.""" for root, _, files in os.walk(directory): for f in files: if f.endswith(extensions): yield os.path.join(root, f)python for subj in subjects: for obj in objects: edges.append((subj, obj))python matrix = [[0] * n for _ in range(n)]Technical Analysis
The scanner recursively processes every supported file beneath the current working directory without enforcing limits on directory depth, file count, individual file size, total input size, unique node count, or generated edge count.
A triple containing parallel subject and object lists produces their Cartesian product. If there are
Ssubjects andOobjects, the parser createsS × Oedge tuples. Once parsing is complete, every unique subject and object becomes a graph node, and the scanner allocates a densen × nadjacency matrix. Consequently, matrix memory usage grows quadratically with the number of unique nodes, even when the graph itself is sparse.An attacker-controlled or unexpectedly large supported document can therefore consume excessive CPU and memory. Recursive directory scanning can amplify this condition when many crafted files are placed anywhere below the execution directory.
Attack Path
- An attacker supplies a supported file, such as a Markdown document, or places it beneath a directory that the user will scan.
- The file contains triples with large parallel subject and object lists, many unique node names, or both.
- The scanner recursively discovers and parses the file.
- Cartesian-product processing creates a large number of in-memory edge tuples.
- The scanner c ...[truncated 877 chars]
- Remediation
View remediation
Remediation Suggestions
-
Enforce configurable upper bounds on:
- Number of scanned files.
- Directory traversal depth.
- Individual and aggregate input size.
- Subject and object list lengths.
- Generated edges.
- Unique graph nodes.
-
Reject oversized input before Cartesian expansion and return a clear error rather than silently continuing.
-
Replace the dense adjacency matrix with a sparse representation, such as a dictionary mapping each source node to a set of target nodes:
python from collections import defaultdict adjacency = defaultdict(set) for source, target in edges: adjacency[source].add(target)-
If matrix output is mandatory, generate and print one row at a time rather than retaining the entire matrix in memory.
-
Deduplicate edges during parsing to reduce unnecessary storage and ensure statistics reflect unique graph edges.
-
Allow callers to specify an explicit input file or constrained input directory instead of always recursively scanning the current working directory.
-
Document safe operating limits and consider operating-system resource restrictions when processing untrusted documents.
-
