T09 · Insecure Skill Coding Practices
- Location
scripts/yotta_mirror.py:235- Finding
Deterministic Unsalted Student Pseudonyms Permit Re-identification and Cross-Report Linkage
- Content
View full analysis
Vulnerability Details
File Location:
scripts/yotta_mirror.py:235-238
Supporting Workflow Documentation:SKILL.md:89,references/privacy.md:17-19
Vulnerability Type: Weak pseudonymization of sensitive student identifiers
Risk Level: MediumVulnerable Code
python def anonymize_id(student): digest = hashlib.sha256(str(student).encode("utf-8")).hexdigest() return "S-" + digest[:8]Technical Analysis
The
--anonymizefeature hashes each student identifier directly with SHA-256, without a secret key or dataset-specific random salt, and retains only the first eight hexadecimal characters. This produces a deterministic 32-bit pseudonym.Student names, enrollment numbers, and other identifiers often have small or predictable input spaces. Anyone who receives a report can enumerate likely identifiers, calculate the same SHA-256 prefixes, and compare them with the report values. Truncating the digest to 32 bits also creates a material collision risk as the number of processed identifiers grows.
Because the transformation is deterministic and not scoped to a particular institution, dataset, or report, identical source identifiers produce identical pseudonyms across reports. This permits cross-report correlation even when the recipient cannot immediately recover the underlying identifier.
The documented workflow recommends using
--anonymizebefore externally sharing reports. Consequently, pseudonymized identifiers may cross from the trusted local environment to external recipients, while the original identifiers are expected to remain confidential.Attack Path
- An authorized user processes a score sheet using the
--anonymizeoption. - The Skill replaces each student identifier with
S-followed by the first eight hexadecimal characters of its unsalted SHA-256 digest. - The resulting report is shared with an external recipient under the assumption that the identifier ...[truncated 1025 chars]
- An authorized user processes a score sheet using the
- Remediation
View remediation
Remediation Suggestions
- Replace direct SHA-256 hashing with HMAC-SHA-256 using a cryptographically random secret key:
python digest = hmac.new(secret_key, student.encode("utf-8"), hashlib.sha256).hexdigest() - Use a separate secret key for each institution or dataset to prevent cross-dataset correlation.
- Retain a longer digest segment, such as at least 128 bits, to reduce collision risk.
- Store the HMAC key separately from exported reports and source datasets, with access restricted to authorized personnel.
- If reproducibility is unnecessary, generate random per-report identifiers and maintain any identifier mapping in a protected local file.
- Describe this feature as pseudonymization rather than irreversible anonymization, and warn that pseudonymized reports remain sensitive.
- Provide a key-management option or explicit dataset-scope parameter so users can control whether identifiers remain linkable across authorized reports.
- Replace direct SHA-256 hashing with HMAC-SHA-256 using a cryptographically random secret key:
