Skip to content

code_intelligence.security: 4 vulnerability classes from the legacy analyzer have no canonical detector (XSS, weak TLS/auth, log info-disclosure, timing attacks) #12587

Description

@mrveiss

Summary

Discovered while converging the 3 duplicated analyzers onto code_intelligence/* in #12362. The legacy code_analysis.src.security_analyzer.SecurityAnalyzer (now a deprecated shim delegating to the canonical code_intelligence.security package) had 10 detection categories via regex. The #712/#9856 modularization that produced the canonical code_intelligence.security package carried over/exceeded most of them (SQL/command injection, hardcoded secrets, weak hash, insecure random, pickle/YAML deserialization, path traversal, debug-mode, missing-input-validation) but 4 categories have no detector in the canonical package today — this is a live production gap, not just a shim-parity issue, since api/code_intelligence.py and tasks/analytics_tasks.py already exclusively use the canonical package.

Also folded WEAK_ENCRYPTION (DES/RC4/3DES/Blowfish) into the canonical analyzer as part of #12362 since it mapped cleanly onto an existing-but-unused enum member — the 4 below do not have that shortcut.

Missing categories

  1. weak_authentication — SSL verification disabled (verify=False), check_hostname=False, disabled auth flags. Legacy regex: verify\s*=\s*False, check_hostname\s*=\s*False. No matching VulnerabilityType enum member exists (MISSING_AUTH_CHECK is semantically about authorization checks, not TLS verification) — needs a new enum member, e.g. INSECURE_TLS_VERIFICATION.
  2. xss_vulnerabilities — XSS_VULNERABILITY enum member exists but is never emitted by any check. Legacy patterns (innerHTML =, document.write, render_template_string with concatenation) were written for a **/*.py-only scan and are largely inapplicable to Python source — a proper implementation should scope to Jinja/render_template_string misuse in Python handlers, not raw JS string matching.
  3. information_disclosure (secrets in logs/print) — SENSITIVE_DATA_LOGGING enum member exists but is never emitted. Legacy regex flagged print/logger.* calls containing password|secret|key|token.
  4. timing_attacks — No enum member exists at all. Legacy regex (==\s*[^=].*?(?:password|secret|token|hash)) was extremely noisy (flags any equality comparison naming a variable hash); a real implementation should detect direct == comparison of two values from AST-traceable sensitive-named variables and recommend hmac.compare_digest, not blind regex on all comparisons.

Also noted (low priority)

code_intelligence/security/patterns.py::COMMAND_INJECTION_PATTERNS is defined and exported from the package __init__.py but never wired into SecurityAnalyzer._regex_analysis() — dead constant. Low priority since AST-based command-injection coverage (eval/exec/os.system+concat/subprocess shell=True) already covers the same ground with better precision; worth removing the dead wiring or actually using it, either is fine.

Why deferred from #12362

#12362 was scoped to converging the 3 duplicated analyzers onto the canonical implementations without capability loss for what's actually reachable from live consumers today. These 4 categories were already gaps in production before that PR (the canonical package has been the sole production security analyzer since #9856) — porting them properly (new enum members + low-noise AST/regex checks + tests, especially for weak_authentication and timing_attacks) is a distinct, non-trivial unit of work deserving its own PR and review rather than being rushed into a consolidation PR.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions