Summary
Discovered while converging the 3 duplicated analyzers onto code_intelligence/* in #12362. The legacy code_analysis.src.security_analyzer.SecurityAnalyzer (now a deprecated shim delegating to the canonical code_intelligence.security package) had 10 detection categories via regex. The #712/#9856 modularization that produced the canonical code_intelligence.security package carried over/exceeded most of them (SQL/command injection, hardcoded secrets, weak hash, insecure random, pickle/YAML deserialization, path traversal, debug-mode, missing-input-validation) but 4 categories have no detector in the canonical package today — this is a live production gap, not just a shim-parity issue, since api/code_intelligence.py and tasks/analytics_tasks.py already exclusively use the canonical package.
Also folded WEAK_ENCRYPTION (DES/RC4/3DES/Blowfish) into the canonical analyzer as part of #12362 since it mapped cleanly onto an existing-but-unused enum member — the 4 below do not have that shortcut.
Missing categories
- weak_authentication — SSL verification disabled (
verify=False), check_hostname=False, disabled auth flags. Legacy regex: verify\s*=\s*False, check_hostname\s*=\s*False. No matching VulnerabilityType enum member exists (MISSING_AUTH_CHECK is semantically about authorization checks, not TLS verification) — needs a new enum member, e.g. INSECURE_TLS_VERIFICATION.
- xss_vulnerabilities —
XSS_VULNERABILITY enum member exists but is never emitted by any check. Legacy patterns (innerHTML =, document.write, render_template_string with concatenation) were written for a **/*.py-only scan and are largely inapplicable to Python source — a proper implementation should scope to Jinja/render_template_string misuse in Python handlers, not raw JS string matching.
- information_disclosure (secrets in logs/print) —
SENSITIVE_DATA_LOGGING enum member exists but is never emitted. Legacy regex flagged print/logger.* calls containing password|secret|key|token.
- timing_attacks — No enum member exists at all. Legacy regex (
==\s*[^=].*?(?:password|secret|token|hash)) was extremely noisy (flags any equality comparison naming a variable hash); a real implementation should detect direct == comparison of two values from AST-traceable sensitive-named variables and recommend hmac.compare_digest, not blind regex on all comparisons.
Also noted (low priority)
code_intelligence/security/patterns.py::COMMAND_INJECTION_PATTERNS is defined and exported from the package __init__.py but never wired into SecurityAnalyzer._regex_analysis() — dead constant. Low priority since AST-based command-injection coverage (eval/exec/os.system+concat/subprocess shell=True) already covers the same ground with better precision; worth removing the dead wiring or actually using it, either is fine.
Why deferred from #12362
#12362 was scoped to converging the 3 duplicated analyzers onto the canonical implementations without capability loss for what's actually reachable from live consumers today. These 4 categories were already gaps in production before that PR (the canonical package has been the sole production security analyzer since #9856) — porting them properly (new enum members + low-noise AST/regex checks + tests, especially for weak_authentication and timing_attacks) is a distinct, non-trivial unit of work deserving its own PR and review rather than being rushed into a consolidation PR.
Summary
Discovered while converging the 3 duplicated analyzers onto
code_intelligence/*in #12362. The legacycode_analysis.src.security_analyzer.SecurityAnalyzer(now a deprecated shim delegating to the canonicalcode_intelligence.securitypackage) had 10 detection categories via regex. The #712/#9856 modularization that produced the canonicalcode_intelligence.securitypackage carried over/exceeded most of them (SQL/command injection, hardcoded secrets, weak hash, insecure random, pickle/YAML deserialization, path traversal, debug-mode, missing-input-validation) but 4 categories have no detector in the canonical package today — this is a live production gap, not just a shim-parity issue, sinceapi/code_intelligence.pyandtasks/analytics_tasks.pyalready exclusively use the canonical package.Also folded
WEAK_ENCRYPTION(DES/RC4/3DES/Blowfish) into the canonical analyzer as part of #12362 since it mapped cleanly onto an existing-but-unused enum member — the 4 below do not have that shortcut.Missing categories
verify=False),check_hostname=False, disabled auth flags. Legacy regex:verify\s*=\s*False,check_hostname\s*=\s*False. No matchingVulnerabilityTypeenum member exists (MISSING_AUTH_CHECKis semantically about authorization checks, not TLS verification) — needs a new enum member, e.g.INSECURE_TLS_VERIFICATION.XSS_VULNERABILITYenum member exists but is never emitted by any check. Legacy patterns (innerHTML =,document.write,render_template_stringwith concatenation) were written for a**/*.py-only scan and are largely inapplicable to Python source — a proper implementation should scope to Jinja/render_template_stringmisuse in Python handlers, not raw JS string matching.SENSITIVE_DATA_LOGGINGenum member exists but is never emitted. Legacy regex flaggedprint/logger.*calls containingpassword|secret|key|token.==\s*[^=].*?(?:password|secret|token|hash)) was extremely noisy (flags any equality comparison naming a variablehash); a real implementation should detect direct==comparison of two values from AST-traceable sensitive-named variables and recommendhmac.compare_digest, not blind regex on all comparisons.Also noted (low priority)
code_intelligence/security/patterns.py::COMMAND_INJECTION_PATTERNSis defined and exported from the package__init__.pybut never wired intoSecurityAnalyzer._regex_analysis()— dead constant. Low priority since AST-based command-injection coverage (eval/exec/os.system+concat/subprocess shell=True) already covers the same ground with better precision; worth removing the dead wiring or actually using it, either is fine.Why deferred from #12362
#12362 was scoped to converging the 3 duplicated analyzers onto the canonical implementations without capability loss for what's actually reachable from live consumers today. These 4 categories were already gaps in production before that PR (the canonical package has been the sole production security analyzer since #9856) — porting them properly (new enum members + low-noise AST/regex checks + tests, especially for weak_authentication and timing_attacks) is a distinct, non-trivial unit of work deserving its own PR and review rather than being rushed into a consolidation PR.