Skip to content

Add challenge 111: Split-KV Attention Reduction (Medium) - #300

Open
claude[bot] wants to merge 1 commit into
mainfrom
challenge/111-split-kv-attention-reduction
Open

claude[bot] wants to merge 1 commit into
mainfrom
challenge/111-split-kv-attention-reduction

Conversation

@claude

@claude claude Bot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Summary

  • New medium challenge: a Flash-Decoding-style merge kernel that combines partial attention outputs from multiple KV-cache chunks into the final attention result.
  • Teaches the log-sum-exp merge trick used to safely combine independently-normalized softmax outputs — a core primitive in production LLM inference stacks (vLLM, FlashInfer) for long-context decoding.
  • Inputs are per-split (partial_out, partial_lse) tuples; the solver computes stable weights, normalizes, and reduces along the split axis.

Validation

  • Verified with run_challenge.py --action submit on Tesla T4 — all tests passed.
  • pre-commit run --all-files passes clean (black, isort, flake8, clang-format, mojo format).

Test plan

  • challenge.py reference matches example output [[2, 3], [3, 2]]
  • Functional tests cover single-split, all-equal-lse, zero partial_out, non-power-of-2 splits/heads/dims, and realistic sizes
  • Performance test (num_splits=128, num_heads=64, head_dim=128) fits well within Tesla T4 memory
  • All 6 starter files present and compile without solving the problem
  • Linting passes

🤖 Generated with Claude Code

Add a Flash-Decoding-style merge kernel that combines per-chunk partial
attention outputs into the final attention result. The solver must apply
the log-sum-exp trick to merge partial (out, lse) tuples across KV splits
without numerical overflow — a core primitive in production LLM inference
stacks (vLLM, FlashInfer) for long-context decoding.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

Copy link
Copy Markdown
Contributor

Thanks for the PR! Challenge ID 111 is already taken on main (challenges/medium/111_*). Please renumber to the next free ID (currently unused near the top of the range: 120+, or any gap such as 95/97/98/99/100/101/102/103/104) and update the directory name + PR title accordingly. Happy to merge once the ID no longer collides.

Copy link
Copy Markdown
Contributor

Follow-up: IDs 98, 102, and 104 were just merged from other PRs — please avoid those. Prefer 95, 97, 99–101, 103, or 120+.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant