Skip to content

ML Course: Solution Articles for New Problems - #5696

Merged
neetcode-gh merged 7 commits into
mainfrom
ml/pr5-solution-articles
Apr 12, 2026
Merged

ML Course: Solution Articles for New Problems#5696
neetcode-gh merged 7 commits into
mainfrom
ml/pr5-solution-articles

Conversation

@ahmadbasyouni10

Copy link
Copy Markdown
Collaborator

Summary

  • Solution articles for all new ML problems across PRs 1-5: weight initialization, word embeddings, RMS normalization, training diagnostics, KV-cache, grouped query attention, dead ReLU detector, tokenization edge cases
  • Updated code-gpt article for W_o output projection rollout

Replaces previously merged chained PRs (#5666, #5668, #5674, #5688, #5689) which merged into each other instead of main.

Test plan

  • Verify each article renders correctly on the problem page
  • Verify math/code blocks display properly

Made with Cursor

ahmadbasyouni10 and others added 7 commits April 12, 2026 14:01
Add solution articles for multi-layer backpropagation, weight
initialization, and batch normalization. These follow the same
format as the existing 27 ML solution articles (Prerequisites,
Concept, Solution with Python tabs, Common Pitfalls, In the GPT
Project, Key Takeaways).

Made-with: Cursor
- multi-headed-self-attention: Add output_proj (W_o) linear layer after
  concatenating heads, matching standard practice
- transformer-block: Add output_proj to inner MultiHeadedSelfAttention
  class, consistent with multi-head attention problem
- weight-initialization: Rewrite check_activations to use raw weight
  matrices (torch.randn * std) instead of nn.Linear for cross-platform
  determinism

Made-with: Cursor
Reduces precision from 4 to 2 decimal places for the
check_activations method to absorb cross-platform floating
point differences in multi-layer matrix operations.

Made-with: Cursor
- Remove softmax from solution, return raw logits
- Add W_o output projection to inner MHA class
- Update all explanatory text: probabilities → logits
- Update shape table and key takeaways

Made-with: Cursor
- rms-normalization.md — RMSNorm vs LayerNorm, numpy implementation, Llama/Mistral context
- training-diagnostics.md — 3-method class (activation stats, gradient stats, diagnose), Karpathy debugging recipe

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two new ML course problems (PR4), bringing total to 34.

Dead ReLU Detector: detect_dead_neurons + suggest_fix with priority-ordered
diagnostic rules. 8 test cases covering all fix branches.

Tokenization Edge Cases: greedy tokenization exposing number inconsistency,
token counting, and fertility scoring. 8 test cases.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@neetcode-gh
neetcode-gh merged commit c5ca0e8 into main Apr 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants