Skip to content
#

reward-functions

Here are 24 public repositories matching this topic...

Stock-Predictor-V4

A reinforcement learning model specialized in stock prediction utilizing deep learning techniques, incorporating reward mechanisms, compatible with any machine equipped with Python.

  • Updated May 18, 2024
  • Python

Group Relative Policy Optimization (GRPO) implementations - NanoAhaMoment, GRPO:Zero, Simple GRPO, and GRPO from Scratch - spanning vLLM + DeepSpeed, custom Transformer stack, Bottle HTTP reference server, and pure PyTorch. Compares generation backends, reference policy strategies, reward designs, and loss functions on GSM8K and Countdown tasks.

  • Updated Jul 22, 2026
  • Python

Improve this page

Add a description, image, and links to the reward-functions topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the reward-functions topic, visit your repo's landing page and select "manage topics."

Learn more