AgentDish directory
reinforcement-learning
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#896
↓ -6
nano-llm-posttraining
Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior. |
AI Development / LLM Training & Fine-tuning | 84 | ↓ -6 | 41 days ago | Details |
|
#1677
↑ +2
GenZ LLM
A small post-trained Qwen2.5-0.5B-Instruct model tuned to write in Gen Z slang, with training and inference notebooks plus dataset files in the repo. |
AI Models / Fine-Tuned LLMs | 79 | ↑ +2 | 123 days ago | Details |
|
#1712
↑ +6
EdotEnv (E.env)
Market-derived RL environments for training agents on quant trading and long-horizon planning under adversarial noise. |
AI/ML / Reinforcement Learning | 78 | ↑ +6 | 37 days ago | Details |
|
#1773
↑ +6
Agora-1: The Multi-Agent World Model
Agora-1 is a multi-agent world model from Odyssey that simulates shared real-time environments for up to four participants, human or AI, with a focus on gaming, robotics, reinforcement learning, and foundation model research. |
AI Research / World Models | 78 | ↑ +6 | 115 days ago | Details |
|
A paper page about FrogNano, a 4B coding agent trained with RL on roughly 1,500 SWE environments using synthetic tasks and no distillation from a larger model. |
Agents / Coding Agent | 74 | ↓ -1 | 2 days ago | Details |
|
A blog post describing a small reinforcement-learning agent trained with PPO to play and beat a Pokelike/Pokerogue-style game, including the input representation, model architecture, and training approach. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 99 days ago | Details |
|
#1983
↑ +1
World Models for Planning Agents
An educational article explaining world models, latent states, dynamics learning, and planning for agents, with examples from gridworld, Dreamer, and MuZero. |
Writing / Copywriting | 73 | ↑ +1 | 117 days ago | Details |
|
arXiv paper on self-speculating LLM agents that predict their next tool call to hide tool latency and improve next-call Hit@1 while preserving task success. |
Research / AI Paper | 72 | ↑ +1 | 44 days ago | Details |