2 results found Sort:

21
329
apache-2.0
6
🌾 OAT: A research-friendly framework for LLM online alignment, including preference learning, reinforcement learning, etc.
Created 2024-10-15
36 commits to main branch, last one 5 days ago
implementation of distributed reinforcement learning with distributed tensorflow
Created 2020-04-07
44 commits to master branch, last one 4 years ago