GitSocial · Repository PassportPARTIAL · 2026-10-01Z

Improving-LLM-Models-with-RLHF-PPO-DPO Mattral · PARTIAL

A modular, production-grade framework for Reinforcement Learning from Human Feedback (RLHF) with Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO).

github.com/Mattral/Improving-LLM-Models-with-RLHF-PPO-DPO · ★ 27 · Forks 6 · Size 667 KB

SUMMARY

Technologies 21
Scored 8
Observed 8
Practices 6
Evidence 23
Skips 13

COVERAGE

Analyzed 63 files · 34 commits · 0 API calls

TECHNOLOGIES & DEPTH

TOML LANGUAGE Depth 70
1 files · PRODUCTION
Markdown LANGUAGE Depth 70
20 files · PRODUCTION
YAML LANGUAGE Depth 70
5 files · PRODUCTION
Python LANGUAGE Depth 70
27 files · PRODUCTION
pip BUILD_TOOL Depth 80
1 files · CONFIGURATION
Poetry BUILD_TOOL Depth 80
1 files · CONFIGURATION
AWS SDK CLOUD Depth —
0 files · config only
Black QUALITY Depth —
0 files · config only
Blinker LIBRARY Depth —
0 files · config only
Click LIBRARY Depth —
0 files · config only
Jinja2 LIBRARY Depth —
0 files · config only
MarkupSafe LIBRARY Depth —
0 files · config only
NumPy LIBRARY Depth —
0 files · config only
Pydantic LIBRARY Depth 58
1 files · PRODUCTION
Requests LIBRARY Depth —
0 files · config only
Werkzeug LIBRARY Depth —
0 files · config only
lxml LIBRARY Depth —
0 files · config only
mypy QUALITY Depth —
0 files · config only
pandas LIBRARY Depth —
0 files · config only
pytest TESTING Depth 45
3 files · TEST
typing_extensions LIBRARY Depth —
0 files · config only

PRACTICES

documentation · observedautomated_tests · observedcontinuous_integration · observedcontainerization · absentlinting · observedformatting · absent

ACTIVITY & OWNERSHIP

First commit 2024-02-02
Last commit 2026-06-06
Active months 6
Commits 34

View on GitHub · GitSocial