faithfulness-judge sanlee-ys · PARTIAL
Can an LLM judge be trusted to catch unsupported claims? Measured against 189 human-labeled claims on public defense text: kappa 0.72-0.75 across two model tiers, 90-98% fabrication recall, and neither axis separates the tiers at this sample size.
github.com/sanlee-ys/faithfulness-judge · ★ 0 · Forks 0 · Size 565 KB
SUMMARY
Technologies 9
Scored 8
Observed 8
Practices 6
Evidence 10
Skips 1
COVERAGE
Analyzed 45 files · 82 commits · 0 API calls
TECHNOLOGIES & DEPTH
Markdown LANGUAGE Depth 70
11 files · PRODUCTION
TOML LANGUAGE Depth 70
1 files · PRODUCTION
YAML LANGUAGE Depth 70
9 files · PRODUCTION
Python LANGUAGE Depth 70
20 files · PRODUCTION
uv BUILD_TOOL Depth 80
1 files · CONFIGURATION
Poetry BUILD_TOOL Depth 80
1 files · CONFIGURATION
Pydantic LIBRARY Depth 58
1 files · PRODUCTION
pytest TESTING Depth 49
4 files · TEST
python-dotenv LIBRARY Depth —
0 files · config only
PRACTICES
documentation · observedautomated_tests · observedcontinuous_integration · observedcontainerization · absentlinting · observedformatting · absent
ACTIVITY & OWNERSHIP
First commit 2026-07-18
Last commit 2026-09-26
Active months 3
Commits 82