corpus-mill cahlen · PARTIAL
Multimodal video annotation pipeline — local GPU end-to-end (audio, vision, OCR, faces, brands, chat, music). Turns any long-form video with people into a time-aligned event corpus for synthetic-data construction, training-set curation, and analysis.
github.com/cahlen/corpus-mill · ★ 3 · Forks 0 · Size 23.2 MB
SUMMARY
Technologies 17
Scored 15
Observed 15
Practices 6
Evidence 18
Skips 2
COVERAGE
Analyzed 201 files · 106 commits · 0 API calls
TECHNOLOGIES & DEPTH
Dockerfile LANGUAGE Depth 70
1 files · PRODUCTION
Shell LANGUAGE Depth 70
1 files · PRODUCTION
JavaScript LANGUAGE Depth 70
7 files · PRODUCTION
JSON LANGUAGE Depth 70
1 files · PRODUCTION
TOML LANGUAGE Depth 70
1 files · PRODUCTION
Markdown LANGUAGE Depth 70
9 files · PRODUCTION
YAML LANGUAGE Depth 70
1 files · PRODUCTION
Python LANGUAGE Depth 70
142 files · PRODUCTION
SQL LANGUAGE Depth 70
1 files · PRODUCTION
Poetry BUILD_TOOL Depth 80
1 files · CONFIGURATION
Click LIBRARY Depth 66
2 files · PRODUCTION, TEST
FastAPI FRAMEWORK Depth 66
2 files · PRODUCTION, TEST
HTTPX LIBRARY Depth —
0 files · config only
NumPy LIBRARY Depth 78
32 files · PRODUCTION, TEST
Pydantic LIBRARY Depth —
0 files · config only
Uvicorn LIBRARY Depth 58
1 files · PRODUCTION
pytest TESTING Depth 74
30 files · PRODUCTION, TEST
PRACTICES
documentation · observedautomated_tests · observedcontinuous_integration · absentcontainerization · observedlinting · observedformatting · absent
ACTIVITY & OWNERSHIP
First commit 2026-04-26
Last commit 2026-05-06
Active months 2
Commits 106