conversation-dataset-generator cahlen · COMPLETED
Generate multi-speaker conversational datasets for LLM fine-tuning. N-persona support, topic variation, conversation continuation, creative briefs with web search, character pools, Docker support. ShareGPT format with configurable role mapping.
github.com/cahlen/conversation-dataset-generator · ★ 17 · Forks 2 · Size 1.1 MB
SUMMARY
Technologies 12
Scored 12
Observed 12
Practices 6
Evidence 13
Skips 0
COVERAGE
Analyzed 65 files · 105 commits · 0 API calls
TECHNOLOGIES & DEPTH
Dockerfile LANGUAGE Depth 70
1 files · PRODUCTION
Markdown LANGUAGE Depth 70
11 files · PRODUCTION
YAML LANGUAGE Depth 70
15 files · PRODUCTION
Python LANGUAGE Depth 70
26 files · PRODUCTION
pip BUILD_TOOL Depth 80
1 files · CONFIGURATION
Apache Kafka MESSAGE_BROKER Depth 69
2 files · CONFIGURATION
Elasticsearch DATABASE Depth 69
2 files · CONFIGURATION
MongoDB DATABASE Depth 69
2 files · CONFIGURATION
NumPy LIBRARY Depth 74
4 files · PRODUCTION, TEST
RabbitMQ MESSAGE_BROKER Depth 69
2 files · CONFIGURATION
Redis CACHE Depth 69
2 files · CONFIGURATION
pytest TESTING Depth 53
8 files · TEST
PRACTICES
documentation · observedautomated_tests · observedcontinuous_integration · absentcontainerization · observedlinting · absentformatting · absent
ACTIVITY & OWNERSHIP
First commit 2025-04-10
Last commit 2026-05-01
Active months 2
Commits 105