I build open-source AI and media tools that turn difficult technical workflows into software people can actually use.
My work focuses on speech synthesis, audiobook automation, local AI, model training, and developer tooling. I enjoy taking ambitious ideas from early experiments through packaging, documentation, cross-platform support, and real-world adoption.
Portfolio · YouTube · ebook2audiobook · Hugging Face · Discord community
- Working in information technology at Adult Swim
- Studying at Georgia State University
- Conducting research with the Alser Lab in GSU's Department of Computer Science
At the Georgia State Undergraduate Research Conference, ebook2audiobook received:
- 1st Place — Applied Research and Entrepreneurship, Poster Presentations
- Global Engagement Award
Project sponsor: Professor Mohammed Alser. View the official GSURC winners.
An open-source platform for converting ebooks into fully chaptered audiobooks locally on CPU or GPU.
- Supports multilingual speech synthesis, voice cloning, OCR, translation, metadata, and multiple ebook and audio formats
- Provides a Gradio interface, headless CLI, Docker workflows, and remote notebook deployments
- Runs across Windows, macOS, Linux, Apple Silicon, NVIDIA CUDA, AMD ROCm, Intel XPU, and other hardware targets
Automatic multi-character audiobook generation built into the ebook2audiobook ecosystem.
- Identifies characters and attributes dialogue to speakers
- Produces voice-tagged SML for distinct narrator and character voices
- Brings the original VoxNovel concept into the actively maintained ebook2audiobook workflow
A unified environment for training and adapting speech models.
- Supports synthetic-data-assisted voice training workflows
- Turns short voice references into larger training datasets
- Enables compact voice models designed for faster, lower-resource deployment
| Project | What it demonstrates |
|---|---|
| JellyDisk | Desktop and CLI media tooling for turning Jellyfin libraries into polished DVD sets |
| Sound Monitor | Real-time environmental noise collection, analysis, visualization, and evidence-based reporting |
| offlineYoutube | Local video downloading, Whisper transcription, embeddings, and semantic search |
| VoxNovel | The legacy predecessor to E2A-SML and an early multi-character audiobook pipeline |
| Doc2 Interview | Document processing and AI-generated interview-style conversations |
| Custom Quote Attribution Pipeline | BERT and DistilGPT-2 experimentation for quotation identification and speaker attribution |
- Languages: Python, Go, JavaScript
- AI and audio: TTS, voice cloning, speech-model training, Whisper, NLP, embeddings
- Product engineering: Docker, Gradio, FFmpeg, CLI tools, cross-platform packaging, automated testing
- Platforms: GitHub Actions, Hugging Face, Google Colab, Kaggle
- Coqui TTS — contributed additional model support
- xtts-finetune-webui — improved functionality and produced a pre-built Docker image
I maintain the ebook2audiobook community and am open to interesting engineering opportunities and collaborations.
Visit my portfolio, watch my tutorials on YouTube, or reach me through the ebook2audiobook Discord community.




