A QLoRA fine-tuned email writing assistant. Adapts a small open LLM (Qwen2.5-1.5B-Instruct) to write short, casual emails instead of formal corporate boilerplate, using PEFT for the LoRA adapters and bitsandbytes for 4-bit quantization. Only about 4.4M of the model's 1.5B parameters (0.28%) are actually trained.
The project goes beyond a single before/after demo: it includes a prompted-base-model baseline, pairwise LLM-judge evaluation, fact preservation and fabrication testing, a generalization check across unseen categories, an inference benchmark, and a FastAPI + MCP serving layer. Every number below comes from a script in eval/ and is logged with full detail in RESULTS.md.
Full numbers, methodology, and honest caveats are in RESULTS.md. The headlines:
Fine-tuning beats a well-prompted base model, not just a raw one. A prompted-base-model baseline was built first, since a good system prompt with a few examples can match a light fine-tune. Pairwise LLM-judge evaluation across 60 unseen topics:
| Comparison | QLoRA win rate |
|---|---|
| QLoRA vs base model | 100% |
| QLoRA vs prompted-base model | 73% |
QLoRA fabricates far less than either baseline. 18 synthetic prompts embedded specific facts (names, days, times, numbers) and checked whether each model preserved them or invented different ones:
| Variant | Fact preservation | Fabrication rate |
|---|---|---|
| base | 81% | 44% |
| prompted-base | 90% | 17% |
| QLoRA | 85% | 6% |
Style holds up across categories it never saw in training. Tested on 60 topics spanning 6 categories (recruiter replies, scheduling, follow-ups, rejections, status updates, cold outreach), none of which overlap the training data. Style consistency scored 3.8 to 4.6 out of 5 across all categories, a spread of 0.8, which supports the model having learned a general tone rather than memorizing a handful of patterns.
The fine-tune is also faster to use. QLoRA emails are ready in about 3 seconds versus about 8 for the base model, because it learned to write short emails instead of rambling to the token limit. VRAM usage is essentially unchanged (the adapter is a few MB on top of the quantized base model).
Two honest caveats worth knowing before trusting these numbers fully: the judge/human agreement sanity check (comparing the LLM judge's picks against a human's) hasn't been run yet, and the training set, while now 264 unique emails, is still small by research standards, so "generalization" here means a modest style-adaptation result rather than a large-scale claim.
Topic: Rescheduling next week's call
Before (base model): "Subject: Rescheduling Next Week's Call. Dear [Recipient's Name], I hope this message finds you well. After careful consideration of our current project deadlines and priorities, we have decided to reschedule the upcoming call... Best regards, [Your Full Name] [Your Position] [Your Contact Information]"
After (QLoRA fine-tuned): "Hi guys, I was wondering if we could reschedule our meeting for thursday at 3pm? I'm pretty sure i got my calendar wrong last time but that is the best slot i can find right now. Thanks!"
More examples in outputs/three_way_comparison.md (base, prompted-base, and QLoRA side by side across all 60 test topics).
The fine-tuned model is also wired up as an MCP tool, invoked directly from Claude Desktop:
Works with Python 3.10+ and a CUDA GPU. A local NVIDIA GPU, Google Colab, Kaggle, or a rented cloud GPU all work the same way. CPU-only execution works but is very slow.
git clone https://github.com/bazzal99/MailTune.git
cd MailTune
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txtVerify the GPU is visible:
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"This should print True and a GPU name. If it prints False, reinstall PyTorch with a CUDA-enabled build matching your driver (see pytorch.org).
For the evaluation scripts that use an LLM judge (Weeks 2 and 4), create a .env file in the repo root with ANTHROPIC_API_KEY=sk-ant-....
Running in Google Colab or Kaggle instead
- Upload/unzip this project into the notebook's working directory.
- Set the runtime/accelerator to a GPU (Colab: Runtime, Change runtime type, T4 GPU. Kaggle: Settings, Accelerator, GPU).
- Run the same
pip install -r requirements.txtand the verification command above in a notebook cell, prefixing shell commands with!. - Run the pipeline scripts the same way, e.g.
!python scripts/train_qlora.py. - To retrieve results, zip the
outputs/folder and download it, or usefrom IPython.display import FileLink; FileLink('outputs/three_way_comparison.md')for a single file.
1. Data preparation. data/emails_raw.txt holds source emails in SUBJECT: ... / EMAIL: ... blocks. The current file has 264 unique synthetic emails, generated by scripts/generate_synthetic_data.py with facts injected programmatically so fabrication is prevented by construction, not just hoped for. Swap in real sent emails for a genuine personal style-transfer result (strip signatures, quoted reply chains, and anything sensitive first).
python scripts/generate_synthetic_data.py # optional, regenerates the synthetic corpus
python scripts/prepare_dataset.pyProduces data/train.jsonl, an instruction/response dataset pairing an email topic with the target email body.
2. Training. QLoRA fine-tuning via PEFT, bitsandbytes, and TRL's SFTTrainer.
python scripts/train_qlora.pySaves the trained adapter to outputs/qlora-adapter/. Takes a few minutes on a T4-class GPU with this dataset size.
3. Evaluation. The full eval suite lives in eval/, run in this order:
python eval/prompted_baseline.py # base vs prompted-base vs QLoRA, all topics
python eval/audit_train_overlap.py # confirms eval topics aren't in the training set
python eval/pairwise_judge.py # LLM-judge win rates
python eval/fact_preservation.py # fact preservation and fabrication rates
python eval/generalization.py # style consistency by category
python eval/inference_benchmark.py # latency, throughput, VRAMEach script appends its results to RESULTS.md automatically.
4. Serving.
python -m uvicorn serve:app --reloadExposes /health and POST /generate ({"topic": "..."} in, {"email": "..."} out). The model loads once at startup.
python mcp_server.pyExposes the same model as an MCP tool (compose_email_tool), connectable from Claude Desktop or Claude Code. Both serve.py and mcp_server.py share one model-loading module (serving/model_service.py) so the model is never loaded twice.
train_qlora.py calls model.print_trainable_parameters(), which reports:
trainable params: 4,358,144 || all params: 1,548,072,448 || trainable%: 0.2815
Under 0.3% of the model's weights are updated. The resulting adapter is a few megabytes rather than a multi-gigabyte model copy, and loads on top of the frozen, 4-bit-quantized base model at inference time.
MailTune/
├── data/
│ ├── emails_raw.txt # source emails (SUBJECT/EMAIL blocks)
│ └── train.jsonl # generated instruction/response dataset
├── scripts/
│ ├── generate_synthetic_data.py # builds the 264-email synthetic corpus
│ ├── prepare_dataset.py # raw text -> instruction/response JSONL
│ ├── train_qlora.py # QLoRA fine-tuning
│ └── inference.py # simple before/after generator
├── eval/
│ ├── topics.py # shared 60-topic test set
│ ├── model_utils.py # shared model loading for eval scripts
│ ├── prompted_baseline.py # base vs prompted-base vs QLoRA
│ ├── audit_train_overlap.py # checks eval topics aren't in training data
│ ├── pairwise_judge.py # LLM-judge win rates
│ ├── fact_preservation.py # fact preservation and fabrication rates
│ ├── generalization.py # style consistency by category
│ └── inference_benchmark.py # latency, throughput, VRAM
├── serving/
│ └── model_service.py # shared model loader for serve.py and mcp_server.py
├── serve.py # FastAPI serving layer
├── mcp_server.py # MCP tool wrapper
├── outputs/ # generated artifacts (adapter, comparisons, raw results)
├── docs/
│ └── mcp_demo.png # screenshot of the MCP tool in Claude Desktop
├── RESULTS.md # full evaluation results, every number traceable to a script
└── requirements.txt
- The training set (264 unique synthetic emails) is still small by research standards. This is a style-adaptation demo with real measured evidence behind it, not a claim of state-of-the-art results.
- The emails in
data/emails_raw.txtare synthetic, generated with facts injected programmatically to prevent fabrication by construction, not any one person's real writing. Swap in real sent emails for a genuine personal style-transfer result. - The judge/human agreement sanity check for the pairwise evaluation hasn't been run yet, so treat the win-rate numbers as strong but not fully independently verified.
- No memorization or canary-string testing was done. This is a known, named gap, not a silent omission.
- If a 1.5B base model doesn't show enough of a style shift for a larger dataset,
Qwen2.5-3B-Instructor aLlama-3.2-3Bvariant are drop-in replacements (updateBASE_MODELin the training and serving scripts), at the cost of more VRAM and training time.
MIT
