Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MailTune

A QLoRA fine-tuned email writing assistant. Adapts a small open LLM (Qwen2.5-1.5B-Instruct) to write short, casual emails instead of formal corporate boilerplate, using PEFT for the LoRA adapters and bitsandbytes for 4-bit quantization. Only about 4.4M of the model's 1.5B parameters (0.28%) are actually trained.

The project goes beyond a single before/after demo: it includes a prompted-base-model baseline, pairwise LLM-judge evaluation, fact preservation and fabrication testing, a generalization check across unseen categories, an inference benchmark, and a FastAPI + MCP serving layer. Every number below comes from a script in eval/ and is logged with full detail in RESULTS.md.

Results

Full numbers, methodology, and honest caveats are in RESULTS.md. The headlines:

Fine-tuning beats a well-prompted base model, not just a raw one. A prompted-base-model baseline was built first, since a good system prompt with a few examples can match a light fine-tune. Pairwise LLM-judge evaluation across 60 unseen topics:

Comparison QLoRA win rate
QLoRA vs base model 100%
QLoRA vs prompted-base model 73%

QLoRA fabricates far less than either baseline. 18 synthetic prompts embedded specific facts (names, days, times, numbers) and checked whether each model preserved them or invented different ones:

Variant Fact preservation Fabrication rate
base 81% 44%
prompted-base 90% 17%
QLoRA 85% 6%

Style holds up across categories it never saw in training. Tested on 60 topics spanning 6 categories (recruiter replies, scheduling, follow-ups, rejections, status updates, cold outreach), none of which overlap the training data. Style consistency scored 3.8 to 4.6 out of 5 across all categories, a spread of 0.8, which supports the model having learned a general tone rather than memorizing a handful of patterns.

The fine-tune is also faster to use. QLoRA emails are ready in about 3 seconds versus about 8 for the base model, because it learned to write short emails instead of rambling to the token limit. VRAM usage is essentially unchanged (the adapter is a few MB on top of the quantized base model).

Two honest caveats worth knowing before trusting these numbers fully: the judge/human agreement sanity check (comparing the LLM judge's picks against a human's) hasn't been run yet, and the training set, while now 264 unique emails, is still small by research standards, so "generalization" here means a modest style-adaptation result rather than a large-scale claim.

Example: before and after

Topic: Rescheduling next week's call

Before (base model): "Subject: Rescheduling Next Week's Call. Dear [Recipient's Name], I hope this message finds you well. After careful consideration of our current project deadlines and priorities, we have decided to reschedule the upcoming call... Best regards, [Your Full Name] [Your Position] [Your Contact Information]"

After (QLoRA fine-tuned): "Hi guys, I was wondering if we could reschedule our meeting for thursday at 3pm? I'm pretty sure i got my calendar wrong last time but that is the best slot i can find right now. Thanks!"

More examples in outputs/three_way_comparison.md (base, prompted-base, and QLoRA side by side across all 60 test topics).

Live demo: MCP tool

The fine-tuned model is also wired up as an MCP tool, invoked directly from Claude Desktop:

MCP demo

Setup

Works with Python 3.10+ and a CUDA GPU. A local NVIDIA GPU, Google Colab, Kaggle, or a rented cloud GPU all work the same way. CPU-only execution works but is very slow.

git clone https://github.com/bazzal99/MailTune.git
cd MailTune
python -m venv venv
source venv/bin/activate        # Windows: venv\Scripts\activate
pip install -r requirements.txt

Verify the GPU is visible:

python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

This should print True and a GPU name. If it prints False, reinstall PyTorch with a CUDA-enabled build matching your driver (see pytorch.org).

For the evaluation scripts that use an LLM judge (Weeks 2 and 4), create a .env file in the repo root with ANTHROPIC_API_KEY=sk-ant-....

Running in Google Colab or Kaggle instead
  1. Upload/unzip this project into the notebook's working directory.
  2. Set the runtime/accelerator to a GPU (Colab: Runtime, Change runtime type, T4 GPU. Kaggle: Settings, Accelerator, GPU).
  3. Run the same pip install -r requirements.txt and the verification command above in a notebook cell, prefixing shell commands with !.
  4. Run the pipeline scripts the same way, e.g. !python scripts/train_qlora.py.
  5. To retrieve results, zip the outputs/ folder and download it, or use from IPython.display import FileLink; FileLink('outputs/three_way_comparison.md') for a single file.

Pipeline

1. Data preparation. data/emails_raw.txt holds source emails in SUBJECT: ... / EMAIL: ... blocks. The current file has 264 unique synthetic emails, generated by scripts/generate_synthetic_data.py with facts injected programmatically so fabrication is prevented by construction, not just hoped for. Swap in real sent emails for a genuine personal style-transfer result (strip signatures, quoted reply chains, and anything sensitive first).

python scripts/generate_synthetic_data.py   # optional, regenerates the synthetic corpus
python scripts/prepare_dataset.py

Produces data/train.jsonl, an instruction/response dataset pairing an email topic with the target email body.

2. Training. QLoRA fine-tuning via PEFT, bitsandbytes, and TRL's SFTTrainer.

python scripts/train_qlora.py

Saves the trained adapter to outputs/qlora-adapter/. Takes a few minutes on a T4-class GPU with this dataset size.

3. Evaluation. The full eval suite lives in eval/, run in this order:

python eval/prompted_baseline.py       # base vs prompted-base vs QLoRA, all topics
python eval/audit_train_overlap.py     # confirms eval topics aren't in the training set
python eval/pairwise_judge.py          # LLM-judge win rates
python eval/fact_preservation.py       # fact preservation and fabrication rates
python eval/generalization.py          # style consistency by category
python eval/inference_benchmark.py     # latency, throughput, VRAM

Each script appends its results to RESULTS.md automatically.

4. Serving.

python -m uvicorn serve:app --reload

Exposes /health and POST /generate ({"topic": "..."} in, {"email": "..."} out). The model loads once at startup.

python mcp_server.py

Exposes the same model as an MCP tool (compose_email_tool), connectable from Claude Desktop or Claude Code. Both serve.py and mcp_server.py share one model-loading module (serving/model_service.py) so the model is never loaded twice.

What PEFT buys here

train_qlora.py calls model.print_trainable_parameters(), which reports:

trainable params: 4,358,144 || all params: 1,548,072,448 || trainable%: 0.2815

Under 0.3% of the model's weights are updated. The resulting adapter is a few megabytes rather than a multi-gigabyte model copy, and loads on top of the frozen, 4-bit-quantized base model at inference time.

Project structure

MailTune/
├── data/
│   ├── emails_raw.txt              # source emails (SUBJECT/EMAIL blocks)
│   └── train.jsonl                 # generated instruction/response dataset
├── scripts/
│   ├── generate_synthetic_data.py  # builds the 264-email synthetic corpus
│   ├── prepare_dataset.py          # raw text -> instruction/response JSONL
│   ├── train_qlora.py              # QLoRA fine-tuning
│   └── inference.py                # simple before/after generator
├── eval/
│   ├── topics.py                   # shared 60-topic test set
│   ├── model_utils.py              # shared model loading for eval scripts
│   ├── prompted_baseline.py        # base vs prompted-base vs QLoRA
│   ├── audit_train_overlap.py      # checks eval topics aren't in training data
│   ├── pairwise_judge.py           # LLM-judge win rates
│   ├── fact_preservation.py        # fact preservation and fabrication rates
│   ├── generalization.py           # style consistency by category
│   └── inference_benchmark.py      # latency, throughput, VRAM
├── serving/
│   └── model_service.py            # shared model loader for serve.py and mcp_server.py
├── serve.py                        # FastAPI serving layer
├── mcp_server.py                   # MCP tool wrapper
├── outputs/                        # generated artifacts (adapter, comparisons, raw results)
├── docs/
│   └── mcp_demo.png                # screenshot of the MCP tool in Claude Desktop
├── RESULTS.md                      # full evaluation results, every number traceable to a script
└── requirements.txt

Scope and limitations

  • The training set (264 unique synthetic emails) is still small by research standards. This is a style-adaptation demo with real measured evidence behind it, not a claim of state-of-the-art results.
  • The emails in data/emails_raw.txt are synthetic, generated with facts injected programmatically to prevent fabrication by construction, not any one person's real writing. Swap in real sent emails for a genuine personal style-transfer result.
  • The judge/human agreement sanity check for the pairwise evaluation hasn't been run yet, so treat the win-rate numbers as strong but not fully independently verified.
  • No memorization or canary-string testing was done. This is a known, named gap, not a silent omission.
  • If a 1.5B base model doesn't show enough of a style shift for a larger dataset, Qwen2.5-3B-Instruct or a Llama-3.2-3B variant are drop-in replacements (update BASE_MODEL in the training and serving scripts), at the cost of more VRAM and training time.

License

MIT

About

QLoRA fine-tuned email assistant with a full evaluation suite: pairwise LLM-judge comparison, fact-preservation testing, generalization checks, and FastAPI/MCP serving.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages