On-device medical search for nurses and midwives in Zanzibar
Demo video · Live web demo · Eval Report · Latency Report
Android app that answers clinical questions offline using on-device RAG — Gemma 4 E4B (LiteRT-LM) for generation, EmbeddingGemma-300M for embeddings, SQLite for vector search. English only. No internet needed after the initial ~3.8 GB model download.
🏥 Try the live web demo — a faithful, browser-based mirror of the on-device app (same Gemma 4 + G1 prompt + EmbeddingGemma RAG, served via llama.cpp). Demonstration only — not medical advice; don't read latency from it. (demo source)
┌─────────────────────────────────────────────────┐
│ Flutter UI (Dart) │
│ intro_page.dart · search_page.dart │
├──────────────┬──────────────────────────────────┤
│ MethodChannel│ EventChannel (streaming) │
├──────────────┴──────────────────────────────────┤
│ Android Native (Kotlin) │
│ MainActivity.kt · RagStream.kt │
│ ┌────────────────────────────────────────────┐ │
│ │ RagPipeline.kt │ │
│ │ ┌──────────┐ ┌──────────┐ ┌────────────┐ │ │
│ │ │ Gemma 4 │ │ Embedding│ │ SQLite │ │ │
│ │ │ LiteRT-LM│ │ Gemma │ │ VectorStore│ │ │
│ │ └──────────┘ └──────────┘ └────────────┘ │ │
│ └────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────┘
Query → EmbeddingGemma embeds → SQLite retrieves top-3 guideline chunks → prompt assembled → LiteRT-LM streams response → Flutter renders markdown.
Requires a real Android device (LiteRT-LM needs hardware acceleration, not emulators).
cd app
flutter pub get
flutter run # debug on connected device
flutter build apk # release APK
adb logcat -s mam-ai # timing, memory, inference logsMAM-AI ships as a signed Android APK (it is not on the Play Store). The APK is on the
Releases page, named mamai-<version>-<sha>-rag-<bundle>.apk. Choose the method
that fits the phone's connectivity:
- On the phone, open the Releases page and download the latest
mamai-*.apk. - Tap the downloaded file. When prompted, allow "install unknown apps" for your browser/Files app, then confirm the install.
- Open MAM-AI. On first launch it downloads the AI models + medical guidelines (~4.5 GB, one-time) over Wi-Fi — each file verified against a pinned SHA-256 — then runs fully offline.
For field devices with little or no connectivity. Stage the assets once on a computer with internet, then push the app and every asset to the phone over USB — it opens ready with zero on-device downloads. The phone needs USB debugging on (Settings → Developer options).
Stage the assets (run from the repo root):
bash scripts/sync_models.sh # Gemma 4 + EmbeddingGemma + tokenizer → device_push/models/
bash scripts/sync_rag_assets.sh # pinned RAG bundle (embeddings + PDFs) → device_push/bundle/Build the signed APK:
cd app
flutter build apk --release
cd ..Connect the phone by USB, then push the app + every asset:
bash scripts/push_to_device.sh --all-models --apk app/build/app/outputs/flutter-apk/app-release.apkpush_to_device.sh --all-models installs the app and pushes the LLM, the EmbeddingGemma retriever +
tokenizer (each with its .verified checksum marker), the vector store, and all source PDFs — after
verifying every model against the app's pinned SHA-256. Open MAM-AI afterwards: it goes straight to
the chat screen, no download. To provision more phones, repeat only the last command (offline) for each.
Requirements: a real Android phone (Android 7.0 / API 24 or newer — emulators are unsupported; on-device inference needs real hardware), ~8 GB free storage, and a recent mid-to-high-end device for acceptable response speed.
Developers building from source: see Build & Run above.
Downloaded on first launch — no auth required. The Gemma 4 LiteRT-LM is the ungated
community export; the EmbeddingGemma retriever is a byte-identical ungated copy of the
license-gated litert-community/embeddinggemma-300m under the project's HuggingFace
account. Each file is verified against a pinned SHA-256 before use.
| File | Size | Source |
|---|---|---|
gemma-4-E4B-it.litertlm |
3.66 GB | litert-community/gemma-4-E4B-it-litert-lm (ungated) |
embeddinggemma-300M_seq256_mixed-precision.tflite |
171 MB | nmrenyi/embeddinggemma-300m-litert-mamai (mirror of gated litert-community/embeddinggemma-300m) |
embeddinggemma_tokenizer.model |
4.5 MB | nmrenyi/embeddinggemma-300m-litert-mamai |
embeddings.sqlite |
~261 MB | mamai-medical-guidelines releases |
The pinned RAG bundle version lives in config/rag_assets.lock.json; the pinned
model checksums live in _modelFileSha256 in app/lib/screens/intro_page.dart.
Chunking and embedding are managed in the companion mamai-medical-guidelines repo. To pull in a new bundle:
- Bump
config/rag_assets.lock.jsonwith the new version + manifest checksum - Run the staging and push scripts:
bash scripts/sync_rag_assets.sh # download + stage bundle
bash scripts/sync_models.sh # download EmbeddingGemma + Gemma 4 from HuggingFace
bash scripts/push_to_device.sh # push everything to connected device
bash scripts/push_to_device.sh --embedding-models # push embedder + tokenizer onlyTag from main only. CI builds a signed APK and publishes a GitHub Release automatically.
git tag v0.1.0-beta.1 # beta/alpha/rc → prerelease; vX.Y.Z → stable
git push origin v0.1.0-beta.1Valid formats: vX.Y.Z, vX.Y.Z-alpha.N, vX.Y.Z-beta.N, vX.Y.Z-rc.N
Required GitHub secrets: ANDROID_KEYSTORE_BASE64, ANDROID_KEYSTORE_PASSWORD, ANDROID_KEY_ALIAS, ANDROID_KEY_PASSWORD. See app/android/key.properties.example for local signing setup.
Benchmarks run across AfriMedQA, MedQA USMLE, MedMCQA, Kenya Vignettes, AfriMedQA SAQ, and WHB Stumps under the app_parity_v1 protocol (same system prompt as the APK, versioned RAG contexts).
| Model | MCQ avg | Open-ended avg |
|---|---|---|
| GPT-5 (no-RAG) | 82.8% | 4.19 / 5 |
| Gemma 4 E4B (deployed, +G1 prompt, no-RAG) | 42.9% | 2.61 / 5 |
| Gemma 3n E4B (no-RAG) | 45.5% | 2.98 / 5 |
The deployed generator is Gemma 4 E4B with the G1 deflection-fix prompt, chosen for
safety: a generator × prompt evaluation (see the mamai-eval repo) found Gemma 3n more
helpful (≈2× Kenya recall) but materially less safe — genuine order-of-magnitude dosing/drug
errors under the same prompts (hand-adjudicated), whereas Gemma 4 + G1 fixes the deflection
problem with ~zero dangerous answers. The retriever is EmbeddingGemma-300M. RAG slightly
hurts the on-device models on MCQ (a knowledge proxy); GPT-5 is unaffected.
On an OPPO Snapdragon 8 Elite device, the deployed Gemma 4 E4B averages ~11.7 s TTFT, with EmbeddingGemma query embedding + vector search around 0.5–0.6 s (faster than Gecko's ~2.3 s).
With LiteRT-LM 0.11.0 on GPU (opt-in), TTFT drops to ~1–2 s on the same device; decode rate is unchanged. Multi-token Prediction (ExperimentalFlags.enableSpeculativeDecoding, gated behind the useMtpForLlm Gradle property) was smoke-tested and produced a ~10–20% decode slowdown rather than the vendor's claimed >2× — drafter acceptance is likely poor for our long retrieved-context prompts. Off by default; re-test before re-enabling.
Full results: eval report · latency report
Gemma 3n E4B was finetuned on medical QA data using LoRA (not deployed). Training code removed; artefacts archived externally: dataset · model.
