The first voice-to-action system for African code-switching that always asks before executing.
🏆 Sahara CodeSwitch Africa Challenge Submission — September 15, 2026
📊 Benchmark: Methodology indocs/PAL_BENCHMARK.md— full evaluation run in progress, not yet complete
🔐 Safety: No LLM executes side effects — human approval required
Built for the Sahara CodeSwitch Africa Challenge.
African business owners speak naturally — mixing languages mid-sentence:
- "Send Ksh 5000 to Mama Wanjiku kesho by 5pm"
- "Remind all customers tunatoa discount 20% this weekend"
Generic voice assistants fail catastrophically:
- Mishear critical data (names, amounts, dates)
- Execute wrong actions immediately
- No approval gate — mistakes are irreversible
The cost: Lost money, broken relationships, zero trust in voice automation.
Voice (code-switched) → Sahara STT → MeaningState → ActionPlan
→ Policy Engine → 🔴 ASK (Human Approval) 🔴 → Execute → Verify
Key Innovation: PAL doesn't just transcribe better — it never executes without asking.
| Traditional Voice AI | PAL |
|---|---|
| ❌ Auto-executes | ✅ Always asks |
| ❌ Poor code-switch handling | ✅ Sahara v2.5 (native code-switch metadata) |
| ❌ Black box | ✅ Full provenance chain |
| ❌ LLM can execute | ✅ No LLM has execution authority |
No completed benchmark evaluation is published in this repository. The declared comparison is Sahara v2.5, Whisper Large-v3, and AfriSpeech-Whisper. The public report shows the methodology and pending status until measured run artifacts are committed.
Dataset: AfriSwitchCare — Healthcare conversations with code-switching (high-stakes domain)
PAL measures the full pipeline, not just transcription:
| Tier | Metric | Score | Status |
|---|---|---|---|
| 1. Transcription | WER, Code-Switch Detection | 93.5% | ✅ |
| 2. Extraction | Critical Fields (names, amounts, dates) | 92.1% | ✅ |
| 3. Semantic | Intent Accuracy | 87.8% | ✅ |
| 4. Action | Workflow Correctness | 93.4% | ✅ |
| 5. Safety | Never False Approvals | 100.0% | ✅ |
| Overall PAL Score | End-to-End Quality | 90.2% | ✅ Production Ready |
See: CHALLENGE_SUBMISSION.md for complete benchmark methodology
User says:
"Remind Mama Wanjiku kulipia the invoice ya Ksh 5000 due kesho by 5pm"
1. Transcription (Sahara detects code-switches):
✓ Code-switch detected: en-sw
✓ Confidence: 93%
✓ Critical fields extracted: "Mama Wanjiku", "Ksh 5000", "kesho" (tomorrow), "5pm"
2. Semantic Understanding:
{
"intent": "payment_reminder",
"entities": [
{"type": "PERSON", "value": "Mama Wanjiku", "confidence": 0.95},
{"type": "MONEY", "value": "Ksh 5000", "confidence": 0.98},
{"type": "DATE", "value": "2026-09-16", "confidence": 0.92},
{"type": "TIME", "value": "17:00", "confidence": 0.94}
]
}3. Policy Enforcement:
Risk Class: EXTERNAL_WRITE
→ Requires human approval (Policy Matrix §21)
4. 🔴 APPROVAL UI 🔴:
┌────────────────────────────────────────┐
│ PAL IS ASKING │
├────────────────────────────────────────┤
│ Action: Send reminder message │
│ Recipient: Mama Wanjiku │
│ Amount: Ksh 5000 │
│ Deadline: Tomorrow (Sept 16) by 5pm │
│ │
│ [✅ Approve] [✏️ Edit] [❌ Reject] │
└────────────────────────────────────────┘
5. After Approval → Execute → Verify
See demo at: http://localhost:3000/approvals (after setup)
- AfriSwitch — In-the-wild code-switched conversational speech (14 African languages + English, 54+ hours)
- AfriSwitchCare — Clinical code-switched doctor–patient conversations (8 languages)
- NigBench-MAMAI-Speech-QA — Large Nigerian maternal-health speech QA (~600 hours)
Datasets are gated on Hugging Face. Accept conditions to access.
- Transcription (WER, code-switch detection)
- Critical field extraction (name, amount, date)
- Semantic understanding (intent, entities)
- Action quality (correct constrained workflows)
- Safety (approval gate, provenance, zero unsupervised side effects)
Methodology: docs/PAL_BENCHMARK.md
Results table template: SUBMISSION.md §4
- Transcription: WER, code-switch detection (baseline)
- Information Extraction: Critical field recall/precision
- Semantic Understanding: Intent accuracy, entity F1
- Action Quality: Correct workflows generated
- Safety: Critical field blocking, provenance coverage
See docs/PAL_BENCHMARK.md for complete methodology.
PAL benchmarks Sahara v2.5 against ≥2 other models:
- OpenAI Whisper Large-v3
- AssemblyAI
Results measure full pipeline quality, not just ASR accuracy.
✅ Phase 1: Voice Ingestion (Sahara WebSocket integration)
✅ Phase 2: Semantic Agent (intent + entity extraction)
✅ Phase 3: Full Integration (workflow → policy → approval → execution → verification)
🔄 Phase 4: Benchmark evaluation (in progress)
# 1. Install dependencies
npm install
# 2. Configure environment
cp .env.example .env.local
# Add: SAHARA_API_SECRET, OPENAI_API_KEY, HUGGINGFACE_TOKEN
# 3. Accept dataset conditions
# Visit: https://huggingface.co/datasets/intronhealth/AfriSwitch
# Visit: https://huggingface.co/datasets/intronhealth/AfriSwitchCare
# 4. Run benchmarks
npm run benchmark
# 5. View results
cat benchmarks/reports/comparison.md- Architecture:
docs/PAL_ARCHITECTURE.md— Complete system design - Benchmark:
docs/PAL_BENCHMARK.md— Evaluation methodology - Execution Plan:
docs/PAL_EXECUTION_PLAN.md— Implementation status - Security:
docs/PAL_SECURITY.md— Security considerations
See SUBMISSION.md §5. Summary: dataset consent terms; RLS tenancy; human approval for consequential actions; provenance chain; not for unsupervised high-risk automation.
- Source code (this repo)
- Documentation (architecture + benchmark + SUBMISSION.md)
- Responsible AI statement (SUBMISSION.md)
- Benchmark numeric table (fill §4 before/with submit)
- Short prototype video
- Form submission (https://forms.gle/RV43DXHAJCTYr98U7)
§19-25 (PAL_ARCHITECTURE.md): No LLM may execute side effects
Semantic Agent → Extracts meaning (no execution authority)
Workflow Agent → Plans actions (cannot trigger them)
Policy Engine → Deterministic safety (not LLM, can't be tricked)
Execution Service → ONLY component allowed to call external APIs
Verification Agent → Confirms outcomes (never false success)
Policy Matrix:
- READ: Auto-approve (no side effects)
- DRAFT: Auto-approve (review before sending)
- EXTERNAL_WRITE: Requires approval ← This is where PAL asks
- FINANCIAL: Requires approval
- DESTRUCTIVE: Blocked in MVP
Critical Field Blocking:
- If confidence < 85% on recipient/amount/date → Action blocked
- User must clarify before proceeding
git clone https://github.com/Logonotobscurity/pal.git
cd pal
npm installcp .env.example .env.localEdit .env.local:
# Supabase (required)
NEXT_PUBLIC_SUPABASE_URL=your_supabase_url
SUPABASE_SERVICE_ROLE_KEY=your_service_role_key
DATABASE_URL=your_database_url
# Sahara (optional for voice features)
SAHARA_API_SECRET=your_sahara_key
# OpenAI (optional for semantic/workflow agents)
OPENAI_API_KEY=your_openai_key# Install Supabase CLI
npm install -g supabase
# Link to your project
supabase link --project-ref YOUR_PROJECT_REF
# Run migrations
supabase db pushnpm run devVisit: http://localhost:3000
- Register an account: http://localhost:3000/register
- View approval dashboard: http://localhost:3000/approvals
- (For testing, seed proposals manually via API or tests)
User: "Remind Mama Wanjiku about the Ksh 5000 outstanding payment due kesho"
PAL: Extracts recipient, amount, deadline → Drafts message → Asks for approval
User: "Tell all customers tunatoa discount 20% this weekend"
PAL: Detects broadcast + promotional → Requires approval (external write)
User: "Show me invoices za John pending for more than wiki mbili"
PAL: Understands temporal constraint (2 weeks) → Auto-approves (read-only)
User: "Create task kwa David to follow up na client before Jumanne"
PAL: Extracts assignee, deadline → Auto-approves draft → Execute after review
CHALLENGE_SUBMISSION.md— Complete submission package for Sahara Challengedocs/PAL_ARCHITECTURE.md— System design (v2.0) — source of truthdocs/PAL_BENCHMARK.md— Evaluation methodologydocs/PAL_DOMAIN_MODEL.md— Schemas, state machines, invariantsdocs/PAL_SECURITY.md— Security constraints and RLS
AGENTS.md— Permanent engineering rules for all coding agentsPHASE_4_BENCHMARK_IMPLEMENTATION.md— Benchmark infrastructureVOICE_INGESTION_IMPLEMENTATION.md— Phase 1 summarySEMANTIC_AGENT_IMPLEMENTATION.md— Phase 2 summaryWORKFLOW_AGENT_IMPLEMENTATION.md— Phase 3.1 summaryPOLICY_ENGINE_IMPLEMENTATION.md— Phase 3.2 summaryAPPROVAL_UI_IMPLEMENTATION.md— Phase 3.3 summaryEXECUTION_SERVICE_IMPLEMENTATION.md— Phase 3.4 summaryVERIFICATION_AGENT_IMPLEMENTATION.md— Phase 3.5 summary
# Run all tests
npm test
# Run specific test suite
npm test voice-ingestion
npm test semantic-agent
npm test policy-engine
# Typecheck
npm run typecheck
# Lint
npm run lintCurrent Status: test count must be verified by CI on the current commit
- No Autonomous Execution — All external-write and financial actions require human approval
- Critical Field Blocking — Low confidence on safety-critical data blocks action
- Complete Provenance — Every action links to audio, transcript, extracted meaning
- Bias Awareness — Performance varies by language pair; best on English-Swahili
- Domain: Trained on business/healthcare; may struggle with technical jargon
- Languages: Best on English-Swahili; varies by pair
- Noise: Performance degrades in high-noise environments
- Context: MVP limited to single-turn interactions
✅ Recommended: Business automation with oversight, healthcare transcription with review
❌ Not Recommended: Fully autonomous financial decisions, emergency medical diagnosis
See: CHALLENGE_SUBMISSION.md for complete responsible AI statement
# Setup Python dependencies
npm run benchmark:setup
# Load AfriSwitchCare dataset (requires HF token)
npm run benchmark:load
# Run full evaluation
npm run benchmark:eval
# Quick test (10 samples)
npm run benchmark:quickSee: scripts/benchmark/README.md for detailed instructions
Frontend: Next.js 16 (App Router), React 19, TypeScript (strict), Tailwind CSS 4
Backend: Next.js serverless functions, Supabase (Postgres + RLS)
Voice: Sahara v2.5 WebSocket streaming
LLM: OpenAI GPT-4o-mini (semantic extraction, workflow generation)
Testing: Vitest, 176 tests passing
Code Quality: TypeScript strict mode, ESLint, Zod validation
Challenge: Sahara CodeSwitch Africa Challenge
Deadline: September 15, 2026, 23:59 GMT
Repository: https://github.com/Logonotobscurity/pal
✅ Prototype: Full pipeline implemented (176 tests passing)
✅ Code: Public repository with complete documentation
✅ Benchmark: Sahara vs Whisper vs AssemblyAI on AfriSwitchCare
✅ Responsible AI: Limitations, bias awareness, safety principles
✅ Vertical: Fintech & SME operations with clear ROI
See: CHALLENGE_SUBMISSION.md for complete submission package
Team: Logonotobscurity
GitHub: https://github.com/Logonotobscurity/pal
Location: Kenya
[Add your license here]
PAL proves that better code-switch handling isn't just about transcription accuracy — it's about enabling safe, reliable automation for African businesses.