Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DualParser

A resume parser that runs classic NER and an LLM side by side, so you can see exactly where they agree — and where they don't.

Every uploaded resume is processed through two independent extraction pipelines — a classic spaCy NER pipeline and a Groq LLaMA 3.3 70B pipeline — normalized into the same schema, and shown next to each other. It can also score a resume against a job description using semantic similarity.


Screenshots

Upload — drag-and-drop a resume and both pipelines run in parallel:

Upload screen

Side-by-side comparison — classic NER vs LLM extraction on the same resume, with per-pipeline latency and cost:

Comparison view

Job matching — paste resume text and a job description to get a semantic match score plus matched/missing skills:

Job match panel


Why hybrid NER + LLM?

Classic NER (spaCy + gazetteers) LLM (Groq LLaMA 3.3 70B)
Cost per resume Free, runs locally Small per-token cost
Latency Slower to load model, fast per-call once warm Fast (Groq inference), network-dependent
New field/domain Needs a new gazetteer or labeled data Works immediately, zero retraining
Consistency Deterministic, same input → same output Can vary slightly between calls
Maintenance Gazetteers need upkeep per domain Prompt/schema upkeep only

Running both against the same schema turns that tradeoff from a guess into something you can actually measure — see /compare and the eval/ notebook.


Architecture

flowchart TD
    A[Resume upload<br/>PDF / DOCX] --> B[Text extraction<br/>pdfplumber / python-docx]
    B --> C{Section splitter}
    C --> D[Classic NER pipeline<br/>spaCy + gazetteers]
    C --> E[LLM pipeline<br/>Groq LLaMA 3.3 70B<br/>JSON schema extraction]
    D --> F[Normalization<br/>clean + dedupe each result]
    E --> F
    F --> G[(SQLite<br/>parsed_resumes)]
    F --> H[POST /parse response<br/>NER result + LLM result]
    H --> I[React frontend<br/>side-by-side comparison]

    J[Resume text + Job description] --> K[Sentence-transformers<br/>embeddings]
    K --> L[Cosine similarity<br/>+ skill overlap]
    L --> M[POST /match response<br/>score + matched/missing skills]
    M --> I
Loading

Project Structure

backend/
  app/
    main.py                     FastAPI app, mounts /parse /compare /match
    config.py                   Settings (env-driven: Groq key, model names, DB url)
    schemas/resume.py           Shared ParsedResume schema — the contract both
                                 pipelines must produce
    db/                         SQLAlchemy models + session (SQLite by default)
    api/
      parse.py                  POST /parse   — upload → both pipelines → normalize → store
      compare.py                POST /compare — run both pipelines on raw text, return metrics
      match.py                  POST /match   — resume vs job description similarity
    services/
      extraction/                PDF (pdfplumber) + DOCX (python-docx) → raw text
      ner/
        spacy_pipeline.py        spaCy NER + gazetteer matching (entry point)
        section_parser.py        Section-splitting + per-entry extraction logic
        gazetteers/               skills.txt, degrees.txt, job_titles.txt, certifications.txt
      llm/groq_extractor.py      Groq structured JSON extraction + cost estimation
      normalization/             Cleans each pipeline's output (email/phone/dedupe)
      matching/jd_matcher.py     Sentence-transformers embeddings + cosine similarity
      evaluation/metrics.py      Precision/recall/F1 (needs labeled ground truth) +
                                 latency/cost comparison + plain-language discussion
frontend/                       React + Tailwind (upload, side-by-side comparison view)
data/
  raw_resumes/                  Uploaded files land here (gitignored)
  labeled/                      Ground-truth labels for evaluation, e.g. {"skills": [...]}
  test_set/                     Held-out resumes for the eval report
eval/                           Evaluation notebook comparing NER vs LLM

Clone & setup

1. Clone

git clone https://github.com/<m-sameerkhan>/dualparser.git
cd dualparser

2. Backend

cd backend
conda create -n dualparser python=3.11
conda activate dualparser
pip install -r requirements.txt
python -m spacy download en_core_web_trf

cp .env.example .env
# then edit .env and set GROQ_API_KEY

uvicorn app.main:app --reload

Backend runs at http://localhost:8000. Check http://localhost:8000/health to confirm it's up.

3. Frontend

In a second terminal:

cd frontend
npm install
npm run dev

Frontend runs at http://localhost:5173 and expects the backend at http://localhost:8000 (override with a VITE_API_BASE_URL env var if needed).

4. Try it

Open http://localhost:5173, drop in a resume (PDF or DOCX), and you'll see both pipelines' output side by side.


API

Endpoint Method Description
/parse POST Upload a resume, run both pipelines, return + store both results
/compare POST Run both pipelines on raw text, return latency/cost + P/R/F1 (if ground truth supplied)
/match POST Score a resume against a job description (semantic similarity + skill overlap)
/health GET Health check

Full request/response shapes are in app/schemas/resume.py, or just check the auto-generated docs at http://localhost:8000/docs once the backend is running.


Extending entity types

  1. Add a labeled term list to app/services/ner/gazetteers/<new_type>.txt (one term per line, lowercase).
  2. Wire it into spacy_pipeline.py alongside _SKILLS_GAZETTEER etc.
  3. Add the same field to the _SYSTEM_PROMPT in groq_extractor.py so the LLM pipeline extracts it too — keeping both pipelines aligned on the same schema is what makes /compare meaningful.
  4. Add labeled examples to data/labeled/ so /compare can report precision/recall/F1 for the new type, not just latency/cost.

Known limitations / TODOs

  • PDF multi-column handling is currently naive (page.extract_text()); true column-aware extraction needs word-level x/y clustering.
  • Skill extraction outside a resume's own "Skills" section (e.g. skills only mentioned in a project description) still relies on a gazetteer, which is bounded by whatever terms are listed in gazetteers/skills.txt.
  • No fine-tuned NER model yet — gazetteers + regex only (see stretch goals).
  • Skill-overlap matching in /match is exact-string based; synonym-aware matching (e.g. "JS" == "JavaScript") is a stretch goal.
  • No batch parsing yet (one resume per /parse call).

Stretch goals

  • Fine-tune a NER model on a labeled resume dataset instead of relying on gazetteers + regex.
  • Batch parsing: multiple resumes in, ranked list out.
  • Explainability: highlight which text span in the source resume each extracted field came from.

About

A resume parser that runs classic NER and an LLM head-to-head, comparing extraction accuracy, latency, and cost plus job-description matching via semantic similarity.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages