I build LLM systems that work in production: agents, RAG, prompt pipelines, and quality measurement.
Currently a prompt engineer at an AI tutor for school students: moderation, subject prompts, text-to-speech, several languages. I measure quality instead of eyeballing it - eval datasets, LLM-as-judge, version-to-version prompt comparisons. Plus answer safety and prompt injection resistance.
Projects
- skillprobe - eval harness for agent skills: does a SKILL.md trigger when it should and does the model follow it, with repeats and confidence intervals across OpenAI, Qwen, DeepSeek, GLM and Kimi
- memento - persistent task memory for Claude Code: the agent maintains task files, decisions and hypotheses itself, outside models verify what was written
- ai-video-factory - decodes viral videos with vision LLMs and generates consistent-character series: FastAPI, Next.js, Docker
- crm-seller-ai - AI sales agent for Bitrix24 and amoCRM: runs the dialog, qualifies leads, moves deals through the pipeline
- summarix-ai - paid AI Telegram bot template: aiogram 3, Telegram Stars, subscriptions, Docker, CI
Stack: Python, FastAPI, LangGraph, LangChain, OpenAI API, RAG, Qdrant, PostgreSQL, Docker, MCP, GigaChat, YandexGPT
Contact: Telegram @andreys_nm · open to remote work and relocation