Skip to content
View bertolucci-rl's full-sized avatar

Organizations

@impossibleG

Block or report bertolucci-rl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bertolucci-rl/README.md

Ricardo Bertolucci — Applied Mathematics, Data Science and AI Engineering. From mathematical reasoning to production AI systems.

Applied Mathematics · Data Science · AI Engineering
Building machine learning and AI systems from mathematical reasoning, careful validation and practical decision-making.

Portfolio LinkedIn Email Hugging Face Impossible G

I'm a Data Scientist and AI Engineer with a B.Sc. and M.Sc. in Mathematics from UNESP. My work spans predictive modeling, statistical analysis and document-based AI, with an emphasis on reproducible experiments, thoughtful validation and understanding how models behave beyond a single score.

What I build — Probabilistic ML, Retrieval and RAG, and AI Systems.

Selected work

RiskPilot AI — Probabilistic machine learning. In development. Open repository.

Credit-risk modeling with probability quality at the center. Logistic Regression, LightGBM and XGBoost compared on a frozen holdout, with log loss, Brier score, calibration diagnostics and paired-bootstrap uncertainty estimates. The broader decision platform is being built incrementally.

61.5k frozen holdout · ROC-AUC 0.762 · ECE 0.0022
Python · scikit-learn · LightGBM · XGBoost · pytest

Repository →   ·   Model comparison report →


Normative RAG — Retrieval and document AI. Independent prototype. Open Hugging Face Space.

Finding relevant evidence in institutional documents. An independent prototype for semantic search over SEST SENAT documents, using multilingual embeddings, source references and retrieval evaluation, with an optional LLM layer.

Semantic retrieval · source grounding · optional LLM layer
Python · Sentence Transformers · ChromaDB · Streamlit

Hugging Face Space →


Payment Risk Modeling — Tabular ML and feature engineering. Case study. Open repository.

Estimating late-payment risk from customer and billing data. A supervised-learning case study covering data integration, feature engineering, preprocessing pipelines and comparison of probabilistic classifiers using ROC-AUC and log loss.

Data integration · feature engineering · probabilistic classifiers
pandas · scikit-learn · Random Forest · XGBoost

Repository →

Open-source organization

Impossible G — open-source, local-first AI infrastructure.

I also build under Impossible G, focused on local-first, self-hosted AI infrastructure. Its projects package open models as usable services for embeddings, OCR, voice and inference through familiar application interfaces.

Organization →   ·   Website →


From reasoning to systems — mathematical structure, statistical validation, machine learning models, and production AI systems.

Technical toolkit

Core tools
Python scikit-learn LightGBM XGBoost FastAPI LangGraph

Modeling & statistics
Probabilistic classification · calibration · validation design · experiment tracking

Retrieval & LLM applications
Embeddings · semantic search · reranking logic · source-grounded answers

Applications & engineering
APIs · Streamlit apps · Git workflows · testing with pytest

Mathematical foundation

B.Sc. (2022) and M.Sc. (2024) in Mathematics — UNESP, Brazil.
My graduate work focused on functional analysis and operator theory. In applied work, I bring that attention to assumptions, mathematical structure and the limits of what the data can support.

I also maintain a Mathematics Self-Study Roadmap: a bilingual, book-based path from mathematical foundations to advanced undergraduate and graduate-level topics.

How I work

Start with a baseline. Establish a reproducible reference before adding complexity.

Validate the right question. Match evaluation to the prediction setting and the decision the model is meant to support.

Make the evidence inspectable. Keep assumptions, experiments, limitations and results documented alongside the code.


Open to Data Science and AI Engineering opportunities and technical collaborations.

Explore my portfolio   ·   Connect on LinkedIn   ·   Get in touch

Popular repositories Loading

  1. math-self-study-roadmap math-self-study-roadmap Public

    A self study roadmap for mathematics in portuguese.

    HTML 5 1

  2. riskpilot-ai riskpilot-ai Public

    Credit Risk Decision Intelligence Platform: probabilistic PD modeling, calibration, decision science, explainability, monitoring, FastAPI and an AI Risk Analyst

    Jupyter Notebook 1

  3. bertolucci-rl bertolucci-rl Public

    Config files for my GitHub profile.

  4. Predicao-de-pagamento-de-mensalidades Predicao-de-pagamento-de-mensalidades Public

    Jupyter Notebook

  5. modelo_pred_alugueis_completo modelo_pred_alugueis_completo Public

    Jupyter Notebook

  6. bertolucci-rl.github.io bertolucci-rl.github.io Public

    HTML