Skip to content
View ashutro's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report ashutro

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ashutro/README.md

MasterHead

Hi πŸ‘‹, I'm ASHUTOSH KUMAR

The world speaks in numbers, I translate their stories. ️ Caffeine-powered data enthusiast, bridging the gap between data and decisions with every cup.

Coding

ashutro

hubbigdata

  • πŸ”­ I’m currently working in Infosys

  • 🌱 I’m currently learning AWS || Airflow || AWS Glue || Kafka

  • πŸ’¬ Ask me about Python|| MySQL || PostgreSQL || Spark || Apache Spark

  • πŸ“« How to reach me iamashutosh.kumar.dev@gmail.com

  • πŸ“„ Know about my experiences Ashutosh Kumar's Resume

Connect with me:

hubbigdata ashutro

Languages and Tools:

android aws azure bash c cplusplus docker gcp git graphql hadoop hive java kafka linux mongodb mssql mysql opencv oracle pandas photoshop postgresql python pytorch scala scikit_learn seaborn sqlite tensorflow


πŸ’Ό Work Experience

Databricks Data Engineer / Team Lead | Infosys - Nuuday Denmark (2026 - Present)
  • Led Databricks data engineering delivery for Nuuday and built Databricks Asset Bundles, reusable notebooks, and job workflows, reducing deployment effort by 70% and release errors by 60%.
  • Developed multi-source ingestion pipelines from SQL databases, REST APIs, ServiceNow, Adobe, ACS, ICH, and other sources, improving onboarding speed for new sources by 80%.
  • Integrated Azure Key Vault-backed secret management for database credentials, API tokens, and service connections, reducing credential maintenance effort by 70%.
  • Migrated existing data pipelines into Databricks Lakeflow Designer, standardizing orchestration, lineage, monitoring, and handover while reducing manual operational checks by 75%.
  • Designed Delta Lake pipelines using hybrid medallion and layered architecture, improving pipeline maintainability by 80% and reducing reprocessing time by 70%.
  • Implemented data quality rules, audit checks, and exception handling across ingestion layers, increasing data validation coverage by 85%.
  • Supported team-lead responsibilities across planning, code review, release coordination, and production issue resolution, improving SLA adherence by 70% across data and BI teams.
Big Data Engineer | Infosys (2022 - 2025)
  • Built and maintained ETL pipelines for large datasets with a strong focus on accuracy, reliability, and data quality rates of 95% or higher.
  • Improved Redshift data integrity by 80% through normalization, partitioning, and optimized SQL query strategies, reducing regional bottlenecks and improving performance.
  • Created complex SQL validation scripts for Redshift and S3-based staging, reducing discrepancies by 70% and improving executive reporting accuracy.
  • Enhanced production scripts using Python, Airflow, and Spark, improving processing efficiency by 80% for millions of records.
  • Automated ETL health checks with Shell scripts on Linux, reducing manual intervention by 70% and improving SLA compliance by 60%.
  • Partnered with BI teams on MSTR dashboard data delivery, reducing report refresh delays by 50% and improving data availability for business users.

πŸ† Awards & Certifications

View Details
Awards:
  • Digital Specialist Engineer - L1 at Infosys for consistent delivery of robust cloud data platforms.
  • AI Builder Skill Tag at Infosys for expertise in AI and Data Engineering.
  • Cloud Data Professional Skill Tag at Infosys for proven expertise in cloud-native data engineering.
  • Insta Award - Certificate of Appreciation from Infosys Data Analytics leadership.
  • Commendation Certificate - Insta Award at Infosys for rapid onboarding and strong ownership.
Certifications:
  • Certified Data Engineer Professional - Databricks (2025)
  • Certified Data Engineer - Associate Databricks (2025)
  • Certified Developer - Associate AWS (2025)
  • Certified Cloud Practitioner AWS (2025)

πŸ“Š GitHub Stats

ashutro stats ashutro streak

ashutro top langs

Pinned Loading

  1. Query-gpt-ollama Query-gpt-ollama Public

    QueryGPT is a Streamlit-based AI assistant that allows you to interact with your SQLite database using natural language. It leverages LLMs like **Ollama (Mistral)** or **OpenAI GPT-4** via LangChai…

    Python 2

  2. dataforge dataforge Public

    πŸš€ Docker-first data engineering platform: Spark, Hadoop (HDFS), Kafka + Zookeeper, Airflow, Postgres, MongoDB, and Jupyter β€” all wired together with Docker Compose.

  3. Analyzing-Movie-Ratings-in-Hive Analyzing-Movie-Ratings-in-Hive Public

    The "Analyzing Movie Ratings" project uses Apache Hive to manage and analyze movie rating data. Hive, built on Hadoop, allows querying and analysis of large datasets in Hadoop's HDFS. This project …

  4. Query-gpt-codellama-7b Query-gpt-codellama-7b Public

    A local LLM-powered natural language to SQL converter using FastAPI and SQLite. This project allows users to input natural language queries and receive corresponding SQL queries, execute them on a …

    Python

  5. learn-apache-airflow learn-apache-airflow Public

    A complete Apache Airflow learning journey from scratch to advanced with step-by-step DAGs, examples, and real-world projects. apache-airflow, data-engineering, etl-pipeline, python, orchestration,…

    Python 1

  6. ashutroKumar.github.io ashutroKumar.github.io Public

    HTML