Modern data warehouse from scratch
Built a modern data warehouse using SQL Server, covering the full ETL pipeline, dimensional data modelling, and analytics-ready layers.
Turning tangled data into clear decisions.
I help teams see what the numbers are really saying, translating dashboards, models and messy spreadsheets into decisions people can act on.
Download CV ↓I sit between the data and the people who need it, asking the sharper question before writing the query, then delivering answers that change what the team does next.
My work spans the full analytics loop: scoping the real business question, wrangling and modeling the data, building dashboards and forecasts, and most importantly communicating the "so what" to stakeholders who don't live in spreadsheets.
I care about clean data, honest charts, and recommendations that hold up in the meeting after the meeting.
Built a modern data warehouse using SQL Server, covering the full ETL pipeline, dimensional data modelling, and analytics-ready layers.
Advanced SQL analytics on top of a data warehouse's gold layer — exploration, ranking, time-series, segmentation and RFM, built with window functions, CTEs and T-SQL reporting views.
End-to-end retail analytics: cleaned a 3,900-row consumer shopping dataset in Python/pandas, then ran SQL Server EDA covering revenue drivers, customer segmentation and product performance.
A deep-learning sarcasm detector on news headlines, combined with VADER sentiment to correct misleading labels — comparing LSTM, Bi-LSTM, Attention-LSTM and Transformer models.
A relational schema for branches, employees, members, books, and issue/return tracking — full CRUD, CTAS summary tables, and stored procedures for issuing/returning books and branch performance reporting.
A simulated university e-commerce platform: 3NF schema design, synthetic data generation, automated GitHub Actions ETL running every 6 hours, and R/Quarto analysis of purchasing behaviour.
A retrieval-augmented generation pipeline answering natural-language questions about Mumbai's railway stations — MiniLM embeddings, FAISS nearest-neighbour search, and GPT-2 grounded generation.
A full NLP pipeline — cleaning, TF-IDF vectorization, and a benchmark of eleven classifiers plus Voting/Stacking ensembles. Multinomial Naive Bayes was selected and serialized for its precision.
A SQL project: database setup, null-handling data cleaning, and business queries covering best-selling months, top 5 customers, and morning/afternoon/evening sales shifts.
15 SQL queries exploring Netflix's movie/TV catalogue — content-type split, top countries and genres, longest titles, and keyword-based content categorisation using CTEs and recursive queries.
Explore the visualizations behind these projects — filterable, published dashboards you can click through yourself.