Portfolio

Projects & Work

A production AI platform I architected end to end, plus 14+ open-source projects across RAG, machine learning, NLP, infra, and web, plus a stack of Kaggle notebooks.

Open Source: 14+ / Flagship: Private
Flagship

Production Work

My biggest project lives in production, not on GitHub. Here's what it does.

Production system Core contributor · RAG, underwriting & LLM orchestration · source private

CRE Loan Underwriting Platform

A production AI platform for commercial real estate lending whose core I built, turning raw financial documents into auditable credit decisions. Multi-LLM orchestration, a RAG document-intelligence pipeline, and a deterministic credit risk engine, deployed to AWS.

01
RAG Document Pipeline

Classification + extraction over 24 CRE document types, OCR routing, dedup, and pgvector retrieval with reranking.

02
Credit Risk Engine

Deterministic, rule-based scoring (DSCR, LTV, NOI, FICO) with tiered decisions and zero LLM in the critical path.

03
Multi-LLM Orchestration

Provider fallback across Bedrock, Gemini & OpenAI with per-provider circuit breakers and graceful degradation.

AWS Bedrock Textract pgvector / HNSW FastAPI ARQ / Redis ECS Fargate GitHub Actions boto3
See the full capability breakdown
Open Source

Open-Source Projects

End-to-end projects, with source on GitHub unless noted

github: @bijay-odyssey / visibility: public

Showing all projects
AI / RAG

Upstream Fix: Qdrant Python Client

Python · Qdrant · open source

Formula evaluation in Qdrant's official Python client accepted zero raised to a negative exponent, which is undefined, instead of rejecting it. Reviewed and merged by a Qdrant maintainer after CI passed.

  • Merged into the client used by everyone running Qdrant from Python, not a fork
View PR #1429
AI / RAG

Smart Kar: Tax Document Analysis

FastAPI · pgvector · Arq · Tesseract · React · source private

A final-year college project and an end-to-end applied AI system for Nepal's SMEs: it ingests tax documents, extracts and classifies them, flags anomalies, and produces IRD-ready VAT and TDS reports. Sole developer of the system, in a two-person team.

  • Bilingual extraction: PyMuPDF for digital PDFs with a Tesseract OCR fallback for scanned input in English and Nepali, table reconstruction from OCR word geometry, and fuzzy vendor matching
  • LLM-assisted classification, anomaly detection and tax-rule extraction, with a passthrough fallback when no provider key is configured, so the pipeline degrades to deterministic behaviour instead of failing
Infra
smart-secretary

Python · FastAPI · Pydantic v2 · pytest

An orchestration layer over a fleet of local services. It routes a natural-language request to whichever project can answer it, starts that service on demand, and says plainly when the capability does not exist instead of improvising one.

  • Capability registry with four transports (HTTP, subprocess, cross-interpreter Python, in-process) over 34 capabilities from 38 registered projects, each declaring its contract in a versioned manifest
  • A path denylist and per-capability egress policy evaluated before any service starts or any prompt leaves the machine, so a private corpus cannot reach a third-party model. 12,000 lines, 215 tests
Machine Learning
scrapify

Python · Apify · pandas

A scraping and analysis pipeline for hiring-market data, built to replace opinion about in-demand skills with measured evidence.

  • Caught a confounder in its own first conclusion before acting on it: one employer's near-identical postings drove 54% of the apparent demand for a single framework and 91–100% of three others. Excluding that employer moved the signal from 26.6% to 14.8% and the rest to noise
NLP
Intelli-oppo

Python · CLI · MIT

A debate engine that argues the opposite of whatever you claim, without ever fabricating to do it. Useful as a devil's-advocate pass over an argument you are too close to.

GitHub
AI / RAG
RAG Precision Enhancement

FAISS · Cross-Encoder · Reranking

Comparative study showing why retrieval alone isn't enough. Two-stage retrieval + reranking pipeline with noise injection to stress-test semantic search performance.

  • Noise injection with 20+ hard-negative documents
  • Cross-encoder reranking (BAAI/bge-reranker-base)
  • Evaluated via Hit Rate@k and MRR metrics
GitHub
Web App
ChurnShield

Flask · Scikit-learn · SHAP · SQLite

Flask web app for real-time churn prediction with user authentication, admin dashboard, CSV export, and SHAP-powered explainability. Automated retention strategy generator included.

  • Real-time prediction with Random Forest pipeline
  • SHAP-based feature importance per prediction
  • Automated retention strategy generator
Machine Learning
IEEE Fraud Detection

XGBoost · SHAP · Imbalanced Learning · Optuna

Transaction classification on 1M+ Kaggle IEEE-CIS records. Achieved ROC-AUC 0.95 and F1 0.66 on the imbalanced fraud class with Optuna hyperparameter tuning.

  • Processed 1M+ transaction records
  • Feature reduction via Spearman correlation
  • Hyperparameter tuning with Optuna
Machine Learning
Time Series Forecasting

LightGBM · Prophet · Feature Engineering

SKU-level supermarket price forecasting on 1.7M+ rows using LightGBM and Prophet with lag, rolling window, and calendar features.

  • Processed 1.7M+ retail records
  • TimeSeriesSplit cross-validation
  • Evaluated via RMSE, MAPE, SMAPE
Machine Learning
Retail Price Optimization

Random Forest · Optuna · SHAP

Competition-aware ML workflow integrating historical sales and competitor pricing. Achieved R² ≈ 0.92 with Random Forest and SHAP explainability for business decisions.

  • Full ML pipeline with EDA
  • Hyperparameter tuning with Optuna
  • Profit and revenue optimization analysis
GitHub
Machine Learning
Anomaly Detection

Isolation Forest · LOF · KMeans · DBSCAN

Multi-method anomaly detection pipeline combining model-based, cluster-based, and statistical approaches on fraud and retail datasets with PCA visualization.

  • Multiple detection algorithms benchmarked
  • Evaluation via ROC-AUC and precision/recall
  • PCA visualization of anomalies
NLP
Text Embeddings Explorer

sentence-transformers · FAISS · Visualization

Scripts and experiments for exploring text embedding models, visualizing embedding spaces, and benchmarking retrieval quality across different encoder models.

GitHub
AI / RAG
Semantic Search with FAISS

FAISS · sentence-transformers · Python

End-to-end semantic search implementation using FAISS for fast approximate nearest neighbor search with sentence-transformers for dense embedding generation.

GitHub
AI / RAG
Qdrant Vector DB Experiments

Qdrant · Python · Vector Search

Experiments with Qdrant vector database: collection management, payload filtering, hybrid search, and performance benchmarks for RAG applications.

GitHub
NLP
Text Chunking Strategies

Python · langchain-text-splitters · Benchmarking

Comparative study of text chunking strategies (fixed-size, sentence-based, recursive, and semantic chunking), evaluating their impact on RAG retrieval quality.

GitHub
Machine Learning
Regression Portfolio

Scikit-learn · XGBoost · LightGBM · Optuna

Collection of regression projects across different domains (housing prices, demand forecasting, and insurance costs), with thorough EDA, feature engineering, and evaluation.

Machine Learning
Customer Segmentation

KMeans · DBSCAN · PCA · Silhouette Analysis

Unsupervised learning project exploring clustering algorithms for customer segmentation. Includes PCA for dimensionality reduction and silhouette analysis for optimal cluster selection.

No projects found for this filter.

Kaggle

Notebooks & Competitions

27 public notebooks · 58 upvotes: community-tested data science