A curated portfolio of 30 unsupervised machine learning projects covering clustering, anomaly detection, topic discovery, recommendation systems, vector retrieval, association-rule mining, dimensionality reduction, behavioral segmentation and graph community detection.
The repository is designed as a practical reference for learners, practitioners, recruiters, and reviewers who want to explore complete unsupervised learning workflows with reproducible code and preserved notebook outputs.
- End-to-end unsupervised learning projects with practical problem statements.
- Clustering, anomaly detection, NLP, recommendation, GenAI, cybersecurity, healthcare, industrial, geospatial, and customer analytics use cases.
- Executed notebook outputs preserved for easier review.
- Reproducible workflows with preprocessing, model selection, internal validation, stability analysis, interpretation, and documented limitations.
- A mix of maintained original projects, selectively refurbished projects, and newly added portfolio projects.
- No prediction target is used during unsupervised model fitting.
- Internal metrics are paired with stability, interpretability, and domain-relevance checks.
- Synthetic or known labels, when available, are used only for retrospective evaluation—not during model training.
- Demo datasets and generated results are deterministic and clearly identified.
- Anomaly scores, clusters, recommendations, and discovered patterns are presented as decision-support signals rather than ground truth.
- Relative paths, reproducible seeds, dependency guidance, saved notebook outputs, and the locked creator branding are preserved.
Tajamul Khan — Data Scientist and AI Engineer
