This repository contains a collection of machine learning projects covering the full ML pipeline — from basic data analysis and model building to advanced techniques such as model validation, feature engineering, and ensemble methods.
The goal of this repository is to demonstrate a structured understanding of machine learning concepts and practical skills in applying them to real-world data.
File: machine_learning_pipeline_basics.ipynb
- Introduction to ML workflow
- Data preprocessing and train/test split
- Baseline models and evaluation
File: linear_regression_regularization.ipynb
- Linear regression and gradient descent
- Overfitting vs underfitting
- Regularization techniques (L1, L2)
- Model evaluation metrics
File: model_validation_and_hyperparameter_tuning.ipynb
- Cross-validation techniques
- Hyperparameter optimization
- Feature selection methods
- Avoiding data leakage
File: classification_models_comparison.ipynb
- Logistic Regression, KNN, Naive Bayes, SVM
- Comparison of classification algorithms
- Evaluation using performance metrics
File: tree_based_models_and_ensembles.ipynb
- Decision Trees (CART)
- Random Forest
- Gradient Boosting
- Comparison of ensemble methods
File: clustering_feature_engineering.ipynb
- K-means, DBSCAN, hierarchical clustering
- Cluster quality evaluation (Elbow, Silhouette)
- Using cluster labels as features in supervised learning
File: dimensionality_reduction_analysis.ipynb
- PCA and matrix factorization
- Manifold learning methods
- Visualization and structure preservation
- Machine Learning fundamentals
- Supervised & Unsupervised Learning
- Model evaluation and validation
- Feature engineering
- Dimensionality reduction
- Ensemble methods
- Python
- NumPy, pandas
- scikit-learn
- matplotlib
- Application of ML to neurobiological and psychological datasets
- Advanced deep learning models
- Research-oriented projects combining ML and cognitive science
Anna Panasenko
GitHub: https://github.com/Nyutapan