Neuro-Screen: A Hybrid Ensemble Framework for Detection of Cognitive Impairment in Insomniac University Students
A CatBoost + ANN hybrid ensemble that detects cognitive impairment in insomniac university students: 95.20% accuracy, 0.982 ROC-AUC.
Thesis that fuses a gradient-boosting classifier with a three-layer neural network by averaging their probability outputs. Trained on 2,237 survey responses from Bangladeshi university students, the ensemble beats every standalone model across all metrics and pinpoints the lifestyle factors that most strongly predict cognitive decline.
[ 01 ]
Research Overview
Cognitive impairment linked to insomnia is widespread among university students but rarely screened for. Existing computational work stops at diagnosing sleep disorders or general mental-health distress. It does not classify the cognitive decline that poor sleep produces.
Neuro-Screen fills that gap with a hybrid ensemble: a CatBoost gradient-boosting classifier, which is strong on categorical, structured survey data, combined with a three-layer ANN, which is strong on non-linear feature interactions. The two paths run in parallel and their probability outputs are averaged into a final Healthy / Impaired decision.
Evaluated on a stratified 80/20 holdout, the ensemble reaches 95.20% accuracy and a 0.982 ROC-AUC, higher than either component alone, and is packaged as a Streamlit screening dashboard with a conversational check-in.
[ 02 ]
Problem Statement
- Most university students get insufficient sleep and a substantial share meet clinical insomnia criteria. Both conditions measurably impair attention, memory, and academic performance.
- Existing ML studies target sleep disorders themselves (insomnia, apnea) or general mental health, not the cognitive decline insomnia causes, and they rely on standalone classifiers rather than hybrid ML-DL designs.
- As a result, no automated framework exists for early detection of cognitive impairment among insomniac university students.
[ 03 ]
Objective
- Develop a hybrid CatBoost + ANN ensemble that fuses both models' probability outputs into a binary cognitive-health classification.
- Benchmark the ensemble against standalone CatBoost and ANN across accuracy, precision, recall, F1-score, and ROC-AUC.
- Identify which behavioral, psychological, and sleep-related features most strongly predict cognitive impairment.
- Provide a practical, data-driven screening approach that helps universities flag at-risk students early.
[ 04 ]
Methodology
- Survey design and collection: 2,237 responses from students at 30+ Bangladeshi universities, capturing demographics, lifestyle and behavioural habits, mental stamina, insomnia indicators, and cognitive-symptom indicators.
- Preprocessing and target engineering: prepared the 21 survey input features (20 categorical + age), then aggregated multiple cognitive-symptom indicators into a single binary Healthy / Impaired target.
- Dual-path training: CatBoost on the raw categorical answers and a three-layer ANN (128-64-1 MLP) on a 104-dimension one-hot encoding of the same features, trained in parallel.
- Ensemble blending: the final decision is the arithmetic mean of both models' predicted probabilities.
- Evaluation: an 80/20 stratified split (1,790 train / 447 test) scored with accuracy, precision, recall, F1-score, and ROC-AUC, plus confusion-matrix and learning-curve analysis.
[ 05 ]
Models Used
CatBoost
Gradient-boosted decision trees with native categorical-feature handling; captures non-linear interactions between survey answers.
ANN (128-64-1 MLP)
Three-layer feed-forward network (128-64-1) over the one-hot encoded survey features.
Hybrid ensemble
Arithmetic-mean blending of both models' probabilities, combining complementary strengths and beating either alone.
[ 06 ]
Dataset
- 2,237 survey responses from students at 30+ Bangladeshi universities.
- 21 input features spanning demographics, caffeine intake, bedtime device use, cognitive load, stress frequency, mental stamina, spacing-out / audio-lag episodes, sleep hours, night awakenings, sleep quality, forgetfulness, reminder reliance, brain fog, missed deadlines, GPA impact, and fatigue.
- A binary target engineered by aggregating multiple cognitive-symptom indicators into Healthy / Impaired.
[ 07 ]
Implementation
- Built in Python on Google Colab with pandas (data manipulation), CatBoost (gradient boosting), and PyTorch (ANN).
- Both modules trained in parallel, then combined with probability-level ensemble blending for the final binary prediction.
- Validated on the 447-response held-out test set with a confusion matrix and ROC-AUC analysis.
- Deployed as a Streamlit screening prototype: a slider-based Quick Check-in and a conversational check-in assistant that both return a 0–100 risk score with model confidence and the driving factors.
[ 08 ]
Key Features
- Parallel CatBoost + ANN architecture with an averaged ensemble decision.
- Target engineering that turns multiple cognitive-symptom signals into one robust binary label.
- Stratified evaluation with confusion matrix, ROC-AUC, and per-model comparison.
- Feature-importance analysis surfacing the strongest predictors of impairment.
- Streamlit prototype with both a form-based and a conversational screening interface.
[ 09 ]
Results
- The hybrid beats both standalones on every metric: CatBoost reaches 94.12% accuracy and the ANN 91.25%, while the ensemble reaches 95.20%.
- Hybrid ROC-AUC of 0.9820 vs. CatBoost 0.9780 and ANN 0.9450. The sharp early trajectory shows high sensitivity with a low false-positive rate across thresholds, which suits screening use cases.
- The insomnia–cognition link is confirmed in the data: roughly 55% of the insomniac group is classified Impaired versus 25% of the non-insomnia group.
[ 10 ]
Baseline comparison
| Model | Accuracy | Delta |
|---|---|---|
| NeuroScreen ensemblewinner | 95.20% | baseline |
| CatBoost standalone | 94.12% | +1.08 vs best |
| ANN standalone | 91.25% | +3.95 vs best |
[ 11 ]
What predicts impairment
01
Mental / Physical Fatigue
importance 0.095
02
Stress Frequency
importance 0.095
03
GPA Impact (Sleep)
importance 0.095
04
Overall Sleep Quality
importance 0.093
05
Average Sleep Hours
importance 0.092
06
Bedtime Device Use
importance 0.082
07
Reminder Reliance
importance 0.066
08
Caffeine Intake
importance 0.063
[ 12 ]
Outcome
- A validated, cost-effective screening framework for university health surveillance that flags at-risk students for academic and psychological intervention.
- Identified mental/physical fatigue, stress frequency, and GPA impact as the strongest predictors (0.095), with sleep quality (0.093) outweighing sleep duration (0.092).
- Positioned as a screening aid rather than a medical diagnostic, following privacy-by-design and responsible-AI principles on anonymized data.
[ 13 ]
Tools & Technologies
[ 14 ]
Challenges
- Self-reported survey data carry response and recall bias.
- A convenience-sampled, self-reported survey cohort limits broader generalizability.
- Binary classification only: no severity levels for cognitive impairment.
- Hyperparameter tuning was constrained by local CPU/RAM availability.
[ 15 ]
Future Improvements
- Multi-institutional and longitudinal data collection to study temporal changes in cognitive health.
- Extend the app's per-prediction SHAP explanations with LIME and clinical validation.
- Add objective physiological biomarkers such as actigraphy or EEG signals.
- Deploy the framework within real campus health-surveillance systems.
[ 16 ]
Related work
Twitter Sentiment Analysis with TF-IDF and Multinomial Naïve Bayes
Course research on the Sentiment140 corpus. 1.6M tweets are balanced down to 200k (100k positive + 100k negative), run through a deep preprocessing chain: lowercasing, URL/@/# removal, punctuation stripping, stopword removal, WordNet synonym substitution, Porter stemming, and lemmatization, then vectorized with TF-IDF and classified by Multinomial Naïve Bayes on a stratified 80/20 split.
Early Warning Model for High-Value Customer Drop-Off
Course research in Data Science. Cleans 541,909 transactions into a 4,338-customer RFM matrix, segments customers with unsupervised clustering compared across three algorithms, then builds a supervised early-warning classifier on month-over-month RFM decay to flag high-value customers at risk of churn.
Explainable Bangla Toxic Comment Detection using BanglaBERT with SHAP and LIME
Data-mining research (IEEE-format paper) that fine-tunes BanglaBERT on Bengali toxicity data using 5-fold stratified cross-validation and adds an explainability layer: SHAP and LIME reveal which words drove each toxic / non-toxic decision, tackling the black-box problem in low-resource Bangla NLP.
[ 17 ]
Get the thesis
The complete thesis report covers the literature review, dataset construction, full feature set, and 16 reference papers. Reach out and I'll share it.
Request the thesis