ab·
← Back to selected work

MSc Data Science / Applied machine learning / 2024

AI-Enabled Digital Twin
for Enhanced Healthcare

Exploring the data science behind combining indoor location, environmental context and simulated patient vitals.

My MSc dissertation investigates a proposed extension to a digital-twin platform: bringing wearable vital-sign data together with real-time location and environmental information. The portfolio focus is the research approach, data preparation and evaluation of machine-learning models.

Synthetic-data researchPatient vitals were simulated in Python because real patient records were unavailable. The work explores a technical approach; it does not establish clinical effectiveness.

The research question

Does including occupancy and environmental context help a model distinguish normal and elevated heart-rate categories within the study dataset?

My technical contribution

  • Proposed a unified data structure connecting indoor location, environmental measurements and wearable vital signs.
  • Used Python, NumPy and pandas for synthetic-data generation and data preparation.
  • Explored class imbalance with SMOTE and compared classification approaches.
  • Evaluated models with and without occupancy to investigate the contribution of that feature.

The data science workflow

  1. 1. Frame and connect the dataDefine how location identifiers, timestamps, environmental readings and vital signs could be brought into a common analytical dataset.
  2. 2. Generate and prepareSimulate patient vitals, use seeded randomness for reproducibility, and prepare numerical features with NumPy and pandas.
  3. 3. Examine class balanceThe dissertation explores SMOTE to address under-represented elevated-heart-rate examples. A reproducible rerun must keep resampling within the training data.
  4. 4. Compare modelsCompare Support Vector Machine and Random Forest classifiers using scikit-learn. The dissertation also discusses CNN and LSTM experiments; the appendix focuses on SVM and Random Forest.
  5. 5. Test the value of contextCompare feature sets with and without occupancy. Other inputs include temperature, humidity, noise level, air quality and duration.
  6. 6. Evaluate beyond accuracyReview precision, recall, F1-score and class support alongside accuracy to understand the different types of classification error.

What the comparison is designed to reveal

Removing occupancy creates a feature-ablation comparison: a way to examine whether that input adds useful information within the experimental dataset. It does not prove that occupancy causes changes in heart rate or that the model generalises to real patients.

Methods and tools

PythonNumPypandasscikit-learnSMOTE / imbalanced-learnMatplotlibClassificationFeature ablation

Research scope and limitations

The data-integration architecture is proposed, and the patient-vital experiments use synthetic data. The reported scores remain historical dissertation results, not independently reproduced portfolio benchmarks. A follow-up implementation would need to examine data-generation assumptions, preprocessing, leakage controls and performance on an untouched test set.

Connecting product thinking with data science

This work connects my digital-twin product experience with hands-on analytical methods: starting with a use case, deciding what data is needed, comparing modelling choices and communicating the limits of the evidence.

Original MSc project dissertation by Arjun Chirayath Balakrishnan, dated 30 September 2024. Shared here as academic work, not as a peer-reviewed publication. The PDF is the original submission and retains its original wording and reported results. This page highlights the methods; no live prediction service is presented.