RAPT: Bridging the Gap in Healthcare Representation via Time-Aware Pre-training
RAPT: Pre-training of Time-Aware Transformer for Learning Robust Healthcare Representation
This paper introduces RAPT (RepresentAtion by Pre-training time-aware Transformer), a novel framework for learning robust healthcare representations from Electronic Health Records (EHR). By combining a Time-Aware Multi-head Attention (TMA) mechanism with three specialized self-supervised pre-training tasks, the model achieves state-of-the-art performance across four pregnancy-related downstream tasks, including Gestational Diabetes and Hypertension prediction.
TL;DR
Predicting pregnancy complications from Electronic Health Records (EHR) is notoriously difficult due to irregular check-up schedules and missing data. This paper presents RAPT, a Time-Aware Transformer meticulously designed to learn robust medical representations through self-supervised pre-training. By training on "surrogate" tasks like masking and similarity matching, RAPT outperforms traditional models by a significant margin, offering both high accuracy and clinical interpretability.
The "Sparse & Irregular" Wall in Healthcare AI
Mining EHR data, specifically in prenatal care, involves navigating a minefield of data quality issues:
- Irregularity: Unlike fixed-rate sensor data, hospital visits occur at irregular intervals (e.g., more frequent near delivery).
- Sparsity: Many patients skip tests, and rare complications provide very few "positive" labels for training.
- Short Sequences: Early-stage prediction essentially means making life-critical decisions based on just 3-5 data points.
Traditional RNNs (like LSTM) or standard Transformers treat visits as an ordered list, often ignoring the actual "days elapsed" between them. RAPT was born from the insight that the tempo of physiological change is just as important as the value of the change.
Methodology: The RAPT Architecture
RAPT's innovation lies in two dimensions: Architecture and Training Strategy.
1. Time-Aware Multi-head Attention (TMA)
Standard self-attention only knows the relative order of tokens. RAPT modifies this by introducing Time-Aware Self-Attention (TSA). It injects the time span directly into the attention score calculation:
This allows the model to "pay more attention" to visits that are clinically more relevant based on the time they occurred.

2. The Three-Pronged Pre-training
To gain "medical intuition" without relying on manual labels, RAPT uses three tasks:
- Similarity Prediction: A Siamese network that forces the model to map similar health profiles (even unlabeled) to nearby points in latent space.
- Masked Prediction: Similar to BERT's MLM, it hides clinical indicators (like blood pressure) and asks the model to reconstruct them from other visits.
- Reasonability Check: The model must detect if a visit sequence has been tampered with (e.g., swapping a 3rd-trimester weight with a 1st-trimester weight), forcing it to learn temporal growth trends.
Experimental Results: SOTA Performance
RAPT was tested against 6 baselines (including HiTANet and Retain) on four tasks:
- Diabetes & Hypertension Prediction: Binary classification tasks.
- Outcome Prediction: Regressing future physiological indicators.
- Risk Period Prediction: Identifying which specific week a patient is in danger.
| Task | RAPT (AUC) | Best Baseline (AUC) | Improvement |
|---|---|---|---|
| Diabetes | 0.867 | 0.813 (HiTANet) | +6.6% |
| Risk Period | 0.985 | 0.965 (Dipole) | +2.1% |

Ablation studies confirmed that the Masked Prediction task was the most critical for diagnostic accuracy, proving that learning to "fill in the blanks" of missing medical tests is a powerful proxy for diagnostic reasoning.
Deep Insight: Beyond Just Numbers
The authors didn't just build a black box; they integrated Sensitivity Analysis. By calculating the gradient of the output with respect to input features, the system can tell a doctor: "I am predicting Hypertension because this patient's Diastolic Pressure is rising faster than the normal pregnancy curve."
Visualizing embeddings with t-SNE: Pre-training (b) effectively clusters healthy vs. anomalous patients even before fine-tuning, compared to a random initialization (a).
Conclusion & Future Outlook
RAPT proves that Transformers, when equipped with temporal awareness and self-supervised objectives, can handle the messy reality of clinical data. While the current study focuses on pregnancy, the methodology is broadly applicable to any chronic disease management involving longitudinal tracking.
Limitations: The model currently focuses on numerical and categorical data. Integrating clinical notes (NLP) and ultrasound images (CV) into this time-aware framework remains an exciting frontier for the next generation of RAPT.
