RAPT: Bridging the Gap in Healthcare Representation via Time-Aware Pre-training

RAPT: Pre-training of Time-Aware Transformer for Learning Robust Healthcare Representation

2021-08-13
Houxing Ren, Jingyuan Wang, Wayne Xin Zhao, Ning Wu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces RAPT (RepresentAtion by Pre-training time-aware Transformer), a novel framework for learning robust healthcare representations from Electronic Health Records (EHR). By combining a Time-Aware Multi-head Attention (TMA) mechanism with three specialized self-supervised pre-training tasks, the model achieves state-of-the-art performance across four pregnancy-related downstream tasks, including Gestational Diabetes and Hypertension prediction.

TL;DR

Predicting pregnancy complications from Electronic Health Records (EHR) is notoriously difficult due to irregular check-up schedules and missing data. This paper presents RAPT, a Time-Aware Transformer meticulously designed to learn robust medical representations through self-supervised pre-training. By training on "surrogate" tasks like masking and similarity matching, RAPT outperforms traditional models by a significant margin, offering both high accuracy and clinical interpretability.

The "Sparse & Irregular" Wall in Healthcare AI

Mining EHR data, specifically in prenatal care, involves navigating a minefield of data quality issues:

  • Irregularity: Unlike fixed-rate sensor data, hospital visits occur at irregular intervals (e.g., more frequent near delivery).
  • Sparsity: Many patients skip tests, and rare complications provide very few "positive" labels for training.
  • Short Sequences: Early-stage prediction essentially means making life-critical decisions based on just 3-5 data points.

Traditional RNNs (like LSTM) or standard Transformers treat visits as an ordered list, often ignoring the actual "days elapsed" between them. RAPT was born from the insight that the tempo of physiological change is just as important as the value of the change.

Methodology: The RAPT Architecture

RAPT's innovation lies in two dimensions: Architecture and Training Strategy.

1. Time-Aware Multi-head Attention (TMA)

Standard self-attention only knows the relative order of tokens. RAPT modifies this by introducing Time-Aware Self-Attention (TSA). It injects the time span directly into the attention score calculation:

This allows the model to "pay more attention" to visits that are clinically more relevant based on the time they occurred.

Model Architecture

2. The Three-Pronged Pre-training

To gain "medical intuition" without relying on manual labels, RAPT uses three tasks:

  1. Similarity Prediction: A Siamese network that forces the model to map similar health profiles (even unlabeled) to nearby points in latent space.
  2. Masked Prediction: Similar to BERT's MLM, it hides clinical indicators (like blood pressure) and asks the model to reconstruct them from other visits.
  3. Reasonability Check: The model must detect if a visit sequence has been tampered with (e.g., swapping a 3rd-trimester weight with a 1st-trimester weight), forcing it to learn temporal growth trends.

Experimental Results: SOTA Performance

RAPT was tested against 6 baselines (including HiTANet and Retain) on four tasks:

  • Diabetes & Hypertension Prediction: Binary classification tasks.
  • Outcome Prediction: Regressing future physiological indicators.
  • Risk Period Prediction: Identifying which specific week a patient is in danger.
TaskRAPT (AUC)Best Baseline (AUC)Improvement
Diabetes0.8670.813 (HiTANet)+6.6%
Risk Period0.9850.965 (Dipole)+2.1%

Experimental Comparison

Ablation studies confirmed that the Masked Prediction task was the most critical for diagnostic accuracy, proving that learning to "fill in the blanks" of missing medical tests is a powerful proxy for diagnostic reasoning.

Deep Insight: Beyond Just Numbers

The authors didn't just build a black box; they integrated Sensitivity Analysis. By calculating the gradient of the output with respect to input features, the system can tell a doctor: "I am predicting Hypertension because this patient's Diastolic Pressure is rising faster than the normal pregnancy curve."

Qualitative Analysis - t-SNE Visualizing embeddings with t-SNE: Pre-training (b) effectively clusters healthy vs. anomalous patients even before fine-tuning, compared to a random initialization (a).

Conclusion & Future Outlook

RAPT proves that Transformers, when equipped with temporal awareness and self-supervised objectives, can handle the messy reality of clinical data. While the current study focuses on pregnancy, the methodology is broadly applicable to any chronic disease management involving longitudinal tracking.

Limitations: The model currently focuses on numerical and categorical data. Integrating clinical notes (NLP) and ultrasound images (CV) into this time-aware framework remains an exciting frontier for the next generation of RAPT.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize self-supervised contrastive learning on longitudinal Electronic Health Records (EHR) beyond RAPT.
  • Which paper originally introduced the Time-Aware LSTM (T-LSTM) for patient subtyping, and how does RAPT's Time-Aware Multi-head Attention differ in its mathematical formulation of time intervals?
  • Explore how self-supervised pre-training methods like RAPT are being adapted for multi-modal healthcare data involving both tabular EHR and medical imaging.
Contents
RAPT: Bridging the Gap in Healthcare Representation via Time-Aware Pre-training
1. TL;DR
2. The "Sparse & Irregular" Wall in Healthcare AI
3. Methodology: The RAPT Architecture
3.1. 1. Time-Aware Multi-head Attention (TMA)
3.2. 2. The Three-Pronged Pre-training
4. Experimental Results: SOTA Performance
5. Deep Insight: Beyond Just Numbers
6. Conclusion & Future Outlook