SeqST-ResNet: Rethinking Sequential Order in Spatio-Temporal Task Prediction
SeqST-ResNet: A Sequential Spatial Temporal ResNet for Task Prediction in Spatial Crowdsourcing
The paper introduces SeqST-ResNet, a novel deep learning architecture for task appearance prediction in Spatial Crowdsourcing (SC). By refining the spatio-temporal residual learning framework to process historical data in a sequential rather than concatenated manner, it achieves State-of-the-Art (SOTA) performance on real-world taxi request datasets.
TL;DR
Spatial Crowdsourcing (SC) platforms like Uber or Deliveroo rely on matching workers to tasks. To maximize efficiency, platforms must "predict the future" to guide idle workers to high-demand areas. SeqST-ResNet improves upon previous SOTA models by replacing simple data concatenation with a sequential residual learning structure, reducing prediction error (RMSE) by over 20% compared to the standard ST-ResNet.
Background & Motivation: The "Concatenation" Flaw
In urban computing, we often treat a city as a grid of "task images." Current deep learning models, such as ST-ResNet, handle temporal trends (hourly, daily, weekly) by stacking these images into a single 3D tensor and feeding them into a CNN.
However, the authors identify a critical logical gap: Concatenation ignores the inherent sequential flow of time. In a standard CNN, the relationship between and is treated similarly to any other feature dimension. SeqST-ResNet posits that a model should "absorb" history step-by-step to capture the true evolution of demand.
Methodology: The Power of Sequential Addition
SeqST-ResNet maintains three branches to capture:
- Interval-level (Closeness): Immediate past hours.
- Day-level (Period): Periodic patterns at the same hour yesterday.
- Week-level (Trend): Long-term trends from the same day last week.
Architecture Decomposition
The core Innovation lies in the Addition Layer. Instead of stacking all time steps and then convolving, SeqST-ResNet processes , passes it through a Residual Unit, and then adds the result to the next time step before it enters its own convolution layer. This creates a chain of dependency similar to an RNN but leverages the spatial power of CNNs.

Fig 1: The SeqST-ResNet Architecture showing the three-tier temporal modeling and the sequential flow within each branch.
Spatial Intuition
Each grid cell in the image isn't independent. Neighboring grids often belong to the same functional zone (e.g., a business district). By using a kernel, the model captures local spatial correlations, while the depth of the ResNet allows it to "see" more distant spatial dependencies.
Experimental Validation
Using the Didi GAIA dataset (11 million taxi orders in Chengdu), the model was tested against several baselines.
Performance Gains
- Smoothing is Key: The best variant,
SeqST-ResNet-3AVG, uses a moving average to filter out "noise" or anomalies in the task data, achieving the lowest RMSE of 12.95. - High-Intensity Regions: The model shines brightest in "Busy" regions. As demand density increases, the gap between SeqST-ResNet and standard ST-ResNet widens, proving that temporal sequencing is more critical when the signal is strong.

Fig 2: Prediction error across regions with different request intensities. Note how SeqST-ResNet (blue) stays significantly lower than ST-ResNet (green) as the region becomes busier.
Critical Insight & Conclusion
The success of SeqST-ResNet suggests that temporal order is an Inductive Bias that cannot be ignored in spatial tasks. While traditional CNNs are excellent for static images, urban tasks are dynamic "videos."
Limitations & Future Work
While powerful, the model relies on a fixed grid partition. Modern research is moving toward Graph Convolutional Networks (GCNs) to model non-Euclidean road networks. Additionally, the authors suggest that an Attention Mechanism could further help the model focus on specific historic "events" (like a sudden festival) that break the standard periodic cycle.
Takeaway for Practitioners: If you are building demand prediction systems, don't just stack your historical features. Respect the arrow of time through sequential modeling architectures.
