Reviving the Past: 3D Reconstruction of Lost Heritage via Machine Learning and Photogrammetry

Processing Historical Film Footage with Photogrammetry and Machine Learning for Cultural Heritage Documentation

2019-10-15
Francesca Condorelli, Fulvio Rinaudo
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an integrated pipeline combining Faster R-CNN for object detection and Structure-from-Motion (SfM) for the 3D metric reconstruction of Cultural Heritage from historical film footage. The study utilizes the open-source COLMAP framework to overcome the technical challenges of low-quality archival videos, achieving a metric accuracy comparable to modern photogrammetric benchmarks.

TL;DR

Researchers at the Polytechnic of Turin have developed a workflow to turn historical "city films" into accurate 3D models. By combining Faster R-CNN for automatic monument detection and COLMAP (SfM) for metric reconstruction, they successfully extracted 3D data from distorted, low-resolution 16mm film with an accuracy of approximately 0.33 cm, proving that archival footage is a viable source for scientific documentation.

Context: Why Historical Films are Data Goldmines

In the field of Cultural Heritage, we often face a tragic gap: monuments are destroyed or altered before they can be digitally surveyed. Historical archives contain the only visual traces of these structures. However, these "data" are trapped in unorganized film reels with:

  • No Metadata: No focal length, sensor size, or lens distortion info.
  • Poor Geometry: Cameras were panned or tilted for aesthetics, not for the "3x3 rules" of photogrammetry.
  • Physical Degradation: Scratches, humidity damage, and low resolution.

Methodology: The Detection-to-Reconstruction Pipeline

The authors proposed a specialized workflow designed to bypass the limitations of commercial black-box software.

1. Automated Detection (Machine Learning)

Instead of manually scrubbing through hours of footage, the team utilized Faster R-CNN implemented in TensorFlow. They focused on the Tour Saint Jacques in Paris—a monument that has been physically moved and altered over time.

  • The dataset: 600 training images including contemporary photos, archival stills, and "negative" images of the Paris skyline to reduce false positives.
  • Performance: The model achieved 100% accuracy in identifying the tower across 10 historical films (1910s–1960s).

2. Overcoming Camera Motion

Most historical videos use Tilting (Up/Down) or Panning (Rolling). In these cases, the "baseline" (the distance between camera positions) is tiny, which usually breaks 3D triangulation.

Camera Motion Schemes Figure: The common camera motions found in archival footage—Tilting, Panning, and Trucking.

3. Open-Source SfM via COLMAP

By using COLMAP, the researchers could control the Simple Radial Camera Model. This is crucial when the internal orientation of the lens is unknown. They used "Sequential Matching," which is optimized for video frames where overlapping features are found in a temporal order.

3D Reconstruction Process Figure: The photogrammetric pipeline—from feature extraction to 3D point cloud reconstruction.

Experiments: Validating Accuracy

The fundamental question was: Is the 3D model scientifically valid? The authors created a benchmark using a high-end Canon 5DS R (50.6 MP) to shoot the Valentino Castle in Turin, mimicking the "Tilting" motion of old films.

MetricHistorical Case Study (16mm)Modern Benchmark (5DS R)
Mean Error (px)0.23 px0.36 px
GSD (Ground Sample Dist.)1.43 cm/px1.2 cm/px
Mean Error (cm)0.33 cm0.1 cm

Surprisingly, the historical footage performed exceptionally well. The error was within the range of one pixel, making it sufficiently accurate for architectural history and restoration planning.

Accuracy Comparison Figure: Deviation analysis between the model generated from film and the "ground truth" 3D mesh.

Critical Insight & Conclusion

The study proves that we don't always need "perfect" data to get professional results. The "intelligence" lies in the selection of the tool—using open-source SfM allows researchers to manually intervene where "closed" software like Metashape might fail due to lack of EXIF data.

Takeaway for the Industry: Archives are no longer just passive repositories; they are active datasets. Using this pipeline, historians can virtually "re-walk" cities that no longer exist, and architects can verify the structural transformations of heritage sites over a century with centimeter-level precision.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning and Structure-from-Motion (SfM) specifically for the 3D reconstruction of "lost" or "destroyed" urban heritage using low-resolution archival sources.
  • Which study first introduced the use of Faster R-CNN for architectural element detection, and how does the approach in this paper modify the training strategy for historical, degraded video frames?
  • Explore if current Neural Radiance Fields (NeRF) or Gaussian Splatting techniques have been applied to sparse, low-quality historical film footages for volumetric reconstruction of Cultural Heritage.
Contents
Reviving the Past: 3D Reconstruction of Lost Heritage via Machine Learning and Photogrammetry
1. TL;DR
2. Context: Why Historical Films are Data Goldmines
3. Methodology: The Detection-to-Reconstruction Pipeline
3.1. 1. Automated Detection (Machine Learning)
3.2. 2. Overcoming Camera Motion
3.3. 3. Open-Source SfM via COLMAP
4. Experiments: Validating Accuracy
5. Critical Insight & Conclusion