Boocture: Bridging the Gap Between MOOCs and eBooks via Hierarchical Indexing

Boocture: Automatic Educational Videos Hierarchical Indexing with eBooks

2020-12-08
Shay Horovitz, Yaniv Ohayon
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Boocture, an automated framework designed to link educational video segments with the hierarchical structure of eBooks. Utilizing NLP, Bayesian optimization, and machine learning models like KNN and Spectral Clustering, it achieves state-of-the-art alignment across chapter, sub-chapter, and sub-topic levels.

TL;DR

Educational videos are engaging but often disorganized; eBooks are thorough but can be "painful" to navigate. Boocture is a novel ML-driven solution that automatically aligns video segments with the specific chapter-level index of a textbook. By using transcripts and hierarchical matching, it allows students to jump from a confusing 10-minute video clip directly to the precise sub-chapter in a book that explains the concept.

Background: The Educational Content Silo

In the current digital learning landscape, MOOCs and textbooks exist as alternative resources rather than complementary ones. A student watching a lecture on "The Brain Structure" often has to manually hunt through a 500-page PDF to find the corresponding reading. Previous research attempted video segmentation using visual cues (slides) or simple keyword matching, but these methods often ignore the hierarchical wisdom embedded in a textbook’s table of contents.

The Boocture Methodology: Two Paths to Alignment

The core of Boocture lies in its ability to take raw video transcripts and map them onto a hierarchical tree. The researchers explored two primary segmentation strategies:

1. Independent Segmentation Method (ISM)

This method views the video as its own data source first. It uses a sliding window to scan transcripts, converts them into TF-IDF vectors, and builds an Adjacency Matrix. By applying Spectral Clustering, the system identifies natural "topic breaks" within the lecture without needing the book's input initially.

2. Dependent Segmentation Method (DSM)

Here, the book acts as the anchor. The video timeline is divided based on which book chapter segments are most similar to. While intuitive, experimental results showed that lecture flow doesn't always mirror a book's structure, making ISM generally more robust.

Model Architecture Figure 1: The Boocture Pipeline – from raw video/book input to final hierarchical mapping.

Hierarchical Matching: The Secret Sauce

Once segments are identified, how do we know if a segment belongs to chapter 2.3 or 2.3.1? Boocture uses a Hierarchical Skeleton:

  • Step 1: Match segment to a main chapter.
  • Step 2: Inherit the parent chapter and search only within its sub-chapters.
  • Keyword Boosting: To ensure accuracy, the titles of chapters are "artificially weighted" (repeated 20x in the corpus) to ensure the vectorizer prioritizes the author’s defined topics.

Adjacency Matrix Figure 2: The Adjacency Matrix used in ISM to identify clusters of similar topics.

Experiments and Insights

The researchers tested the system in two scenarios:

  • Explicit Relation: The video was designed specifically for the book (e.g., MIT OpenCourseWare).
  • Implicit Relation: The video and book are from different sources but cover the same topic.

Key Findings:

  • KNN Smoothing Wins: Using a K-Nearest Neighbor approach for matching (smoothing the "votes" of adjacent segments) yielded an 87% hit ratio, outperforming simple majority voting.
  • Human Consensus: The study used "Normalized Entropy" to measure how much humans agreed on a segment's topic. When humans were certain (low entropy), Boocture was nearly 100% accurate.

Performance Data Table 1: Comparison of Matching Methods. ISM combined with KNN consistently provided superior results.

Critical Analysis & Future Outlook

Boocture proves that the "Table of Contents" of a book is a powerful inductive bias for organizing unstructured video data. However, the reliance on TF-IDF means the system might struggle with highly paraphrased content where the lecturer uses vastly different terminology than the author.

Takeaway: The future of personalized learning lies in "Cross-Medium Navigation." Systems like Boocture could soon be integrated into YouTube or Kindle, allowing for a seamless "Learn by Watching, Deepen by Reading" experience.

Limitation: The current system relies on text transcripts. Future iterations could benefit from Visual Grounding—recognizing diagrams in the video that also appear in the book to further anchor the alignment.

Find Similar Papers

Try Our Examples

  • Find recent papers addressing automatic video-to-text alignment in educational contexts using Large Language Models (LLMs) instead of traditional TF-IDF.
  • What were the seminal works in "Topic Segmentation of Lecture Videos," and how does the Boocture hierarchical approach differ from their linear segmentation methods?
  • Are there studies that apply Boocture-like hierarchical indexing to non-textbook sources, such as aligning news videos with technical documentation or whitepapers?
Contents
Boocture: Bridging the Gap Between MOOCs and eBooks via Hierarchical Indexing
1. TL;DR
2. Background: The Educational Content Silo
3. The Boocture Methodology: Two Paths to Alignment
3.1. 1. Independent Segmentation Method (ISM)
3.2. 2. Dependent Segmentation Method (DSM)
4. Hierarchical Matching: The Secret Sauce
5. Experiments and Insights
6. Critical Analysis & Future Outlook