Beyond the Obvious: Diversifying Venture Capital via Indirect Association Mining
Company Investment Recommendation Based on Data Mining Techniques
This paper introduces an unsupervised, non-personalized recommendation approach for identifying alternative company investments using data mining on a massive Knowledge Graph of 7.5 million entities. By employing Indirect Association Rules Mining (IARM) and transfer learning, the system suggests investment opportunities that allow investors to diversify portfolios beyond their usual industry constraints.
TL;DR
Sifting through millions of startups to find the next "unicorn" is a needle-in-a-haystack problem for investors. This paper presents an unsupervised recommendation system built on a massive 7.5-million-company Knowledge Graph. By leveraging Indirect Association Rules Mining (IARM), the system identifies hidden investment patterns that human experts might miss, achieving an F1-score of 0.958 in recommending high-potential diversification opportunities.
The Problem: Information Sparsity and Human Limitation
In the world of Venture Capital (VC), the most promising innovators are often the ones with the least amount of public data. Human experts are limited by "bounded rationality"—they can only track specific geographies or technologies.
Current recommendation systems in finance usually fail here because:
- Personalization Bias: They require a rich history of an investor's past moves, which is useless for new funds.
- Data Sparsity: Startups lack the decades of financial reporting seen in public markets.
- The "Silo" Effect: Algorithms often suggest more of the same, preventing the very diversification that protects an investment portfolio.
Methodology: The Power of Indirect Links
The core innovation lies in the move from "Frequent Patterns" (what usually goes together) to Indirect Associations.
1. The Knowledge Graph Foundation
The authors fused five commercial datasets into a unified RDF Knowledge Graph stored in GraphDB. This allowed them to treat investments not just as rows in a table, but as a complex network of relationships between companies, investors, and funding events.

2. Indirect Association Rules Mining (IARM)
Instead of finding companies that the same investor always buys (Direct Association), the algorithm looks for pairs that are rarely in the same portfolio but are both highly dependent on a Mediator Set .
- Logic: If two companies move in similar circles but aren't yet "discovered" together, they represent a latent investment trend or a perfect alternative for diversification.
3. Strategy Generation via CN2
After clustering companies based on features like funding amount and RDF rank, the authors used CN2 rule induction. This generates readable "If-Then" rules that explain why a certain cluster of companies is a good fit for an investor's current profile.
Experiments & Results
The model focused on 322,445 companies active in the last three years. By filtering the massive graph down to recent startups, the researchers ensured the recommendations were relevant to current market trends.
| Metric | Score |
|---|---|
| Precision | 0.959 |
| Recall | 0.958 |
| F1-Measure | 0.958 |
The evaluation results (highlighted in the table below) show that the density-based clustering and JRip classification provided extremely stable performance across different company clusters.

Key Takeaway from Results: The system doesn't just suggest another "Software" company if you invested in one; it might suggest a "Fintech" startup with similar growth metrics and investor backing, facilitating genuine portfolio diversification.
Critical Insight & Future Outlook
This work highlights a shift in FinTech from simple predictive modeling to Knowledge-Based Recommendation. By using an unsupervised approach, the system bypasses the "cold start" problem that plagues personalized recommenders.
Limitations & Next Steps: While the statistical accuracy is high, the authors admit this is a pre-selection tool. A human expert is still required for the "final mile" of due diligence. Furthermore, the model currently struggles with qualitative data; the next frontier involves using Large Language Models (LLMs) to encode company descriptions into these knowledge graphs to capture the "vibe" and mission of a startup, not just its balance sheet.
Conclusion
By mapping the "hidden" associations in global investment data, we can move beyond human networking and local biases. This methodology provides a mathematically rigorous way to discover the market disruptors of 2026 and beyond.
