Norms2Onto: Bridging the Gap in Financial Knowledge Modeling via Machine Learning

A New Automatic Ontology Construction Method Based on Machine Learning Techniques: Application on financial corpus

2021-11-01
Amani Drissi, Ahmed Khemiri, Salma Sassi, Richard Chbeir, R. Chbeir
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Norms2Onto, a semi-automatic ontology learning framework designed to model complex financial domains using IFRS and IAS standards. By combining NLP preprocessing with supervised machine learning, it achieves a high-fidelity representation of financial knowledge, significantly outperforming manual expert classification in scale and concept coverage.

TL;DR

The complexity of International Financial Reporting Standards (IFRS) often creates a "knowledge silo" difficult for even experts to navigate. This paper presents Norms2Onto, a pipeline that automates the transition from messy financial text to structured, interactive ontologies. Using Random Forest and NLP, it extracts 2,000 relevant concepts where human experts previously managed only 50, boasting a 0.93 precision rate.

Problem & Motivation: The Complexity of Global Finance

Since 2005, IFRS has been the global language of business. However, the sheer volume of these standards makes manual modeling—creating "ontologies" to shared meaning—a bottleneck.

The authors identify three critical challenges:

  1. Extraction: How to filter relevant financial terms from dense text?
  2. Enrichment: How to automatically assign precise definitions to these terms?
  3. Linkage: How to visualize the intricate web of relationships between diverse standards like IFRS 7 and IAS 32?

Previous works were either entirely manual or lacked a systematic approach to term extraction and relationship identification.

Methodology: The Norms2Onto Pipeline

The architecture is divided into three functional modules:

1. Preprocessing (The Cleaner)

Text is tokenized and tagged using Part-of-Speech (POS) and NP Chunking. This is vital because financial terms are often compound nouns (e.g., "Financial Instrument" vs "Financial").

2. Learning (The Brain)

The system identifies candidate concepts using TF-IDF. It then employs four supervised algorithms to predict the category and subcategory of each term:

  • Linear SVC
  • Logistic Regression
  • Random Forest (The best performer)
  • Multinomial Naive Bayes

3. Visualization (The Interface)

The output isn't just a list; it’s an interactive graph where concepts (blue), standards (red), and subcategories (green) are interconnected.

Overall Architecture of Norms2Onto

Experiments & Performance

The researchers tested the models on five core standards (IFRS 7, 13, IAS 1, 7, 32).

SOTA Comparison: ML vs. Human

The results were striking. When compared to a manual classification by an accounting expert, Norms2Onto demonstrated a massive leap in scale:

  • Manual: 50 concepts.
  • Norms2Onto: 2,000 concepts (with 1,900 validated as relevant by the IFRS Glossary).

Algorithmic Efficiency

The Random Forest algorithm proved to be the most robust for this domain-specific task, outperforming simpler models like Naive Bayes by a significant margin.

Performance Comparison Table

Critical Insight & Conclusion

The true value of Norms2Onto lies in its hybrid resource approach. By utilizing both the unstructured text of the norms and structured knowledge from external financial dictionaries (Investopedia and IFRS Glossary), the system mimics an expert's external lookup process.

Takeaway: While semi-automatic, this method proves that machine learning can drastically reduce the "knowledge engineering" tax in fintech.

Future Work: The authors suggest expanding the corpus to the entire IFRS/IAS catalog. However, a limitation to note is the current reliance on expert validation for the final "relevant concept" transition—future iterations might look toward LLMs (Large Language Models) to further automate this validation step.

Find Similar Papers

Try Our Examples

  • Search for recent studies that use Deep Learning or Transformer-based models for automatic financial ontology construction beyond classical machine learning.
  • Which paper originally proposed the six layers of Ontology Learning (Terms to Rules), and how does Norms2Onto specifically innovate on the 'Concept Formation' layer compared to that seminal work?
  • Explore how these automated financial ontologies are currently being integrated into Knowledge Graph Retrieval-Augmented Generation (KG-RAG) systems for financial auditing.
Contents
Norms2Onto: Bridging the Gap in Financial Knowledge Modeling via Machine Learning
1. TL;DR
2. Problem & Motivation: The Complexity of Global Finance
3. Methodology: The Norms2Onto Pipeline
3.1. 1. Preprocessing (The Cleaner)
3.2. 2. Learning (The Brain)
3.3. 3. Visualization (The Interface)
4. Experiments & Performance
4.1. SOTA Comparison: ML vs. Human
4.2. Algorithmic Efficiency
5. Critical Insight & Conclusion