SimoOn: Bridging Semantic Gaps in Software Benchmarking via Ontology-Based Similarity
Ontology-Based Similarity Measurement in Software Projects through SimReq Framework
This paper introduces an ontology-based expansion of the SimReq framework to perform similarity measurements across software projects. By leveraging the ISBSG (International Software Benchmarking Standards Group) repository, the authors enable multi-attribute matching (e.g., Functional Size, Work Effort) to support software reuse and project planning.
TL;DR
Determining how "similar" two software projects are is a cornerstone of effective project management, yet traditional keyword searches often miss the mark. This paper introduces an ontology-driven approach using the SimReq framework and a new application called SimoOn. By testing against the massive ISBSG repository, the authors prove that conceptual similarity can unlock hidden insights for software reuse and planning.
The "Keyword" Trap in Software Engineering
Project managers often ask: "Have we done something like this before?" To answer this, they usually query databases for labels or specific values. However, software projects are multi-dimensional. A project might be similar in Functional Size but vastly different in Data Quality.
Prior work often failed because:
- Lack of Context: Traditional IR systems don't understand that "Financial System" and "Accounting App" might share structural similarities.
- Rigid Queries: Users are often forced into binary "exact matches" rather than "degree of similarity."
The authors' insight is to use Ontology—a formal specification of shared conceptualization—to provide a structured backbone for these comparisons.
Methodology: The SimReq Framework & SimoOn
The core of the study is the transition of the SimReq framework from purely textual requirements to a more robust, multi-attribute system.
1. The SimoOn Architecture
The authors developed SimoOn, a prototype that allows for a three-step similarity discovery process:
- Selection: Dynamic choosing of attributes (e.g., Work Effort, Project Duration).
- Weighting/Specification: Entering specific target values.
- Thresholding: Defining the "similarity percentage" (e.g., "Find projects 90% similar to X").
2. The Mathematical Foundation
The framework references several similarity models, including the Tversky model, which moves beyond simple geometric distance to account for common and distinctive features:
This allows the system to prioritize what the user deems important—whether it's the overlap of features or the absence of specific risks.
Figure 1: The modified SimReq Framework focusing on attribute ontology.
Experiments: Sorting through 5,000+ Projects
To validate the approach, the authors used the ISBSG Data Repository (Release 11). This is a gold standard in the industry, containing data from over 5,000 projects worldwide across 100+ attributes.
Query Scenario
The researchers set a rigorous test case seeking projects with:
- Function Size of 1,000 (95% similarity).
- Summary Work Effort of 5,000 hours (90% similarity).
- Data Quality Rating of 'B'.
The Result
Out of 4,744 valid projects, the multi-stage similarity filter narrowed the pool down to just 2 final projects. These results (shown below) provided the "blueprint" for how a manager might proceed with a new project of similar scale.
Figure 2: The execution of the SimoOn application through similarity thresholds.
| Functional Size | Work Effort | Data Quality | Application Type | Development Tech |
|---|---|---|---|---|
| 979 | 4997 | B | Financial / Accounting | Waterfall |
| 990 | 4618 | A | Transaction System | Data Modelling / JAD |
| 1049 | 4768 | B | Management Info System | Process Modelling / JAD |
| Table 1: Highly similar candidate projects identified by the framework. |
Critical Insight & Conclusion
The real value of this ontology-based approach isn't just "finding a match"—it's predictive guidance. By finding highly similar past projects, a manager can decide:
- Should we use Waterfall or Agile?
- Is New Development or Enhancement more efficient for this size?
Limitations: The current prototype is limited to a few specific attributes and numeric data. Future Outlook: The authors plan to integrate more complex data types like full-length project descriptions and multimedia, potentially moving toward a "Full-Spectrum Similarity" engine for software engineering.
