Scenarios to MSCs: Bridging the Gap Between Natural Language and Formal Models

Scenarios: Identifying Missing Objects and Actions by Means of Computational Linguistics

2007-10-01
Leonid Kof
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a computational linguistics framework designed to transform natural language scenarios into Message Sequence Charts (MSCs). The core innovation is a stack-based algorithm that automatically identifies and reconstructs missing communicating objects and actions often omitted by human authors in industrial requirements documents.

TL;DR

Requirements engineering often struggles with "deletion"—the human tendency to omit obvious details in scenarios. This paper presents a specialized computational linguistics approach that uses a stack-based interaction model and ontology mapping to transform messy natural language scenarios into formal Message Sequence Charts (MSCs), successfully filling in the blanks where authors forgot to mention who is sending what to whom.

The Problem: The "Obvious" is Often Missing

In industrial settings, system behavior is typically described in natural language sequences called scenarios. However, because authors (typically engineers) find certain facts "obvious," they frequently omit critical details. A sentence like "The instrument cluster turns on" lacks a clear sender. This leads to:

  • Imprecise specifications that cause downstream development errors.
  • Incomplete behavioral models that cannot be simulated or verified.
  • Manual extraction fatigue, where analysts spend hours re-verifying implicit logic.

Methodology: The Logic of the Message Stack

The author’s solution hinges on two pillars: Domain Ontologies and a Situation Stack.

1. Linguistic Decomposition

Each sentence is parsed to find the verb, subject, and objects. Using an ontology extracted from the document, the system identifies which nouns are "communicating objects" (e.g., Driver, Car, Instrument Cluster).

2. The Interaction Stack

To solve the "missing object" problem, the system doesn't just look at one sentence at a time—it maintains a stack of active messages.

  • Push: If a message starts a new interaction.
  • Pop: If a message acts as a reply to a previous one.
  • Inference: If a sentence has no receiver, the system looks at the top of the stack. If the current sender was the receiver of the last message, the new receiver is likely the original sender (a reply).

The Logic of Stack Management

Figure: Rules for identifying missing messages and managing the stack state.

Experimental Evidence

The approach was tested on an instrument cluster specification (a car dashboard).

Manual Rephrasing vs. Automatic Extraction

The authors found that while strictly "raw" text (full of passive voice and compound sentences) was hard for the parser, minimal rephrasing (taking about 2 hours for 37 scenarios) allowed the system to generate highly accurate MSCs.

SentenceExtracted SenderExtracted ReceiverStack Action
The driver switches on the car.drivercarpush
The instrument cluster turns on.ins. clust.[Inferred: car]push/pop

Manual Translation Benchmark

Figure: Comparing a manual scenario mental model to the formal MSC output.

Ontology Validation

A fascinating side effect of this method is Ontology Debugging. If the generated MSC looked wrong (e.g., identifying "input speed" as a physical sender), it immediately flagged that the domain ontology was cluttered with non-hardware concepts. Refining this ontology took only 2 hours and significantly improved model quality.

Critical Analysis & Conclusion

The Takeaway

The true value of this work lies in its Heuristic Intuition. By treating a requirements scenario as a "discourse" where objects have "focus," the author moves beyond simple keyword matching and into the realm of semantic behavioral reconstruction.

Limitations

  1. Linguistic Constraints: The current tool struggles with passive voice ("The message is sent") and complex subordinate clauses.
  2. Preprocessing: It still requires a human to "clean up" the grammar before the magic happens.
  3. Implicit Repetition: Loops and timers (e.g., "stays active for 30s") are still difficult to extract without deeper semantic reasoning.

Future Outlook

As we move toward LLM-driven requirements engineering, the stack-based logic proposed here remains highly relevant. It provides a formal framework to verify if an AI's interpretation of a requirement is logically sound and "closed" in terms of message passing.

Final Extracted MSC Output

Find Similar Papers

Try Our Examples

  • Find recent papers that automate the generation of Message Sequence Charts or UML diagrams from natural language requirements using Large Language Models (LLMs).
  • Which paper originally proposed the "Grosz situation stack" or Centering Theory for discourse analysis, and how has it been applied to software engineering beyond this work?
  • Search for studies that evaluate the accuracy of Part-of-Speech (POS) tagging and dependency parsing in the specific context of industrial technical specifications versus general corpora.
Contents
Scenarios to MSCs: Bridging the Gap Between Natural Language and Formal Models
1. TL;DR
2. The Problem: The "Obvious" is Often Missing
3. Methodology: The Logic of the Message Stack
3.1. 1. Linguistic Decomposition
3.2. 2. The Interaction Stack
4. Experimental Evidence
4.1. Manual Rephrasing vs. Automatic Extraction
4.2. Ontology Validation
5. Critical Analysis & Conclusion
5.1. The Takeaway
5.2. Limitations
5.3. Future Outlook