PlanMiner: Bridging Machine Learning and Symbolic Planning through Expression Discovery

Discovering relational and numerical expressions from plan traces for learning action models

2021-03-22
José Á. Segura-Muros, Raúl Pérez, Juan Fernández-Olivares
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces PlanMiner, a novel domain learning system that automatically induces PDDL action models from partially observed plan traces. It uniquely combines symbolic regression with inductive rule learning to discover both relational and numerical expressions, achieving F-Scores above 0.85 across various IPC benchmarks.

TL;DR

Knowledge engineering for automated planning is notoriously difficult and error-prone. PlanMiner is a sophisticated domain learner that bridges the gap between raw data (plan traces) and formal PDDL models. Unlike its predecessors, it doesn't just learn logic; it discovers the arithmetic and relational laws (like fuel consumption or distance constraints) that govern a world, even when the data is riddled with missing values.

Background & Positioning

In the world of Automated Planning (AP), the "Domain Model" is the engine. If the model is wrong, the plan is useless. While prior work like ARMS and LAMP successfully learned STRIPS (logic-only) models, they hit a wall with Numerical Fluents. PlanMiner positions itself as a SOTA solution that treats domain learning as a hybrid problem: a mix of Symbolic Regression (to find formulas) and Classification (to find logic).

The "Why": Why Numeric Discovery is Hard

Most machine learning algorithms are "black boxes." If you want to know why a rover can't move, a neural network might give you a probability, but PDDL requires a hard constraint (e.g., (>= (energy ?r) (dist ?w1 ?w2))). The authors identified two main hurdles:

  1. Data Incompleteness: We rarely know the state of every object at every second.
  2. Expression Expressivity: How do you "guess" an arithmetic formula from a sequence of numbers without trying every infinite combination?

Methodology: The PlanMiner Pipeline

The core innovation lies in a multi-stage translation process that turns noisy traces into clean PDDL code.

1. Dataset Extraction & Schematization

Raw grounded actions (e.g., goto rover1 waypoint A) are generalized into schema forms (goto ?arg1 ?arg2). The system then creates an attribute-value matrix where rows are intermediate states and columns are predicates.

2. Symbolic Regression via Heuristic Search

This is the "secret sauce." To find how a variable like energy changes, PlanMiner uses an informed graph search. It starts with an empty expression and adds operators (+, -, *, /) and operands (constants or attributes), guided by a Mean Absolute Percentage Error (MAPE) heuristic to prune the search space.

Model Architecture and Symbolic Search Logic Figure 1: High-level overview of the PlanMiner architecture, showing the transition from plan traces to PDDL.

3. Inductive Rule Learning (NSLV)

Once the formulas are discovered, they are added back into the dataset as new features. The NSLV algorithm then searches for the most descriptive "white-box" rules that separate a "pre-state" (before an action) from a "post-state" (after an action).

Experimental Insights

The authors tested PlanMiner on benchmarks from the International Planning Competition (IPC), including the famous Rovers and ZenoTravel domains.

Key Metrics:

  • Resilience: At 50% incompleteness, PlanMiner still maintains an F-Score above 0.85 in most domains.
  • Validity: In domains like Depots and Satellite, the learned models were 100% valid (able to generate successful plans) even with missing data.

Experimental Results Comparison Figure 2: Performance comparison across different domains. Notice the stability of PlanMiner (indicated in the charts) versus reference algorithms as incompleteness increases.

Critical Analysis & Discussion

While PlanMiner is a significant step forward, it isn't perfect. The authors candidly point out "Over-information errors." Because the system is so good at finding relations, it sometimes finds "spurious" ones—true but redundant statements (e.g., if A is next to B, it also learns B is next to A, leading to cluttered models).

Furthermore, the system relies on a closed-world assumption for logic but handles missing values using an open-world approach, which is a clever but complex balancing act.

Takeaway and Future Outlook

PlanMiner proves that we don't need a perfectly labeled dataset to learn the fundamental laws of a planning environment. By combining the interpretability of Symbolic Regression with the robustness of Inductive Logic Programming, we move closer to autonomous systems that can "observe" a human performing a task and instantly generate a formal, machine-readable manual for that task.

Future Work: The authors aim to tackle durative actions (actions that take time) and noise (erroneous data), which would bring this tech even closer to real-world robotic applications.

Find Similar Papers

Try Our Examples

  • Find recent papers published after 2021 that extend PDDL domain learning to include PDDL+ features like continuous time and durative actions.
  • Which paper first proposed the NSLV (Non-Symmetric Learning algorithm with Vision) algorithm, and how does PlanMiner modify its rule extraction for planning contexts?
  • Search for studies that utilize Deep Reinforcement Learning or Transformer-based models to discover symbolic action models from state transition traces.
Contents
PlanMiner: Bridging Machine Learning and Symbolic Planning through Expression Discovery
1. TL;DR
2. Background & Positioning
3. The "Why": Why Numeric Discovery is Hard
4. Methodology: The PlanMiner Pipeline
4.1. 1. Dataset Extraction & Schematization
4.2. 2. Symbolic Regression via Heuristic Search
4.3. 3. Inductive Rule Learning (NSLV)
5. Experimental Insights
5.1. Key Metrics:
6. Critical Analysis & Discussion
7. Takeaway and Future Outlook