
Seville, Spain, 31 August – 5 September 2025,
Springer LNCS, Vol. 16044, 2026.
Keywords: Information Storage and Retrieval, Data Structures and Information Theory, Data Mining and Knowledge Discovery, Algorithm Analysis and Problem Complexity, e-Commerce/e-business, Computer Applications
About this book
This book constitutes the refereed proceedings of the 23rd International Conference on Business Process Management, BPM 2025, which took place in Seville, Spain, in September 2025. The 30 full papers included in this book were carefully reviewed and selected from 132 submissions. They were organized in topical sections on foundation, engineering, and management.
Rome, Italy, 23-27 October 2023,
IEEE Proceedings, 2023.
Keywords: Data Mining and Knowledge Discovery, Business Process Management, Machine Learning, Computer Applications, Health Informatics
About this book
This volume constitutes the revised selected papers of several workshops which were held in conjunction with the 5th International Conference on Process Mining, ICPM 2023, held in Rome, Italy, in October 23–27, 2023.
The 38 revised full papers presented in this book were carefully reviewed and selected from 85 submissions. The book also contains one invited talk.
Bozen-Bolzano, Italy, October 23-28, 2022,
Lecture Notes in Business Information Processing 468, 2023.
Keywords: Data Mining and Knowledge Discovery, Business Process Management, Machine Learning, Computer Applications, Health Informatics
About this book
This open access book constitutes revised selected papers from the International Workshops held at the 4th International Conference on Process Mining, ICPM 2022, which took place in Bozen-Bolzano, Italy, during October 23–28, 2022.
The conference focuses on the area of process mining research and practice, including theory, algorithmic challenges, and applications. The co-located workshops provided a forum for novel research ideas. The 42 papers included in this volume were carefully reviewed and selected from 89 submissions.
Manufacturing & Service Operations Management (MSOM) [FT50, ABDC A*, Q1], 2025.
Keywords: Predictive process monitoring; Inter-case predictions; Knowledge-driven encoding; Data-driven encoding
Abstract
Problem Definition: Much of the focus of queueing theory (QT) is on performance evaluation that supports comparative analytics, i.e., comparing performance measures under different intervention. However, analytical results are very sensitive to model assumptions. We develop a data-driven Structured Causal Model adapted to queueing systems (SCQM) that automatically adapts to the data generating process, finds causal relations, and supports comparative analytics. Numerical experiments show that the accuracy of SCQM is competitive with QT even for examples where analytical queueing solutions are available.
Methodology: We use structured causal modeling methodology to develop a non-queueing simulator without prior knowledge of the system. We employ Machine Learning (ML) models for identifying the parent sets and causal relations. We then provide intervention analysis using Monte Carlo simulation.
Managerial Implications: We develop an accurate self-adapting data-driven performance evaluator for congested systems that requires no prior knowledge of the system dynamics. Using this method, companies can perform comparative analytics of interventions for queueing systems that may not be analytically solvable.
Manufacturing & Service Operations Management (MSOM) [FT50, ABDC A*, Q1], 2025.
Keywords: Predictive process monitoring; Inter-case predictions; Knowledge-driven encoding; Data-driven encoding
Abstract
Problem definition: Motivated by technological advances in real-time data collection about customers location in service systems, we study the effect of partial visibility of customers on waiting time prediction. We consider systems where the predictor observes only a subset of the customers interacting with the system while serving all customers indiscriminately.
Methodology/results: We formulate a novel model of a partially visible queue and analyze the waiting time prediction problem, deriving a closed-form expression for the optimal prediction. This facilitates quantifying the performance loss of arbitrary prediction methods because of partial visibility and their inherent limitations (i.e., bias and variance). We compare the performance of a wide range of commonly used predictive methods and examine how partial visibility along with other system parameters affects their performance. We further extend these numerical analyses to queueing systems that exhibit characteristics that are common in practice and that were studied in the service operations literature.
Managerial implications: Our analysis shows that the phenomenon of invisible customers profoundly impacts the ability to accurately predict waiting times and should, therefore, be considered an important factor in the development of prediction tools. Such tools cannot be effectively deployed if technological barriers or operational limitations prevent a sufficiently high level of data integrity. This work provides specific insights into the effectiveness of various commonly used prediction methods, some of which are shown to be highly sensitive to partial visibility and other queueing systems characteristics. Our findings suggest that machine learning methods that use carefully chosen features offer the most effective generic solution for waiting time prediction in the presence of invisible customers and explain the mechanisms through which partial visibility deteriorates the performance of prediction methods.
Information Systems [Q1], 127: 102447, 2025.
Abstract
Business Process Simulation (BPS) is an approach to analyze the performance of business processes under different scenarios. For example, BPS allows us to estimate the impact of adding one or more resources on the cycle time of a process. The starting point of BPS is a process model annotated with simulation parameters (a BPS model). BPS models may be manually designed, based on information collected from stakeholders and from empirical observations, or automatically discovered from historical execution data. Regardless of its provenance, a key question when using a BPS model is how to assess its quality. In particular, in a setting where we are able to produce multiple alternative BPS models of the same process, this question becomes: How to determine which model is better, to what extent, and in what respect? In this context, this article studies the question of how to measure the quality of a BPS model with respect to its ability to accurately replicate the observed behavior of a process. Rather than pursuing a one-size-fits-all approach, the article recognizes that a process covers multiple perspectives. Accordingly, the article outlines a framework that can be instantiated in different ways to yield quality measures that tackle different process perspectives. The article defines a number of concrete quality measures and evaluates these measures with respect to their ability to discern the impact of controlled perturbations on a BPS model, and their ability to uncover the relative strengths and weaknesses of two approaches for automated discovery of BPS models. The evaluation shows that the proposed measures not only capture how close a BPS model is to the observed behavior, but they also help us to identify the sources of discrepancies.
INFORMS Journal on Computing [Q1], 36(3): 766–786, 2024.
Abstract
We apply supervised learning to a general problem in queueing theory: using a neural net, we develop a fast and accurate predictor of the stationary system-length distribution of a GI/GI/1 queue—a fundamental queueing model for which no analytical solutions are available. To this end, we must overcome three main challenges: (i) generating a large library of training instances that cover a wide range of arbitrary interarrival and service time distributions, (ii) labeling the training instances, and (iii) providing continuous arrival and service distributions as inputs to the neural net. To overcome (i), we develop an algorithm to sample phase-type interarrival and service time distributions with complex transition structures. We demonstrate that our distribution-generating algorithm indeed covers a wide range of possible positive-valued distributions. For (ii), we label our training instances via quasi-birth-and-death(QBD) that was used to approximate PH/PH/1 (with phase-type arrival and service process) as labels for the training data. For (iii), we find that using only the first five moments of both the interarrival and service times distribution as inputs is sufficient to train the neural net. Our empirical results show that our neural model can estimate the stationary behavior of the GI/GI/1—far exceeding other available methods in terms of both accuracy and runtimes.
Information Systems [Q1], 53, 2015, 278–295.
Abstract
Information systems have been widely adopted to support service processes in various domains, e.g., in the telecommunication, finance, and health sectors. Information recorded by systems during the operation of these processes provides an angle for operational process analysis, commonly referred to as process mining. In this work, we establish a queueing perspective in process mining to address the online delay prediction problem, which refers to the time that the execution of an activity for a running instance of a service process is delayed due to queueing effects. We present predictors that treat queues as first-class citizens and either enhance existing regression-based techniques for process mining or are directly grounded in queueing theory. In particular, our predictors target multi-class service processes, in which requests are classified by a type that influences their processing. Further, we introduce queue mining techniques that derive the predictors from event logs recorded by an information system during process execution. Our evaluation based on large real-world datasets, from the telecommunications and financial sectors, shows that our techniques yield accurate online predictions of case delay and drastically improve over predictors neglecting the queueing perspective.

International Conference on Process Mining (ICPM), 2025. Best paper award
Abstract
With recent technological advances, process logs, which were traditionally deterministic in nature, are being captured from non-deterministic sources, such as uncertain sensors or machine learning models (that predict activities using cameras). In the presence of stochastically-known logs, logs that contain probabilistic information, the need for stochastic trace recovery increases, to offer reliable means of understanding the processes that govern such systems. We design a novel deep learning approach for stochastic trace recovery, based on Diffusion Denoising Probabilistic Models (DDPM), which makes use of process knowledge (either implicitly by discovering a model or explicitly by injecting process knowledge in the training phase) to recover traces by denoising. We conduct an empirical evaluation demonstrating state-of-the-art performance with up to a 25% improvement over existing methods, along with increased robustness under high noise levels.
International Joint Conference on Artificial Intelligence (IJCAI), IJCAI2025 (CORE A*)
Abstract
Proactive scheduling creates robust offline schedules that optimize resource utilization and minimize job flow times. This work addresses scheduling challenges in business processes, often encountered in service systems, which differ from traditional applications like manufacturing due to inherent uncertainties in activity durations, and human resource availability. We model the business process scheduling problem (BPSP) as a variation of stochastic resource-constrained multi-project scheduling (RCMPSP), and apply process mining to infer unknown parameter values from historical event data. To overcome the randomness in activity durations, we transform the problem into its deterministic counterpart, and prove that the latter provides a lower bound on the Makespan of the stochastic problem. Our approach integrates data-driven Monte Carlo simulation with constraint programming to generate proactive schedules that guarantee, with high probability, that the Makespan remains below a predefined threshold. We evaluate our approach using synthetic datasets with varying levels of uncertainty and size. In addition, we apply the approach to a real-world dataset from an outpatient cancer hospital, demonstrating its effectiveness in optimizing the process Makespan by an average of 5% to 14%.
AAAI Conference on Artificial Intelligence, AAAI2025 (CORE A*)
Abstract
Scheduling is adopted in various domains to assign jobs to resources, such that an objective is optimized. While schedules enable the analysis of the underlying system, publishing them also incurs a privacy risk. Recently, privacy attacks on schedules have been proposed, which may reveal sensitive information on the jobs by solving an inverse scheduling problem. In this work, we study the protection against such attacks. We formulate the problem of privacy-and-utility preservation of schedules, which bounds both, the privacy leakage and the loss in the utility of the schedule due to obfuscation. We address the problem based on a set of perturbation functions for schedules, study their instantiations for standard scheduling problems, and implement privacy-and-utility-aware publishing of a schedule using constraint programming. Experiments with synthetic and real-world schedules demonstrate the feasibility, robustness, and effectiveness of our mechanism.
AAAI Conference on Artificial Intelligence, AAAI2023 (CORE A*)
Abstract
Schedules define how resources process jobs in diverse domains, reaching from healthcare to transportation, and, therefore, denote a valuable starting point for analysis of the underlying system. However, publishing a schedule may disclose private information on the considered jobs. In this paper, we provide a first threat model for published schedules, thereby defining a completely new class of data privacy problems. We then propose distance-based measures to assess the privacy loss incurred by a published schedule, and show their theoretical properties for an uninformed adversary, which can be used as a benchmark for informed attacks. We show how an informed attack on a published schedule can be phrased as an inverse scheduling problem. We instantiate this idea by formulating the inverse of a well-studied single-machine scheduling problem, namely minimizing the total weighted completion times. An empirical evaluation for synthetic scheduling problems shows the effectiveness of informed privacy attacks and compares the results to theoretical bounds on uninformed attacks.
ACM SIGOPS Symposium on Operating Systems Principles (SOSP), 2021 (CORE A*)
Abstract
Decision-making in large-scale compute clouds relies on accurate workload modeling. Unfortunately, prior models have proven insufficient in capturing the complex correlations in real cloud workloads. We introduce the first model of large-scale cloud workloads that captures long-range inter-job correlations in arrival rates, resource requirements, and lifetimes. Our approach models workload as a three-stage generative process, with separate models for: (1) the number of batch arrivals over time, (2) the sequence of requested resources, and (3) the sequence of lifetimes. Our lifetime model is a novel extension of recent work in neural survival prediction. It represents and exploits inter-job correlations using a recurrent neural network. We validate our approach by showing it is able to accurately generate the production virtual machine workload of two real-world cloud providers.
International Conference on Automated Planning and Scheduling (ICAPS), ICAPS2019 (CORE A*)
Abstract
A significant challenge in declarative approaches to scheduling is the creation of a model: the set of resources and their capacities and the types of activities and their temporal and resource requirements. In practice, such models are developed manually by skilled consultants and used repeatedly to solve different problem instances. For example, in a factory, the model may be used each day to schedule the current customer orders. In this work, we aim to automate the creation of such models by learning them from event data. We introduce a novel methodology that combines process mining, timed Petri nets (TPNs), and constraint programming (CP). The approach learns a sub-class of TPN from event logs of executions of past schedules and maps the TPN to a broad class of scheduling problems. We show how any problem of the scheduling class can be converted to a CP model. With new instance data (e.g., the day’s orders), the CP model can then be solved by an off-the-shelf solver. Our approach provides an end-to-end solution, going from event logs to model-based optimal schedules. To demonstrate the value of the methodology we conduct experiments in which we learn and solve scheduling models from two types of data: logs generated from job-shop scheduling benchmarks and real-world event logs from an outpatient hospital.
AAAI Conference on Artificial Intelligence, AAAI2019 (16% acceptance rate) (CORE A*)
Abstract
Time prediction is an essential component of decision making in various Artificial Intelligence application areas, including transportation systems, healthcare, and manufacturing. Predictions are required for efficient resource allocation and scheduling, optimized routing, and temporal action planning. In this work, we focus on time prediction in congested systems, where entities share scarce resources. To achieve accurate and explainable time prediction in this setting, features describing system congestion (e.g., workload and resource availability), must be considered. These features are typically gathered using process knowledge, (i.e., insights on the interplay of a system’s entities). Such knowledge is expensive to gather and may be completely unavailable. In order to automatically extract such features from data without prior process knowledge, we propose the model of congestion graphs, which are grounded in queueing theory. We show how congestion graphs are mined from raw event data using queueing theory based assumptions on the information contained in these logs. We evaluate our approach on two real-world datasets from healthcare systems where scarce resources prevail: an emergency department and an outpatient cancer clinic. Our experimental results show that using automatic generation of congestion features, we get an up to 23% improvement in terms of relative error in time prediction, compared to common baseline methods. We also detail how congestion graphs can be used to explain delays in the system.
Business Process Management (BPM), BPM2017 (13% acceptance rate; Best Paper Award) (CORE A)
Abstract
This book constitutes the proceedings of the 15th International Conference on Business Process Management, BPM 2017, held in Barcelona, Spain, in September 2017.The 19 revised full papers papers presented were carefully reviewed and selected from 116 initial submissions. The topics selected by the authors demonstrate an increasing interest of the research community in the area of process mining, resonated by an equally fast-growing uptake by different industry sectors. The papers are organized in topical sections on process modeling; process mining; assorted BPM topics; decisions and understanding; and process knowledge.