simOutrank clusters the traces of an event log by an
outranking model of similarity. Instead of collapsing every
perspective into a single distance, it scores each pair of traces on
several problem-specific criteria – activities, transitions, duration,
case attributes – and aggregates them with ELECTRE-III-style concordance
and discordance. The result is a credibility matrix S that
a spectral or hierarchical algorithm then partitions.
This vignette walks through the whole pipeline on the bundled
illustrative_log, a 25-case customer-service log from
Delias et al. (2023).
library(simOutrank)
head(illustrative_log)
#> case_id activity timestamp status satisfaction
#> 1 1 B 2021-01-01 00:01:00 GOLD High
#> 2 1 E 2021-01-01 00:02:00 GOLD High
#> 3 2 B 2021-01-01 00:01:00 GOLD High
#> 4 2 E 2021-01-01 00:02:00 GOLD High
#> 5 3 B 2021-01-01 00:01:00 GOLD High
#> 6 3 E 2021-01-01 00:02:00 GOLD Highas_traces() turns an event log into a
traces object: the time-ordered, integer-encoded activity
sequence of every case, plus the case-level attribute table.
traces <- as_traces(illustrative_log,
case_id = "case_id", activity = "activity",
timestamp = "timestamp")
traces
#> <traces>: 25 cases, 5 distinct activities
#> trace length: min 2, median 5, max 7
#> case attributes: status, satisfaction
summary(traces)[1:5, ]
#> case_id n_events n_distinct
#> 1 1 2 2
#> 2 2 2 2
#> 3 3 2 2
#> 4 4 2 2
#> 5 5 2 2Criteria are built with the crit_* helpers. Each carries
a direction (similarity or dissimilarity), a weight, and ELECTRE-III
thresholds (indifference, similarity, and an optional veto). Here we
mirror the four criteria of the source paper: what activities occur, in
what order, and two case attributes.
criteria <- list(
crit_activity_profile(weight = 0.2, indifference = 0.7, similarity = 0.8,
veto = 0.4),
crit_edit_distance(weight = 0.2, similarity = 2, indifference = 3, veto = 6),
crit_nominal("status", weight = 0.3),
crit_nominal("satisfaction", weight = 0.3)
)
criteria[[1]]
#> <criterion 'activity_profile'>: direction = similarity, weight = 0.2
#> indifference = 0.7, similarity = 0.8, veto = 0.4Thresholds may be fixed numbers or quantiles of the criterion’s own
value distribution via as_quantile(); the parameter-tuning vignette explores that
choice.
outrank_similarity() resolves any quantile thresholds,
normalises the weights, and aggregates the per-criterion indices into
S.
eigengap() plots the smallest eigenvalues of the
normalized Laplacian; a gap after the k-th value suggests k groups.
The paper targets four natural groups (tier x satisfaction). Spectral clustering uses a seed for reproducible k-means.
clust <- cluster_traces(sim, k = 4, seed = 42)
clust
#> <outrank_clust>: 25 cases in 4 clusters (spectral)
#> cluster sizes: 1=6, 2=8, 3=5, 4=6
split(names(clust$memberships), clust$memberships)
#> $`1`
#> [1] "11" "12" "13" "14" "15" "24"
#>
#> $`2`
#> [1] "16" "17" "18" "19" "20" "21" "22" "23"
#>
#> $`3`
#> [1] "6" "7" "8" "9" "10"
#>
#> $`4`
#> [1] "1" "2" "3" "4" "5" "25"Join memberships back onto the event log, inspect a cluster’s attribute profile, and compute validity indices.
augmented <- augment_log(clust, illustrative_log)
#> Joining on case-id column 'case_id'.
head(augmented)
#> case_id activity timestamp status satisfaction .cluster
#> 1 1 B 2021-01-01 00:01:00 GOLD High 4
#> 2 1 E 2021-01-01 00:02:00 GOLD High 4
#> 3 2 B 2021-01-01 00:01:00 GOLD High 4
#> 4 2 E 2021-01-01 00:02:00 GOLD High 4
#> 5 3 B 2021-01-01 00:01:00 GOLD High 4
#> 6 3 E 2021-01-01 00:02:00 GOLD High 4
plot(clust, type = "profile", attribute = "status")Delias, P., Doumpos, M., Grigoroudis, E. and Matsatsinis, N. (2023). Improving the non-compensatory trace-clustering decision process. International Transactions in Operational Research, 30(3), 1387–1406. doi:10.1111/itor.13062 ```