moabb.analysis.plotting.plot_critical_difference#

moabb.analysis.plotting.plot_critical_difference(data, pipelines=None, alpha=0.05, higher_is_better=True, figsize=None)[source]#

Plot average pipeline ranks and Nemenyi critical-difference groups.

Session identity is checked before aggregation: within each dataset-and-subject block, every pipeline must contain the same sessions. Scores are then macro-averaged over sessions per subject and averaged over subjects within each dataset. This prevents a missing session from being silently averaged away while still giving every subject equal weight. Each dataset therefore contributes one comparable score per pipeline and one rank to the cross-dataset comparison.

Parameters:
  • data (pandas.DataFrame) – Output of Results.to_dataframe(). Must contain dataset, pipeline, subject, and score columns.

  • pipelines (list of str | None) – Pipelines to include. If None, all pipelines are included.

  • alpha (float) – Significance level for the Nemenyi post-hoc critical difference.

  • higher_is_better (bool) – If True, larger scores receive better (lower) ranks. If False, smaller scores receive better ranks.

  • figsize (tuple of (float, float) | None) – Figure size. If None, the width and height scale with the number of pipelines.

Returns:

fig – Critical-difference diagram. Horizontal bars connect maximal groups of pipelines whose mean-rank differences do not exceed the Nemenyi critical difference.

Return type:

matplotlib.figure.Figure

Notes

The comparison uses a Friedman test followed by the Nemenyi critical difference for complete blocks (the same pipelines evaluated on every dataset). Incomplete dataset-by-pipeline score matrices are rejected rather than silently changing the set of benchmark datasets per pair. If an evaluation column is present, all rows must belong to the same evaluation protocol; protocol identity is never averaged away. Learning-curve results must also be filtered to a single data_size; permutations at that fixed size are treated as repeated measurements within each session, and every pipeline must contain the same permutation identities per block.

References

[1]

Demšar, J. (2006). Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, 7, 1–30. https://www.jmlr.org/papers/v7/demsar06a.html