moabb.datasets.Ma2022#

class moabb.datasets.Ma2022(subjects=None, sessions=None, *, return_all_modalities=False)[source]#

Bases: BaseBIDSDataset

[source]

Dataset Snapshot

Ma2022

Imagery, 2 classes (left_hand vs right_hand)

AuthorsJun Ma, Banghua Yang, Wenzheng Qiu, Yunzhe Li, Shouwei Gao, Xinxing Xia

πŸ‡¨πŸ‡³β€‚Shanghai University, CNΒ·2022
Imagery Code: Ma-edf2022 25 subjects 5 sessions 32 ch 250 Hz 2 classes 4.0 s trials CC BY 4.0

Class Labels: left_hand, right_hand

Overview

Cross-session motor imagery dataset (SHU) from Ma et al. 2022.

Dataset from

The SHU dataset contains EEG recordings from 25 healthy, BCI-naive subjects (13 males, 12 females, aged 20-24 years) performing cued left- vs right-hand grasping motor imagery. Each subject completed five independent sessions recorded on five different days, 2 to 3 days apart, which makes the dataset specifically suited for studying cross-session variability in motor imagery BCIs.

Each session was designed with 100 trials (50 left-hand, 50 right-hand, randomized order); the data paper's Table 2 reports "90 to 100" trials per session. The released files retain 74 to 100 trials per session after the source-side bad-segment rejection described in the data paper, for 11,988 trials in total. Signals were recorded from 32 EEG channels (10-10 according to the paper, called 10-20 in the source sidecar; unipolar reference on M1, ground on AFz) at 250 Hz. Only the 4 s motor imagery window is stored (1000 samples per trial), so the analysis interval spans the full stored window. The source sidecar describes an 8 s protocol whereas the paper text states 7.5 s; missing preparation and cue segments are not reconstructed.

This loader reads the authors' EDF release, hosted in BIDS form on NEMAR (nm000288). Each EDF contains the retained 4 s windows concatenated in time, not continuous amplifier recordings. BIDS DatasetType: raw describes the deposit layout, not its processing state. Events and bad-channel flags are read from the BIDS sidecars; MNE converts the EDF microvolt calibration to SI volts. Zero-duration EDGE boundary annotations mark the joins between stored windows, so MNE's default filtering does not mix neighboring, discontinuous trials.

Nine sessions contain a channel numerically near zero in the authors' MATLAB release, with blank EDF calibration fields. The deposit assigns a small replacement physical range and marks the channel bad; this is not a recovery of the missing calibration. The untouched EDF is preserved under sourcedata/. Bad channels remain flagged, without interpolation or deletion. Since MOABB's default EEG selection excludes bads, full-dataset analyses must use a common good-channel set (exclude F3, T6 and A2). Explicitly picking a bad channel includes its replacement-calibrated signal. Channel positions are template estimates from standard_1020 (called colin27_1020 in newer MNE versions), not measured locations.

Citation & Impact

Stimulus Protocol
../_images/Ma2022.svg

4s task window per trial Β· 2-class imagery paradigm Β· 1 runs/session across 5 sessions

HED Event Tags
HED tags2/2 events annotated

Source: MOABB BIDS HED annotation mapping.

Agent-action
2
Sensory-event
2
left_hand
Sensory-eventAgent-action
right_hand
Sensory-eventAgent-action

HED tree view

Tree Β· left_hand
β”œβ”€ Sensory-event
β”‚  β”œβ”€ Experimental-stimulus
β”‚  └─ Visual-presentation
└─ Agent-action
   └─ Imagine
      β”œβ”€ Move
      └─ Left
         └─ Hand
Tree Β· right_hand
β”œβ”€ Sensory-event
β”‚  β”œβ”€ Experimental-stimulus
β”‚  └─ Visual-presentation
└─ Agent-action
   └─ Imagine
      β”œβ”€ Move
      └─ Right
         └─ Hand
Channel Summary
Total channels32
EEG32 (Ag/AgCl)
Montagestandard_1020
Sampling250 Hz
ReferenceM1 (unipolar)
Notch / line50 Hz

This diagram is automatically generated from MOABB metadata. Please consult the original publication to confirm the experimental protocol details.

Cross-session motor imagery dataset (SHU) from Ma et al. 2022.

Dataset from [1].

The SHU dataset contains EEG recordings from 25 healthy, BCI-naive subjects (13 males, 12 females, aged 20-24 years) performing cued left- vs right-hand grasping motor imagery. Each subject completed five independent sessions recorded on five different days, 2 to 3 days apart, which makes the dataset specifically suited for studying cross-session variability in motor imagery BCIs.

Each session was designed with 100 trials (50 left-hand, 50 right-hand, randomized order); the data paper’s Table 2 reports β€œ90 to 100” trials per session. The released files retain 74 to 100 trials per session after the source-side bad-segment rejection described in the data paper, for 11,988 trials in total. Signals were recorded from 32 EEG channels (10-10 according to the paper, called 10-20 in the source sidecar; unipolar reference on M1, ground on AFz) at 250 Hz. Only the 4 s motor imagery window is stored (1000 samples per trial), so the analysis interval spans the full stored window. The source sidecar describes an 8 s protocol whereas the paper text states 7.5 s; missing preparation and cue segments are not reconstructed.

Important

The released data is preprocessed: bad segments were removed, the baseline was corrected, and a 0.5-40 Hz FIR band-pass filter was applied by the authors before disclosure.

This loader reads the authors’ EDF release, hosted in BIDS form on NEMAR (nm000288). Each EDF contains the retained 4 s windows concatenated in time, not continuous amplifier recordings. BIDS DatasetType: raw describes the deposit layout, not its processing state. Events and bad-channel flags are read from the BIDS sidecars; MNE converts the EDF microvolt calibration to SI volts. Zero-duration EDGE boundary annotations mark the joins between stored windows, so MNE’s default filtering does not mix neighboring, discontinuous trials.

Nine sessions contain a channel numerically near zero in the authors’ MATLAB release, with blank EDF calibration fields. The deposit assigns a small replacement physical range and marks the channel bad; this is not a recovery of the missing calibration. The untouched EDF is preserved under sourcedata/. Bad channels remain flagged, without interpolation or deletion. Since MOABB’s default EEG selection excludes bads, full-dataset analyses must use a common good-channel set (exclude F3, T6 and A2). Explicitly picking a bad channel includes its replacement-calibrated signal. Channel positions are template estimates from standard_1020 (called colin27_1020 in newer MNE versions), not measured locations.

Note

NEMAR is the only download source for this EDF/BIDS loader, including when the provider is set to upstream. Failures propagate: there is no fallback to the scientifically different MATLAB representation. data_path returns the five EDF paths; sessions are numbered "0" to "4" in MOABB. The code Ma-edf2022 isolates downloads, caches and evaluation results from the former MATLAB loader.

This dataset is from the same laboratory as Yang2025 (WBCIC-SHU, a distinct 2025 multi-day recording) and is unrelated to Ma2020 (different team, different recording).

References

[1]

J. Ma, B. Yang, W. Qiu, Y. Li, S. Gao, and X. Xia, β€œA large EEG dataset for studying cross-session variability in motor imagery brain-computer interface,” Scientific Data, vol. 9, p. 531, 2022. DOI: 10.1038/s41597-022-01647-1

from moabb.datasets import Ma2022
dataset = Ma2022()
data = dataset.get_data(subjects=[1])
print(data[1])

Dataset summary

#Subj

25

#Chan

32

#Classes

2

#Trials / class

239.76

Trials length

4 s

Freq

250 Hz

#Sessions

5

#Runs

1

Total_trials

11988

Participants

  • Population: healthy

  • Age: 22.52 (range: 20-24) years

  • BCI experience: naive

Equipment

  • Amplifier: Wuhan Greentech 32-channel Ag/AgCl cap, Brickcom wireless amplifier

  • Electrodes: Ag/AgCl

  • Montage: standard_1020

  • Reference: M1 (unipolar)

Preprocessing

  • Data state: preprocessed

  • Bandpass filter: 0.5-40 Hz

  • Steps: bad segment rejection (EEGLAB amplitude > 100 uV, manually confirmed), baseline removal, 0.5-40 Hz FIR band-pass filter, epoching to the 4 s motor imagery window

  • Notes: Preprocessing was applied by the data authors before disclosure; the released trials are the 4 s motor imagery segments only.

Data Access

Experimental Protocol

  • Paradigm: imagery

  • Task type: motor_imagery_grasping

  • Feedback: none

  • Stimulus: visual cue

Notes

Added in version 1.8.0.

__init__(subjects=None, sessions=None, *, return_all_modalities=False)[source]#

Initialize function for the BaseDataset.

property all_subjects#

Full list of subjects available in this dataset (unfiltered).

convert_to_bids(path=None, subjects=None, overwrite=False, format='EDF', verbose=None, generate_figures=False)[source]#

Convert the dataset to BIDS format.

Saves the raw EEG data in a BIDS-compliant directory structure. Unlike the caching mechanism (see CacheConfig), the files produced here do not contain a processing-pipeline hash (desc-<hash>) in their names, making the output a clean, shareable BIDS dataset.

Parameters:
  • path (str | Path | None) – Directory under which the BIDS dataset will be written. If None the default MNE data directory is used (same default as the rest of MOABB).

  • subjects (list of int | None) – Subject numbers to convert. If None, all subjects in subject_list are converted.

  • overwrite (bool) – If True, existing BIDS files for a subject are removed before saving. Default is False.

  • format (str) – The file format for the raw EEG data. Supported values are "EDF" (default), "BrainVision", and "EEGLAB".

  • verbose (str | None) – Verbosity level forwarded to MNE/MNE-BIDS.

  • generate_figures (bool) – If True, generate interactive neural signature HTML figures in {bids_root}/derivatives/neural_signatures/. Requires plotly (pip install moabb[interactive]). Default is False.

Returns:

bids_root – Path to the root of the written BIDS dataset.

Return type:

pathlib.Path

Examples

>>> from moabb.datasets import AlexMI
>>> dataset = AlexMI()
>>> bids_root = dataset.convert_to_bids(path="/tmp/bids", subjects=[1])

Notes

Use CacheConfig to configure caching for get_data(). Use moabb.datasets.bids_interface.get_bids_root to get the BIDS root path.

Added in version 1.5.

data_path(subject, path=None, force_update=False, update_path=None, verbose=None)[source]#

Get path to local copy of a subject data.

Parameters:
  • subject (int) – Number of subject to use

  • path (None | str) – Location of where to look for the data storing location. If None, the environment variable or config parameter MNE_DATASETS_(dataset)_PATH is used. If it doesn’t exist, the β€œ~/mne_data” directory is used. If the dataset is not found under the given path, the data will be automatically downloaded to the specified folder.

  • force_update (bool) – Force update of the dataset even if a local copy exists.

  • update_path (bool | None Deprecated) – If True, set the MNE_DATASETS_(dataset)_PATH in mne-python config to the given path. If None, the user is prompted.

  • verbose (bool, str, int, or None) – If not None, override default verbose level (see mne.verbose()).

Returns:

path – Local path to the given data file. This path is contained inside a list of length one, for compatibility.

Return type:

list of str

download(subject_list=None, path=None, force_update=False, update_path=None, accept=False, verbose=None)[source]#

Download BIDS EDFs, not the original sourcedata distribution.

get_additional_metadata(subject, session, run)[source]#

Load additional metadata for a specific subject, session, and run.

Parameters:
  • subject (str) – The identifier for the subject.

  • session (str) – The identifier for the session.

  • run (str) – The identifier for the run.

Returns:

A DataFrame containing the additional metadata if available, otherwise None.

Return type:

None | pandas.DataFrame

get_block_repetition(paradigm, subjects, block_list, repetition_list)[source]#

Select data for all provided subjects, blocks and repetitions.

subject -> session -> run -> block -> repetition

See also

get_data

Parameters:
  • subjects (List of int) – List of subject number

  • block_list (List of int) – List of block number

  • repetition_list (List of int) – List of repetition number inside a block

Returns:

data – dict containing the raw data

Return type:

Dict

get_data(subjects=None, cache_config=None, process_pipeline=None, n_jobs=1)[source]#

Return the data corresponding to a list of subjects.

The returned data is a dictionary with the following structure:

data = {"subject_id": {"session_id": {"run_id": run}}}

subjects are on top, then we have sessions, then runs. A sessions is a recording done in a single day, without removing the EEG cap. A session is constitued of at least one run. A run is a single contiguous recording. Some dataset break session in multiple runs.

Processing steps can optionally be applied to the data using the *_pipeline arguments. These pipelines are applied in the following order: raw_pipeline -> epochs_pipeline -> array_pipeline. If a *_pipeline argument is None, the step will be skipped. Therefore, the array_pipeline may either receive a mne.io.Raw or a mne.Epochs object as input depending on whether epochs_pipeline is None or not.

Parameters:
  • subjects (List of int) – List of subject number

  • cache_config (dict | CacheConfig) – Configuration for caching of datasets. See CacheConfig for details.

  • process_pipeline (sklearn.pipeline.Pipeline | None) – Optional processing pipeline to apply to the data. To generate an adequate pipeline, we recommend using moabb.make_process_pipelines(). This pipeline will receive mne.io.BaseRaw objects. The steps names of this pipeline should be elements of StepType. According to their name, the steps should either return a mne.io.BaseRaw, a mne.Epochs, or a numpy.ndarray. This pipeline must be β€œfixed” because it will not be trained, i.e. no call to fit will be made.

  • n_jobs (int) – Number of jobs to run in parallel over subjects (passed to joblib.Parallel). Default 1 (sequential). Per-subject processing (reading, filtering, resampling, epoching) is independent, so this gives a near-linear speedup for datasets with many subjects.

Returns:

data – dict containing the raw data

Return type:

Dict

property metadata[source]#

Return structured metadata for this dataset.

Returns the DatasetMetadata object from the centralized catalog, or None if metadata is not available for this dataset.

Returns:

The metadata object containing acquisition parameters, participant demographics, experiment details, and documentation. Returns None if no metadata is registered for this dataset.

Return type:

DatasetMetadata | None

Examples

>>> from moabb.datasets import BNCI2014_001
>>> dataset = BNCI2014_001()
>>> dataset.metadata.participants.n_subjects
9
>>> dataset.metadata.acquisition.sampling_rate
250.0
sourcedata_path(subject=None, path=None, force_update=False, verbose=None)[source]#

Get the dataset’s original pre-BIDS distribution from NEMAR.

Where data_path() fetches the original files from the upstream host, this fetches the copy NEMAR mirrors under sourcedata/. The files keep their upstream names, so the two are interchangeable in content – but NEMAR stays reachable when the upstream host is slow, rate-limited, behind a bot gate, or retired.

Parameters:
  • subject (int | str | None) – Restrict the download to one subject, resolved through the deposit’s sourcedata_provenance.json. Deposits enriched before that manifest recorded subjects fall back to the whole tree with a warning. When None the whole tree is fetched.

  • path (None | str) – Base path where MOABB stores datasets.

  • force_update (bool) – Re-fetch even when a local copy is present.

  • verbose (bool, str, int, or None) – If not None, override default verbose level.

Returns:

Local path to the sourcedata directory.

Return type:

str

Raises:
  • ValueError – If the dataset declares no nemar_id.

  • moabb.datasets.download.NemarDownloadError – If the download fails or the deposit publishes no sourcedata/.