EgoBrain-Mini / README.md
lin-nie's picture
README: Known data issues as a table
3e90086 verified
|
Raw History Blame Contribute Delete
23.4 kB
metadata
license: cc-by-nc-4.0
pretty_name: EgoBrain-Mini
tags:
  - eeg
  - egocentric-video
  - brain-computer-interface
  - multimodal
  - imu
  - motion-capture
  - audio
  - neuroscience
  - video
  - timeseries
language:
  - en
size_categories:
  - n<1K
configs:
  - config_name: default
    data_files:
      - split: train
        path: browse/segments.parquet
extra_gated_fields:
  First Name: text
  Last Name: text
  Email: text
  Affiliation: text
  Position (PhD/Master/Faculty/Industry): text
  Country: country
  Personal or Lab Website (optional): text
  How did you learn about the EgoBrain-Mini project?: text
  What do you plan to use EgoBrain-Mini for?: text
  Why are you interested in our EgoBrain-Mini Project?: text
  I confirm this dataset will only be used for research purposes: checkbox
  I agree not to redistribute the dataset: checkbox
  I agree not to attempt re-identification of any participant: checkbox
  I agree to properly cite the EgoBrain paper and dataset: checkbox

[ICLR 2026]
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding

EgoBrain-Mini — lightweight, fast, and fully synchronized

Project Page Paper PDF Demo

東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia

Nie Lin · Yansen Wang · Dongqi Han · Weibang Jiang · Jingyuan Li · Ryosuke Furuta · Yoichi Sato* · Dongsheng Li* · *(Co-corresponding authors)*

[Teaser Figure]

arXiv License CC BY-NC-4.0

This is EgoBrain-Mini, a lightweight companion release of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action Understanding" — the egocentric subset (first-person video + EEG + IMU + audio), meant as a quick, no-request-needed way to explore the data and its alignment before diving into the full EgoBrain release.

Nie (Elon) Lin, Yansen Wang, Dongqi Han, Weibang Jiang, Jingyuan Li, Ryosuke Furuta, Yoichi Sato*, Dongsheng Li* (*co-corresponding authors). "EgoBrain: Synergizing Minds and Eyes For Human Action Understanding", ICLR 2026.

🗺️ Contents

📢 News

  • [2026-10-04]: Released Subject P0023 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-10-04]: Released Subject P0022 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-10-04]: Updated P0008: 15 of its 39 sub-actions now carry EEG quality flags (a loose electrode, twice) — see Known data issues.
  • [2026-10-04]: Released Subject P0021 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-10-04]: Released Subject P0020 egocentric video + EEG + IMU + audio (Phase 2). ⚠️ In 32 of its 39 sub-actions an EEG reference electrode had lost contact — see Known data issues.
  • [2026-10-03]: Released Subject P0019 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-10-03]: Released Subject P0018 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-27]: Released Subject P0017 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-21]: Released Subject P0016 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-20]: Released Subject P0015 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-16]: Released Subject P0014 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-14]: Released Subject P0013 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-12]: Released Subject P0012 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-09]: Released Subject P0011 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-07]: Released Subject P0010 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-06]: Released Subject P0009 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-06]: Released Subject P0008 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-04]: Released Subject P0007 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-03]: Released Subject P0006 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-02]: Released Subject P0005 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-09-02]: Released Subject P0004 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-08-20]: Released Subject P0003 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-08-19]: Released Subject P0002 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-08-19]: Released Subject P0001 egocentric video + EEG + IMU + audio (Phase 2).
  • [2026-08-18]: Created the Hugging Face repository for ut-vision/EgoBrain-Mini.

📦 What's actually in the box

23 subjects (P0001–P0023), 897 real-world sub-actions (39 per subject), four modalities — first-person video, EEG, a body-worn IMU, and audio — all recorded from the same person doing the same thing at the same instant, and individually verified, not sampled.

EEG cap + first-person view of the desk setup

Modality File Format Shape Rate / Size Notes
🧠 EEG (raw) eeg/raw.npy float32 array, µV [32, T] 32 channels @ 256 Hz Channel-selected only, otherwise unprocessed
🧠 EEG (clean) eeg/clean.npy float32 array, standardized [32, T] 32 channels @ 200 Hz Detrended, 0.1–75 Hz bandpass + 50 Hz notch, RANSAC bad-channel interpolation, average reference, exponential-moving standardization — see Processing
🖐️ IMU imu/imu.parquet Parquet, long format [T, 20] per device 4 body-worn sensors, rate varies by subject/session — see meta.json's imu_hz, don't assume a fixed number Left/right wrist + left/right upper arm — see IMU columns
🎙️ Audio audio/ego.wav WAV, mono PCM16 [T] 16 kHz The Ego camera's own microphone
🎥 Video video/ego.mp4 H.264 / MP4 — 854×480 @ ~30 fps First-person; a copy of its own mic audio is embedded too, but audio/ego.wav is the higher-fidelity source
📄 Provenance video/ego.json JSON — — Exact raw source chapter + frame range this clip was cut from (for traceability; raw video itself is not distributed)

T is the total sample count for that one sub-action's clip — duration_sec × sample_rate — not a separate dimension on top of the rate. For example, the 76.2 s setup-calibration sub-action has eeg/raw.npy.shape == (32, 19508) (76.2 × 256 Hz) and eeg/clean.npy.shape == (32, 15240) (76.2 × 200 Hz exactly). Since duration varies per sub-action (from ~1.5 s to ~17 min across the 897 of them), T is different for every file — it's never a fixed constant you can hard-code.

The array itself carries no metadata — it's just numbers. To go the other way (shape → duration), either read meta.json's duration_sec directly (the authoritative source, already used in Quick start), or divide back out: duration_sec ≈ T / sample_rate (e.g. 15240 / 200 = 76.2), since the sample rate per file is fixed and documented above.

Every one of the 897 sub-actions has all six of these. Nothing here is a placeholder or a partial modality.

📂 Dataset Structure

EgoBrain-Mini/
├── assets/
├── P0001/
│   ├── manifest.json                      ← index of all 39 segments + metadata
│   ├── segments/
│   │   └── phase2/
│   │       ├── P0001_phase2_000_setup-calibration/
│   │       │   ├── meta.json               ← action label, timestamps, duration, modality flags,
│   │       │   │                             and eeg_quality_flags where EEG quality is compromised
│   │       │   ├── eeg/
│   │       │   │   ├── raw.npy
│   │       │   │   └── clean.npy
│   │       │   ├── imu/
│   │       │   │   └── imu.parquet
│   │       │   ├── audio/
│   │       │   │   └── ego.wav
│   │       │   └── video/
│   │       │       ├── ego.mp4
│   │       │       ├── ego.json
│   │       │       └── redaction_windows.json  ← only present where needed, see Privacy below
│   │       └── ...                             (39 segments total, same layout)
│   └── reports/                             ← bundled, offline-first HTML report, see below
├── P0002/                                   ← same layout, 39 segments
├── ...                                      ← more subjects
└── README.md

<segment_id> looks like P0001_phase2_022_drink-water-b — subject, phase, sequence number, and a human-readable action slug.

IMU columns

imu.parquet is long-format: one row per sample, a device column (imu_1…imu_4) selecting which sensor, and a ts column (Unix timestamp):

Column Meaning
AccX_g, AccY_g, AccZ_g Acceleration X/Y/Z (g)
GyroX_dps, GyroY_dps, GyroZ_dps Angular velocity X/Y/Z (degrees/second)
AngleX_deg, AngleY_deg, AngleZ_deg Orientation angle X/Y/Z (degrees)
MagX_uT, MagY_uT, MagZ_uT Magnetic field X/Y/Z (µT)
QuatW, QuatX, QuatY, QuatZ Orientation quaternion
Temperature_C Sensor temperature (°C)
Battery_pct Battery level (%)

imu_1 = left wrist, imu_2 = left upper arm, imu_3 = right wrist, imu_4 = right upper arm.

The IMU's own sample rate is not the same across subjects* — P0001 and P0022 onward logged at ~200 Hz (the intended setting), but P0002–P0021 logged at ~10 Hz for their entire sessions (a ~20x difference; same hardware, a different logging configuration — P0002's raw column names were also formatted differently, see above). Don't hard-code a rate: each segment's meta.json has an imu_hz field computed directly from that file's own ts column, and that's the number to use for any T ↔ duration conversion involving imu.parquet.

At ~200 Hz, timestamps come in batches. The sensors send samples to the logger in small Bluetooth batches, and the logger stamps every sample of a batch with the batch's arrival time — so most rows (80–90%) share the previous row's ts, with batches ~30–40 ms apart. The samples themselves are real (the values change from row to row); only the per-sample time is coarse. If you need one time per sample, spread each batch evenly up to the next batch, per device:

def respace(ts, imu_hz):
    """ts: one device's timestamps, sorted. Spread each batch evenly up to the next batch."""
    ts = np.asarray(ts, float); out = ts.copy()
    starts = np.flatnonzero(np.r_[True, np.diff(ts) > 0])
    ends = np.r_[starts[1:], len(ts)]
    for a, b in zip(starts, ends):
        nxt = ts[b] if b < len(ts) else ts[a] + (b - a) / imu_hz
        out[a:b] = ts[a] + (nxt - ts[a]) * np.arange(b - a) / (b - a)
    return out

dev = imu[imu["device"] == "imu_1"].sort_values("ts", kind="stable")
t = respace(dev["ts"].to_numpy(), meta["imu_hz"])   # strictly increasing, ~imu_hz apart

* During early data collection (P0002–P0021), the IMU devices weren't configured to their maximum output rate (the team was still getting familiar with the hardware) — this is why the main EgoBrain release doesn't include IMU for those subjects at all. EgoBrain-Mini releases it anyway rather than dropping it, since a lower-rate signal is still real, usable data — just be aware of it via imu_hz rather than assuming parity with P0001 and P0022 onward. We're still looking into whether this can be backfilled or otherwise compensated for in a future update.

⏱️ How the modalities are aligned

Multimodal synchronization is achieved through an EEG-anchored, cross-modal calibration procedure: EEG is treated as the ground-truth clock, and every other modality's timeline is corrected to match it — not aligned pairwise against each other. The alignment for this subject was independently re-verified against 100% of the published sub-actions (not a sample) before release, including an EEG interval-marker cross-check baked directly into the bundled report's EEG chart (see below).

🔍 Explore it interactively

Each subject's own <subject>/reports/ (e.g. P0001/reports/, P0002/reports/) is a self-contained, offline-first HTML report — no server, no internet connection, no dependencies. Download this repo, then just open <subject>/reports/index.html in a browser:

  • A gallery of that subject's 39 sub-actions with thumbnails
  • Per-action pages with synced video + live EEG / IMU / audio charts (drag the video, the charts scroll with it, and vice versa — click any chart to jump the video there)
  • Plain-language explanations of what each chart actually shows and why

(HF's own file viewer won't execute the JavaScript in these pages — you do need the files on your own machine. That's by design: it means the report works identically whether you're online or not.)

Prefer to browse without downloading first? The Dataset Viewer above (the "Data Studio" / table tab on this page) shows every sub-action across every subject as one row — thumbnail, subject, action label, phase, duration, and which modalities it has — before you commit to downloading anything. It's generated from browse/segments.parquet, a lightweight index (thumbnails only, no video/EEG/IMU payload) built purely for this at-a-glance browsing; the real per-segment data still lives under each subject's own segments/phase2/<segment_id>/ folder as described above.

✅ Verification

Every one of the 897 sub-actions — not a sample — was individually checked:

  • EEG integrity: all 32 channels finite, no flat/dead channels
  • Alignment: cross-checked against an independent, EEG-hardware-logged marker timestamp
  • Coverage: 897 / 897, no segments skipped or excluded from verification
  • EEG quality (every released subject): two checks over the whole session — the headset's own contact-quality readings for its two ear reference electrodes, and the raw amplitude (a large artifact shared across channels, sustained above ~1 mV). Affected spans are flagged per segment — see Known data issues

⚠️ Known data issues

Known problems in the released data. The files themselves are left unaltered; every affected range is listed in that segment's meta.json (eeg_quality_flags) and marked in the bundled report.

Subject Modality Issue Affected range Segments Recommendation
P0008 EEG An electrode came loose twice; the signal is dominated by a 1–3 mV artifact shared across channels (normal EEG: ~0.05–0.15 mV) phase2_021 1 s → phase2_025 249 s, and phase2_027 1 s → phase2_036 185 s 15 of 39 Exclude from EEG analyses
P0020 EEG An ear reference electrode (CMS/DRL) lost contact; all 32 channels are affected (alpha power ~85% lower; blinks and chewing still visible) phase2_007 111.3 s → phase2_038 36.4 s (contact briefly back in phase2_019, 154.5–217.5 s) 32 of 39 Exclude from band-power / amplitude-based analyses

Video, audio and IMU are unaffected in both cases.

How they were found

  • P0008 — the raw amplitude jumps at the start of each episode and falls back when an experimenter re-attaches the electrode (tagged as not-task data in phase2_025 and phase2_036; phase2_026, in between, is normal). The headset's signal-quality reading is mostly 0 over both episodes. The reference electrodes also briefly lost contact inside them (in phase2_024, phase2_025, phase2_036); those ranges are flagged too.
  • P0020 — an experimenter re-attaches the subject's right-ear electrode ~18 s into phase2_038_closing-calibration (that segment is also tagged as an interaction). The headset's own contact-quality channels show the contact had been lost since phase2_007, and the EEG shows a sharp transient at both the loss and the re-attachment.

Using the flags. Segments without eeg_quality_flags have no known EEG issue. rel_start / rel_end are seconds from the segment start; issue is signal_quality_poor or reference_contact_lost:

"eeg_quality_flags": [
  {"rel_start": 0.0, "rel_end": 154.5, "issue": "reference_contact_lost",
   "description": "...", "evidence": "..."}
]

To keep only unflagged EEG samples:

flags = meta.get("eeg_quality_flags", [])
fs = 200                                           # clean.npy rate
keep = np.ones(eeg_clean.shape[1], bool)
for f in flags:
    keep[int(f["rel_start"] * fs):int(f["rel_end"] * fs)] = False
eeg_ok = eeg_clean[:, keep]                         # samples outside every flagged range

🔒 Privacy

99 of the 897 sub-actions briefly show something that needed redacting — most often a laptop's own lock screen or an app window, a hand-held phone, or earlier pages of a shared notebook being leafed through; occasionally a bystander incidentally in frame. Rather than blank out the whole picture for something covering a fraction of it, redaction is region-limited blur: only the specific area of the frame containing the sensitive detail is blurred (strong enough that text, faces, and names are unreadable), for only the seconds it's actually visible — the rest of the frame, and the rest of the clip, is untouched. A small number of windows also mute audio, when the audio itself (not just the picture) needed it. This is baked into the video at the source, not a runtime toggle. video/redaction_windows.json documents exactly which seconds and which region of the frame were affected, for transparency. Nothing else in the release is redacted.

If you come across any content in this release that you believe may reveal a subject's private information, please contact the maintainers immediately at nielin@iis.u-tokyo.ac.jp and it will be re-processed and re-uploaded.

🚫 What's not in this release

  • Raw, unprocessed video — not distributed at all; video/*.json retains exact provenance (source file + frame range) in case higher-fidelity access is ever needed for approved use.
  • Only Phase 2 (real-world action) sub-actions are included in this subject's data.

🚀 Quick start

import json
import numpy as np
import pandas as pd

root = "P0001/segments/phase2/P0001_phase2_022_drink-water-b"

meta = json.load(open(f"{root}/meta.json"))
eeg_clean = np.load(f"{root}/eeg/clean.npy")        # [32, T] @ 200 Hz
imu = pd.read_parquet(f"{root}/imu/imu.parquet")     # long format, 4 devices

print(meta["action"], meta["duration_sec"], "sec")
print("EEG:", eeg_clean.shape)
print("IMU devices:", imu["device"].unique())

⚙️ Processing

eeg/clean.npy is produced by (in order): linear + 6th-order detrending → 0.1–75 Hz bandpass filter → 50 Hz notch filter → resample to 200 Hz → RANSAC-based bad-channel detection and interpolation → average reference → exponential-moving-average standardization. eeg/raw.npy skips all of this — it's the channel-selected signal at its native 256 Hz, in µV.

🌟 Acknowledgement

We sincerely thank all participants who contributed their time and effort to the data collection process.

This project was initiated and developed during Nie Lin's research internship at Microsoft Research Asia (MSRA), and later continued at the Institute of Industrial Science (IIS), The University of Tokyo. We thank all collaborators and mentors at MSRA and IIS for their valuable guidance and support.

We also acknowledge the support from our advisors, collaborators, and the community.

📜 License

EgoBrain-Mini is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.

By using this dataset, you agree to:

  • Use the dataset only for non-commercial research purposes
  • Not redistribute the data
  • Not attempt to identify the participant
  • Comply with applicable ethical and data protection regulations

If you have questions regarding licensing or usage, please contact the maintainers at nielin@iis.u-tokyo.ac.jp.

License details: https://creativecommons.org/licenses/by-nc/4.0/

CC BY-NC 4.0