Download README.md from ut-vision/EgoBrain-Mini: direct link, hf CLI and curl.
- Browser
- Download file 23.4 kB
-
https://huggingface.co/datasets/ut-vision/EgoBrain-Mini/resolve/main/README.md
- Command line
-
hf download hf://datasets/ut-vision/EgoBrain-Mini/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/datasets/ut-vision/EgoBrain-Mini/resolve/main/README.md
license: cc-by-nc-4.0
pretty_name: EgoBrain-Mini
tags:
- eeg
- egocentric-video
- brain-computer-interface
- multimodal
- imu
- motion-capture
- audio
- neuroscience
- video
- timeseries
language:
- en
size_categories:
- n<1K
configs:
- config_name: default
data_files:
- split: train
path: browse/segments.parquet
extra_gated_fields:
First Name: text
Last Name: text
Email: text
Affiliation: text
Position (PhD/Master/Faculty/Industry): text
Country: country
Personal or Lab Website (optional): text
How did you learn about the EgoBrain-Mini project?: text
What do you plan to use EgoBrain-Mini for?: text
Why are you interested in our EgoBrain-Mini Project?: text
I confirm this dataset will only be used for research purposes: checkbox
I agree not to redistribute the dataset: checkbox
I agree not to attempt re-identification of any participant: checkbox
I agree to properly cite the EgoBrain paper and dataset: checkbox
[ICLR 2026]
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
EgoBrain-Mini — lightweight, fast, and fully synchronized
東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia
Nie Lin · Yansen Wang · Dongqi Han · Weibang Jiang · Jingyuan Li · Ryosuke Furuta · Yoichi Sato* · Dongsheng Li* · *(Co-corresponding authors)*
This is EgoBrain-Mini, a lightweight companion release of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action Understanding" — the egocentric subset (first-person video + EEG + IMU + audio), meant as a quick, no-request-needed way to explore the data and its alignment before diving into the full EgoBrain release.
Nie (Elon) Lin, Yansen Wang, Dongqi Han, Weibang Jiang, Jingyuan Li, Ryosuke Furuta, Yoichi Sato*, Dongsheng Li* (*co-corresponding authors). "EgoBrain: Synergizing Minds and Eyes For Human Action Understanding", ICLR 2026.
🗺️ Contents
- 📢 News
- 📦 What's actually in the box
- 📂 Dataset Structure
- ⏱️ How the modalities are aligned
- 🔍 Explore it interactively
- ✅ Verification
- ⚠️ Known data issues
- 🔒 Privacy
- 🚫 What's not in this release
- 🚀 Quick start
- ⚙️ Processing
- 🌟 Acknowledgement
- 📜 License
📢 News
- [2026-10-04]: Released Subject P0023 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-10-04]: Released Subject P0022 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-10-04]: Updated P0008: 15 of its 39 sub-actions now carry EEG quality flags (a loose electrode, twice) — see Known data issues.
- [2026-10-04]: Released Subject P0021 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-10-04]: Released Subject P0020 egocentric video + EEG + IMU + audio (Phase 2). ⚠️ In 32 of its 39 sub-actions an EEG reference electrode had lost contact — see Known data issues.
- [2026-10-03]: Released Subject P0019 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-10-03]: Released Subject P0018 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-27]: Released Subject P0017 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-21]: Released Subject P0016 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-20]: Released Subject P0015 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-16]: Released Subject P0014 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-14]: Released Subject P0013 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-12]: Released Subject P0012 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-09]: Released Subject P0011 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-07]: Released Subject P0010 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-06]: Released Subject P0009 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-06]: Released Subject P0008 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-04]: Released Subject P0007 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-03]: Released Subject P0006 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-02]: Released Subject P0005 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-09-02]: Released Subject P0004 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-08-20]: Released Subject P0003 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-08-19]: Released Subject P0002 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-08-19]: Released Subject P0001 egocentric video + EEG + IMU + audio (Phase 2).
- [2026-08-18]: Created the Hugging Face repository for ut-vision/EgoBrain-Mini.
📦 What's actually in the box
23 subjects (P0001–P0023), 897 real-world sub-actions (39 per subject), four modalities — first-person video, EEG, a body-worn IMU, and audio — all recorded from the same person doing the same thing at the same instant, and individually verified, not sampled.
| Modality | File | Format | Shape | Rate / Size | Notes |
|---|---|---|---|---|---|
| 🧠 EEG (raw) | eeg/raw.npy |
float32 array, µV | [32, T] |
32 channels @ 256 Hz | Channel-selected only, otherwise unprocessed |
| 🧠 EEG (clean) | eeg/clean.npy |
float32 array, standardized | [32, T] |
32 channels @ 200 Hz | Detrended, 0.1–75 Hz bandpass + 50 Hz notch, RANSAC bad-channel interpolation, average reference, exponential-moving standardization — see Processing |
| 🖐️ IMU | imu/imu.parquet |
Parquet, long format | [T, 20] per device |
4 body-worn sensors, rate varies by subject/session — see meta.json's imu_hz, don't assume a fixed number |
Left/right wrist + left/right upper arm — see IMU columns |
| 🎙️ Audio | audio/ego.wav |
WAV, mono PCM16 | [T] |
16 kHz | The Ego camera's own microphone |
| 🎥 Video | video/ego.mp4 |
H.264 / MP4 | — | 854×480 @ ~30 fps | First-person; a copy of its own mic audio is embedded too, but audio/ego.wav is the higher-fidelity source |
| 📄 Provenance | video/ego.json |
JSON | — | — | Exact raw source chapter + frame range this clip was cut from (for traceability; raw video itself is not distributed) |
T is the total sample count for that one sub-action's clip — duration_sec × sample_rate — not a
separate dimension on top of the rate. For example, the 76.2 s setup-calibration sub-action has
eeg/raw.npy.shape == (32, 19508) (76.2 × 256 Hz) and eeg/clean.npy.shape == (32, 15240) (76.2 × 200 Hz
exactly). Since duration varies per sub-action (from ~1.5 s to ~17 min across the 897 of them), T is different
for every file — it's never a fixed constant you can hard-code.
The array itself carries no metadata — it's just numbers. To go the other way (shape → duration), either
read meta.json's duration_sec directly (the authoritative source, already used in Quick start),
or divide back out: duration_sec ≈ T / sample_rate (e.g. 15240 / 200 = 76.2), since the sample rate
per file is fixed and documented above.
Every one of the 897 sub-actions has all six of these. Nothing here is a placeholder or a partial modality.
📂 Dataset Structure
EgoBrain-Mini/
├── assets/
├── P0001/
│ ├── manifest.json ← index of all 39 segments + metadata
│ ├── segments/
│ │ └── phase2/
│ │ ├── P0001_phase2_000_setup-calibration/
│ │ │ ├── meta.json ← action label, timestamps, duration, modality flags,
│ │ │ │ and eeg_quality_flags where EEG quality is compromised
│ │ │ ├── eeg/
│ │ │ │ ├── raw.npy
│ │ │ │ └── clean.npy
│ │ │ ├── imu/
│ │ │ │ └── imu.parquet
│ │ │ ├── audio/
│ │ │ │ └── ego.wav
│ │ │ └── video/
│ │ │ ├── ego.mp4
│ │ │ ├── ego.json
│ │ │ └── redaction_windows.json ← only present where needed, see Privacy below
│ │ └── ... (39 segments total, same layout)
│ └── reports/ ← bundled, offline-first HTML report, see below
├── P0002/ ← same layout, 39 segments
├── ... ← more subjects
└── README.md
<segment_id> looks like P0001_phase2_022_drink-water-b — subject, phase, sequence number, and a
human-readable action slug.
IMU columns
imu.parquet is long-format: one row per sample, a device column (imu_1…imu_4) selecting which
sensor, and a ts column (Unix timestamp):
| Column | Meaning |
|---|---|
AccX_g, AccY_g, AccZ_g |
Acceleration X/Y/Z (g) |
GyroX_dps, GyroY_dps, GyroZ_dps |
Angular velocity X/Y/Z (degrees/second) |
AngleX_deg, AngleY_deg, AngleZ_deg |
Orientation angle X/Y/Z (degrees) |
MagX_uT, MagY_uT, MagZ_uT |
Magnetic field X/Y/Z (µT) |
QuatW, QuatX, QuatY, QuatZ |
Orientation quaternion |
Temperature_C |
Sensor temperature (°C) |
Battery_pct |
Battery level (%) |
imu_1 = left wrist, imu_2 = left upper arm, imu_3 = right wrist, imu_4 = right upper arm.
The IMU's own sample rate is not the same across subjects* — P0001 and P0022 onward logged at
~200 Hz (the intended setting), but P0002–P0021 logged at ~10 Hz for their entire sessions (a ~20x
difference; same hardware, a different logging configuration — P0002's raw column names were also
formatted differently, see above). Don't hard-code a rate: each segment's meta.json has an imu_hz
field computed directly from that file's own ts column, and that's the number to use for any
T ↔ duration conversion involving imu.parquet.
At ~200 Hz, timestamps come in batches. The sensors send samples to the logger in small Bluetooth
batches, and the logger stamps every sample of a batch with the batch's arrival time — so most rows
(80–90%) share the previous row's ts, with batches ~30–40 ms apart. The samples themselves are real
(the values change from row to row); only the per-sample time is coarse. If you need one time per
sample, spread each batch evenly up to the next batch, per device:
def respace(ts, imu_hz):
"""ts: one device's timestamps, sorted. Spread each batch evenly up to the next batch."""
ts = np.asarray(ts, float); out = ts.copy()
starts = np.flatnonzero(np.r_[True, np.diff(ts) > 0])
ends = np.r_[starts[1:], len(ts)]
for a, b in zip(starts, ends):
nxt = ts[b] if b < len(ts) else ts[a] + (b - a) / imu_hz
out[a:b] = ts[a] + (nxt - ts[a]) * np.arange(b - a) / (b - a)
return out
dev = imu[imu["device"] == "imu_1"].sort_values("ts", kind="stable")
t = respace(dev["ts"].to_numpy(), meta["imu_hz"]) # strictly increasing, ~imu_hz apart
* During early data collection (P0002–P0021), the IMU devices weren't configured to
their maximum output rate (the team was still getting familiar with the hardware) — this is why the
main EgoBrain release doesn't include IMU for those subjects at all. EgoBrain-Mini releases it anyway
rather than dropping it, since a lower-rate signal is still real, usable data — just be aware of it via
imu_hz rather than assuming parity with P0001 and P0022 onward. We're still looking into whether this can be backfilled
or otherwise compensated for in a future update.
⏱️ How the modalities are aligned
Multimodal synchronization is achieved through an EEG-anchored, cross-modal calibration procedure: EEG is treated as the ground-truth clock, and every other modality's timeline is corrected to match it — not aligned pairwise against each other. The alignment for this subject was independently re-verified against 100% of the published sub-actions (not a sample) before release, including an EEG interval-marker cross-check baked directly into the bundled report's EEG chart (see below).
🔍 Explore it interactively
Each subject's own <subject>/reports/ (e.g. P0001/reports/, P0002/reports/) is a self-contained,
offline-first HTML report — no server, no internet connection, no dependencies. Download this repo,
then just open <subject>/reports/index.html in a browser:
- A gallery of that subject's 39 sub-actions with thumbnails
- Per-action pages with synced video + live EEG / IMU / audio charts (drag the video, the charts scroll with it, and vice versa — click any chart to jump the video there)
- Plain-language explanations of what each chart actually shows and why
(HF's own file viewer won't execute the JavaScript in these pages — you do need the files on your own machine. That's by design: it means the report works identically whether you're online or not.)
Prefer to browse without downloading first? The Dataset Viewer above (the "Data Studio" /
table tab on this page) shows every sub-action across every subject as one row — thumbnail, subject,
action label, phase, duration, and which modalities it has — before you commit to downloading anything.
It's generated from browse/segments.parquet, a lightweight index (thumbnails only, no video/EEG/IMU
payload) built purely for this at-a-glance browsing; the real per-segment data still lives under each
subject's own segments/phase2/<segment_id>/ folder as described above.
✅ Verification
Every one of the 897 sub-actions — not a sample — was individually checked:
- EEG integrity: all 32 channels finite, no flat/dead channels
- Alignment: cross-checked against an independent, EEG-hardware-logged marker timestamp
- Coverage: 897 / 897, no segments skipped or excluded from verification
- EEG quality (every released subject): two checks over the whole session — the headset's own contact-quality readings for its two ear reference electrodes, and the raw amplitude (a large artifact shared across channels, sustained above ~1 mV). Affected spans are flagged per segment — see Known data issues
⚠️ Known data issues
Known problems in the released data. The files themselves are left unaltered; every affected range is
listed in that segment's meta.json (eeg_quality_flags) and marked in the bundled report.
| Subject | Modality | Issue | Affected range | Segments | Recommendation |
|---|---|---|---|---|---|
| P0008 | EEG | An electrode came loose twice; the signal is dominated by a 1–3 mV artifact shared across channels (normal EEG: ~0.05–0.15 mV) | phase2_021 1 s → phase2_025 249 s, and phase2_027 1 s → phase2_036 185 s |
15 of 39 | Exclude from EEG analyses |
| P0020 | EEG | An ear reference electrode (CMS/DRL) lost contact; all 32 channels are affected (alpha power ~85% lower; blinks and chewing still visible) | phase2_007 111.3 s → phase2_038 36.4 s (contact briefly back in phase2_019, 154.5–217.5 s) |
32 of 39 | Exclude from band-power / amplitude-based analyses |
Video, audio and IMU are unaffected in both cases.
How they were found
- P0008 — the raw amplitude jumps at the start of each episode and falls back when an
experimenter re-attaches the electrode (tagged as not-task data in
phase2_025andphase2_036;phase2_026, in between, is normal). The headset's signal-quality reading is mostly 0 over both episodes. The reference electrodes also briefly lost contact inside them (inphase2_024,phase2_025,phase2_036); those ranges are flagged too. - P0020 — an experimenter re-attaches the subject's right-ear electrode ~18 s into
phase2_038_closing-calibration(that segment is also tagged as an interaction). The headset's own contact-quality channels show the contact had been lost sincephase2_007, and the EEG shows a sharp transient at both the loss and the re-attachment.
Using the flags. Segments without eeg_quality_flags have no known EEG issue. rel_start /
rel_end are seconds from the segment start; issue is signal_quality_poor or
reference_contact_lost:
"eeg_quality_flags": [
{"rel_start": 0.0, "rel_end": 154.5, "issue": "reference_contact_lost",
"description": "...", "evidence": "..."}
]
To keep only unflagged EEG samples:
flags = meta.get("eeg_quality_flags", [])
fs = 200 # clean.npy rate
keep = np.ones(eeg_clean.shape[1], bool)
for f in flags:
keep[int(f["rel_start"] * fs):int(f["rel_end"] * fs)] = False
eeg_ok = eeg_clean[:, keep] # samples outside every flagged range
🔒 Privacy
99 of the 897 sub-actions briefly show something that needed redacting — most often a laptop's own
lock screen or an app window, a hand-held phone, or earlier pages of a shared notebook being leafed
through; occasionally a bystander incidentally in frame. Rather
than blank out the whole picture for something covering a fraction of it, redaction is region-limited
blur: only the specific area of the frame containing the sensitive detail is blurred (strong enough
that text, faces, and names are unreadable), for only the seconds it's actually visible — the rest of
the frame, and the rest of the clip, is untouched. A small number of windows also mute audio, when the
audio itself (not just the picture) needed it. This is baked into the video at the source, not a
runtime toggle. video/redaction_windows.json documents exactly which seconds and which region of the
frame were affected, for transparency. Nothing else in the release is redacted.
If you come across any content in this release that you believe may reveal a subject's private information, please contact the maintainers immediately at nielin@iis.u-tokyo.ac.jp and it will be re-processed and re-uploaded.
🚫 What's not in this release
- Raw, unprocessed video — not distributed at all;
video/*.jsonretains exact provenance (source file + frame range) in case higher-fidelity access is ever needed for approved use. - Only Phase 2 (real-world action) sub-actions are included in this subject's data.
🚀 Quick start
import json
import numpy as np
import pandas as pd
root = "P0001/segments/phase2/P0001_phase2_022_drink-water-b"
meta = json.load(open(f"{root}/meta.json"))
eeg_clean = np.load(f"{root}/eeg/clean.npy") # [32, T] @ 200 Hz
imu = pd.read_parquet(f"{root}/imu/imu.parquet") # long format, 4 devices
print(meta["action"], meta["duration_sec"], "sec")
print("EEG:", eeg_clean.shape)
print("IMU devices:", imu["device"].unique())
⚙️ Processing
eeg/clean.npy is produced by (in order): linear + 6th-order detrending → 0.1–75 Hz bandpass filter →
50 Hz notch filter → resample to 200 Hz → RANSAC-based bad-channel detection and interpolation → average
reference → exponential-moving-average standardization. eeg/raw.npy skips all of this — it's the
channel-selected signal at its native 256 Hz, in µV.
🌟 Acknowledgement
We sincerely thank all participants who contributed their time and effort to the data collection process.
This project was initiated and developed during Nie Lin's research internship at Microsoft Research Asia (MSRA), and later continued at the Institute of Industrial Science (IIS), The University of Tokyo. We thank all collaborators and mentors at MSRA and IIS for their valuable guidance and support.
We also acknowledge the support from our advisors, collaborators, and the community.
📜 License
EgoBrain-Mini is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.
By using this dataset, you agree to:
- Use the dataset only for non-commercial research purposes
- Not redistribute the data
- Not attempt to identify the participant
- Comply with applicable ethical and data protection regulations
If you have questions regarding licensing or usage, please contact the maintainers at nielin@iis.u-tokyo.ac.jp.
License details: https://creativecommons.org/licenses/by-nc/4.0/
