Load HDF5 / Zarr (OME-Zarr)(HDF5/Zarr 数据加载器)
Load HDF5 / Zarr (OME-Zarr) - User Manual
Dragonfly Prototype Apps · Load HDF5 / Zarr (OME-Zarr)...
版本 Version 1.0 · 2026-07-14
第一部分 中文手册
目录
1. 简介
2. 适用场景
3. 安装与启用
4. 运行环境与首次配置
5. 界面说明
6. 使用步骤
7. 参数说明
8. 输出结果
9. 常见问题与故障排除
10. 注意事项与已知限制
11. 参考资料
1. 简介
Load HDF5 / Zarr (OME-Zarr) 用于浏览*分块(chunked)容器*中的数据集——一个 HDF5 文件(.h5/.hdf5/.he5/.nxs/...)或一个 Zarr / OME-Zarr / N5 文件夹——预览选中的数据集,并把它(可降采样)导入为 Dragonfly 的图像 Channel。
它封装了两个开源读取库:h5py(HDF5 绑定,BSD 许可)与 zarr v3(Zarr / OME-Zarr / N5,MIT 许可),并用 numpy、Pillow 处理数组与缩略图。全部为预编译 wheel——无需 Java、无需 MSVC、无需 GPU。对 OME-Zarr,体素尺寸从顶层 group 的 multiscales 元数据(coordinateTransformations 的 scale 与 axes[].unit)读取并写入 Channel 的 spacing。
输入是磁盘上的文件/文件夹;输出是一个新的 Dragonfly Channel。它面向 Bio-Formats 与 ITK 加载器不擅长处理的大型/下一代数据(光片显微、电镜、云端 OME-NGFF)。
它不是分割、重建或通用显微厂商格式加载器。它只从 HDF5/Zarr 容器中读取数组数据——CZI/LIF/DICOM 等请用 Bio-Formats / ITK 加载器,TIFF 图像堆栈请用 Dragonfly 原生导入。
2. 适用场景
- 打开以 HDF5 或 OME-Zarr 存储的光片显微体数据。
- 加载云端/下一代 OME-NGFF(OME-Zarr)多尺度金字塔,从下拉框中选取某个分辨率层级或通道数组。
- 读取以 HDF5/NeXus 数据集保存的电镜或仿真结果。
- 通过在完整读取前设置 Downsample 系数,快速导入超大体数据的低分辨率预览。
- 在决定导入什么之前,先查看容器内容(路径、形状、dtype、chunk 大小)。
3. 安装与启用
本插件随 Prototype Apps 完整包(Full Package) 一同发布。默认关闭,需在安装时勾选。
1. 将完整包解压到一个较短的目录路径(例如 C:\PrototypeApps),以避免 Windows MAX_PATH 问题。
2. 运行 Install_FullPackage.bat。
3. 在安装对话框中勾选 "Load HDF5 / Zarr (OME-Zarr)"(所有插件默认均未勾选)。
4. 点击 Install。文件安装到 %LOCALAPPDATA% 下——无需管理员权限。
5. 完全重启 Dragonfly——菜单仅在启动时扫描。
重启后它出现在 Prototype Apps > Load HDF5 / Zarr (OME-Zarr)...,位于 Reconstruction & Imaging(重建与成像) 分组下。之后可在 Developer > Prototype Labs... > Menu Item Manager 中启用/停用;每次切换后需重启 Dragonfly。
4. 运行环境与首次配置
本插件使用独立的 Python 虚拟环境(env.kind = venv_in_code)。Dragonfly 进程本身从不导入 h5py 或 zarr——它们通过子进程(runner/hz_runner.py)在 venv 中运行,采用基于文件的 IPC。
在 Setup 标签页点击 Setup Environment (h5py + zarr)。它会用你的 Base Python 创建 venv,并 pip 安装 runner/requirements.txt 中列出的依赖:numpy、Pillow、h5py、zarr。下载量很小(数十 MB 预编译 wheel)。仅首次配置需要联网,之后可完全离线运行。
- GPU:不需要,也不使用。
- Java / MSVC / WSL / 外部工具:均不需要(依赖全为预编译 wheel)。
- venv 位置:安装代码目录内的
venv文件夹(...\GenericMenuItems\HDF5ZarrLoader\venv)。解析出的 python 路径会保存到 Analysis venv python 字段 /hz_config.json。 - Base Python 字段:留空则自动检测——依次尝试 Dragonfly 自带 Python、
py启动器、PATH 上的python/python3,以及常见 Miniconda/Anaconda/Python 安装位置。若自动检测选错,可用 Browse... 指定具体python.exe。 - Job root:每次运行的作业文件夹(config/status/results/缩略图)的工作目录,默认
C:\HDF5ZarrLoaderJobs。
5. 界面说明
面板以可停靠窗口打开,含标题行、两个标签页(Setup、Load)、绿色状态行,以及底部只读日志。
Setup 标签页
- Analysis venv python —— venv 解释器路径;由 Setup Environment 自动填入(占位符:*set by Setup Environment*)。
- Base Python(+ Browse...)—— 用于创建 venv 的解释器;留空 = 自动检测 Dragonfly/系统 Python。
- Job root —— 作业输出文件夹;默认
C:\HDF5ZarrLoaderJobs。 - Setup Environment (h5py + zarr) —— 创建 venv 并 pip 安装依赖。
Load 标签页 —— "Container" 分组
- HDF5 file / Zarr folder 输入框,配 File...(打开 HDF5 文件;筛选
*.h5 *.hdf5 *.he5 *.nxs)与 Folder...(打开 Zarr / OME-Zarr / N5 文件夹)。选取任一后会清空之前的预览并自动执行 List datasets。 - Dataset 下拉框 + List datasets 按钮 —— 枚举容器内每个数组(每项显示路径、形状与 dtype)。
- Downsample 数值框 —— 范围 1-64,默认 1;提示:*Read every Nth voxel along the spatial axes (1 = full).*
- Spacing unit 下拉框 —— *as-is / from OME*(默认)、*micrometer (µm)*、*millimeter (mm)*、*meter (m)*。
- Preview 按钮 —— 读取中间切片/缩略图并填充元数据表。
Load 标签页 —— 预览区
- 缩略图窗格(左)—— 显示渲染的预览图,或 *No preview*。
- 元数据表(右,Property/Value)—— Container、Dataset、Shape、Dimensions、Data type、Chunks、OME scale (z,y,x)、OME unit、Preview downsample。
Load 标签页 —— "Import" 分组
- Output Channel title —— 所建 Channel 的名称;留空默认为
file [dataset]。 - Import -> Publish as Channel(蓝色按钮)—— 读取选中数据集(可降采样)并发布为 Dragonfly Channel。
6. 使用步骤
1. 仅首次:在 Setup 标签页点击 Setup Environment (h5py + zarr),等待状态行显示 *Environment ready*。
2. 切换到 Load 标签页。点击 File...(HDF5 文件)或 Folder...(Zarr / OME-Zarr / N5 文件夹)。数据集会被自动列出。
3. 在 Dataset 下拉框中选择要加载的数组(OME-Zarr 请选取想要的分辨率层级/通道数组)。
4. 可选:设置 Downsample 系数以快速低分辨率导入,并设置 Spacing unit。
5. 点击 Preview 查看缩略图与元数据(形状、dtype、chunks、OME scale)。
6. 可选:输入 Output Channel title。
7. 点击 Import -> Publish as Channel。数据集会在 venv 中读取并发布为 Channel;状态行报告所建 Channel 名称。
7. 参数说明
这是一个加载器,可调数值参数很少——核心是容器/数据集的选择,加上以下选项:
参数 | 默认值 | 范围 | 说明 |
Downsample | 1 | 1-64(整数) | 步长:沿最后 <=3 个(空间)轴每隔 N 个体素读取一次。越大越快、分辨率越低。 |
Spacing unit | as-is / from OME | none / um / mm / m | 体素 spacing 的单位。 |
Output Channel title | (空) | 自由文本 | Channel 名称;留空 = |
Base Python | (空) | 路径 | 用于创建 venv 的解释器;留空 = 自动检测。 |
Job root | C:\HDF5ZarrLoaderJobs | 路径 | 每次运行作业输出的文件夹。 |
8. 输出结果
Import 会创建一个 Dragonfly 图像 Channel(经 createChannelFromNumpyArray,带 Z 轴)。它像任何导入的体数据一样出现在对象列表 / Data Properties 中。
- 命名:使用 Output Channel title,留空时为
<文件名> [<dataset>](顶层数组的 dataset 记为root)。 - 维度处理:维度 >3 的数组(如 OME-Zarr
t,c,z,y,x)通过对前导轴取索引 0 归约为 (Z,Y,X);2D 数组变为 (1,Y,X)。 - 数据类型:uint8/uint16/int8/int16/float32 原样保留;其他 dtype 转为 float32。
- Spacing:对 OME-Zarr,体素 scale 换算为米(ORS 约定)并乘以 Downsample 系数,使物理尺寸保持正确;HDF5(无内嵌 scale)则沿用 Dragonfly 默认值,除非提供带单位的 OME 源。
- Origin:设为 (0, 0, 0)。
9. 常见问题与故障排除
- "Run Setup first, or set the analysis venv python." —— venv 尚未创建。到 Setup 标签页点击 Setup Environment。
- "Select a valid HDF5 file or Zarr folder first." —— 未选择容器或路径已不存在。用 File.../Folder... 重新选取。
- "List datasets and choose one first." —— 导入前请点击 List datasets 并在 Dataset 下拉框中选中一项。
- Setup 失败 —— 通常是没有可用的 Base Python 或首次配置无网络。用 Browse... 显式指定 python.exe 并检查网络;日志会显示 pip 输出。
- 数据集列表为空 —— 容器可能只含 <2D 的数组(列表仅保留 ndim >= 2 的数组),或格式不被 h5py/zarr 读取。N5 经 zarr 为尽力支持。
- Import 时内存不足 —— 超大单块数据集会被整体读入内存。加大 Downsample 系数,或导入更低分辨率的 OME-Zarr 层级。
10. 注意事项与已知限制
- 每次仅导入一个数据集。多通道/多尺度 OME-Zarr 需逐个数组导入——分别从下拉框选取各层级/通道。
- 前导轴被丢弃:对 >3D 数据仅导入时间/通道轴的索引 0(Z,Y,X);其他时间点/通道需分别导入各自的数组。
- 体素 spacing 仅从 OME-Zarr 的
multiscales元数据读取;普通 HDF5/Zarr 数组在此不含 spacing。 - N5 经 zarr 为尽力支持,可能无法打开每一种 N5 变体。
- 超大单块读取仍会整体载入内存;降采样能缓解但无法消除。
- 已通过离线往返测试验证 HDF5(嵌套数据集)与 Zarr / OME-Zarr(带 multiscales scale 的 5D);在 Dragonfly 内的实机验证尚待完成。
11. 参考资料
- h5py(HDF5 for Python,BSD 许可)—— https://www.h5py.org/
- zarr-python v3(Zarr / N5,MIT 许可)—— https://zarr.readthedocs.io/
- OME-NGFF / OME-Zarr 规范 —— https://ngff.openmicroscopy.org/
- NumPy(https://numpy.org/)与 Pillow(https://python-pillow.org/),用于数组处理与缩略图。
- 插件的
README.md与DESIGN.md,位于DragonflyPlugins/HDF5ZarrLoader-Plugin/。
Part II English Manual
Contents
1. Introduction
2. Use cases
3. Installation & enabling
4. Runtime environment & first-run setup
5. Interface
6. How to use
7. Parameters
8. Output
9. FAQ & troubleshooting
10. Notes & known limitations
11. References
1. Introduction
Load HDF5 / Zarr (OME-Zarr) browses the datasets inside a *chunked container* - an HDF5 file (.h5/.hdf5/.he5/.nxs/...) or a Zarr / OME-Zarr / N5 directory - lets you preview a chosen dataset, and imports it (optionally downsampled) as a Dragonfly image Channel.
It wraps two open-source readers: h5py (HDF5 bindings, BSD license) and zarr v3 (Zarr / OME-Zarr / N5, MIT license), with numpy and Pillow for array handling and thumbnails. All are prebuilt wheels - no Java, no MSVC, no GPU. For OME-Zarr the voxel size is read from the top group's multiscales metadata (coordinateTransformations scale + axes[].unit) and written to the Channel spacing.
Input is a file/folder on disk; output is a new Dragonfly Channel. It is aimed at big / next-generation data (light-sheet, EM, cloud OME-NGFF) that the Bio-Formats and ITK loaders don't handle well.
This is not a segmentation, reconstruction, or general microscopy-vendor loader. It reads array data out of HDF5/Zarr containers only - use the Bio-Formats / ITK loaders for CZI/LIF/DICOM etc., and Dragonfly's native import for TIFF stacks.
2. Use cases
- Opening light-sheet microscopy volumes stored as HDF5 or OME-Zarr.
- Loading cloud / next-generation OME-NGFF (OME-Zarr) multiscale pyramids, picking a specific resolution level or channel array from the dropdown.
- Reading electron-microscopy or simulation results saved as HDF5/NeXus datasets.
- Quickly importing a low-resolution preview of a very large volume by setting a Downsample factor before the full read.
- Inspecting a container's contents (paths, shapes, dtypes, chunk sizes) before deciding what to import.
3. Installation & enabling
This plugin ships in the Prototype Apps Full Package. It is disabled by default and must be checked at install time.
1. Unzip the Full Package to a short directory path (e.g. C:\PrototypeApps) to avoid Windows MAX_PATH issues.
2. Run Install_FullPackage.bat.
3. In the installer dialog, check "Load HDF5 / Zarr (OME-Zarr)" (all plugins are unchecked by default).
4. Click Install. Files are copied under %LOCALAPPDATA% - no administrator rights are needed.
5. Fully restart Dragonfly - the menu is scanned only at startup.
After restart it appears under Prototype Apps > Load HDF5 / Zarr (OME-Zarr)... in the Reconstruction & Imaging section. You can later enable/disable it from Developer > Prototype Labs... > Menu Item Manager; restart Dragonfly after each toggle.
4. Runtime environment & first-run setup
The plugin uses an isolated Python virtual environment (env.kind = venv_in_code). The Dragonfly process itself never imports h5py or zarr - they run inside the venv via a subprocess (runner/hz_runner.py) with file-based IPC.
On the Setup tab, click Setup Environment (h5py + zarr). This builds a venv from your Base Python and pip-installs the exact dependencies from runner/requirements.txt: numpy, Pillow, h5py, zarr. The download is small (tens of MB of prebuilt wheels). Internet is required only for this first setup; afterwards the plugin runs fully offline.
- GPU: not required, not used.
- Java / MSVC / WSL / external tools: none required (all deps are prebuilt wheels).
- Venv location: a
venvfolder inside the installed code directory (...\GenericMenuItems\HDF5ZarrLoader\venv). The resolved python path is saved to the Analysis venv python field /hz_config.json. - Base Python field: leave blank to auto-detect Dragonfly's own Python, then the
pylauncher,python/python3on PATH, and common Miniconda/Anaconda/Python install locations. Use Browse... to point at a specificpython.exeif auto-detect picks the wrong one. - Job root: working directory for per-run job folders (config/status/results/thumbnail), default
C:\HDF5ZarrLoaderJobs.
5. Interface
The panel opens as a dockable window with a header line and two tabs (Setup, Load), a green status line, and a read-only log at the bottom.
Setup tab
- Analysis venv python - path to the venv interpreter; filled in automatically by Setup Environment (placeholder: *set by Setup Environment*).
- Base Python (+ Browse...) - interpreter used to build the venv; blank = auto-detect Dragonfly/system Python.
- Job root - folder for job outputs; default
C:\HDF5ZarrLoaderJobs. - Setup Environment (h5py + zarr) - builds the venv and pip-installs the requirements.
Load tab - "Container" group
- HDF5 file / Zarr folder field with File... (opens an HDF5 file; filter
*.h5 *.hdf5 *.he5 *.nxs) and Folder... (opens a Zarr / OME-Zarr / N5 directory). Picking either clears the previous preview and auto-runs List datasets. - Dataset dropdown + List datasets button - enumerates every array in the container (each entry shows path, shape and dtype).
- Downsample spin box - range 1-64, default 1; tooltip *Read every Nth voxel along the spatial axes (1 = full).*
- Spacing unit dropdown - *as-is / from OME* (default), *micrometer (µm)*, *millimeter (mm)*, *meter (m)*.
- Preview button - reads a middle slice/thumbnail and fills the metadata table.
Load tab - preview area
- Thumbnail pane (left) - shows the rendered preview image, or *No preview*.
- Metadata table (right, Property/Value) - Container, Dataset, Shape, Dimensions, Data type, Chunks, OME scale (z,y,x), OME unit, Preview downsample.
Load tab - "Import" group
- Output Channel title - name for the created Channel; blank defaults to
file [dataset]. - Import -> Publish as Channel (blue button) - reads the chosen dataset (optionally downsampled) and publishes it as a Dragonfly Channel.
6. How to use
1. First run only: on the Setup tab, click Setup Environment (h5py + zarr) and wait for *Environment ready* in the status line.
2. Switch to the Load tab. Click File... (for an HDF5 file) or Folder... (for a Zarr / OME-Zarr / N5 folder). The datasets are listed automatically.
3. In the Dataset dropdown, choose the array to load (for OME-Zarr, pick the resolution level / channel array you want).
4. Optionally set a Downsample factor for a fast low-resolution import, and a Spacing unit.
5. Click Preview to check the thumbnail and metadata (shape, dtype, chunks, OME scale).
6. Optionally type an Output Channel title.
7. Click Import -> Publish as Channel. The dataset is read in the venv and published as a Channel; the status line reports the created Channel name.
7. Parameters
This is a loader, so it has few numeric tunables - the key controls are the container/dataset selection plus these options:
Parameter | Default | Range | Meaning |
Downsample | 1 | 1-64 (integer) | Stride: read every Nth voxel along the last <=3 (spatial) axes. Larger = faster, lower resolution. |
Spacing unit | as-is / from OME | none / um / mm / m | Unit for the voxel spacing. |
Output Channel title | (blank) | free text | Channel name; blank = |
Base Python | (blank) | path | Interpreter to build the venv; blank = auto-detect. |
Job root | C:\HDF5ZarrLoaderJobs | path | Folder for per-run job outputs. |
8. Output
Import creates one Dragonfly image Channel (via createChannelFromNumpyArray, with a Z axis). It appears in the object list / Data Properties like any imported volume.
- Naming: the Output Channel title, or
<filestem> [<dataset>]when left blank (datasetrootfor a top-level array). - Dimensionality: arrays with >3 dimensions (e.g. OME-Zarr
t,c,z,y,x) are reduced to (Z,Y,X) by taking index 0 on the leading axes; a 2D array becomes (1,Y,X). - Data type: uint8/uint16/int8/int16/float32 are kept as-is; any other dtype is cast to float32.
- Spacing: for OME-Zarr the voxel scale is converted to meters (ORS convention) and multiplied by the Downsample factor so the physical size stays correct; HDF5 (no embedded scale) uses Dragonfly defaults unless you supply a unit-carrying OME source.
- Origin: set to (0, 0, 0).
9. FAQ & troubleshooting
- "Run Setup first, or set the analysis venv python." - the venv hasn't been built. Go to the Setup tab and click Setup Environment.
- "Select a valid HDF5 file or Zarr folder first." - no container is chosen or the path no longer exists. Re-pick with File.../Folder...
- "List datasets and choose one first." - click List datasets and select an entry in the Dataset dropdown before importing.
- Setup fails - usually no usable Base Python or no internet on first setup. Set Base Python explicitly (Browse... to a python.exe) and check the connection; the log shows the pip output.
- Empty dataset list - the container may hold only <2D arrays (listing keeps arrays with ndim >= 2), or the format isn't readable by h5py/zarr. N5 is best-effort via zarr.
- Out of memory on Import - a very large single-chunk dataset is read fully into memory. Increase the Downsample factor, or import a lower-resolution OME-Zarr level.
10. Notes & known limitations
- One dataset per import. Multi-channel / multiscale OME-Zarr is imported one array at a time - pick each level/channel from the dropdown separately.
- Leading axes are dropped: for >3D data only index 0 of the time/channel axes is imported (Z,Y,X); other timepoints/channels require separate imports of their own arrays.
- Voxel spacing is read only from OME-Zarr
multiscalesmetadata; plain HDF5/Zarr arrays carry no spacing here. - N5 is best-effort (via zarr) and may not open every N5 variant.
- Very large single-chunk reads still load into RAM; downsampling mitigates but doesn't eliminate this.
- Validated offline by round-tripping HDF5 (nested datasets) and Zarr / OME-Zarr (5D with multiscales scale); live in-Dragonfly verification is pending.
11. References
- h5py (HDF5 for Python, BSD license) - https://www.h5py.org/
- zarr-python v3 (Zarr / N5, MIT license) - https://zarr.readthedocs.io/
- OME-NGFF / OME-Zarr specification - https://ngff.openmicroscopy.org/
- NumPy (https://numpy.org/) and Pillow (https://python-pillow.org/) for array handling and thumbnails.
- Plugin
README.mdandDESIGN.mdinDragonflyPlugins/HDF5ZarrLoader-Plugin/.