Measurements & AnalysisChinese & English

Cluster Measurements (UMAP/HDBSCAN)

Cluster Measurements (UMAP/HDBSCAN) is a Dragonfly plugin that performs unsupervised clustering of a MultiROI's per-label measurement table. Once you have computed per-label measurements (volume, surface area, sphericity

Updated 2026-07-09User manual

测量聚类 (UMAP/HDBSCAN) 插件用户手册

Cluster Measurements (UMAP/HDBSCAN) - User Manual

Dragonfly Prototype Apps · Cluster Measurements (UMAP/HDBSCAN)...

版本 Version 1.0 · 2026-07-09


第一部分 中文手册

目录

1. 简介

2. 适用场景

3. 安装与启用

4. 运行环境与首次配置

5. 界面说明

6. 使用步骤

7. 参数说明

8. 输出结果

9. 常见问题与故障排除

10. 注意事项与已知限制

11. 参考资料

1. 简介

测量聚类 (UMAP/HDBSCAN) 是 Dragonfly 的一个插件工具,用于对一个 MultiROI(多标签 ROI) 的逐标签测量表进行无监督聚类。当您在 Dragonfly 中对成百上千个分割标签(颗粒、孔隙、晶粒、细胞等)计算了体积、表面积、球度、形状因子、取向等逐标签测量后,这些测量会以标量槽(scalar slot)的形式保存在 MultiROI 上。本插件读取您勾选的标量槽,把每个标签看作一个特征向量,自动把这些标签归入若干形态或成分相近的群组(cluster),并把聚类结果写回为一个新的 MultiROI 标量槽,同时生成一张二维散点图 PNG。

整个计算流程是:读取勾选的标量槽 → 得到一个特征矩阵(每行一个标签)→ 可选标准化(StandardScaler) → 用 UMAP / t-SNE / PCA 把特征降到二维 → 用 KMeans(指定簇数) 或 HDBSCAN(指定最小簇大小) 进行聚类 → 把每个标签的聚类编号写回一个新的标量槽 → 保存一张按聚类着色的二维散点图。之后您只要在 Dragonfly 里按这个新标量槽为 MultiROI 着色,就能直观地看到各个群组。

底层引擎与算法(均为宽松开源许可): 降维与聚类由 scikit-learn(PCA、t-SNE、KMeans、StandardScaler)、umap-learn(UMAP 降维)、hdbscan(HDBSCAN 密度聚类)完成;数值计算基于 numpy;二维散点图由 matplotlib(Agg 无界面后端)绘制。

许可证要点:插件代码遵循 DragonflyPrototypeLabs 仓库的许可条款;所用计算引擎全部为宽松许可 —— scikit-learn(BSD-3-Clause)、umap-learn(BSD-3-Clause)、hdbscan(BSD-3-Clause)、matplotlib(PSF/BSD 风格)、numpy(BSD)。

2. 适用场景

本插件适用于材料科学与生命科学中的颗粒 / 孔隙 / 晶粒 / 细胞分群:依据每个标签的测量值(体积、表面积、形状因子、球度、取向等),把成千上万个标签自动归类为少数几个形态或成分相近的群体,用于分类统计、异常检测与结构表征。

  • 颗粒 / 孔隙分群: 把 CT 分割得到的孔隙或颗粒按体积、球度、形状因子分成大孔/小孔、球形/片状等群组,做定量分类统计。
  • 晶粒表征: 依据晶粒的形态测量把晶粒分组,识别异常晶粒或不同的晶粒族群。
  • 细胞分群: 在生命科学图像里,依据细胞体积、形状等测量把细胞分为若干形态类别。
  • 异常检测: 使用 HDBSCAN 时,少数不属于任何密集群组的标签会被标记为噪声(noise,编号 -1),便于挑出离群的异常结构。
  • 降维可视化: 即使不做严格分群,也可以借助二维散点图直观查看高维测量在二维空间中的分布结构。

3. 安装与启用

本插件作为 Prototype Apps 的一部分随 Full Package(完整安装包) 分发。安装步骤如下:

1. 将 Full Package 压缩包解压到任意较短的目录(例如 C:\PL\,不要放在很深的下载目录或 OneDrive 重定向的桌面下)。

2. 双击 `Install_FullPackage.bat`。

3. 在弹出的对话框中选择核心安装模式(Fresh 全新 / Compatible 兼容),并在 Prototype Apps 列表中勾选需要的应用。

4. 点击 Install,等待控制台完成。

5. 完全退出并重启 Dragonfly(菜单只在启动时扫描)。

本插件默认未勾选。 在安装器的应用列表中,轻量菜单项默认开启,所有插件默认关闭——请务必手动勾选 “Cluster Measurements (UMAP/HDBSCAN)” 才会安装。

重启 Dragonfly 后,菜单出现在:Prototype Apps ▸ Cluster Measurements (UMAP/HDBSCAN)...(位于该菜单的 “Measurements & Analysis(测量与分析)” 分组下)。点击它会打开一个可停靠的浮动面板。

以后修改勾选: 最方便的方式是在 Dragonfly 内操作——打开 Developer ▸ Prototype Labs... ▸ Menu Item Manager,在底部的 “Prototype Apps (Full Package)” 列表里勾选=部署本插件,取消勾选=移除菜单项;修改后重启 Dragonfly 生效。停用从不删除插件已搭好的环境,重新启用立即可用。也可以随时重跑安装器(上次的选择即为默认值)。

4. 运行环境与首次配置

本插件的重计算(scikit-learn / umap-learn / hdbscan / matplotlib)运行在一个专用虚拟环境(venv)中,而不会污染 Dragonfly 自带的 Python。首次使用前必须在面板里点一次 Setup Environment(搭建环境)。

Setup Environment 做什么

点击 Setup Environment 后,插件会:(1)以 Dragonfly 自带的 Python 为基础解释器(默认),在已安装的插件代码目录内新建一个 venv;(2)通过 pip 从 PyPI 安装 scikit-learn + umap-learn + hdbscan + numpy + matplotlib;(3)把可用的 venv Python 路径记录到配置中,面板 “Status(状态)” 一栏显示为 “Ready:<路径>”。这些依赖的版本未固定(pip 会自动匹配基础 CPython 的合适 wheel)。

联网、GPU、磁盘与路径

  • 需要联网一次: 仅在 Setup Environment 从 PyPI 下载依赖时需要联网;搭建完成后再运行聚类不再需要联网。
  • 不需要 GPU: 全部为纯 CPU 计算(cp310 上均为纯 wheel,无 CUDA)。
  • 不需要 WSL、不依赖任何外部软件。
  • 安装位置: venv 建在已安装插件的代码目录内,即 %LOCALAPPDATA%\comet\<Dragonfly版本>\pythonUserExtensions\GenericMenuItems\ClustersPlotter\venv。重装或更新插件代码时,该 venv 会被保留(不会被清除)。
  • 下载体积: 取决于 scikit-learn、scipy、numpy、matplotlib、umap-learn、hdbscan 及其依赖的 wheel(通常为数十至一两百 MB 量级);实际大小随基础 Python 版本的 wheel 而定,手册不给出精确数字。

Base Python(基础解释器)与失败时的替代方案

面板的 Base Python 字段用于覆盖搭建 venv 所用的基础解释器。留空时使用当前 Dragonfly 自带的 Python;也可以填一个绝对路径(指向 CPython 3.10+ 的 python.exe),或填一个启动器命令(例如 py -3.11)。如果用 Dragonfly 自带 Python 搭建失败(例如 pip 解析 wheel 出错),可在此填另一个你机器上已装的 CPython 3.10+ 重试。

若 Setup Environment 失败,面板日志会打印完整错误信息(traceback)。常见原因是首次搭建时无法联网访问 PyPI,或所选基础 Python 版本没有对应的预编译 wheel。请确认网络可达 PyPI,或改用 Base Python 指向另一个 3.10+ 解释器后重试。

5. 界面说明

面板是一个可停靠的浮动窗口,自上而下由以下几个分组构成。下面逐一说明每个控件(与代码完全一致):

Input MultiROI(输入 MultiROI)

  • MultiROI 下拉框: 列出当前 Dragonfly 会话中带测量标量槽的 MultiROI。每一项显示为 “标题(N labels, M slots)”,即标签数与标量槽数;若当前选中了对象会优先列出。若没有任何带测量的 MultiROI,会显示 “(no MultiROI with measurements - create one first)”。
  • Refresh(刷新)按钮: 重新扫描会话,更新 MultiROI 下拉框。

Feature scalar slots(特征标量槽,勾选参与聚类的测量)

  • 标量槽列表: 一个带复选框的列表,列出所选 MultiROI 的每一个标量槽(显示其测量描述名)。默认全部勾选。请勾选想要作为聚类特征的测量项。
  • Select all(全选)/ Select none(全不选)按钮: 一键勾选或取消全部标量槽。

Clustering parameters(聚类参数)

  • Standardize features (StandardScaler) 复选框: 是否对特征做零均值、单位方差的标准化。默认勾选(开启)。当各标量槽单位/量级差异较大时建议开启。
  • Reducer(降维方法)下拉框: 可选 UMAP / t-SNE / PCA,默认 UMAP。用于把特征降到二维——既是散点图坐标,也是聚类实际所在的空间。
  • Clusterer(聚类算法)下拉框: 可选 KMeans / HDBSCAN,默认 KMeans。
  • KMeans n_clusters(簇数)数字框: KMeans 的目标簇数,范围 1–50,默认 3。仅当 Clusterer 选 KMeans 时可编辑。
  • HDBSCAN min_cluster_size(最小簇大小)数字框: HDBSCAN 判定为一个簇所需的最少标签数,范围 2–1000,默认 5;小于该规模的标签会被判为噪声(-1)。仅当 Clusterer 选 HDBSCAN 时可编辑。

切换 Clusterer 下拉框时,插件会自动启用/禁用对应的数字框:选 KMeans 时只有 n_clusters 可编辑,选 HDBSCAN 时只有 min_cluster_size 可编辑。

Environment(计算 venv)

  • Base Python 输入框: 覆盖搭建 venv 的基础解释器(留空=当前 Dragonfly 自带 Python;或填一个路径 / py -3.11)。
  • Status(状态)标签: 显示环境是否就绪——就绪时显示 “Ready:<venv python 路径>”;未就绪时提示先点 Setup Environment。

操作按钮与输出区

  • Setup Environment 按钮: 搭建 venv 并安装依赖(一次性)。
  • Cluster Measurements 按钮: 用当前选择的 MultiROI、勾选的标量槽和参数运行聚类。
  • 摘要标签: 计算完成后显示一行摘要(样本数、特征数、实际所用降维/聚类方法、簇数、噪声数)。
  • 日志文本框(只读): 显示运行过程与结果信息,包括写回的新标量槽编号和散点图 PNG 的保存路径。

6. 使用步骤

输入要求: 需要一个已经计算好逐标签测量的 MultiROI —— 也就是它上面至少有一个标量槽(scalar slot)。若 MultiROI 没有任何测量标量槽,插件会提示 “This MultiROI has no scalar (measurement) slots. Compute per-label measurements first.”,请先在 Dragonfly 中对该 MultiROI 计算逐标签测量。

端到端操作流程

1. 打开 Prototype Apps ▸ Cluster Measurements (UMAP/HDBSCAN)...。

2. (首次) 若 Status 显示未就绪,先点 Setup Environment 并等待日志出现 “Environment ready:...”。

3. 在 MultiROI 下拉框选择目标 MultiROI(若没看到,点 Refresh;并确认它带测量槽)。

4. 在 Feature scalar slots 列表中勾选要参与聚类的测量项(可用 Select all / Select none 快速切换)。

5. 视需要设置参数:是否 Standardize(默认开启)、Reducer(默认 UMAP)、Clusterer(默认 KMeans);KMeans 设 n_clusters(默认 3),HDBSCAN 设 min_cluster_size(默认 5)。

6. 点击 Cluster Measurements 开始计算;可在日志中观察进度。

7. 得到什么: 计算完成后,插件会把每个标签的聚类编号写回一个新的 MultiROI 标量槽(描述名为 “Cluster id”),日志给出该新槽的编号和散点图 PNG 的保存路径,摘要标签给出样本/特征/簇数等统计。

8. 回到 Dragonfly,把该 MultiROI 按这个新的 “Cluster id” 标量槽着色,即可看到各群组;打开保存的散点图 PNG 可查看二维分布。

当样本(标签)数量太少或某个可选库不可用时,插件会自动降级:UMAP / t-SNE 会退回 PCA,HDBSCAN 会退回 KMeans。摘要中的 “reducer_used / clusterer_used” 会如实反映实际使用的方法,可能与你选择的不同。

7. 参数说明

参数

默认值

说明

MultiROI

会话中第一个可用项

输入的多标签 ROI,必须带测量标量槽(逐标签测量)。

Feature scalar slots(特征标量槽)

全部勾选

勾选哪些标量槽作为聚类特征;每个勾选的槽是特征矩阵的一列。至少勾选一个。

Standardize features(标准化)

开启(勾选)

对每个特征做零均值、单位方差(StandardScaler);槽单位/量级不同时建议开启。

Reducer(降维方法)

UMAP

降到二维的方法:UMAP / t-SNE / PCA。样本太少或库不可用时自动退回 PCA。

Clusterer(聚类算法)

KMeans

聚类算法:KMeans(固定簇数)或 HDBSCAN(密度聚类,可产生噪声)。HDBSCAN 不可用时退回 KMeans。

KMeans n_clusters(簇数)

3

KMeans 的目标簇数,范围 1–50。仅 KMeans 生效;实际簇数会被限制在样本数以内。

HDBSCAN min_cluster_size(最小簇大小)

5

HDBSCAN 判定为一个簇所需的最少标签数,范围 2–1000;更小的群被判为噪声(-1)。仅 HDBSCAN 生效。

Base Python(基础解释器)

空(=Dragonfly 自带 Python)

搭建计算 venv 所用的基础 Python;可填绝对路径或 py -3.11 之类的启动器命令。

补充说明: t-SNE 需要至少 5 个样本才会启用(否则退回 PCA),其困惑度(perplexity)会依样本数自动取值;UMAP 的邻居数 n_neighbors 也会依样本数自动取(上限 15)。特征矩阵中的 NaN / inf 会在计算前被自动清零(置为 0)。这些都是内部自适应行为,无需手动设置。

8. 输出结果

一次成功的聚类会产生以下结果:

  • 新的 MultiROI 标量槽 “Cluster id”: 每个标签(1..N,背景标签 0 不参与)对应一个整数聚类编号(≥0 为某个簇;-1 表示噪声,只有 HDBSCAN 可能出现)。这个新槽被追加到 MultiROI 现有标量槽之后。插件写入后会设置该槽的窗宽/范围并发布(publish),因此新槽会即时出现在 Dragonfly 中。
  • 二维散点图 PNG: 一张按聚类着色的二维散点图(matplotlib 绘制),噪声点用灰色 “x” 标记,各簇用不同颜色;当簇数不超过 12 时附带图例。文件保存在临时作业目录中,日志会打印其完整路径。
  • 摘要信息: 一行文本摘要,包含样本数、特征数、实际所用的降维方法、实际所用的聚类算法、簇数与噪声数。

如何查看结果

1. 在 Dragonfly 中选中该 MultiROI,把它的着色标量槽切换为新写入的 “Cluster id” 槽——此时不同群组会以不同数值着色,在 2D/3D 视图里即可区分各群组。

2. 打开日志中给出的散点图 PNG 路径,查看二维空间中各群组的分布与分离程度。

3. 参考摘要行判断聚类是否合理(例如簇数是否符合预期、噪声是否过多),必要时调整参数重跑。

聚类编号本身没有 “大小/好坏” 的物理含义——它只是群组的标识。若要解释每个群组代表什么形态或成分,请结合原始测量(体积、球度等)对各群组做统计。

9. 常见问题与故障排除

问:MultiROI 下拉框是空的 / 显示 “no MultiROI with measurements”。

答:说明当前会话没有带测量标量槽的 MultiROI。请先在 Dragonfly 中对目标 MultiROI 计算逐标签测量(生成标量槽),再回到面板点 Refresh。

问:点 Cluster Measurements 报错 “This MultiROI has no scalar (measurement) slots”。

答:所选 MultiROI 上没有任何测量标量槽。本插件对已有测量做聚类,并不自己计算体积/球度等测量——请先在 Dragonfly 里计算逐标签测量。

问:提示环境未搭建 / “environment not set up”。

答:首次使用需先点 Setup Environment(需联网一次)。若已点过但仍提示未就绪,查看日志中的错误信息;可在 Base Python 填另一个 3.10+ 的 Python 后重试。

问:我选了 UMAP / HDBSCAN,但摘要里 reducer_used / clusterer_used 显示成 PCA / KMeans?

答:这是自动降级的正常行为。当样本(标签)数太少,或 umap-learn / hdbscan 库在当前环境不可用时,UMAP/t-SNE 会退回 PCA、HDBSCAN 会退回 KMeans。若确需 UMAP/HDBSCAN,请确认 Setup Environment 成功装上了对应的库,并保证有足够多的标签。

问:HDBSCAN 把大部分标签都标成了噪声(-1)。

答:说明数据中没有足够密集的群组,或 min_cluster_size 设得偏大。可减小 min_cluster_size,或改用 KMeans 指定固定簇数;也可尝试开启/关闭标准化、更换降维方法。

问:Setup Environment 下载依赖失败。

答:多为无法访问 PyPI(网络/代理问题)或所选基础 Python 没有对应 wheel。确认网络可达 PyPI,或在 Base Python 填另一个 CPython 3.10+ 解释器后重试。

10. 注意事项与已知限制

  • 只对已有测量聚类: 插件不计算测量本身,输入 MultiROI 必须已经带有逐标签测量标量槽。
  • 背景标签排除: 标签 0(背景)不参与聚类;聚类编号按标签 1..N 顺序写回新槽的对应位置。
  • 结果写回为新槽: 每次运行都会追加一个名为 “Cluster id” 的新标量槽,不会覆盖已有测量;多次运行会累积多个新槽。
  • 自动降级不可关闭: 样本太少或库缺失时会静默退回 PCA/KMeans,请以摘要中的 reducer_used / clusterer_used 为准。
  • 二维嵌入: 降维固定降到二维(便于绘图和聚类),不提供更高维嵌入选项。
  • 首次需联网: 仅 Setup Environment 阶段需要联网;之后离线可用。不需要 GPU、不需要 WSL、不依赖外部软件。
  • 随机性: 内部使用固定随机种子以尽量复现;但 t-SNE / UMAP 等方法对参数与样本量较敏感,不同数据下结果可能不同。
  • 菜单发现: 启用/停用或安装后必须完全重启 Dragonfly,菜单才会出现或消失。

11. 参考资料

  • scikit-learn(PCA / t-SNE / KMeans / StandardScaler):https://scikit-learn.org
  • UMAP(umap-learn):https://umap-learn.readthedocs.io
  • HDBSCAN(hdbscan):https://hdbscan.readthedocs.io
  • matplotlib:https://matplotlib.org
  • NumPy:https://numpy.org
  • 安装与启用总说明:参见 Full Package 根目录的 UserManual_用户手册.docx 总手册,以及 Developer ▸ Prototype Labs... ▸ Menu Item Manager。


Part II English Manual

Contents

1. Overview

2. Use Cases

3. Installation and Enabling

4. Runtime Environment and First-Run Setup

5. Interface Guide

6. Usage Steps

7. Parameter Reference

8. Outputs

9. FAQ and Troubleshooting

10. Notes and Known Limitations

11. References

1. Overview

Cluster Measurements (UMAP/HDBSCAN) is a Dragonfly plugin that performs unsupervised clustering of a MultiROI's per-label measurement table. Once you have computed per-label measurements (volume, surface area, sphericity, shape factor, orientation, ...) on the hundreds or thousands of labels in a MultiROI, those measurements live on the MultiROI as scalar slots. This plugin reads the slots you tick, treats each label as a feature vector, automatically sorts the labels into a handful of morphologically or compositionally similar clusters, writes the result back as a NEW MultiROI scalar slot, and saves a 2D scatter PNG.

The pipeline is: read the ticked scalar slots into a feature matrix (one row per label) -> optionally standardize (StandardScaler) -> embed to 2D with UMAP / t-SNE / PCA -> cluster with KMeans (fixed number of clusters) or HDBSCAN (minimum cluster size) -> write each label's cluster id back as a new scalar slot -> save a 2D scatter colored by cluster. Afterwards, simply color the MultiROI by that new slot in Dragonfly to see the groups.

Underlying engines and algorithms (all permissively licensed): dimensionality reduction and clustering use scikit-learn (PCA, t-SNE, KMeans, StandardScaler), umap-learn (UMAP reduction) and hdbscan (HDBSCAN density clustering); numerics rely on numpy; the 2D scatter is drawn with matplotlib (Agg head-less backend).

Licensing: the plugin code follows the DragonflyPrototypeLabs repository's terms; every compute engine is permissive -- scikit-learn (BSD-3-Clause), umap-learn (BSD-3-Clause), hdbscan (BSD-3-Clause), matplotlib (PSF/BSD-style), numpy (BSD).

2. Use Cases

The plugin is aimed at grouping particles, pores, grains, or cells in materials and life science: using per-label measurements (volume, surface area, shape factor, sphericity, orientation, ...) it automatically sorts thousands of labels into a few similar groups for classification, outlier detection and structural characterization.

  • Particle / pore grouping: sort CT-segmented pores or particles by volume, sphericity, and shape factor into large/small or spherical/plate-like groups for quantitative classification.
  • Grain characterization: group grains by their morphological measurements to spot anomalous grains or distinct grain families.
  • Cell grouping: in life-science imagery, group cells into morphological categories from volume, shape, etc.
  • Outlier detection: with HDBSCAN, labels that belong to no dense group are flagged as noise (id -1), making it easy to pick out anomalous structures.
  • Dimensionality-reduction visualization: even without strict grouping, the 2D scatter lets you see how the high-dimensional measurements distribute in 2D.

3. Installation and Enabling

The plugin ships as part of Prototype Apps in the Full Package. To install:

1. Unzip the Full Package to a short folder (e.g. C:\PL\), not a deep Downloads path or a OneDrive-redirected Desktop.

2. Double-click `Install_FullPackage.bat`.

3. In the dialog, pick the core install mode (Fresh / Compatible) and tick the Prototype Apps you want.

4. Click Install and wait for the console to finish.

5. Fully restart Dragonfly (menus are discovered only at startup).

This plugin is unchecked by default. In the installer's app list, the light menu items are ON but all plugins are OFF -- you must tick "Cluster Measurements (UMAP/HDBSCAN)" for it to be installed.

After restarting, the menu appears at Prototype Apps ▸ Cluster Measurements (UMAP/HDBSCAN)... (under the "Measurements & Analysis" section of that menu). Clicking it opens a dockable floating panel.

Changing your choices later: the easiest way is inside Dragonfly -- open Developer ▸ Prototype Labs... ▸ Menu Item Manager; the "Prototype Apps (Full Package)" list at the bottom has a checkbox per app: tick = deploy this plugin, untick = remove the menu entry. Restart Dragonfly to apply. Disabling never deletes a plugin's environment, so re-enabling is instant. You can also re-run the installer at any time (your previous choices become the defaults).

4. Runtime Environment and First-Run Setup

The heavy compute (scikit-learn / umap-learn / hdbscan / matplotlib) runs in a dedicated virtual environment (venv), never touching Dragonfly's own Python. Before first use you must click Setup Environment once in the panel.

What Setup Environment does

When you click Setup Environment, the plugin (1) uses Dragonfly's own Python as the base interpreter (by default) to build a venv inside the installed plugin code directory; (2) pip-installs scikit-learn + umap-learn + hdbscan + numpy + matplotlib from PyPI; (3) records the resulting venv Python path so the panel's "Status" line reads "Ready: <path>". The dependency versions are left unpinned (pip resolves wheels matching the base CPython).

Internet, GPU, disk and paths

  • Internet needed once: only while Setup Environment downloads the dependencies from PyPI; running clustering afterwards needs no internet.
  • No GPU required: everything is pure CPU (pure wheels on cp310, no CUDA).
  • No WSL and no external application are needed.
  • Install location: the venv is built inside the installed plugin code directory, i.e. %LOCALAPPDATA%\comet\<Dragonfly version>\pythonUserExtensions\GenericMenuItems\ClustersPlotter\venv. Reinstalling or updating the plugin code keeps that venv (it is never clobbered).
  • Download size: depends on the wheels for scikit-learn, scipy, numpy, matplotlib, umap-learn and hdbscan (typically on the order of tens to a couple hundred MB); the exact size varies with the base Python's wheels, so no precise figure is given here.

Base Python and fallback if setup fails

The panel's Base Python field overrides the base interpreter used to build the venv. Blank = this Dragonfly's own Python; otherwise enter an absolute path (to a CPython 3.10+ python.exe) or a launcher command (e.g. py -3.11). If building the venv with Dragonfly's own Python fails (e.g. pip cannot resolve a wheel), point Base Python at another CPython 3.10+ installed on your machine and retry.

If Setup Environment fails, the panel log prints the full traceback. The usual causes are no internet access to PyPI on the first build, or the chosen base Python having no matching pre-built wheel. Make sure PyPI is reachable, or switch Base Python to another 3.10+ interpreter and retry.

5. Interface Guide

The panel is a dockable floating window, laid out top to bottom in the following groups. Each control is described below, matching the code exactly:

Input MultiROI

  • MultiROI dropdown: lists the MultiROIs in the current session that carry measurement scalar slots. Each entry reads "title (N labels, M slots)" -- its label count and slot count; a currently selected object is listed first. If none carry measurements, it shows "(no MultiROI with measurements - create one first)".
  • Refresh button: rescans the session and updates the MultiROI dropdown.

Feature scalar slots (tick which to cluster on)

  • Slot list: a checkbox list of every scalar slot on the selected MultiROI (showing each measurement's description). All ticked by default. Tick the measurements you want to use as clustering features.
  • Select all / Select none buttons: tick or untick every slot at once.

Clustering parameters

  • Standardize features (StandardScaler) checkbox: zero-mean, unit-variance per feature. Checked (on) by default. Recommended when slots have different units/scales.
  • Reducer dropdown: UMAP / t-SNE / PCA, default UMAP. The 2D embedding used both for the scatter and as the space clustering runs in.
  • Clusterer dropdown: KMeans / HDBSCAN, default KMeans.
  • KMeans n_clusters spinbox: target number of KMeans clusters, range 1-50, default 3. Editable only when Clusterer is KMeans.
  • HDBSCAN min_cluster_size spinbox: the smallest group of labels that counts as a cluster, range 2-1000, default 5; smaller groups become noise (-1). Editable only when Clusterer is HDBSCAN.

Switching the Clusterer dropdown automatically enables/disables the matching spinbox: with KMeans only n_clusters is editable, with HDBSCAN only min_cluster_size is editable.

Environment (clustering venv)

  • Base Python field: overrides the base interpreter for the venv (blank = this Dragonfly's Python; or a path / py -3.11).
  • Status label: shows whether the environment is ready -- "Ready: <venv python path>" when set up, otherwise a prompt to click Setup Environment first.

Action buttons and output area

  • Setup Environment button: builds the venv and installs dependencies (one time).
  • Cluster Measurements button: runs clustering with the current MultiROI, ticked slots and parameters.
  • Summary label: after a run, shows a one-line summary (sample count, feature count, the reducer/clusterer actually used, number of clusters, number of noise labels).
  • Log text box (read-only): shows progress and results, including the new scalar-slot index written back and the saved scatter PNG path.

6. Usage Steps

Input requirement: a MultiROI that already has per-label measurements -- that is, at least one scalar slot. If the MultiROI has no measurement slots, the plugin reports "This MultiROI has no scalar (measurement) slots. Compute per-label measurements first.", so compute per-label measurements on it in Dragonfly beforehand.

End-to-end workflow

1. Open Prototype Apps ▸ Cluster Measurements (UMAP/HDBSCAN)....

2. (First time) if Status shows not ready, click Setup Environment and wait for the log to read "Environment ready: ...".

3. Pick the target MultiROI in the MultiROI dropdown (click Refresh if it is not listed, and make sure it has measurement slots).

4. In Feature scalar slots, tick the measurements to cluster on (use Select all / Select none to toggle quickly).

5. Set parameters as needed: Standardize (on by default), Reducer (UMAP by default), Clusterer (KMeans by default); set n_clusters (default 3) for KMeans or min_cluster_size (default 5) for HDBSCAN.

6. Click Cluster Measurements to start; watch progress in the log.

7. What you get: on completion, the plugin writes each label's cluster id back as a NEW MultiROI scalar slot (described as "Cluster id"); the log reports the new slot index and the scatter PNG path, and the summary label reports sample/feature/cluster statistics.

8. Back in Dragonfly, color the MultiROI by this new "Cluster id" slot to see the groups; open the saved scatter PNG to view the 2D distribution.

When there are too few samples (labels) or an optional library is unavailable, the plugin falls back automatically: UMAP / t-SNE fall back to PCA, and HDBSCAN falls back to KMeans. The summary's "reducer_used / clusterer_used" faithfully reflect what was actually used, which may differ from what you chose.

7. Parameter Reference

Parameter

Default

Description

MultiROI

first available in session

The input multi-label ROI; must carry measurement scalar slots (per-label measurements).

Feature scalar slots

all ticked

Which scalar slots to use as clustering features; each ticked slot is one column of the feature matrix. Tick at least one.

Standardize features

on (checked)

Zero-mean, unit-variance per feature (StandardScaler); recommended when slots have different units/scales.

Reducer

UMAP

The 2D embedding method: UMAP / t-SNE / PCA. Falls back to PCA when samples are too few or the library is unavailable.

Clusterer

KMeans

Clustering algorithm: KMeans (fixed count) or HDBSCAN (density, can produce noise). Falls back to KMeans when HDBSCAN is unavailable.

KMeans n_clusters

3

Target number of KMeans clusters, range 1-50. Applies only to KMeans; the effective count is capped at the sample count.

HDBSCAN min_cluster_size

5

Smallest group of labels that counts as a cluster, range 2-1000; smaller groups are labelled noise (-1). Applies only to HDBSCAN.

Base Python

blank (= Dragonfly's own Python)

Base interpreter for building the compute venv; can be an absolute path or a launcher command like py -3.11.

Additional notes: t-SNE requires at least 5 samples to run (otherwise it falls back to PCA), and its perplexity is chosen automatically from the sample count; UMAP's n_neighbors is also chosen automatically (capped at 15). NaN / inf values in the feature matrix are cleaned to 0 before computing. These are internal adaptive behaviors that need no manual tuning.

8. Outputs

A successful clustering run produces the following:

  • New MultiROI scalar slot "Cluster id": each label (1..N; background label 0 is excluded) gets an integer cluster id (>=0 for a cluster; -1 for noise, only possible with HDBSCAN). The new slot is appended after the MultiROI's existing scalar slots. After writing, the plugin sets the slot's window/range and publishes it, so the new slot appears in Dragonfly immediately.
  • 2D scatter PNG: a 2D scatter colored by cluster (drawn with matplotlib), with noise points shown as grey "x" markers and each cluster in a distinct color; a legend is included when there are at most 12 clusters. The file is saved in a temporary job directory and its full path is printed to the log.
  • Summary: a one-line text summary with the sample count, feature count, the reducer actually used, the clusterer actually used, and the number of clusters and noise labels.

How to view the results

1. Select the MultiROI in Dragonfly and switch its coloring scalar slot to the newly written "Cluster id" slot -- the different groups are then colored by value and can be distinguished in the 2D/3D views.

2. Open the scatter PNG at the path shown in the log to see how well-separated the groups are in 2D.

3. Read the summary line to judge whether the clustering is reasonable (e.g. whether the cluster count matches expectations, whether there is too much noise), and re-run with adjusted parameters if needed.

The cluster id itself has no physical "size/quality" meaning -- it is only a group label. To interpret what each group represents (morphology or composition), analyze each group's original measurements (volume, sphericity, etc.).

9. FAQ and Troubleshooting

Q: The MultiROI dropdown is empty / shows "no MultiROI with measurements".

A: The session has no MultiROI with measurement scalar slots. Compute per-label measurements on your MultiROI in Dragonfly first (to create scalar slots), then click Refresh in the panel.

Q: Cluster Measurements errors with "This MultiROI has no scalar (measurement) slots".

A: The selected MultiROI carries no measurement scalar slots. This plugin clusters existing measurements; it does not compute volume/sphericity itself -- compute per-label measurements in Dragonfly first.

Q: It says the environment is not set up.

A: First use requires Setup Environment (internet needed once). If you already ran it but it still reports not ready, check the log for the error; you can set Base Python to another 3.10+ Python and retry.

Q: I chose UMAP / HDBSCAN, but the summary shows reducer_used / clusterer_used as PCA / KMeans?

A: That is the expected automatic fallback. When there are too few samples (labels), or umap-learn / hdbscan is unavailable in the environment, UMAP/t-SNE fall back to PCA and HDBSCAN falls back to KMeans. If you truly need UMAP/HDBSCAN, make sure Setup Environment installed those libraries and that there are enough labels.

Q: HDBSCAN labels most of the labels as noise (-1).

A: The data has no sufficiently dense groups, or min_cluster_size is set too large. Lower min_cluster_size, or switch to KMeans with a fixed cluster count; you can also try toggling standardization or changing the reducer.

Q: Setup Environment fails to download the dependencies.

A: Usually PyPI is unreachable (network/proxy) or the chosen base Python has no matching wheel. Make sure PyPI is reachable, or set Base Python to another CPython 3.10+ interpreter and retry.

10. Notes and Known Limitations

  • Clusters existing measurements only: the plugin does not compute the measurements; the input MultiROI must already carry per-label measurement scalar slots.
  • Background excluded: label 0 (background) does not participate; cluster ids are written back for labels 1..N in order.
  • Result written as a new slot: each run appends a new "Cluster id" scalar slot and never overwrites existing measurements; repeated runs accumulate multiple new slots.
  • Automatic fallback cannot be disabled: with too few samples or a missing library it silently falls back to PCA/KMeans, so trust the summary's reducer_used / clusterer_used.
  • 2D embedding: reduction is fixed to 2D (for plotting and clustering); no higher-dimensional embedding option is offered.
  • Internet needed once: only during Setup Environment; afterwards it runs offline. No GPU, no WSL, no external application required.
  • Randomness: a fixed random seed is used for reproducibility, but methods like t-SNE / UMAP are sensitive to parameters and sample size, so results can vary across datasets.
  • Menu discovery: after enabling/disabling or installing, Dragonfly must be fully restarted before the menu appears or disappears.

11. References

  • scikit-learn (PCA / t-SNE / KMeans / StandardScaler): https://scikit-learn.org
  • UMAP (umap-learn): https://umap-learn.readthedocs.io
  • HDBSCAN (hdbscan): https://hdbscan.readthedocs.io
  • matplotlib: https://matplotlib.org
  • NumPy: https://numpy.org
  • General install/enable guidance: see the overview UserManual_用户手册.docx in the Full Package root, and Developer ▸ Prototype Labs... ▸ Menu Item Manager.
You’ve reached the end of this manual.Explore the library →