Measurements & AnalysisChinese & English

Segmentation QA: Compare Masks in 3D

Segmentation QA (3D) answers one concrete question: how good is this segmentation? It treats one published MultiROI or ROI as the ground truth and another as the segmentation to evaluate, then reports the standard agreem

Updated 2026-08-06User manual

Segmentation QA (3D)(三维分割质量评估)

Segmentation QA (3D) - User Manual

Dragonfly Prototype Apps · Segmentation QA: Compare Masks in 3D...

版本 Version 1.0 · 2026-07-31


第一部分 中文手册

目录

1. 简介

2. 两个层次的指标,以及为什么不报准确率

3. 配对方式:贪心还是最优

4. 边界距离(可选)

5. 操作步骤

6. TP/FP/FN 叠加结果

7. 环境需求与限制

1. 简介

Segmentation QA (3D) 用来回答一个很具体的问题:我这次分割做得有多好? 它把一个已发布的 MultiROI 或 ROI 当作标准答案(ground truth),把另一个当作待评估的分割结果,然后给出一整套业界通用的一致性指标 —— 并且是在真三维下计算的。

本插件重新实现了 ImageJ 插件 MiC(MIT 许可,https://github.com/MultimodalImagingCenter/MiC)的指标体系。MiC 的 README 写着「目前仅支持 2D」,但那句话相对它自己的代码已经过时 —— MiC 实际注册了一个「Mask instant Comparator 3D」菜单项,所以不要把「三维」当成差异化卖点。本插件真正的价值:跑在 Dragonfly 内部、直接处理已发布的 MultiROI/ROI、结果接进 Dragonfly 自己的表格/叠加/报告,并且多给了世界单位的边界距离(ASSD / Hausdorff,MiC 不计算)。「逐层比较」选项保留,便于与 ImageJ 原版对照。

两个输入必须来自同一图像网格(形状完全一致)。如果不一致,插件会直接报错并告诉你两个形状分别是多少 —— 而不是悄悄给你一组错的数字。

2. 两个层次的指标,以及为什么不报准确率

体素层面:TP(两边都认为是前景)、FP(只有待评估结果认为是前景,即多分割)、FN(只有标准答案认为是前景,即漏分割),再由此得出 Precision = TP/(TP+FP)、Recall = TP/(TP+FN)、Jaccard(IoU) = TP/(TP+FP+FN)、F1 = 2TP/(2TP+FP+FN)(F1 与 Dice 系数是同一个数)。

这里故意不报「准确率」和「特异度」。 它们都需要 TN(两边都认为是背景的体素),而在一个体数据里背景体素往往占 99% 以上 —— 于是不管分割多糟糕,含 TN 的指标都会显示 0.99 以上。报这种数字只会误导人,所以本插件不提供。

对象层面:先把标准答案里的每个物体和待评估结果里的每个物体按 IoU 做一对一配对,然后在一串 IoU 阈值上统计:配对成功且 IoU ≥ 阈值的算 TP,没配上的预测物体算 FP,没配上的标准答案物体算 FN。默认阈值是 0.50 到 0.95、步长 0.05(COCO 与 StarDist/Cellpose 的惯例)。由此给出 AP = TP/(TP+FP+FN)、其在整个阈值区间上的平均 mAP,以及全景质量 PQ =(配对 IoU 之和)/(TP + FP/2 + FN/2)。

mAP 只在两边都是带标签的 MultiROI 时才有意义。 如果输入的是二值 ROI,整个前景会被当成一个物体,此时只看体素层面的指标。

3. 配对方式:贪心还是最优

默认按 IoU 从高到低的贪心配对(也是 ImageJ 原版 MiC 的做法),当 IoU 阈值 ≥ 0.5 时贪心可证明是最优的 —— 一个预测物体不可能同时与两个标准答案物体的 IoU 都超过 0.5。结果面板会写出这次用的是哪种规则。

「最大化 IoU 总和(Hungarian)」默认关闭,且不建议开启:它优化的是 IoU 的总和,而本页指标数的是超过阈值的配对个数 —— 两者目标不同。实测过一个例子:IoU 分别为 0.5556 / 0.2857 / 0.2857 时,最大化总和会舍弃 0.5556 那一对(因为 0.2857+0.2857 更大),于是在 IoU 0.50 上报出 TP/FP/FN = 0/2/2、mAP 0.0000,而正确答案非零。

4. 边界距离(可选)

勾选后额外计算 ASSD(平均对称表面距离)、95% Hausdorff 和最大 Hausdorff,并且按体素间距换算成世界单位(各向异性的体素也算得对)。

为什么值得算:两张掩膜可以有 0.9 的 Dice,却处处偏了整整一个体素 —— 体积类指标看不出来,边界距离一眼就看出来。这一项 ImageJ 原版没有,是三维评估里特别值得看的一项。需要 scipy;任一掩膜为空时不会给出 0,而是明确报告「未计算」。

5. 操作步骤

1. 先把两个要比较的分割结果发布(Publish)到会话里 —— 插件只列出已发布的对象。

2. 打开 Prototype Apps ▸ Segmentation QA: Compare Masks in 3D...,在第 1 页选好「标准答案」和「待评估的分割结果」。若列表是空的,点「刷新」。

3. 需要的话调整 IoU 阈值区间,勾选「逐层比较」「边界距离」「发布叠加结果」,然后点「开始比较」。

4. 第 2 页看结果:上方一行是概览(体素 Dice、Jaccard、物体个数、mAP、配对方式),下面是完整指标表和各阈值下的对象指标表。

5. 第 3 页选一个输出文件夹,导出 Excel/CSV/JSON 和曲线图。

6. TP/FP/FN 叠加结果

勾选「完成后发布 TP/FP/FN 叠加结果为 MultiROI」时,插件会新建并发布一个 MultiROI,里面有三个标签:TP(一致)绿色、FP(多分割)红色、FN(漏分割)蓝色。

这是整个插件最实用的部分:数字告诉你「差多少」,叠加结果告诉你「差在哪」。红色成片出现在物体边缘 → 分割整体偏胖;蓝色出现在小目标位置 → 模型漏掉了小目标;两个相邻物体之间没有蓝色分界 → 模型把它们粘连成了一个。

叠加结果与源对象共用体素间距和原点,因此在 2D/3D 视图里与原图严格对齐。若两者完全一致(没有任何 FP/FN 体素),插件会明确告知而不是发布一个空对象。

7. 环境需求与限制

无需 GPU、无需联网、无需管理员权限,也不需要安装任何环境:全部在 Dragonfly 内部用自带的 numpy 运行。最优配对与边界距离需要 scipy,Excel 导出需要 openpyxl,曲线图需要 matplotlib —— 这些 Dragonfly 都已自带。

限制:两个输入必须同形状;时间序列请逐个时间点比较(插件读取当前时间点);对象层面的指标依赖标签,二值 ROI 只能做体素层面的评估。


Part II English Manual

Contents

1. Introduction

2. Two levels of metric, and why accuracy is not reported

3. Matching: greedy or optimal

4. Boundary distances (optional)

5. How to use it

6. The TP/FP/FN overlay

7. Requirements and limits

1. Introduction

Segmentation QA (3D) answers one concrete question: how good is this segmentation? It treats one published MultiROI or ROI as the ground truth and another as the segmentation to evaluate, then reports the standard agreement metrics - computed over the whole 3D volume.

This reimplements the metric set of the ImageJ plugin MiC (MIT, https://github.com/MultimodalImagingCenter/MiC). MiC’s README says it works only in 2D, but that line is stale relative to its own code — MiC ships a registered “Mask instant Comparator 3D” entry, so 3D is not the differentiator. What this plugin actually offers is running inside Dragonfly on published MultiROI/ROI objects, with results wired into Dragonfly’s own tables, overlays and reports, plus boundary distances in world units (ASSD / Hausdorff), which MiC does not compute. The per-slice mode is kept so results can be compared with the ImageJ original.

The two inputs must come from the same image grid (identical shape). A mismatch is reported as an explicit error naming both shapes - never as silently wrong numbers.

2. Two levels of metric, and why accuracy is not reported

Voxel level: TP (both call it foreground), FP (only the evaluated mask does - over-segmentation), FN (only the ground truth does - missed), then precision = TP/(TP+FP), recall = TP/(TP+FN), Jaccard (IoU) = TP/(TP+FP+FN) and F1 = 2TP/(2TP+FP+FN), which is the same number as the Dice coefficient.

Accuracy and specificity are deliberately NOT reported. Both need TN, the voxels both sides call background - and in a volume that is often over 99% of the data, so any TN-based score reads above 0.99 no matter how bad the segmentation is. Reporting it would only mislead.

Object level: every ground-truth object is matched one-to-one with a predicted object by IoU, then counted across a sweep of IoU thresholds: a matched pair with IoU >= threshold is a TP, an unmatched prediction is an FP, an unmatched ground-truth object is an FN. The default sweep is 0.50 to 0.95 in steps of 0.05 (the COCO and StarDist/Cellpose convention). From those counts: AP = TP/(TP+FP+FN), its mean over the sweep mAP, and panoptic quality PQ = (sum of matched IoU) / (TP + FP/2 + FN/2).

mAP is only meaningful when BOTH inputs are labelled MultiROIs. With a binary ROI the whole foreground is one object, so only the voxel-level metrics apply.

3. Matching: greedy or optimal

Matching is GREEDY by descending IoU by default (what the ImageJ MiC original does), and at IoU thresholds >= 0.5 greedy is provably optimal - one prediction cannot exceed 0.5 IoU with two ground-truth objects at once. The results panel always states which rule produced the numbers.

"Maximise TOTAL IoU (Hungarian)" is off by default and not recommended: it optimises the SUM of IoU, while these metrics count PAIRS ABOVE THE THRESHOLD - a different objective. Measured case: with IoUs 0.5556 / 0.2857 / 0.2857 it discards the 0.5556 pair (0.2857+0.2857 is a larger sum) and reports TP/FP/FN = 0/2/2 and mAP 0.0000 at IoU 0.50, where the correct answer is non-zero.

4. Boundary distances (optional)

When ticked, the plugin also reports ASSD (average symmetric surface distance), 95% Hausdorff and maximum Hausdorff, scaled by the voxel spacing so the numbers are world units even for anisotropic voxels.

Why it is worth having: two masks can share a Dice of 0.9 and still be a whole voxel off everywhere. Volume-overlap metrics cannot see that; a boundary distance shows it immediately. Upstream MiC has no equivalent. Needs scipy; if either mask is empty the result is reported as "not computed" rather than as 0.

5. How to use it

1. Publish both segmentations into the session - the plugin lists published objects only.

2. Open Prototype Apps > Segmentation QA: Compare Masks in 3D... and pick the ground truth and the prediction on tab 1. Press Refresh if the lists are empty.

3. Adjust the IoU sweep if needed, tick per-slice comparison / boundary distances / publish overlay, then press Compare now.

4. Tab 2 shows a one-line headline (voxel Dice, Jaccard, object counts, mAP, matching method), the full metric table, and the per-threshold object table.

5. Tab 3 exports Excel/CSV/JSON and the plots to a folder of your choice.

6. The TP/FP/FN overlay

With "Publish a TP/FP/FN overlay MultiROI when done" ticked, a new MultiROI is created and published with three labels: TP (agree) green, FP (extra) red, FN (missed) blue.

This is the most useful part of the plugin: the numbers say how much is wrong, the overlay says where. Red rims around every object means the segmentation is systematically too fat; blue where the small objects are means they were missed; no blue boundary between two touching objects means the model merged them.

The overlay inherits the source spacing and origin, so it lines up exactly in the 2D/3D views. If the two segmentations agree everywhere (no FP or FN voxel at all) the plugin says so instead of publishing an empty object.

7. Requirements and limits

No GPU, no internet, no admin rights and nothing to install: everything runs inside Dragonfly on the bundled numpy. Optimal matching and boundary distances need scipy, Excel export needs openpyxl, plots need matplotlib - all shipped with Dragonfly.

Limits: the two inputs must have the same shape; for a time series compare one time point at a time (the plugin reads the current one); object-level metrics need labels, so a binary ROI supports voxel-level evaluation only.

You’ve reached the end of this manual.Explore the library →