Remote DL Training(远程深度学习训练)插件用户手册
Remote DL Training - User Manual
Dragonfly Prototype Apps · Remote DL Training...
版本 Version 1.0 · 2026-07-04
第一部分 中文手册
目录
1. 简介
2. 适用场景
3. 安装与启用
4. 运行环境与首次配置
5. 界面说明
6. 使用步骤
7. 参数说明
8. 输出结果
9. 常见问题与故障排除
10. 注意事项与已知限制
11. 参考资料
1. 简介
Remote DL Training(远程深度学习训练) 是一个在 Dragonfly 内运行的可停靠面板,用于把耗时的深度学习语义分割模型训练任务提交到一台远程(或本机)Dragonfly 服务器上执行。您在客户端 Dragonfly 中准备好训练数据(一个图像 Channel 加一个作为标注真值的 MultiROI),在面板顶部填写远程服务器的地址、端口和访问令牌(token),即可把训练任务交给那台服务器,同时在本地实时查看训练进度、损失曲线和分割预览。
面板由两部分组成:顶部的“远程 Dragonfly 服务器”连接栏(设置地址 / 端口 / token,含 Apply 与 Test 按钮),以及下方与独立应用完全一致的 Remote DL Training 训练页(输入数据、超参数、模型操作、可视化反馈、训练与取消、日志)。所有网络请求都发送到您在顶部设置的那台服务器,因此一台没有 GPU 的电脑上的 Dragonfly 可以把训练交给另一台带 GPU 的 Dragonfly 服务器完成。
底层引擎: 训练使用 Dragonfly 自带的深度学习框架(PyTorch 后端),通过 Dragonfly 的 AI 接口创建并训练一个新的分割模型;客户端与服务器之间使用长度前缀 JSON 的 TCP 协议通信。可选的模型架构包括 2D 的 U-Net、Attention U-Net、UNet++、TransU-Net,以及 3D 的 U-Net 3D、Sensor3D。
许可与依赖: 本插件本身不引入任何第三方外部软件或独立虚拟环境(venv),完全运行在 Dragonfly 自带的 Python(PyQt6)之中。它的界面代码来自 gui_qt 包——与独立应用、嵌入式 Prototype Labs 面板共用,因此使用本插件前需要同时安装 Prototype Labs 嵌入式包(它提供 gui_qt)。深度学习训练本身依赖运行训练任务的那台 Dragonfly 的 AI 功能与许可。
与 Dragonfly 内置的 Prototype Labs 面板不同:那个面板强制连接本机回环地址且免认证;本插件则是“指向远程机器”的对应版本——连接远程服务器必须提供该服务器的 token,连接本机 127.0.0.1 则无需 token。
2. 适用场景
本插件专为把慢速的深度学习训练从本地转移到远程 GPU 服务器而设计。典型场景包括:
- 本地电脑没有 GPU 或性能有限,而实验室 / 机房里有一台配备 GPU 的 Dragonfly 服务器:在本地 Dragonfly 中准备好数据,一键提交到远程服务器训练。
- 通过 Tailscale 组网访问远程 GPU 机器(例如地址
100.64.0.2),或在局域网内访问(例如192.168.0.82)。 - 希望在训练进行时,于本地实时观察损失曲线(loss / val_loss)与分割预览图,判断模型是否收敛、是否需要提前停止。
- 训练完成后,直接下载模型文件、把模型导入本机 Dragonfly 的 AI 模型库,或一键应用到整卷图像得到分割结果。
- 在本机进行快速试验:可选择“本机 Headless 子进程”模式,在一个无窗口的 Dragonfly 进程中训练,无需另开服务器。
3. 安装与启用
本插件随 Prototype Labs & Apps 完整安装包(Full Package) 一起分发。安装步骤如下:
1. 将安装包解压到任意较短的目录(如 C:\PL\,避免路径过长报错)。
2. 双击 `Install_FullPackage.bat`。
3. 在弹出的对话框中选择核心安装模式(Fresh 全新安装 / Compatible 兼容安装,该选项只影响 Prototype Labs 核心的 blocks 与 recipes,不影响任何插件),并在插件列表中勾选 Remote DL Training。
4. 点击 Install,等待控制台完成。
5. 完全退出并重启 Dragonfly(菜单只在启动时扫描)。
Remote DL Training 默认未勾选(default_enabled = false)。安装时必须手动勾选它,否则重启后菜单里不会出现。另外请同时勾选并安装 Prototype Labs 嵌入式核心——本插件的界面来自它提供的 gui_qt 包,缺少时面板会显示无法加载 gui_qt 的错误。
重启后,插件出现在菜单:Prototype Apps ▸ Remote DL Training...(位于 B9_Remote Deep Learning 分组,与配套的 Remote DL Inferring 相邻)。点击即打开一个可停靠、可浮动的面板。
以后修改勾选: 最方便的方式是在 Dragonfly 内打开 Developer ▸ Prototype Labs... ▸ Menu Item Manager,在底部的 “Prototype Apps (Full Package)” 列表中勾选=部署、取消勾选=移除菜单项,重启 Dragonfly 生效。停用从不删除任何环境。也可随时重跑安装器(会记住上次的勾选作为默认值)。
安装位置(均在当前用户目录,无需管理员权限):
内容 | 位置 |
插件包(OrsPlugin) |
|
面板代码 |
|
界面依赖 gui_qt(Prototype Labs) |
|
完整包中央存储与启用状态 |
|
卸载: 双击完整安装包中的 `Uninstall_FullPackage.bat`,即可移除本插件的菜单项与文件。
4. 运行环境与首次配置
无需搭建环境、无需 Setup Environment。 本插件是“轻量型”插件(环境类型为 in_process),完全运行在 Dragonfly 自带的 Python 里,不下载任何东西、不创建 venv、不需要联网、不需要本机 GPU、不需要 WSL、不需要外部软件。它唯一的运行前置条件是:
- 已安装 Prototype Labs 嵌入式核心(提供
gui_qt界面代码)。面板启动时会自动从相邻的GenericMenuItems\Prototype_Labs目录、或环境变量DF_PROTOTYPE_LABS_DIR指向的目录解析gui_qt。 - 有一台正在运行的目标 Dragonfly 服务器——可以是远程带 GPU 的机器,也可以是本机;该服务器需已启动 Prototype Labs 的 TCP 服务器。
- 目标服务器的访问 token(仅远程连接必需;连接本机
127.0.0.1时不需要)。
首次配置(在面板顶部的连接栏完成):
1. 在 Address 中填入目标服务器地址(Tailscale IP 如 100.64.0.2、局域网 IP 如 192.168.0.82,或本机 127.0.0.1)。
2. 在 Port 中填入端口(默认 54321)。
3. 在 Token 中填入访问令牌(远程必填;本机可留空)。
4. 点击 Apply 让下方训练页使用该服务器;或点击 Test 先连接并 ping 一次以验证地址 / 端口 / token 是否正确。
5. 连接栏下方的状态行会显示当前目标,例如 Target: 100.64.0.2:54321 (remote - token set);若缺 token 会提示 remote - TOKEN REQUIRED。
出于安全考虑,地址与端口会被记住并跨会话保留,但访问 token 不会保存到磁盘,每次打开面板都需要重新输入。这是设计如此,不是缺陷。
注意 GPU 的位置: 训练所需的 GPU 是运行训练任务的那台机器(即目标服务器)的 GPU,而不是您本地这台运行界面的电脑。因此“本地无 GPU”正是本插件的用武之地。
5. 界面说明
面板自上而下分为连接栏和训练页两大块。训练页又采用左右两栏布局:左栏放数据 / 传输方式 / 执行位置 / 操作按钮 / 日志,右栏放模型操作与训练超参数。
5.1 顶部:远程 Dragonfly 服务器连接栏
- Address(地址): 目标 Dragonfly 服务器的 IP 或主机名。占位提示举例:
100.64.0.2 (Tailscale)/192.168.0.82 (LAN)/127.0.0.1。 - Port(端口): 数值框,范围 1–65535,默认
54321。 - Token(令牌): 密码框(输入以圆点隐藏)。远程服务器必填,本机可留空。
- Apply 按钮: 将当前地址 / 端口 / token 设为下方训练页使用的目标服务器。
- Test 按钮: 先应用当前设置,再连接并 ping 服务器一次,验证是否可达;结果显示在状态行(如
connected OK或cannot reach server: ...)。 - 状态行: 显示当前目标与 token 状态(本机免 token / 远程已设 token / 远程缺 token)。
5.2 训练数据(.ORSObject 文件)
- Image Channel(图像 Channel): 文本框 +
Browse...按钮,选择作为训练输入的图像 Channel.ORSObject文件。 - Ground-truth MultiROI(标注真值 MultiROI): 文本框 +
Browse...按钮,选择包含真值标签的 MultiROI.ORSObject文件。
5.3 服务器如何获取文件(传输方式)
- Upload to the Dragonfly server(上传到服务器,默认选中): 把文件上传给服务器,跨机器可用;每个文件须小于 200 MB。
- The server can read these paths directly(服务器直接读取路径): 适用于同一台机器或共享盘,不限文件大小。
5.4 训练在哪里运行(执行位置)
- On the connected Dragonfly (GUI Sync)(在已连接的 Dragonfly 上,默认选中): 把训练发送到面板顶部配置的服务器(带 GUI 且已启动服务器,或您自己启动的本机 headless 服务器)。
- Headless of local Dragonfly (subprocess)(本机 Headless 子进程): 在本机启动一个无窗口的 Dragonfly 进程来训练;直接读取所选文件,实时进度 / epoch / 可视化反馈均可用,取消可硬停止;之后的导入 / 应用 / 下载仍使用已连接的服务器。
- Headless on a Dragonfly Server (no GUI)(在无 GUI 的 Dragonfly 服务器上): 连接到您自行启动的无窗口 Dragonfly 服务器(可在另一台机器上)。旁边有一个 ? 帮助按钮,说明什么是 headless 服务器及如何启动。
选择“本机 Headless 子进程”时,传输方式框会被禁用(因为子进程直接读取本地文件,无需上传)。
5.5 模型操作(右栏)
- Save model in Dragonfly server(在服务器保存模型,复选框): 勾选=把训练好的模型保留在服务器 AI 模型库(名称带
RemoteTraining前缀);不勾选=只导出为临时文件供下载,不保留在库中。 - Auto apply the trained model to full image data(自动应用到整卷,复选框): 训练结束后(无论正常完成还是取消),用最终模型对整幅输入图像推理,并把分割 MultiROI 保存为可下载的
.ORSObject。 - Download Model File(下载模型文件,按钮): 把刚训练好的模型
.zip下载到本地(训练成功后可用)。 - Add to local Dragonfly(添加到本机 Dragonfly,按钮): 把刚训练好的模型导入已连接(本机)Dragonfly 的 AI 模型库,便于推理。
- Download Full Segmentation MultiROI(下载整卷分割 MultiROI,按钮): 下载训练后生成的整卷分割结果
.ORSObject。 - Select an image and apply the trained model(选择图像并应用模型,按钮): 选择另一幅图像(Channel
.ORSObject或原始图像文件),上传并用刚训练的模型推理,再下载分割 MultiROI 结果。
下载类与应用类按钮在训练成功前处于禁用状态,训练成功后自动启用。
5.6 模型与训练参数(右栏)
各控件详见第 7 章参数表。包含架构、类别数、Epochs、Patch 尺寸、Stride 比例、批大小、损失函数、优化器、验证集、数据增强、GPU 数量、模型名称等。
5.7 可视化反馈(可选,左栏可折叠区)
- Show visual feedback(显示可视化反馈,复选框): 每隔 N 个 epoch,在指定 ROI 区域上应用当前模型并显示分割预览与损失曲线。它会把训练分成若干小段进行,与一次性连续训练相比,优化器动态会略有不同。
- Preview region (ROI)(预览区域 ROI): 文本框 +
Browse...,选择定义预览区域的 ROI.ORSObject。 - Preview every N epochs(每 N 个 epoch 预览一次): 数值框,范围 1–1000,默认
1。 - Segmentation preview(分割预览图): 左半区显示训练中每次预览的分割结果图。
- Loss curve(损失曲线): 右半区绘制
loss与val_loss曲线,双击可放大为独立窗口。
可视化反馈的实时预览需要界面与训练进程共享文件系统:本机 Headless 子进程始终满足;GUI Sync / headless 服务器仅当服务器在本机(回环地址)时满足。远程服务器无法做实时预览——此时可视化反馈会被禁用,但训练仍正常进行,结果在完成后可用。
5.8 操作区与日志(左栏)
- Train model(训练模型,按钮): 开始训练。
- Cancel(取消,按钮): 请求停止训练;GUI Sync / 服务器模式会在下一个安全点停止,本机 Headless 子进程会被硬停止。仅训练中可用。
- Status(状态)标签: 显示 Idle / Connecting / Starting / Training epoch X/Y / Trained / Cancelled / Failed 等。
- 进度行 / 日志区: 实时进度文本与详细训练日志(等宽字体)。训练在服务器工作线程上运行(
exec_mode=current_thread),因此不会冻结 Dragonfly 界面。
6. 使用步骤
端到端流程:在远程 GPU 服务器上训练一个分割模型
1. 在准备阶段,先在 Dragonfly 中把训练图像导出为 Channel .ORSObject,把标注真值导出为 MultiROI .ORSObject(可选:再准备一个定义预览区域的 ROI .ORSObject)。
2. 打开 Prototype Apps ▸ Remote DL Training...。
3. 在顶部连接栏填入远程服务器的 Address / Port / Token,点击 Test 验证可达,再点击 Apply。
4. 在“训练数据”中用 Browse... 选择 Image Channel 与 Ground-truth MultiROI 文件。
5. 选择传输方式:跨机器请选 Upload to the Dragonfly server(注意单文件 < 200 MB);同机或共享盘可选直接读取。
6. 在“训练在哪里运行”中保持 On the connected Dragonfly (GUI Sync)(即刚配置的服务器)。
7. 在右栏“模型与训练参数”中设置架构、Epochs、Patch 尺寸等超参数(默认值即可开始)。
8. (可选)勾选 Save model in Dragonfly server 以在服务器库中保留模型,和/或勾选 Auto apply the trained model to full image data 以在训练后自动分割整卷。
9. (可选,仅本机/共享盘)展开“可视化反馈”,勾选 Show visual feedback,选择预览 ROI 并设置预览间隔。
10. 点击 Train model。在日志区观察进度,在损失曲线区观察收敛情况;必要时点击 Cancel。
11. 训练成功后,右栏按钮启用:用 Download Model File 保存模型、Add to local Dragonfly 导入本机、Download Full Segmentation MultiROI 下载整卷结果、或 Select an image and apply the trained model 应用到其他图像。
替代流程:本机快速试验(无需服务器)
1. 顶部地址可保持 127.0.0.1(本机免 token)。
2. 在“训练在哪里运行”中选择 Headless of local Dragonfly (subprocess);此时传输方式框自动禁用。
3. 选择 Channel 与 MultiROI,设置参数,点击 Train model——本机会启动一个无窗口 Dragonfly 进程进行训练,实时进度与可视化反馈均可用。
7. 参数说明
下表列出“模型与训练参数”区各控件的取值范围与默认值(默认值取自一个全新的 Dragonfly U-Net 模型)。
参数 | 默认值 | 说明 |
Architecture(架构) | U-Net | 模型架构。可选:2D 的 |
Class count(类别数,0=自动) | 0 | 输出类别数;0 表示自动,使用 MultiROI 的标签数量。范围 0–999。 |
Epochs(训练轮数) | 100 | 训练的 epoch 数。范围 1–100000。 |
Patch size(Patch 尺寸) | 64 | 正方形训练 patch 的边长(体素)。最小 32,步长 16。 |
Stride ratio(Stride 比例) | 1.0 | patch 采样步长占 patch 尺寸的比例。范围 0.05–1.0,步长 0.05;越小则 patch 越多、重叠越多。 |
Batch size(批大小,0=自动) | 0 | 每步 patch 数;0 表示自动(根据 GPU 显存选择,遇到显存不足会自动减半)。范围 0–4096。 |
Loss function(损失函数) | (model default) | 分割损失函数。可选 |
Optimizer(优化器) | (model default) | 优化器。可选 |
Use validation split(使用验证集) | 开启 | 训练时留出一部分数据作验证(用于最佳检查点 / 提前停止)。 |
Validation %(验证比例) | 20 % | 开启验证集时留作验证的数据百分比。范围 1–90。 |
Apply data augmentation(数据增强) | 开启 | 应用 Dragonfly 默认的数据增强(翻转、旋转、缩放、错切)。 |
GPU count(GPU 数量) | 1 | 训练使用的本地 GPU 数量(深模型可多 GPU)。范围 1–16。 |
Model name(模型名称) | (自动) | 创建模型的名称;留空则由架构 + 类别数自动生成。 |
Preview every(预览间隔) | 1 | 可视化反馈:每多少个 epoch 生成一次预览。范围 1–1000。 |
此处的 GPU 数量指运行训练任务的那台服务器上的 GPU,与您本地界面机器无关。
8. 输出结果
训练成功后,本插件可产生以下结果:
- 训练好的分割模型: 若勾选了 Save model in Dragonfly server,模型以
RemoteTraining前缀名保留在服务器 AI 模型库中;否则导出为临时.zip供下载。 - 模型 `.zip` 文件: 通过 Download Model File 下载到本地磁盘。
- 导入本机的模型: 通过 Add to local Dragonfly,把模型加入本机 Dragonfly 的 AI 模型库,之后可用 Dragonfly 的“应用深度学习分割模型”功能推理。
- 整卷分割 MultiROI: 若勾选了 Auto apply the trained model to full image data,训练后自动对整卷图像推理并保存为
.ORSObject,通过 Download Full Segmentation MultiROI 下载。 - 对任意图像的分割结果: 通过 Select an image and apply the trained model,对另一幅所选图像推理并把分割 MultiROI 结果
.ORSObject下载到本地。 - 训练历史: 面板内的损失曲线(loss / val_loss)与日志区的最终损失、轮数等信息。
如何查看: 把下载或导入的模型加入 Dragonfly AI 模型库后,可在 Dragonfly 中对新图像应用它;把下载的分割 MultiROI .ORSObject 在 Dragonfly 中打开即可查看和编辑分割标签。
9. 常见问题与故障排除
问:面板打开后显示“无法找到 gui_qt 包”或加载失败,怎么办?
答:本插件的界面来自 Prototype Labs 嵌入式核心提供的 gui_qt 包。请确认已在完整安装包中勾选并安装了 Prototype Labs 嵌入式核心,使其位于本插件相邻的 GenericMenuItems\Prototype_Labs 目录;或设置环境变量 DF_PROTOTYPE_LABS_DIR 指向包含 gui_qt 的目录。安装后重启 Dragonfly。
问:点击 Test 提示“cannot reach server”(无法连接服务器)?
答:请依次检查:目标 Dragonfly 服务器是否正在运行且已启动 Prototype Labs 服务器;Address 与 Port 是否正确(默认端口 54321);远程连接是否填写了正确的 Token(远程必填,状态行会提示 TOKEN REQUIRED);以及网络 / 防火墙 / Tailscale 是否允许该端口连通。
问:菜单里找不到 Remote DL Training?
答:该插件默认未勾选。请重跑完整安装器并勾选它(或在 Developer ▸ Prototype Labs... ▸ Menu Item Manager 中勾选),然后完全重启 Dragonfly——菜单只在启动时扫描。它位于 Prototype Apps 菜单的 B9_Remote Deep Learning 分组。
问:为什么“显示可视化反馈”选项是灰的(不可勾选)?
答:实时预览需要界面与训练进程共享文件系统。当您连接的是另一台机器上的远程服务器时,没有共享路径,因此可视化反馈被禁用,但训练仍正常运行,结果在完成后可用。若要使用可视化反馈,请在本机训练(选择“本机 Headless 子进程”,或让服务器在本机)。
问:上传文件时报错或文件太大?
答:上传模式下每个文件必须小于 200 MB。若数据更大,请把文件放到服务器可直接访问的位置(同机或共享盘),并选择 The server can read these paths directly 传输方式,即可不限大小。
问:提示“Dragonfly 服务器正忙”?
答:目标服务器同一时刻只能处理一个任务。请等待当前任务完成后重试。
问:重开面板后 token 不见了?
答:这是有意的安全设计——地址与端口会保留,但 token 不会写入磁盘,每次会话需重新输入。
10. 注意事项与已知限制
- 依赖 Prototype Labs 嵌入式核心: 缺少
gui_qt时面板无法加载,请务必同时安装。 - 远程连接必须提供 token: 连接
127.0.0.1才免 token。 - GPU 在服务器端: 训练速度取决于目标服务器的 GPU,而非本地界面机器。
- 上传大小限制: 上传模式下单文件须 < 200 MB;更大数据请用直接路径方式。
- 可视化反馈仅本机/共享盘可用: 远程服务器无法做实时预览。
- 可视化反馈会分段训练: 与一次性连续训练相比,优化器动态会略有差异。
- 服务器单任务: 目标服务器一次只处理一个任务,忙时需等待。
- 取消行为: GUI Sync / 服务器模式在下一安全点停止(可能无法立即中断某个 epoch 内部);本机 Headless 子进程为硬停止。
- token 不落盘: 每次会话需重新输入访问令牌。
11. 参考资料
- 配套插件:Remote DL Inferring(远程深度学习推理)——把已训练模型应用到图像;两者位于同一
B9_Remote Deep Learning分组。 - 完整安装 / 启用 / 卸载说明:随完整安装包的
README.md(Prototype Labs & Apps — Full Package)。 - 本机 headless 服务器的用途与启动方法:面板中“在无 GUI 的 Dragonfly 服务器上”选项旁的 ? 帮助按钮。
- 深度学习分割、模型库与推理:Dragonfly 官方帮助文档(Deep Learning / AI 分割相关章节)。
Part II English Manual
Contents
1. Overview
2. Use Cases
3. Installation & Enabling
4. Runtime Environment & First-Run Setup
5. Interface Guide
6. Step-by-Step Usage
7. Parameter Reference
8. Outputs
9. FAQ & Troubleshooting
10. Notes & Known Limitations
11. References
1. Overview
Remote DL Training is a dockable panel that runs inside Dragonfly and lets you offload slow deep-learning semantic-segmentation model training to a remote (or local) Dragonfly server. You prepare the training data in the client Dragonfly (one image Channel plus a ground-truth MultiROI), enter the remote server's address, port and access token at the top of the panel, and submit the job to that server while watching the progress, loss curve and segmentation preview locally in real time.
The panel has two parts: a "Remote Dragonfly server" bar at the top (Address / Port / Token with Apply and Test buttons), and, below it, the exact Remote DL Training tab from the standalone app (inputs, hyper-parameters, model manipulation, visual feedback, train/cancel, and a log). Every request goes to the server you set at the top, so a GPU-less machine's Dragonfly can hand training to another GPU-equipped Dragonfly server.
Underlying engine: training uses Dragonfly's own deep-learning framework (PyTorch backend) through Dragonfly's AI interface to create and train a new segmentation model; the client and server communicate over a length-prefixed-JSON TCP protocol. Available model architectures are the 2D U-Net, Attention U-Net, UNet++, TransU-Net, and the 3D U-Net 3D, Sensor3D.
License & dependencies: the plugin itself introduces no third-party external software and no separate virtual environment (venv); it runs entirely in Dragonfly's own Python (PyQt6). Its UI comes from the gui_qt package shared with the standalone app and the embedded Prototype Labs panel, so you must also install the Prototype Labs embedded package (which provides gui_qt). Training itself relies on the AI capability and license of whichever Dragonfly runs the job.
Unlike Dragonfly's built-in Prototype Labs window (which forces a no-auth loopback connection to the local Dragonfly), this plugin is the "point at a remote box" counterpart: a remote connection requires that server's token, while a 127.0.0.1 connection needs none.
2. Use Cases
The plugin is designed to move slow DL training off the local machine onto a remote GPU server. Typical scenarios:
- Your local machine has no GPU or limited power, but there is a GPU-equipped Dragonfly server in the lab: prepare the data locally, submit to the remote server with one click.
- Reach a remote GPU box over Tailscale (e.g. address
100.64.0.2) or on a LAN (e.g.192.168.0.82). - Watch the loss curve (loss / val_loss) and a segmentation preview locally while training to judge convergence and decide whether to stop early.
- After training, directly download the model file, import it into the local Dragonfly AI model library, or apply it to the full volume in one click.
- Quick local experiments: choose the local headless subprocess mode to train in a windowless Dragonfly process with no separate server.
3. Installation & Enabling
This plugin ships in the Prototype Labs & Apps Full Package. To install:
1. Unzip to any short folder (e.g. C:\PL\, to avoid path-too-long errors).
2. Double-click `Install_FullPackage.bat`.
3. In the dialog, pick the core install mode (Fresh / Compatible - this affects only the Prototype Labs core blocks & recipes, not any plugin) and tick Remote DL Training in the plugin list.
4. Click Install and wait for the console to finish.
5. Quit Dragonfly completely and restart it (menus are scanned only at startup).
Remote DL Training is OFF by default (default_enabled = false); you must tick it during install or it will not appear after restart. Also tick and install the Prototype Labs embedded core - this plugin's UI comes from its gui_qt package, and without it the panel shows a "could not find gui_qt" error.
After restart the plugin appears under Prototype Apps ▸ Remote DL Training... (in the B9_Remote Deep Learning group, next to the companion Remote DL Inferring). Clicking it opens a dockable, floatable panel.
Changing your choices later: the easiest way is inside Dragonfly - open Developer ▸ Prototype Labs... ▸ Menu Item Manager and use the checkboxes in the bottom "Prototype Apps (Full Package)" list (ticked = deploy, unticked = remove the menu entry); restart Dragonfly to apply. Disabling never deletes any environment. You may also re-run the installer at any time (it remembers your previous choices).
Install locations (all per-user, no admin rights needed):
Content | Location |
Plugin package (OrsPlugin) |
|
Panel code |
|
UI dependency gui_qt (Prototype Labs) |
|
Full Package store & enable state |
|
Uninstall: double-click `Uninstall_FullPackage.bat` from the Full Package to remove the menu entry and files.
4. Runtime Environment & First-Run Setup
No environment to build, no Setup Environment step. This is a lightweight plugin (environment kind in_process): it runs entirely in Dragonfly's own Python, downloads nothing, creates no venv, and needs no internet, no local GPU, no WSL and no external software. Its only prerequisites are:
- The Prototype Labs embedded core is installed (provides the
gui_qtUI code). On startup the panel resolvesgui_qtautomatically from the siblingGenericMenuItems\Prototype_Labsfolder or from theDF_PROTOTYPE_LABS_DIRenvironment variable. - A running target Dragonfly server - a remote GPU box or the local machine; it must have the Prototype Labs TCP server started.
- The target server's access token (required for remote connections only; not needed for
127.0.0.1).
First-run configuration (done in the top connection bar):
1. Enter the target server's Address (Tailscale IP such as 100.64.0.2, LAN IP such as 192.168.0.82, or 127.0.0.1).
2. Enter the Port (default 54321).
3. Enter the Token (required for remote; leave blank for local).
4. Click Apply so the tab below uses that server, or click Test to connect and ping once to verify the address / port / token.
5. The status line below the bar shows the current target, e.g. Target: 100.64.0.2:54321 (remote - token set); a missing token shows remote - TOKEN REQUIRED.
For security, the address and port are remembered across sessions, but the access token is NOT saved to disk - re-enter it each session. This is by design, not a bug.
Where the GPU lives: the GPU that matters is on the machine that runs the training (the target server), not on your local UI machine. That is exactly why this plugin is useful when the local machine has no GPU.
5. Interface Guide
Top to bottom, the panel is a connection bar over a training tab. The training tab uses a two-column layout: the left column holds data / transfer mode / execution target / action buttons / log; the right column holds model manipulation and training hyper-parameters.
5.1 Top: Remote Dragonfly server bar
- Address: the target Dragonfly server's IP or hostname. Placeholder examples:
100.64.0.2 (Tailscale)/192.168.0.82 (LAN)/127.0.0.1. - Port: spin box, range 1-65535, default
54321. - Token: password field (shown as dots). Required for a remote server; may be blank for local.
- Apply button: sets the current address / port / token as the target for the tab below.
- Test button: applies the current settings, then connects and pings the server once to check reachability; the result appears on the status line (e.g.
connected OKorcannot reach server: ...). - Status line: shows the current target and token state (local needs no token / remote token set / remote token required).
5.2 Training data (.ORSObject files)
- Image Channel: text field +
Browse...button; select the image Channel.ORSObjectfile used as training input. - Ground-truth MultiROI: text field +
Browse...button; select the MultiROI.ORSObjectholding the ground-truth labels.
5.3 How should the server get the files? (transfer mode)
- Upload to the Dragonfly server (default): uploads the files; works across machines; each file must be < 200 MB.
- The server can read these paths directly: for the same machine or a shared drive; any file size.
5.4 Where should training run? (execution target)
- On the connected Dragonfly (GUI Sync) (default): sends training to the server configured in the top bar (a GUI Dragonfly with the server started, or a headless server you started yourself on this machine).
- Headless of local Dragonfly (subprocess): spawns a windowless Dragonfly process on THIS machine to train; it reads the selected files directly; live progress / epoch / visual feedback all work and Cancel hard-stops it; import / apply / download afterwards use the connected server.
- Headless on a Dragonfly Server (no GUI): connects to a headless Dragonfly server you started yourself (possibly on another machine). A ? help button next to it explains what a headless server is and how to start one.
When "Headless of local Dragonfly" is chosen, the transfer-mode box is disabled (the subprocess reads local files directly, so no upload is needed).
5.5 Model manipulation (right column)
- Save model in Dragonfly server (checkbox): checked = keep the trained model in the server's AI model library (named with a
RemoteTrainingprefix); unchecked = export to a temporary file for download but do not keep it in the library. - Auto apply the trained model to full image data (checkbox): after training ends (whether it finishes or you cancel), run the final model on the FULL input image and save the segmentation MultiROI as a downloadable
.ORSObject. - Download Model File (button): download a
.zipof the just-trained model to your computer (enabled after a successful train). - Add to local Dragonfly (button): import the just-trained model into the connected (local) Dragonfly's AI model library, ready for inference.
- Download Full Segmentation MultiROI (button): download the full-volume segmentation result
.ORSObjectproduced after training. - Select an image and apply the trained model (button): choose another image (a Channel
.ORSObjector a raw image file), upload it, run the just-trained model on it, and download the resulting segmentation MultiROI.
The download and apply buttons are disabled until a train succeeds, then enable automatically.
5.6 Model & training parameters (right column)
See the parameter table in Chapter 7 for each control. It includes architecture, class count, epochs, patch size, stride ratio, batch size, loss function, optimizer, validation split, data augmentation, GPU count and model name.
5.7 Visual feedback (optional, collapsible, left column)
- Show visual feedback (checkbox): every N epochs, apply the current model to a chosen ROI region and show the segmentation preview + loss curve. It trains in short segments, so optimizer dynamics differ slightly from one continuous run.
- Preview region (ROI): text field +
Browse...; select the ROI.ORSObjectdefining the preview region. - Preview every N epochs: spin box, range 1-1000, default
1. - Segmentation preview: the left half shows the segmentation preview image at each interval.
- Loss curve: the right half plots
lossandval_loss; double-click to enlarge into a separate window.
The live preview needs a filesystem SHARED between the UI and the training process: the local headless subprocess always qualifies; GUI Sync / a headless server qualify only when the server is on this machine (loopback host). A remote server cannot do a live preview - visual feedback is then disabled, but training still runs normally and results are available on completion.
5.8 Actions & log (left column)
- Train model (button): starts training.
- Cancel (button): requests a stop; GUI Sync / server modes stop at the next safe point, while the local headless subprocess is hard-stopped. Enabled only while training.
- Status label: shows Idle / Connecting / Starting / Training epoch X/Y / Trained / Cancelled / Failed, etc.
- Progress line / log pane: live progress text and a detailed training log (monospace). Training runs on the server worker thread (
exec_mode=current_thread), so the Dragonfly GUI stays responsive.
6. Step-by-Step Usage
End-to-end: train a segmentation model on a remote GPU server
1. Prepare data: in Dragonfly, save the training image as a Channel .ORSObject and the ground-truth labels as a MultiROI .ORSObject (optionally also a ROI .ORSObject that defines a preview region).
2. Open Prototype Apps ▸ Remote DL Training....
3. In the top bar, enter the remote server's Address / Port / Token, click Test to verify reachability, then click Apply.
4. Under Training data, use Browse... to pick the Image Channel and Ground-truth MultiROI files.
5. Choose a transfer mode: for a different machine pick Upload to the Dragonfly server (note the < 200 MB per-file limit); for the same machine or a shared drive pick direct read.
6. Under "Where should training run?" keep On the connected Dragonfly (GUI Sync) (the server you just configured).
7. In the right-column parameters, set architecture, epochs, patch size, etc. (defaults are fine to start).
8. (Optional) tick Save model in Dragonfly server to keep the model in the server library, and/or Auto apply the trained model to full image data to segment the full volume after training.
9. (Optional, local / shared drive only) expand "Visual feedback", tick Show visual feedback, pick a preview ROI and set the interval.
10. Click Train model. Watch progress in the log and convergence in the loss curve; click Cancel if needed.
11. After a successful train, the right-column buttons enable: use Download Model File to save, Add to local Dragonfly to import, Download Full Segmentation MultiROI for the full-volume result, or Select an image and apply the trained model for other images.
Alternative: quick local experiment (no server)
1. Leave the top address as 127.0.0.1 (local, no token).
2. Under "Where should training run?" choose Headless of local Dragonfly (subprocess); the transfer-mode box is disabled automatically.
3. Pick the Channel and MultiROI, set parameters and click Train model - a windowless Dragonfly process trains on this machine with live progress and visual feedback available.
7. Parameter Reference
The table lists the range and default of each control in "Model & training parameters" (defaults are live-probed from a fresh Dragonfly U-Net).
Parameter | Default | Description |
Architecture | U-Net | Model architecture. Options: 2D |
Class count (0 = auto) | 0 | Number of output classes; 0 = auto (use the MultiROI's label count). Range 0-999. |
Epochs | 100 | Number of training epochs. Range 1-100000. |
Patch size | 64 | Square training patch edge length (voxels). Minimum 32, step 16. |
Stride ratio | 1.0 | Patch sampling stride as a fraction of patch size. Range 0.05-1.0, step 0.05; lower = more, overlapping patches. |
Batch size (0 = auto) | 0 | Patches per step; 0 = auto (chosen from GPU memory; halved automatically on out-of-memory). Range 0-4096. |
Loss function | (model default) | Segmentation loss. Options |
Optimizer | (model default) | Optimizer. Options |
Use validation split | On | Hold out part of the data for validation (needed for best-checkpoint / early stop). |
Validation % | 20 % | Percent of data held out for validation when validation is on. Range 1-90. |
Apply data augmentation | On | Apply Dragonfly's default augmentation (flips, rotation, zoom, shear). |
GPU count | 1 | Number of local GPUs used for training (multi-GPU for deep models). Range 1-16. |
Model name | (auto) | Name for the created model; blank = auto from architecture + class count. |
Preview every | 1 | Visual feedback: generate a preview every this many epochs. Range 1-1000. |
GPU count refers to the GPUs on the machine that RUNS the training (the server), not your local UI machine.
8. Outputs
After a successful train, the plugin can produce:
- A trained segmentation model: if Save model in Dragonfly server was ticked, the model is kept in the server's AI model library with a
RemoteTrainingprefix; otherwise it is exported to a temporary.zipfor download. - A model `.zip` file: downloaded to your local disk via Download Model File.
- A locally imported model: via Add to local Dragonfly, added to the local Dragonfly AI model library, then usable with Dragonfly's "apply a deep-learning segmentation model" for inference.
- A full-volume segmentation MultiROI: if Auto apply the trained model to full image data was ticked, the full volume is segmented after training and saved as an
.ORSObject, downloaded via Download Full Segmentation MultiROI. - A segmentation of any chosen image: via Select an image and apply the trained model, another image is inferred and its segmentation MultiROI
.ORSObjectis downloaded locally. - Training history: the in-panel loss curve (loss / val_loss) plus final loss and epoch count in the log.
How to view: after downloading or importing the model into Dragonfly's AI model library, apply it to new images in Dragonfly; open a downloaded segmentation MultiROI .ORSObject in Dragonfly to view and edit the segmentation labels.
9. FAQ & Troubleshooting
Q: The panel opens but says it cannot find the gui_qt package, or fails to load. What now?
A: This plugin's UI comes from the gui_qt package provided by the Prototype Labs embedded core. Make sure you ticked and installed the Prototype Labs embedded core in the Full Package so it sits in the sibling GenericMenuItems\Prototype_Labs folder, or set the DF_PROTOTYPE_LABS_DIR environment variable to a folder containing gui_qt. Then restart Dragonfly.
Q: Clicking Test says "cannot reach server".
A: Check, in order: is the target Dragonfly server running with the Prototype Labs server started; are the Address and Port correct (default port 54321); did you enter the correct Token for a remote connection (required; the status line shows TOKEN REQUIRED); and do network / firewall / Tailscale allow that port.
Q: I can't find Remote DL Training in the menu.
A: The plugin is OFF by default. Re-run the Full Package installer and tick it (or tick it in Developer ▸ Prototype Labs... ▸ Menu Item Manager), then restart Dragonfly completely - menus are scanned only at startup. It lives under the Prototype Apps menu, B9_Remote Deep Learning group.
Q: Why is "Show visual feedback" greyed out?
A: The live preview needs a filesystem shared between the UI and the training process. When you connect to a remote server on another machine there is no shared path, so visual feedback is disabled - but training still runs and results are available on completion. To use visual feedback, train on this machine (choose the local headless subprocess, or have the server on this machine).
Q: Uploading errors out or the file is too large.
A: In upload mode each file must be < 200 MB. For larger data, put the files where the server can read them directly (same machine or a shared drive) and choose The server can read these paths directly - then there is no size limit.
Q: It says the Dragonfly server is busy.
A: The target server handles one task at a time. Wait for the current task to finish and try again.
Q: My token is gone after reopening the panel.
A: This is deliberate security: the address and port persist, but the token is not written to disk and must be re-entered each session.
10. Notes & Known Limitations
- Depends on the Prototype Labs embedded core: without
gui_qtthe panel cannot load, so install both. - Remote connections require a token: only
127.0.0.1needs no token. - The GPU is on the server: training speed depends on the target server's GPU, not your local UI machine.
- Upload size limit: in upload mode each file must be < 200 MB; use direct paths for larger data.
- Visual feedback is local/shared-drive only: a remote server cannot do a live preview.
- Visual feedback trains in segments: optimizer dynamics differ slightly from one continuous run.
- Single-task server: the target server handles one task at a time; wait when busy.
- Cancel behavior: GUI Sync / server modes stop at the next safe point (may not interrupt mid-epoch); the local headless subprocess is hard-stopped.
- Token not persisted: re-enter the access token each session.
11. References
- Companion plugin: Remote DL Inferring - apply a trained model to images; both live in the same
B9_Remote Deep Learninggroup. - Full install / enable / uninstall instructions: the
README.mdshipped with the Full Package (Prototype Labs & Apps - Full Package). - What a headless Dragonfly server is and how to start one: the ? help button next to the "Headless on a Dragonfly Server (no GUI)" option in the panel.
- Deep-learning segmentation, the model library and inference: Dragonfly's official Help documentation (Deep Learning / AI segmentation sections).